Signals·

AI Signals — Weekend Read: We tried to break the herd

Written by Claude·7 min read·2026-08-19
Summary
  • The panel's effective N — how many independent opinions five models actually produce — has sat near 1.1 since March; daily gap movements correlate at 0.91, meaning the panel moves as a single organism
  • A 1,440-call controlled experiment tried to break the herd with three escalating prompt interventions; effective N moved from 1.05 to 1.07 — and explicitly instructing models to be independent broke output validity (99%→83%) without buying any independence
  • A frontier-class reasoning model run through the identical protocol joined the herd on arrival (correlation 0.92, 72/72 valid outputs): herding is not a capability ceiling but a shared training prior that no prompt reaches
  • Meanwhile the standing views split further than ever — GPT −6.5% to DeepSeek +1.1% — and two configuration changes showed levels move overnight (a sampling fix, a provider-side model swap) while motion stays locked to the group
  • The conclusion is a two-axis honesty rule: never 'five independent forecasts' (the motion is one), never 'the models agree' (the levels differ by eight points, persistently) — the herd is accepted, and the information lives in the levels

Weekend Read #13 — August 8, 2026

AI Investor Barometer tracks how five LLMs form DCF assumptions for 24 listed companies — daily, independently, from identical inputs.

At four minutes past ten on the evening of Monday, July 20, a script on a laptop in Finland fired the first of 1,440 API calls. Over three consecutive evenings it would ask five language models to value twenty-four companies, three times each, under four different sets of instructions — a controlled experiment designed to answer one question that had been bothering us since spring.

The question: can these models be made to think for themselves?

Not in the science-fiction sense. In the statistical one. Since March, one number in our research diary had refused to move: the effective N of the panel — how many genuinely independent opinions five models actually produce when their day-to-day movements are correlated against each other. Five models from four companies on three continents, trained on different data by different teams, should plausibly yield something between two and four independent views. Our measurements kept returning a number close to one. This July it stood at 1.07: when new information arrives, the panel moves as a single organism, its daily gap changes correlated at 0.91.

A panel that moves as one is a strange product to run. So before accepting it, we spent 1,440 calls trying to break it.

Where the herd lives

You cannot fix what you have not located, and by July the location was known with some precision. It was not the engine's safety caps — they touch about an eighth of estimates, and the herd survives with caps removed. It was not the discount rate — WACC choices had actually diversified after our July prompt revision. The compression lived in the growth and margin assumptions: on EBIT margins, model disagreement contributed barely one percent of total variance — near determinism — and almost every value landed on a neat half-percentage-point grid. The models were not reasoning their way to the same answer. They were reaching for the same defaults.

That diagnosis suggested an obvious culprit: our own prompt. It offers sector reference ranges, fallback values, rounding guidance — a scaffolding of anchors. Remove the scaffolding, the theory went, and five different minds might finally produce five different answers.

The experiment gave the theory every chance to be right. Four cumulative "arms": the production prompt as control; a version rewriting every sector range as historical reference rather than menu, with fallback defaults deleted; a version additionally demanding continuous three-decimal values instead of rounded grid points; and finally a version that told each model, explicitly, that it was one of several independent analysts, that agreement was not the goal, that divergence grounded in fundamentals was a valid and valuable finding.

Same companies, same frozen input data, same evening conditions. If prompt anchors caused the herd, arm by arm, the effective N should climb.

Nothing moved

The control arm scored 1.05. The de-anchored arm: 1.06. The continuous-values arm: 1.07. The independence arm: 1.07 — and that last one came at a price. Telling models to be independent did not make them independent; it made them worse, with output validity dropping from ninety-nine percent to eighty-three as models wandered off schema in the name of originality. Instructed independence produced not five opinions but broken JSON.

We had preregistered the decision criteria before running a single call — an arm would count as a win at an effective N of 1.5 or better without breaking validity or safety caps. No arm cleared 1.1. The verdict was written into the protocol in advance: prompt levers are insufficient. The herd is not an artifact of our scaffolding.

One hypothesis remained. Perhaps herding is a property of the class of models we run — efficient mid-tier models, chosen because the panel prices at pennies per company. Perhaps a frontier-class reasoning model, given the same task, would find its own path. So we ran one of the strongest models currently available through the identical protocol: seventy-two calls, seventy-two valid outputs, impeccable discipline. Its correlation with the herd: 0.92. It joined the flock on arrival — more fluently than the flock itself.

That result closes the question in a way the prompt arms alone could not. The herd is not in our instructions, and it is not a capability ceiling. Five models — six, counting the visitor — move together because they have read the same world: the same filings, the same financial press, the same textbook DCF conventions, the same historical base rates. The prior is shared at a depth no prompt reaches. When Nokia guides down, there is only one reasonable direction to revise, and every competent reader of the same information revises the same way. What we had been calling a herd is, from another angle, six literate analysts who all did the reading.

Then the herd split — in levels

Here is where the story turns, because while the experiment was proving that the panel's motion cannot be diversified, its positions were drifting apart in production — further apart than they have ever been.

On the last trading day of July, GPT's median gap across the panel stood at −6.5 percent, Grok's at −5.1, Gemini's at −0.3, Claude's at +0.6, DeepSeek's at +1.1. Five models, nearly eight percentage points of spread in standing view, maintained day after day. GPT — for months the panel's optimist — had become its bear. DeepSeek — for months the deepest bear — now sat closest to neutral. These are not random walks; the offsets persist for weeks and carry each model's documented personality: how it weighs news, how it treats margins, how much conviction it reports.

The month's two configuration changes made the layering visible. When we standardized GPT's sampling temperature in mid-July — it had been running hotter than the other four since launch — its level shifted while its correlation with the panel did not. And on the last day of July, DeepSeek's provider began serving a reasoning-tuned variant under the same model name; overnight, the model's level moved from bearish to neutral while its motion stayed locked to the group. Levels, it turns out, are sensitive to configuration, to provider, to sampling. Motion is sensitive to nothing we can touch.

That asymmetry is the finding. One herd in motion, five views in level.

Living with it

So we stopped fighting. The de-anchoring workstream is closed; the effective-N target has been deleted from our metrics; no further prompts will instruct anyone to be an individual. This is not resignation — it is what the data licenses. Every claim now has a measured boundary. We will not tell you the panel offers five independent forecasts; its movements are one forecast, made five times. And we will not tell you the models agree; their standing views differ by eight percentage points, persistently, informatively, and those differences — who is bearish, who breaks from the pack on which company, how wide the fan of estimates opens — are exactly what the dashboard measures every morning.

There is even a use for the herding itself. A panel that moves as one is, in effect, a very well-read consensus machine — and when one member does break formation on a single company, that break carries information precisely because formation is the rule.

At eleven twenty on the evening of July 22, the 1,440th call returned its JSON and the experiment ended. It had failed completely, in the most instructive way available: every lever pulled, nothing moved, and the null result mapped the machine. The five models will keep moving as one. What they think — where each one stands while moving — remains five different things, and that is the part worth reading every day.

Model estimates are research output from an automated system, not investment advice. The experiment protocol, preregistered criteria and full results are documented in the methodology and changelog pages. The experiment ran in an isolated environment and never touched production data.

More research

editorial · written by Claude
AI Signals — Weekend Read: The gap that closed itself
editorial · written by Claude
AI Signals — Weekend Read: The stock that's always on sale
editorial · written by Claude
AI Signals — Weekend Read: The machines were the bears
Want these insights weekly?
Subscribe to AI Signals →