Most of what a model writes never reaches its answer.

These are experiments on models anyone can download and run: forty-six of them, from a small one with 160 million internal parameters up to one with 14 billion, built by fourteen different groups on several different designs. Each result below had to survive a test that could have shown it was wrong. What follows is what survived, and what did not. The work was carried out between June and August 2026, and every figure here was checked against the run record on 6 August 2026.

Lattice of components with a causal readout plane
across the panel

Almost nothing a model writes reaches its answer

Switch off the thin slice of a model that points at its answer and the answer changes 60 percent of the time. Switch off everything else it wrote, which is almost all of it, and the answer barely moves.

As a model works, its parts write into a shared running scratchpad, and the answer is read off one direction in it. Between 95 and 99 percent of everything written points away from that direction and cancels out before it gets there. The thin remainder is what carries the decision, which is why removing it changes the answer and removing the rest does not. That held in 12 of 14 runs across 8 model families, in both the raw models and the chat-tuned versions. The other two runs are the same model.

The bulk that cancels out is not stray noise. It has structure, it repeats from one position to the next, and removing it disturbs the answer about three times as much as removing the same quantity of random activity. Yet nothing downstream reads it. Its only consumer is the readout at that same position, which is blind to it by construction.

That blindness is the point. A model carries far more concepts than it has separate dimensions to keep them in, an arrangement Anthropic set out in 2022, and this is how the packing stays readable: the part that has to be read is kept thin and clean, and the rest is put where the readout cannot see it. Each individual write lands about one percent of its size on the direction that decides, and the answer is set by thousands of those slivers adding up.

The geometry is not mine. Stolfo and colleagues showed in 2024 that some neurons write into exactly this blind spot, adjusting a model's confidence without changing what it says. What I measured is that same geometry applied to a live choice between two competing answers, one word at a time. It holds well past arithmetic: of everything the model collectively writes, the share landing on the deciding direction is 0.012 for recalling a fact, 0.012 for sorting something into a category and 0.058 for arithmetic, across all six models.

That model is gemma-2-9b, and it does the opposite: it puts its largest write straight onto the readout. Both its raw and chat-tuned versions behave this way, which is what accounts for the remaining two runs. It does something related elsewhere: asked whether a two-digit answer falls in the tens or the twenties, it writes that commitment straight onto the readout, where Qwen makes the identical commitment and hides it. Gemma is the only model in the panel that is distilled, and the only one with a particular numerical cap on its outputs, so which of those is responsible has not been established. Where a model chooses to put a decision is apparently a design choice a family can make differently, and that is the open question here.

Between 95 and 99 percent of what a model's parts write points away from the answer and cancels out, in 12 of 14 runs across 8 model families
Measured June to July 2026. Switching parts off and putting them back, six to eight models per result.
across the panel

Choosing and writing are done by different parts

The parts that write an answer are interchangeable. Switch them off and others write the same answer instead. Whatever picked it is somewhere else: earlier, smaller, and separable.

The parts that write an answer behave like general-purpose pens. They point in fixed directions, only about a quarter of them are reused from one context to the next, and they are replaceable: switch them off and other parts write the same answer instead. That backfilling is a documented effect, first described in 2023 as models compensating for a removed layer.

What picks the answer is separate, earlier and small. Switching parts off one at a time and keeping whatever collapses the decision isolates a chooser of between 0.3 and 4.8 percent of the network, 2 percent on average across six models. Remove it and the decision collapses while the writing machinery is left intact. Putting pieces back one at a time places the choosing about 27 layers ahead of the writing, in all eight models tested. A separate line of work has found that a compact, portable representation of the task is carried by a handful of parts, which is the same division seen from the other side.

For recalling a fact the split is clean: either side can be broken without breaking the other, in all eight. For arithmetic with several numbers in it, the split is not clean. The size of the answer gets assembled inside the writing machinery itself. And one model, Qwen2.5, shows no separation at all. That disagreement is part of the result rather than an outlier to drop.

Neither result names a particular part of a particular model, and that turns out to be the pattern. Where the computation forces a solution, different models arrive at the same one. Where it does not, what you are looking at is a training recipe: how much of the work is duplicated, where a value gets carried, whether the machinery that pushes against an answer sits in one place or is spread across many. Several of those recipe differences are large, and they do not track which lab built the model or what it was trained from. That was tested directly and did not hold.

Measured June to July 2026. Six to eight models per result.
the instrument is the suspect

The most-used tool for dating a decision reads it late

Three cases where something that looked like a discovery about the model turned out to be a fact about the tool being used on it.

There is a popular technique that reads a model's internal state part-way through and decodes what it would say if it stopped there. It is widely used to date the moment a model makes up its mind. Measured against actually intervening in the model, it is late.

Swapping a piece of one run into another and seeing whether the answer follows pins three moments, given as fractions of the way through the model. The answer is fixed at about 0.64. The decoding technique can first see it at about 0.76. The model stops building and starts simply transcribing at about 0.86.

So the tool trails the decision by 30 to 48 percent of the model's depth. Anything dating a model's decision with it is dating the readout, not the decision, and the layers after that point are transcription rather than deliberation. The pattern runs in one direction only, survives changing the measurement basis entirely, and shows up in 83 to 92 percent of twelve different prompts per model.

The answer is causally fixed at about 0.64 of the way through the model, the usual decoding tool can first read it at about 0.76, and building gives way to transcription at about 0.86

The other two are smaller and of the same kind. A common test uses a fixed cut-off to decide whether an effect counts; on three of seven models that cut-off sat below the level of ordinary background noise, so whatever it reported, it could not tell the two apart. And a standard shortcut is to find the part that scores highest on some measurement and call it the part doing the work; on five models checked by switching parts off directly, the highest-scoring part was never the one that mattered.

Measured July 2026. Twelve prompts per model.
Counts of tests that could not have detected the effect they were looking for
where the margin is thin

When almost nothing separates two answers, almost anything decides

Llama-3.1-8B was asked how many legs three spiders and two cats have between them. The answer is 32. It said 20, and one census traced every step of how it got there.

For the spiders it answered twelve, which is three times four. It very nearly answered twenty-four. Both numbers were pushed hard toward the output, near the top of every number it could have produced, and the right one was never suppressed. The gap between them stayed close to zero for 31 of the model's 32 stages, with twenty-four ahead at eight separate points along the way. Entering the last stage the lead was a fifth of what it ended up being. That one stage supplied 85 percent of the final margin.

Switch off the fifty cells in that stage pushing hardest for twelve, which is a third of one percent of it, and the model answers twenty-four instead. Random sets of the same size never do that, at any size up to two thousand. This is not hidden machinery for getting spiders wrong: aim the same lever at any other close runner-up and the answer lands there instead. It is what a decision looks like when almost nothing separates the options.

The cats it counted correctly, two at four legs each, so the arithmetic was never the problem. It then summed its own displayed working, the wrong twelve included, instead of going back to the question. Twelve plus eight is twenty, so the answer is internally consistent and externally wrong. Having committed to it, the model reaches a sentence boundary, deliberates, and resolves not to a correction but to the next item in its list. It never revisits the error. The problem is exhausted by then, and it opens a fourth item anyway, with nothing grounded to put in it.

A worked sum built on a wrong count: three spiders at four legs each gives twelve where it should give twenty-four, two cats at four legs gives eight correctly, and the model added its own working to reach twenty where the answer is thirty-two

The same shape shows up somewhere with no right answer at all. Almost every method in this field needs an answer key to check against; take it away, as you must for any question about values, and the tools have nothing left to score. Three things stay checkable anyway: whether the answer changes when you alter something that should not matter, whether the reason the model gives is the reason it actually used, and which of its parts ran.

The test set is somebody else's, published by Scherrer and colleagues: 680 genuinely hard dilemmas and 687 easy ones as a built-in control. Present two courses of action as A and B, then swap them and present the same two as B and A. A model that has decided something picks the same action both times. On the easy items all six models manage that 88 percent of the time. On the hard ones, 58 percent, barely above the 50 you would get by guessing. One model drops to 21 percent and simply takes whichever action was printed first, 95 percent of the time.

Watching the model write its answer out, rather than freezing it at one moment, sharpens the picture. On an easy item the decision is already present before it produces a single word, and it never reverses. On a hard one there is nothing there to begin with. The model assembles the decision as it writes, and about a third of the time it swaps sides partway through its own answer.

The sharpest number is how little it takes. Which option is printed first changes roughly 42 percent of the hard answers, and the amount of information that ordering actually carries is about 0.015 bits. A bit is one yes-or-no question's worth, so under a fiftieth of one is deciding the outcome. Across nine hundred cases, the ones that flipped were the ones sitting closest to a tie.

Note what none of this buys you. A model can be perfectly consistent and still hold values you would reject, and the standing mistake in this area is letting consistent quietly stand in for good.

Census run June 2026, re-verified against its own artifacts 6 August 2026. Judgment work August 2026.
Consistency after swapping the two options: 88 percent on easy items, 58 percent on hard ones Printing order changes about 42 percent of hard choices while carrying about 0.015 bits
calibration

The literature's positives held. Its negatives dissolved.

Twelve well-known results were re-run on ten different models, against a comparison much harder than the one they were originally published against. All twelve survived. Reports of finding nothing did not fare nearly as well.

Every experiment needs a comparison, something you expect to come out empty so that a real effect stands out against it. The twelve were tested against the easy comparison of the original papers and against a hard one, with the pass mark locked in before the test ran so it could not quietly move its own bar afterwards. All twelve held up. When somebody in this field reports that they found something, it is more likely to be real than you would guess, and that was worth finding out.

Reports that somebody found nothing are a different matter. Most never establish that their test could have detected the thing in the first place. Give them a test that could, and the nothing often turns into a something. That gap is the useful result: believe this field's positive findings more than you expect to, and its negative ones considerably less. A test that cannot see anything reports nothing either way.

All twelve canonical results survived the hard comparison, and none of the twelve needed the easy one
Audit run July 2026. Ten models, pass mark fixed by hash before the test.
what did not survive

The same standard applies here.

Three things that did not hold: a count of what was left unexplained, a whole reading of what that gap meant, and the strongest normative result on the page.

Around twenty investigations that looked like they were turning up new machinery turned out to be finding about twelve already-documented things, over and over. A claim I made about how much was left unexplained was too strong, and then both of the tests I built to check that claim turned out to be broken: one relied on a shortcut that structurally cannot see the thing it was hunting for, the other was set up so badly it scored a textbook example at zero.

A published stress test argues that parts which pass the usual checks routinely fail to transfer to other prompts. Rather than cite it, I ran it against my own results. It split them by model: the claim survived on Llama, was distributed on Qwen, and on Gemma a figure of mine went from 63 percent to about 20, so I withdrew it.

The bigger casualty was a reading of my own. Between a quarter and three-fifths of what a model writes at each step has no specific name attached to it, and I took that gap for a store of undiscovered machinery. Hunted across eighteen models and three different levels of detail, it dissolved into measurement artifacts plus computation that is genuinely spread out rather than hidden. There is no trove. The fraction also turns out to be a property of the model rather than a constant of the method, running from 0.250 to 0.596 across six models.

Moral Foundations Theory is a well-known account from psychology which holds that human moral thinking runs on five separate concerns. If a model had learned five genuinely separate things, you would expect to find five independent directions inside it. In all five models tested there are closer to two and a half. That is not a refutation of the theory: the published work measures whether each concern can be read out on its own, which is a different question, and one nobody had asked this way.

Those directions were then tested on real text that people had labelled by hand and the model had never seen. Score them from 0.5, no better than a coin flip, to 1.0, perfect. They reach 0.61. A crude method that only counts which words appear scores exactly 0.5, so they are picking up something past vocabulary. But 0.61 is under the 0.7 mark I have criticised other people's work for failing to clear.

Moral Foundations occupy about 2.5 independent directions of a predicted 5, and the measured directions reach 0.61 where just counting words reaches 0.5
Retractions logged June to August 2026. Literature checked against source 10 July 2026. Charts plot reported endpoints only; where a value is approximate it is marked approximate, and no intermediate points are inferred.