Mini-lesson 03 / toward the global-workspace paper · prev: what the lens reads
From an instrument to a claim
Lesson 2 gave us a lens that turns a hidden activation into words. Finding word-like readouts is suggestive — but a picture is not a mechanism. Four functional tests are what promote “the lens finds verbalizable representations” into “these representations form a global workspace.”
The Jacobian lens shows that certain mid-stack activations decode into clean, nameable tokens — Italy, euro, a concept mid-formation. Tempting to stop there and declare the model “thinks in words.” But a decoder can find structure that the model itself never uses. The readout might be an artifact of the lens, not a load-bearing part of the computation.
The question is not “can we read a word off this activation?” but “does the model act on that word — report it, be steered by it, reason over it, generalize it?”
That distinction — between a representation you can observe and one the system actually functions with — is exactly the line that global workspace theory draws in cognitive science. A global workspace isn’t just information that’s present; it’s information that’s broadcast — available to be reported, to drive downstream processes, and to be recombined. The paper’s move is to take those functional criteria and test the lens’s representations against them, one by one.
Each test asks whether a lens-readable representation has one of the properties a workspace item should have. Click through them — they climb from “the model can say it” to “the model can reason with it in situations it never saw.”
A representation earns “workspace” status only if it passes as a thing the model uses — not just a thing we can see.
The tests aren’t a checklist you tick in any sequence — they’re a ladder of increasing evidential weight. Each rung rules out a cheaper explanation the rung below it can’t.
the shape of the argument
Notice this is the same escalation a skeptic would demand of any claim that an internal variable is “real”: show me you can read it, then that it does something, then that it composes, then that it transfers. Pass all four and “the model represents X” stops being a figure of speech.
Clear the four tests and the lens’s readouts are no longer just visualizations. They’re functional units the model reports on, is steered by, reasons over, and reuses — which is precisely the operational definition of a global workspace borrowed from consciousness research. The claim in the paper’s title, “verbalizable representations form a global workspace in language models,” is exactly this bundle of four passes, not a metaphor.
And this is where the trail opens onto live research. If the workspace holds word-like items the model can report, steer, and recombine, then the natural next question is which items live there. Concrete ones — Italy, euro — are the easy case. But could a more abstract axis, something like beauty, satisfy the same four tests? That would make an aesthetic judgment not a vague vibe but a reportable, steerable, reusable workspace variable — and the four tests are exactly the protocol you’d run to find out.
Two questions to lock in why four tests, and in this order.
Check 1 — report vs. modulation
The model can produce the word “euro” when asked about Italy’s currency (passes verbal report). Why isn’t that alone enough to call the representation causal?
Because report shows availability, not influence. The model could be reading the answer off some other pathway while the lens-readable representation just rides along as a correlate. Only directed modulation — intervening on that representation and watching the output move in the predicted direction — shows it’s actually doing causal work. Correlation passes test 1; causation is what test 2 demands.
Check 2 — the load-bearing rung
Of the four, which test most directly earns the word “workspace” — and why?
Flexible generalization. A workspace item’s defining property is that it’s broadcast and reusable — available to processes and contexts beyond the one that produced it. Report, modulation, and reasoning can all, in principle, be true of a narrow representation memorized for specific prompts. Only transfer to unseen contexts shows the representation is a general-purpose unit the rest of the model can pick up and use — which is what “global” means. (The other three rule out cheaper stories; this one delivers the actual claim.)
end of the trail
You now hold the paper’s whole spine: a Jacobian (lesson 1) becomes a lens (lesson 2), and four functional tests (this lesson) promote its readouts into a global workspace. Read the paper ↗ — it should read like a home you already know.