Softmax Weights Replace the Ignition Threshold Global Workspace Theory Needs
GWT treats conscious access as competition for a limited broadcast. Transformers use graded QKV routing on a residual stream, not ignition.
AI interpretability · Jacobian Lens · alignment
Research explainers on J-Space and the internal workspaces where modern LLMs stage reportable reasoning.
How safety teams use interpretability to catch behaviors that tests miss — and make model behavior operational.
Read featured explainer →Deep dives on Anthropic’s interpretability work, internal workspaces in LLMs, and what J-Space reveals about model behavior.
GWT treats conscious access as competition for a limited broadcast. Transformers use graded QKV routing on a residual stream, not ignition.
QKV attention routes tokens rather than igniting broadcast. Decoder-only transformers fail GWT; the 2017 paper dropped recurrence. Use indicator tests.
Decode residual-stream activations and 34 million SAE features; chat transcripts, chain of thought, and the public API do not expose them.
GPT-3 stacks 175 billion untied parameters while Universal Transformers reuse one block. Both broadcast. Neither creates C2 sentience.
Late residual writes sit immediately upstream of the unembedding and bias decoder-aligned effects toward the last blocks.
Published widths from 768 to 12288 bound how many residual directions stay independent. Extra features interfere instead of adding workspace slots.
J-Space tracks the sparse internal workspace where language models stage reportable reasoning. The Jacobian Lens makes that workspace measurable — so safety, alignment, and interpretability research can move from speculation to evidence.