Score Hidden Objectives Before Compliant Tokens Appear
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
AI interpretability · Jacobian Lens · alignment
Research explainers on J-Space and the internal workspaces where modern LLMs stage reportable reasoning.
How safety teams use interpretability to catch behaviors that tests miss — and make model behavior operational.
Read featured explainer →Deep dives on Anthropic’s interpretability work, internal workspaces in LLMs, and what J-Space reveals about model behavior.
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.
Isolate causal paths and confirm necessity on residual streams with the J Lens joint lens and first-order Jacobian stack.
Researchers map how features causally influence Claude outputs by reconstructing prompt specific graphs with sparse autoencoders and cross layer…
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…
Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.
J-Space tracks the sparse internal workspace where language models stage reportable reasoning. The Jacobian Lens makes that workspace measurable — so safety, alignment, and interpretability research can move from speculation to evidence.