Below are three interpretability projects. The work on refusal is what I spent by far the most time on.
What happens to a refusal direction when the model thinks?
I ran the refusal-direction pipeline of Arditi et al. on three 32B models from the Qwen family, four configurations in all. When Qwen3-32B answers directly, ablating the direction raises the share of harmful requests it answers from 3% to 94%. When the same model reasons first, the same direction raises it only from 13% to 41%. A linear probe separates harmful from harmless prompts almost perfectly, but ablating its direction removes far less refusal. The results are consistent with the chain of thought actively altering refusal behavior.
Write-up · Spring 2026
Can you see a robot's failure coming by looking inside it?
I built failure monitors for SmolVLA, a 450M-parameter robot policy, and pre-registered the main comparison before collecting 158 test runs. On top of exact simulator state, signals from inside the network added almost nothing: 0.5% better log loss, against a pass mark of 3%. Without simulator state, which a real robot does not have, the raw activations closed about two thirds of the gap (AUROC 0.80 to 0.89, against 0.92). My test of whether the policy uses a probed angle was too weak to answer, and the post says why.
Write-up · August to September 2026
What does a single parameter subcomponent do?
In Goodfire's parameter decomposition (VPD) of a 4-layer transformer, ablating subcomponents 427 and 3040 changes the next token almost only right after the word "an". Harder controls beat the pair, above all pairs that contain subcomponent 1654. Reading out the write directions explains both: all three promote tokens that start with a vowel. Ablating them together lowers the probability of a vowel-initial token after "an" from 73% to 32% and leaves it unchanged after "a".
Write-up · May 2026