Autointerp Simulation Scores Track Token Naming, Not Explanation Truth

Abstract

Autointerp simulation scoring is used as a quality signal for sparse autoencoder latent descriptions, a use that presumes the score tracks whether a description is correct. We test that presumption with a within-latent 2x2 over whether a description names the latent’s trigger token and whether its semantic gloss is true, on 16 token-driven latents of a Gemma-2-2B layer-12 GemmaScope SAE. Naming the token raises the score by 0.305, and does so for all 16 latents. Conditional on naming, swapping a true gloss for a false one changes it by 0.021 (95% CI -0.11 to +0.15). A judge-free baseline that marks tokens whose surface form appears in the description reproduces the effect. The results suggest that a high score can reflect naming the right token rather than truth of the description.

Publication
NeurIPS 2026 Workshop on InterpScience (submission)
Prince Modi
Prince Modi
Master’s Student, LLM Systems (Inference)

LLM Systems (Inference) @ UCSD