Blogs

2026

Circuits, Gradients, and Tool Trust

Published:

Scaling mechanistic interpretability to a real safety task: when a model must decide whether to trust a wrong tool response, attribution graphs read by a circuit oracle reach 67% accuracy, while a single backward pass at the correct-tool prompt reaches 85%.