Blogs
2026
Circuits, Gradients, and Tool Trust
Published:
Scaling mechanistic interpretability to a real safety task: when a model must decide whether to trust a wrong tool response, attribution graphs read by a circuit oracle reach 67% accuracy, while a single backward pass at the correct-tool prompt reaches 85%.
