AI Safety North East
5 May 2026•Centre for AI Safety, Newcastle University
How reliable are current interpretability methods? Recent work on CoT monitoring and SAEs (co-presented with Toby Pullan)

Theo Farrell, Patrick Leask, Noura Al Moubayed
Mechanistic Interpretability Workshop at ICML 2026
View on OpenReview
Theo Farrell, Patrick Leask, Noura Al Moubayed
ResponsibleFM Workshop at NeurIPS 2025
View on OpenReviewHow reliable are current interpretability methods? Recent work on CoT monitoring and SAEs (co-presented with Toby Pullan)