Skip to main content
Theo Farrell

Research

Reviewing

  • ICML 2026 Mechanistic Interpretability Workshop
  • NeurIPS 2025 Mechanistic Interpretability Workshop
  • NeurIPS 2025 ResponsibleFM Workshop

AI Safety North East

5 May 2026Centre for AI Safety, Newcastle University

How reliable are current interpretability methods? Recent work on CoT monitoring and SAEs (co-presented with Toby Pullan)