
AI Safety Research and Field-building
Theo Farrell
First-class MSci Natural Sciences (Computer Science & Philosophy) from Durham University. Founder and co-organiser of Durham AI Safety Initiative. Previously I worked on relational composition and SAEs (i.e. mechinterp) but now I’m more interested in model organisms, emergent misalignment and inoculation prompting. I’m also increasingly concerned about risks from non-frontier open weight models. I’m currently in London for LASR Labs, investigating evaluation of model organisms (mentored by Satvik Golechha). Outside of AI Safety I play bass guitar 🎸 and dabble with piano and viola 🎹🎻!
Recent Publications
- Sparse Autoencoders Can Learn Graded Latents for Relational Composition
Theo Farrell, Patrick Leask, Noura Al Moubayed
Mechanistic Interpretability Workshop at ICML 2026
- Order by Scale: Relative-Magnitude Relational Composition in Attention-Only Transformers
Theo Farrell, Patrick Leask, Noura Al Moubayed
ResponsibleFM Workshop at NeurIPS 2025
Field-building
- Durham AI Safety Initiative
I founded DAISI in my second year of university and grew weekly attendance from 5 to 20. Supported by the Pathfinder fellowship, the group helps funnel top-university talent into AI Safety.
- Advisor · February 2026 — Ongoing
- Lead Organiser · October 2023 — February 2026