Back to Resources
A natural language processing researcher who studies how to understand, improve, and precisely control language models. His work applies mechanistic interpretability, causal analysis, and cognitive-science-inspired evaluation to questions of model robustness and efficiency, mentoring students who lead publications on topics such as concept unlearning and model steering.
Last verified: August 22, 2026
Suggest a correction