NeurIPS 2026 · Sydney
Interpretability as a Science
Toward rigorous foundations for understanding LLMs
About
As large language models grow in capability, interpretability asks how and why they behave as they do. Yet the field has not converged on notions of explanations at varying levels of abstraction, what evidence supports a claim, or how to design experiments that rule out alternative explanations. This workshop asks what it would take to ground interpretability as a rigorous empirical science—drawing lessons from disciplines that have long studied complex systems.
We focus on the following questions:
- What does it mean to understand an LLM?
- What standards, benchmarks, or evaluation criteria the field could adopt for measurement, causal claims, and falsifiability?
- What can interpretability learn from neuroscience, statistics, and causal representation learning?
A distinctive feature of this workshop is its interactive format. The workshop will host multiple breakout sessions, each led by a facilitator from a relevant discipline, who will give a short lightning talk highlighting important questions related to a specific sub-theme, then moderate a discussion connecting these ideas to interpretability. This format is designed to encourage genuine dialogue and engagement, inviting attendees to collectively shape a shared scientific foundation for interpretability.