Interpretability will reliably detect most model problems by 2027.
Dario Amodei, Anthropic · made Apr 1, 2025 · target Dec 2027 · Open
Anthropic is doubling down on interpretability, and we have a goal of getting to “interpretability can reliably detect most model problems” by 2027.
The prediction
- Who
- Dario Amodei, CEO, Anthropic
- Made
- Apr 1, 2025
- Target
- Dec 31, 2027 (stated as "by 2027")
- Status
- Openin 15 months
- Source
- The Urgency of Interpretability Primary: their own words · original recording
- Topics
- Science
More from Dario Amodei
Related predictions
| Prediction | ||||
|---|---|---|---|---|
| Mar 20, 2026 | Within a decade AI will be able to do much of the work math students now do.Open · Unverified · Terence Tao, UCLA · target Mar 2036 | Mar 2036within a decade | Science | Open |
| November 2025 | Medical superintelligence will arrive in the next few years.Open · Mustafa Suleyman, Microsoft · target Nov 2028 | Nov 2028in the next few years | Science | Open |
| Oct 28, 2025 | OpenAI will have a fully automated AI researcher by 2028.Open · Unverified · Sam Altman, OpenAI · target Dec 2028 | Dec 2028by 2028 | Agents | Open |
| Oct 28, 2025 | OpenAI will have an intern-level AI research assistant by September 2026.Open · Unverified · Sam Altman, OpenAI · target Sep 2026 | Sep 2026by September 2026 | Agents | Open |
| September 2025 | Dec 2030by 2030 | Science | Open | |
| September 2025 | Dec 2030by 2030 | Science | Open |
Sources: the source linked on each prediction, checked weekly; outcomes judged on the evidence linked. Logos via logo.dev; trademarks belong to their owners.