Anthropic's own Alignment Science Lead just put a number on the worst-case outcome of AI development.
Evan Hubinger, who leads alignment science at Anthropic, said he personally believes there is a greater than 10% chance AI could "kill all humans" within the next decade. He said this after another Anthropic researcher resigned over concerns about the pace of development, and added that the company still has no clear solution for aligning a superintelligent AI system.
Why this matters
This isn't a critic or an outside commentator — it's a senior safety researcher at one of the AI labs building the most capable systems today. When the people closest to the technology are the ones raising the alarm, it's a different kind of signal than the usual AI-doom commentary from people with no stake in the outcome.
What This Means for Learners
You don't need to resolve the alignment debate to take one practical lesson from it: the companies building these systems openly admit they don't fully understand or control them yet. That's a reason to stay informed about what a tool can and can't be trusted to do unsupervised, especially as AI agents take on more autonomous, unattended work in day-to-day tools.