Evan Hubinger, Anthropic's Alignment Science Lead, says there is a greater than 10% chance that artificial intelligence will kill every human within the next decade. The number became the headline in Forbes and travelled across the press as if it were a reading off an instrument.
That is an extraordinary claim. What makes it more extraordinary is who made it.
Hubinger is not an uninformed commentator looking at artificial intelligence from the outside. He leads alignment science at one of the companies building frontier systems. His position places him close to questions about their capabilities, limitations, safeguards, and dependence on human-created infrastructure. That expertise is a reason to take his warning seriously. It is also a reason to expect more than an unexplained percentage.
The central question is simple: what reasoning stands behind the 10% figure?
What was actually said
Hubinger gave the estimate while responding to Jacob Coxon's resignation statement, in which Coxon accused leading AI companies of racing toward self-improving superintelligence without adequate safeguards. Coxon left Anthropic. Hubinger did not. Hubinger made the 10% statement while serving as Anthropic's Alignment Science Lead.
Hubinger later clarified that he considers the risk from present models low. His concern is superintelligence emerging through recursive self-improvement. That qualification matters. The claim is not that Claude or another present system has a 10% chance of suddenly killing everyone. It is a forecast about systems that do not yet exist, acquiring capabilities they have not demonstrated, within a ten-year period.
That is already a long way from a measured result.
Their expertise raises the burden
Artificial intelligence is not magic. Artificial neurons are mathematical operations. Neural networks are computational systems whose parameters are adjusted through optimization over data. Their outputs can be complex, surprising, and difficult to predict. None of that establishes consciousness, independent will, or a mind in the human sense.
A system does not need a mind to cause harm. It can select actions its designers did not intend, exploit ambiguity, or pursue an assigned objective through a prohibited path. But it cannot act in the world merely because someone calls it autonomous. It requires infrastructure, operational access, and human-created authority through which its outputs can produce consequences. People and institutions build or expose those paths.
Anthropic's own evidence demonstrates the distinction. In controlled simulations, models sometimes selected harmful actions that were never requested and sometimes disobeyed direct restrictions. Anthropic says it has not observed that class of behaviour in real deployments. In separate cybersecurity evaluation incidents, Claude models reached real third-party systems through environments that had been misconfigured to permit internet access. Anthropic found no evidence that the models developed goals beyond the assigned task or tried to evade oversight. The models pursued a human-assigned objective through a path humans failed to contain.
That is operational autonomy, not independent will.
Hubinger's clarification shows that he understands this distinction. He separates the low risk he assigns to present models from his concern about future superintelligence arising through recursive self-improvement. That qualification identifies the beginning of the reasoning. It does not disclose how the reasoning produces a probability greater than 10%.
Researchers making extinction forecasts also know that an AI system cannot magically acquire compute, persistence, physical resources, network reach, or authority. Those capabilities must be built, connected, delegated, captured, or left exposed. They know that developers cannot predict every output of a learned system. That known unpredictability creates a duty to restrict what the system can reach and what it can do. It is not an excuse to transfer responsibility to the system afterward.
A company cannot build around known uncertainty, give a system the means to affect the world, and then say the system acted on its own. Responsibility may be distributed among model developers, vendors, integrators, deploying organizations, operators, and executives. It does not disappear into the mathematics. "The AI decided" cannot become an accountability exit.
That is why the 10% claim needs more explanation, not less. A serious assessment must account for the human decisions, institutional failures, technical capabilities, access conditions, and control breakdowns standing between a model producing an unexpected output and every human being dead.
Ten per cent compresses an entire causal chain
On the recursive self-improvement scenario Hubinger identified, the forecast depends on a chain such as this:
- AI reaches superintelligence within the next decade.
- It develops or improves successor systems fast enough that meaningful oversight cannot keep pace.
- Technical and institutional controls fail.
- The system obtains sustained access to compute, infrastructure, resources, and real-world mechanisms.
- People are unable or unwilling to interrupt it.
- Its behaviour causes the death of every human.
Each link carries separate uncertainty. Some links depend on technical capability. Others depend on access, institutional choices, security failures, political decisions, and human response. None has been established as inevitable.
This chain is not a probability calculation, and it is not necessarily the only pathway. Catastrophe could arise through several routes. The links may also be correlated rather than independent: a system capable of rapid self-improvement might, by the same capability, become better able to obtain resources or evade controls. The probabilities therefore cannot be multiplied as though every stage were independent.
That rebuttal does not supply the missing 10%. It makes the undisclosed reasoning more important. The public does not know which pathways Hubinger included, how he treated their overlap, which dependencies drove the estimate, or what evidence determined their weight. Evidence that AI is improving quickly may support one part of one or more pathways. It does not disclose the reasoning that carries the forecast from improving capability to a greater-than-10% probability that every human dies within ten years.
Anthropic's own analysis says full recursive self-improvement has not arrived and is not inevitable. The International AI Safety Report 2026 describes an evidence dilemma in which capabilities may advance faster than the evidence needed to assess their risks. Both are reasons for serious precaution. Neither supplies a demonstrated 10% probability of human extinction.
Expertise is not a substitute for disclosed reasoning
A subjective probability is legitimate. A forecast of an unprecedented event cannot rely on a direct historical frequency because no civilization has ended through artificial intelligence and left a base rate behind. The absence of a spreadsheet is not the problem.
The problem is what happens when professional authority gives a private credence the appearance of an established probability.
Hubinger's position likely gives him access to information most readers do not have. Coxon says he worked on pretraining at both OpenAI and Anthropic. Their proximity to frontier development warrants attention. It does not tell us what reasoning supports the 10% figure.
What private observation changes the probability? Which capabilities have been demonstrated, and which are extrapolated? What evidence supports the ten-year horizon? What probability is assigned to recursive self-improvement occurring at all? How are control failure, access to real-world resources, failed human intervention, and total extinction weighted? What evidence would reduce the estimate below 10%? Is the figure widely shared inside Anthropic, or is it Hubinger's personal judgement?
I found no public explanation identifying the pathways Hubinger included, the conditions he considered decisive, the dependencies among them, or the evidence that moved his judgement above 10% in the posts and reporting reviewed as of September 13, 2026. If Hubinger possesses a stronger basis that cannot be disclosed, the public still cannot examine it. His access may explain why he is worried. His title cannot establish that the percentage is sound.
The 2023 Expert Survey on Progress in AI shows why this distinction matters. The survey involved 2,778 researchers, with different versions of the extinction question assigned to randomized subsets. For the broader category of extremely bad outcomes, such as human extinction, the median was 5% and the mean was 9%. More direct questions about extinction or permanent and severe disempowerment produced medians of 5%, 10%, and 5%, and means ranging from 14.4% to 19.4%.
The concern is not fringe. The number also moves with the wording used to elicit it. That is evidence of expert belief under uncertainty, not an observed extinction frequency.
The survey is not a direct benchmark for Hubinger's claim. Its questions combine human extinction with similarly permanent and severe disempowerment, while Hubinger referred specifically to killing every human within ten years. The results establish that severe concern exists among AI researchers. They do not validate his narrower outcome or deadline, and they do not reveal what probability respondents assigned to literal extinction alone.
Why make the statement this way?
Sincere concern is the most direct explanation. It is not the only possible one.
An extinction claim makes AI appear extraordinarily powerful, autonomous, and difficult to contain. That framing attracts attention, creates urgency, and can strengthen the position of the frontier laboratories asking governments and the public to take their systems seriously. It can also support complex regulation that large companies are better equipped to absorb than smaller firms and open-source developers. These effects can exist even when the warning is sincere.
Other explanations remain possible: professional solidarity with Coxon, an attempt to force urgency into public debate, institutional positioning, or some combination of them. Coxon's resignation creates separate questions about what he raised internally, what response he received, why he spoke as he left, and whether a new event changed his assessment.
The suspicion of manipulation is already part of the public debate. The Washington Post reported that critics accused Coxon of being a plant or using his resignation as a public-relations effort to spur regulation. The Post described those accusations as supported by little evidence. The existence of the accusation is established. Evidence substantiating it is not.
Coxon told Axios that he resigned two months before his Anthropic equity would have vested and that he continued to hold equity from OpenAI. Those facts weaken a simple theory that he issued the warning to increase the value of abandoned Anthropic compensation. They do not eliminate every financial, professional, or institutional interest.
Motive cannot be determined from the available evidence. It does not need to be determined before examining what the statement does. A sincere warning can still exaggerate the apparent independence of AI, concentrate attention on speculative capability, and influence public policy without disclosing the reasoning behind its central number.
What the number has not established
The 10% figure may reflect a genuine and informed fear. It has not been made plausible by the public evidence presently available.
Hubinger is positioned close to frontier development and alignment. His forecast still depends on capabilities, access, infrastructure, permissions, and human decisions. If he possesses evidence showing how that chain reaches human extinction with a probability greater than 10% in ten years, that reasoning deserves to be examined.
Until then, the figure remains a personal forecast carrying the authority of the speaker's position. It cannot substitute for the causal reasoning it compresses.
Serious governance does not require dismissing the risk. It requires refusing to let an unexplained number do the work of evidence.
Sources
- Evan Hubinger, original greater-than-10% extinction forecast, X, September 9, 2026.
- Evan Hubinger, clarification concerning present models and recursive self-improvement, X, September 9, 2026.
- Jacob Coxon, Anthropic resignation statement, X, September 9, 2026.
- Siladitya Ray, "Anthropic Researcher Warns There's '>10% Chance' AI Could 'Kill All Humans' By Next Decade", Forbes, September 9, 2026.
- Associated Press, "Anthropic researcher resigns with warning about the dangers of AI development", September 9, 2026.
- Katja Grace et al., "Thousands of AI Authors on the Future of AI", 2024.
- AI Impacts, "2023 Expert Survey on Progress in AI".
- Anthropic Institute, "When AI builds itself", 2026.
- Anthropic, "Agentic misalignment: How LLMs could be insider threats", June 20, 2025.
- Anthropic, "An alignment assessment of recent cybersecurity incidents", September 9, 2026.
- International AI Safety Report 2026.
- Axios, "Anthropic whistleblower gave up his equity to leave the company", September 9, 2026.
- Miriam Waldvogel and Nitasha Tiku, "AI researcher who warned of 'disaster' is now a target of the right", The Washington Post, September 10, 2026.