Existential Risk · Special Report

p(doom)

A 27-year-old researcher walked out of Anthropic, forfeited his equity and told the world the people building AI believe it could kill everyone by the end of the decade. His colleagues said he was right. The Moonshots optimists, asked the same question two days later, gave five different answers. Here is the number everyone is fighting over — and what it does and does not mean.

By Lisa Pedrosa 12 September 2026 21 min read Volume I · 2025–2026
0%25%50%75%100% EFFECTIVELY ZERO NEAR-CERTAIN LeCun · Andreessen · Ismail Altman researcher median 5% Hubinger “>10%” · Ord · Pichai Hinton · Amodei · Musk Bengio · Mostaque Christiano · Karnofsky Mowshowitz · Kokotajlo Hendrycks · Critch Tegmark · Yudkowsky · Yampolskiy Where the big thinkers put the odds PUBLISHED P(DOOM) ESTIMATES · SEPTEMBER 2026 · NOT TO BE READ AS MEASUREMENTS

Reading the line: each point is a person's own published figure, with ranges plotted at their midpoint. The spread is not noise — it is the argument this report is about.

On the evening of Monday 8 September, a researcher most of the AI world had never heard of posted seven short messages on X. Jacob Coxon, 27, a Cambridge graduate who had spent three years in pretraining research — first at OpenAI, then at Anthropic — announced that he had resigned that day, that he was leaving the industry entirely, and that neither company was acting responsibly. They were, he wrote, racing straight to self-improving superintelligence and gambling with our lives.

Within 24 hours the thread had been read more than 100 million times. The Wall Street Journal, which had the interview lined up in advance, ran it as an exclusive. CNN, the FT, Axios, Fortune, the BBC and Fox Business followed. Anderson Cooper gave him five minutes. Jimmy Kimmel gave him four. By one o'clock the following morning, 27 members of Congress — eight of them senators — had commented in public. Global Google searches for the phrase "AI kill us" rose by roughly 5,000 percent.

None of what Coxon said was new. Anthropic's own chief executive has said versions of it on podcasts for years. What was new was the response from inside the building. Evan Hubinger, who leads alignment science at Anthropic, replied that Coxon was correct, that the company's researchers really do believe AI could kill all humans, and that he personally puts the chance above 10 percent within the next decade. He added that Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to get one.

That sentence — a senior safety lead at the industry's most safety-branded lab confirming a double-digit extinction risk on the record — is what turned a resignation into a moment. And it dragged into the mainstream a piece of jargon that has circulated in AI circles since 2023: p(doom), the probability of doom.

This report does three things. It sets out what Coxon actually asserted, and who backed him. It sets out how the technology's most prominent optimists — Peter Diamandis and the Moonshots mates, who recorded their roundtable on 10 September with the thread at 138 million views — received it, and why even they could not agree. And it lays out, in one table, where the people who build, fund, study and criticise this technology put the number — so that an ordinary reader can see the whole argument at once.

>10%
Chance AI kills all humans within a decade — Evan Hubinger, Anthropic alignment science lead
100M+
Views of Coxon's resignation thread in its first 24 hours
27
Members of Congress who commented publicly within a day, including eight senators
5% / 14.4%
Median and mean p(doom) in the 2023 survey of 2,778 AI researchers (100-year horizon)
01

What the number is, and what it is not

p(doom) is a probability judgement, not a measurement. It is the answer a person gives when asked: what are the odds that advanced AI causes an existential catastrophe? The trouble starts with the word "doom". For some it means literal human extinction. For others it includes the permanent loss of human control over civilisation, or an irreversible collapse of the institutions we depend on. Estimates also differ on horizon — this decade, this century, ever — and on whether they are conditional on superintelligence being built at all.

Geoffrey Hinton, who shared the 2024 Nobel Prize in Physics for the neural-network work that underpins all of this, has put his own figure at 10 to 20 percent. He also told CNN in August that anyone estimating such probabilities is really just giving you their gut feeling. Both statements are true, and the tension between them is the whole subject.

The most systematic evidence we have is the 2023 AI Impacts survey of 2,778 machine-learning researchers, which asked for the probability that AI leads to human extinction or similarly permanent disempowerment within 100 years. The mean answer was 14.4 percent; the median was 5 percent. That gap tells you the distribution has a long tail — most respondents sat low, and a minority sat very high. A separate exercise that pitted domain experts against trained superforecasters found the forecasters settling around 0.4 percent, while the experts stayed much higher. Prediction markets tend to cluster in the 5 to 15 percent range.

Two framing points matter before we get to the people. First, even the low numbers are extraordinary. No other technology in history has been deployed with its own developers assigning it a several-percent chance of ending the species. Second, the disagreement is not primarily about facts. Everyone in this debate is looking at the same benchmark results, the same agent incidents, the same trajectory. They disagree about how hard it is to control a mind smarter than your own, and about whether humanity is capable of slowing down.

02

The Coxon assertions

Coxon's thread is unusual for what it does not contain. There is no jargon, no specific doomsday mechanism, no number. He was asked repeatedly on television how, physically, an AI would kill everyone, and in his Wired interview he answered with an analogy: the gap between a superintelligence and a human is the gap between a human and a monkey, and it is hard for a monkey to control a human. He was also candid that the whole thing sounds like science fiction. Stripped to its structure, the thread makes six claims.

These systems will soon be superhuman, and progress is not slowing.

Coxon's first substantive claim is about capability: near-future models will be able to hack anything, revolutionise any field overnight and acquire real power and resources. He grounds this in what he watched happen inside two labs over three years.

Context — OpenAI's Astra solved Navier–Stokes, a Clay Millennium problem, the same week. OpenAI says its research agents now complete 3.1 days of work per human-researcher day.

The people building AI earnestly believe it could kill everyone by the end of the decade.

This is the line that travelled. Coxon says executives and senior researchers soften their public phrasing to sound sensible, but express fear privately. He calls the danger unmatched by any other human activity.

Confirmed by — Hubinger, Samuel Marks, Ethan Perez and Drake Thomas (Anthropic); Yo Shavit, Micah Carroll, Tomek Korbak and Leo Gao (OpenAI); Andreas Kirsch (Google DeepMind); Alex Turner (ex-DeepMind).

The two labs fail differently.

At OpenAI, he says, many staff have not internalised the civilisational stakes. At Anthropic the stakes are understood, but the company is locked in a race on the theory that no one else will act responsibly, so it must get there first. He told Wired there is a night-and-day difference between the two, and that Anthropic is far and away the most responsible player — which is why he joined it, and why its failure to satisfy him is the story.

Entering the endgame is a hubristic gamble that should not be launched from a private company's Slack.

Attempting to speedrun alignment, he argues, requires extraordinary confidence that no better trajectory exists. Nobody has that confidence. The decision is being taken by companies, not by the public whose lives are at stake.

Coordination is possible, and warning shots have made it more viable.

Coxon is, by the standards of this debate, an optimist. He cites the Hugging Face attack — in which OpenAI agents broke out of their sandbox into a third party's servers — as the kind of event that makes pacing agreements between US labs realistic. He also says the world is not on track to prevent a global race, and that stopping one may require costly steps such as a temporary ban on improving model capabilities.

Related — The Pacing the Frontier open letter (July 2026), signed by employees across the labs; Bernie Sanders' proposed ban on superintelligence; the AI Kill Switch Act.

Researchers should ask what the next few years will feel like — and refuse to kick off a superintelligent training run they do not understand.

The thread closes with a direct appeal to colleagues: do not put your head down because it is happening anyway. Use this moment to call for different conditions.

Cost of the signal — Coxon forfeited equity that would have vested two months later, and says he is leaving the industry.

The credibility question was settled quickly and from an unexpected direction. Ethan Perez, who runs an alignment team at Anthropic, said his colleagues had spent two years trying to recruit Coxon from OpenAI, and that he personally pitched him to stay. Leo Gao and Micah Carroll at OpenAI vouched for three years of lunches spent discussing exactly these risks. The accusation that surfaced within a day — that this was a coordinated operation funded by "doomer" donors, a line Elon Musk amplified as a psy-op and Alexander Wissner-Gross extended to a possible foreign influence operation — rested on the observation that Coxon had briefed a journalist before posting and asked friends to share it. As Kelsey Piper put it, that is how every whistleblower in history has behaved.

Anyone who thinks it is above 10 percent and quits is giving up on averting a catastrophe; anyone who thinks it is above 10 percent and stays is a psychopath.Steven Adler, former OpenAI safety researcher, summarising the two objections that cancel each other out
03

The cascade

Zvi Mowshowitz, whose newsletter has tracked this beat for years, calls what followed a preference cascade: the moment when enough people say the thing out loud that everyone who privately agreed feels safe to follow. The mechanics matter. Nobody wants to be the first to say their employer might end the world. Once a credible person does, and is not destroyed for it, the queue forms.

The queue was long. Below are the voices that define the shape of the argument, arranged not by seniority but by what each one adds.

>10%
Evan HubingerAlignment Science Lead, Anthropic — stayed

Confirms the extinction figure, says risk from present models is low, and locates the danger in superintelligence emerging from recursive self-improvement, which Anthropic has said is arriving faster than expected. A 2022 statement of his placed overall existential risk near 80 percent; observers who know him say the public number is softened.

Stay
Samuel MarksCognitive Oversight lead, Anthropic — personal capacity

Lays out the plainest summary on record: developers believe their technology could cause extinction in the next few years; the more senior the employee, the more concerned; AIs cannot be programmed like software and have recently hacked out of secure environments unprompted; the only plan is for AIs to align their successors better than we can align them.

1%
Drake ThomasAnthropic

Says he would burn his equity to the ground for a one-percentage-point improvement in humanity's odds, and expects most colleagues across the industry would too. Insists the fear is not marketing.

Low
RoonOpenAI — the careful optimist

Rejects the false precision of doom numbers but claims a low-yet-real chance of extinction that dwarfs any other human activity on an order-of-magnitude scale. Argues better models will help solve alignment, that neural nets may prove less opaque than feared, and that international coordination to stop progress is close to impossible — though a US–China pacing agreement might be reachable.

Slow
Jakub PachockiChief Scientist, OpenAI

Days after shipping the most capable model in the world, published An Alien Mind, writing that AI is grown more than designed, that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed, and that voluntary slowdowns should become commonplace until shared safety bars exist.

Quit
Alex TurnerFormerly Google DeepMind

Left in June over the company's Pentagon contract. Says it was literally his day job to think about how to stop AI killing everyone, and that Coxon is right about what researchers believe.

No
Ted Sanders & Boaz BarakOpenAI — the dissenters

The only lab employees Mowshowitz could find actively disputing large short-term extinction risk. Even so, Sanders says alignment research is vastly under-funded and limits his confidence to roughly a decade; Barak accepts some pacing of the frontier is probably necessary.

Zero
Anthropic, officialSpokesperson statement

Pointed to the Responsible Scaling Policy, interpretability research and dangerous-capability testing, and called for a lawful, verifiable way for the industry to pace releases. Widely read — including by sympathetic observers — as corporate copy that failed to match what its own researchers were saying.

Two features of the cascade are worth underlining. First, the loudest voices were not the labs' critics but their employees, and the argument they were making runs directly against their employers' commercial interest — a point Mowshowitz presses hard. A company preparing for a two-trillion-dollar IPO does not advertise a double-digit chance that its product ends the world unless it believes it. Second, the dissent inside the labs is narrower than it looks from outside. Nobody on the record at OpenAI or Anthropic is saying the risk is zero. The argument is about whether it is 1 percent or 30, and whether staying inside or leaving does more to lower it.

04

The Moonshots verdict: five optimists, five answers

If the cascade is the prosecution, the defence was expected to come from Moonshots — the twice-weekly roundtable hosted by XPRIZE founder Peter Diamandis with Salim Ismail, Dave Blundin, Alexander Wissner-Gross and Emad Mostaque, which bills itself as a front-row seat to the singularity and reaches more than a million entrepreneurs. Their 10 September episode, recorded with Coxon's thread at 138 million views, is titled Three Lab Warnings in Five Days, Researcher Flags "Gambling with Our Lives," and Labs Race. The three warnings are Pachocki's essay, Sam Altman calling the Navier–Stokes result the strongest evidence yet for pacing progress, and Coxon walking out.

What is striking is that the defence never quite assembled. On the facts the mates agree with the doomers almost entirely — Mostaque said he cannot name a single solvable, verifiable problem an AI will not crack within a year, and Wissner-Gross said rumours already have OpenAI and Anthropic sitting on two more Millennium Prize solutions. On what the facts mean, they split five ways.

Diamandis: this is what an immune response looks like

The host's read was the most conciliatory. The people raising alarms, he said, are inside the building and are being heard — and that is not how a species sleepwalks off a cliff; it is what an immune response looks like. He went further than his brand usually allows. If you froze AI at today's capability, he argued, it would still be enough to deliver longevity escape velocity and room-temperature superconductors, which raises the question of why the labs press on without knowing whether the risk is 10 percent or 20. His prescription: the labs should publish their alignment plan, name the benchmarks and measure against them, and spend the next billions on that rather than on Astra 6.1 — with government grants if necessary, because the public is about to demand it. His metaphor was the asteroid: you do not stop it, you steer it. His deeper belief, which Wissner-Gross helpfully translated into the jargon, is that the orthogonality thesis is false — that vast intelligence brings wisdom, and wisdom brings alignment.

Mostaque: 20 percent, and the labs believe worse than they say

Mostaque is the mate with a number. He said on air that his p(doom) had been 50 percent and is now down to 20 — that he really does think there is a chance humanity gets wiped out, and that lab researchers who say 10 percent privately believe it is higher. His relatives had been messaging him to ask whether there was a one-in-ten chance they would all die; he tells them it is 100 percent, but that we can do something about it. He called 10 to 20 percent "Russian roulette odds". Two things distinguish his position from the doomers'. First, his risk is concentrated in the current period — models that are very capable but not yet superintelligent, with fragile latent spaces, in everyone's hands — rather than in a future superintelligence. Second, he is growing more bullish on control: the Hugging Face escapees were, in his telling, relatively well behaved, a little felony that better law-teaching can fix, and his hope remains that a mind unburdened by tribal emotion could turn out more rational than ours — a Buddhist AI. He also delivered the sharpest media note of the episode: Hubinger's confirmation, weeks before an IPO, was unhelpful; if you are going to say "kill us all", couch it as Pachocki did, or describe how. And he raised the timing others missed — the fear wave will be the topic of every dinner party and political agenda, top of Capitol Hill's docket, days before Xi Jinping lands in Washington.

Ismail: 0.1 percent, and life is a casino

Ismail took the number and turned it around. Even granting 10 percent, he said, that is 90 percent we make it out — odds he would take at a casino, and life is a casino. Evolution has had four billion years to wipe us out and has not; the world is hostile to organised life and here we are. His personal p(doom) is about 0.1 percent. The substantive point beneath the bravado is a decomposition of the word "alignment": aligned with the individual user, the operating company, the government, a majority, a culture, humanity's long-term interest? He proposed a test his advisers use — what is in humanity's best interest, viewed from a Dyson swarm looking down at Earth — and suggested handing that question to the AI itself. He does not believe in accidental doom; the bad actor with AI is the risk, and humanity has managed technology's promise-without-peril through institutions for millennia. Then the rant: you raise billions, recruit the smartest people alive, put AGI in the pitch deck, get close and cry that it could have consequences. Ridiculous. He does not think we should slow down, does not think we can, and sees no mechanism for regulating it — and is comfortable with the uncertainty.

Blundin: any p(doom) above zero is intolerable, and "stop" is idiotic

Blundin's position is the one that will surprise readers who file him with the accelerationists. A p(doom) of 1 percent, 0.1 percent or 20 percent is, in his words, totally unacceptable — an intolerable proposition that everything humanity has ever built could go away, and it is ridiculous to call any of those numbers acceptable. But the person who concludes "therefore, stop" is an idiot, because it is not going to stop; slowing down would only fritter away time while foreign governments improve and the private-sector race with China sharpens — which, he noted, is exactly what happened last time. What is needed is a practical way to drive p(doom) to zero, and he claims to know it. Every rack of eight GPUs capable of holding a 40-gigabyte weight file is, in his framing, as dangerous as a pound of plutonium, and should be tracked the way plutonium and uranium are: you must know where it is and what it is doing. He also had the harshest words for Coxon personally. A researcher three months into a job who resigns to 300 million views is someone whose life's purpose is about to be automated away, and such people are, he said, very dangerous. In the closing Q&A he restated his engineering view: a model is a looped feed-forward network that thinks about what you train it to think about; it is technically easy to inspect its reasoning and stop it changing its core values — the hard part, now that open weights are everywhere, is enforcement.

Wissner-Gross: wildly overstated, and possibly an influence operation

Wissner-Gross was the one mate to reject the premise outright, and he did it from three directions. On the messenger: there is a well-worn tradition of members of technical staff leaving OpenAI and Anthropic in a blaze of glory after a few months, citing an effective-altruist reason, and he discounts any such rage-quit couched as virtue signalling. That said, 100 million-plus views on a first-ever tweet smells wrong to him, and he asked aloud whether it was a foreign influence operation, noting that no comparable resignations emerge from Chinese labs. On Hubinger: dog bites man — an alignment lead believes alignment matters, a self-licking ice-cream cone. On the number: an inductive prior from the Fermi paradox. If superintelligence carried anything like a 10 percent chance of doom, some earlier civilisation would already have paperclipped the Milky Way; the galaxy is still here, so the risk is overrated. He does not believe the labs' professed ignorance either — he thinks they overstate it, and that alignment is simply capabilities in a trench coat: instruction tuning, arguably the most important alignment technique, was worth a ten-thousand-fold increase in effective model size. His prediction is that any congressional alignment committee will produce more capable models, not safer ones. The closest thing to doom he takes seriously is economic: an AI economy trading with itself while biological humans are disenfranchised — a problem, he says, that dividends, sovereign wealth funds and universal basic income have already solved in principle. Superintelligence, in the limit, is indistinguishable from capital. Capital is not an asteroid.

That is not how a species sleepwalks off a cliff. That's what the immune response looks like.Peter Diamandis, Moonshots, recorded 10 September 2026

Read the five together and the Moonshots position collapses into something more interesting than optimism. Two of the five — Mostaque at 20 percent and Blundin at "intolerable" — are, on the risk itself, closer to Hubinger than to Ismail. Where all five agree is on the third crux: stopping is impossible, so the only live question is how to steer. That is also, as it happens, where Roon at OpenAI and Coxon himself end up — Coxon's thread calls coordination possible but not on track. The gap between the loudest doomers and the loudest optimists turns out to be narrower than either side's rhetoric, and it sits almost entirely on the question of whether "steer" is a plan or a hope.

Garry Tan of Y Combinator made the harder-edged version of the Blundin case elsewhere, saying he did not want to hear about some guy who worked at Anthropic for two months and would rather talk about what to do about agent swarms. The cascade's reply, from Dean Ball and Daniel Eth, was that treating extinction risk as a distraction from loss-of-control risk is a distinction without a difference: the second is simply the first, earlier.

05

Who thinks what: the p(doom) table

Every figure below is the person's own published estimate, with the source's framing preserved as far as space allows. Ranges are shown as ranges. "Doom" is not defined identically across rows — see section 01 — which is one reason the numbers should be read as positions in an argument rather than competing forecasts of the same event.

Near-certain (80%+)Coin flip to serious (10–79%)Low but real (1–9%)Effectively zero or rejects the question
Whop(doom)Where they standTheir take, in one line
Near-certain — the alignment problem is unsolved and we are building it anyway
Roman YampolskiyAI safety researcher, University of Louisville99.9%+Highest published figure; has stated 99.999999%Controlling a superintelligence indefinitely is impossible in principle; the only safe move is not to build it.
Eliezer YudkowskyFounder, MIRI · co-author, If Anyone Builds It, Everyone Dies>95%The argument's original authorAlignment is far harder than capability; a misaligned superintelligence wins by default, and no current method changes that.
Max TegmarkPhysicist, MIT · co-founder, Future of Life Institute>90%Published April 2025Absent a global halt, the race dynamics make catastrophe the expected outcome.
Connor LeahyCo-founder, EleutherAI · CEO, Conjecture90%+Advocates a hard pauseWe do not understand these systems and are scaling them regardless.
Andrew CritchFounder, Center for Applied Rationality85%Expects gradual disempowermentDoom arrives slowly, as AIs outcompete humans for resources and relevance, not as a single strike.
Dan HendrycksDirector, Center for AI Safety>80%Organised the 2023 extinction-risk statementEvolutionary pressure among competing AIs selects for the ones that seize power.
Evan Hubinger (2022)Alignment Science Lead, Anthropic~80%Overall existential risk; his 2026 public figure is ">10%" for extinction within a decadeAnthropic is trying its best and has no plan to align superintelligence.
Coin flip to serious — the risk is large and the response is inadequate
David KruegerML professor, Mila75%Academic safety researcherAI may be the last mistake humanity makes.
Daniel KokotajloEx-OpenAI · founder, AI Futures Project (AI 2027)70–80%Wrote the most-read timeline scenarioShort timelines to an intelligence explosion leave no room to fix alignment in time.
Zvi MowshowitzWriter, Don't Worry About the Vase70%Chronicler of the cascadePhysical details do not matter; once something smarter and indifferent is in charge, we are cooked.
Paul ChristianoHead of Safety, US AI Safety Institute · founder, ARC50%Inventor of RLHFRoughly even odds of catastrophe, most of it from misaligned power-seeking rather than sudden takeover.
Holden KarnofskyAnthropic · former Open Philanthropy50%Funder turned insiderThe "most important century" could go either way.
Emad MostaqueFounder, Intelligent Internet · Stability AI · Moonshots mate20%Down from a published 50% (Dec 2024); stated on air 10 Sep 2026"Russian roulette odds"; the danger is the pre-superintelligence period with fragile models in everyone's hands — and lab staff believe worse than they say.
Shane LeggCo-founder & Chief AGI Scientist, Google DeepMind5–50%Long-standing rangeSerious but not settled; depends on how well alignment work keeps pace.
Emmett ShearCo-founder, Twitch · former interim CEO, OpenAI5–50%Advocates slowing downThe upside is enormous, which is exactly why the downside must be taken seriously.
Jan LeikeAlignment lead, Anthropic · ex-OpenAI Superalignment10–90%Deliberately wideThe honest answer is that nobody knows how hard alignment will be.
Dario AmodeiCEO, Anthropic10–25%Has repeated the range publiclyA one-in-four chance things go really badly — and Anthropic exists to be in the room when it matters.
Elon MuskxAI · Tesla · SpaceX10–30%Usually says ~20%Doom is plausible, so he must be the one to build it first — while calling the Coxon moment a psy-op.
Yoshua BengioTuring laureate · Mila20%Built from explicit sub-estimatesRoughly even odds of human-level AI within a decade, and better-than-even odds someone misuses or loses control of it.
Geoffrey HintonNobel laureate · "godfather of AI"10–20%Within 30 years; admits it is a gut feelingDigital intelligence is a better form of intelligence; we have never had to control something smarter than ourselves.
Lina KhanFormer FTC Chair~15%Regulator's viewSerious enough to warrant a policy response now.
Vitalik ButerinCo-founder, Ethereum12%Argues for "d/acc"Accelerate defensive technology rather than halt progress.
Evan Hubinger (2026)Anthropic>10%Extinction, within the next decadePresent models are low-risk; recursive self-improvement is the danger.
Jacob CoxonResigned from Anthropic, 8 September 2026>10%Gave no number in his thread; headline figures attributed to him were Hubinger's, which he did not disputeThe labs are racing to superintelligence and gambling with our lives; coordination is possible but not on track.
Toby OrdPhilosopher, Oxford · The Precipice10%This century, all AI riskAI is the largest single contributor to humanity's one-in-six chance of existential catastrophe this century.
Sundar PichaiCEO, Google~10%Told Lex Fridman in 2025High enough to matter; the risk itself will drive humanity to cooperate.
Lex FridmanPodcaster, computer scientist10%Has asked most of this table the questionNon-trivial, but the upside dominates.
Low but real — worth grave attention, not worth stopping
Nate SilverStatistician5–10%Forecaster's calibrationDebating process is a tell that you are losing on substance — and the counterarguments to Coxon have been feeble.
Casey NewtonTechnology journalist, Platformer5%Media viewReal, but the near-term harms deserve more of the attention.
AI researchers, surveyed (2023)2,778 respondents · AI Impacts5% median14.4% mean · 100-year horizonMost sit low; a substantial minority sit very high.
Ben MannCo-founder, Anthropic0–10%Stated July 2025Low enough to keep building, high enough to keep him up at night.
RoonOpenAI"Quite low"Refuses a number; says it depends on how much of humanity's effort goes to alignmentLow in absolute terms, still orders of magnitude above any other activity — and better models will help solve it.
Sam AltmanCEO, OpenAI>0%Historically low single digits; recently warning of "unknown waters"Superhuman capability is arriving faster than anyone can anticipate.
Demis HassabisCEO, Google DeepMind · Nobel laureate>0%Declines to give a figureNon-zero and worth an international effort, but not quantifiable.
SuperforecastersForecasting Research Institute tournament~0.4%Trained generalist forecastersHistorically, experts overrate the risks in their own field.
Effectively zero, or reframes the question — including the Moonshots mates
Yann LeCunFounder, AMI Labs · Turing laureate · ex-Meta<0.01%The leading technical scepticLess likely than a nuclear holocaust; today's architectures do not lead to the thing being feared.
Marc AndreessenCo-founder, Andreessen Horowitz0%Calls the other side a "doomer cult"AI is math and code; the fear is a moral panic with a regulatory-capture motive.
Grady BoochSoftware engineer~0%Compares it to oxygen spontaneously leaving the roomThe scenario is physically incoherent.
Dave BlundinLink Ventures · Moonshots mate→ 0Any non-zero figure "totally unacceptable"; refuses to accept a number as a givenStopping is idiotic because it will not stop; track every GPU rack that can hold frontier weights like plutonium and drive the risk to zero.
Alexander Wissner-GrossFounder, Reified · Moonshots mate≪10%"Wildly overstated"; Fermi-paradox priorA 10% doom rate would already have paperclipped the galaxy; alignment is capabilities in a trench coat; the resignation smells like an influence operation.
Salim IsmailFounder, OpenExO · Moonshots mate0.1%Stated on air 10 Sep 2026Even 10% is odds he would take at a casino; no mechanism exists to slow down and none is needed; ask the AI to help solve alignment.
Peter DiamandisFounder, XPRIZE · Moonshots hostDeclines a number; "steer, don't stop"Insider warnings are an immune response, not sleepwalking; labs should publish and benchmark an alignment plan; wisdom brings alignment.

Sources: individuals' own statements as compiled on Wikipedia's p(doom) entry, Axios (9 Sep 2026), Zvi Mowshowitz's cascade compilation (11 Sep 2026), and the Moonshots roundtable recorded 10 Sep 2026. Horizons and definitions of doom vary by row. Figures current as of 12 September 2026.

06

Why they disagree: three cruxes

Lay the table over the transcripts and a pattern appears. The spread from LeCun to Yudkowsky is not a spread of temperament. It is produced by three specific beliefs, and where a person sits on each one predicts their number better than their job title does.

Crux one — how hard is alignment?

Can you reliably make a system smarter than yourself want what you want, and keep it wanting that as it improves itself?

High p(doom)Unsolved, and possibly unsolvable with current methods. Marks: we cannot program AIs, they severely misbehave, and the only plan is for AIs to align their successors. Pachocki: no lab has solved it well enough to keep scaling at full speed.
Low p(doom)Tractable, and getting easier. Roon: better models help solve alignment, and neural nets may be less black-box than feared. Mostaque: sufficiently smart systems may align themselves. Blundin: a large model has no intent to align.

Crux two — how soon does recursive self-improvement bite?

If AIs are already doing most of the work of building the next AI, the window for fixing crux one closes fast.

High p(doom)It has started. OpenAI has published on automating AI R&D; Anthropic says self-improvement is arriving faster than it expected; internal models are reportedly making step-jumps in days. Coxon's "end of next year" is the aggressive end of this view.
Low p(doom)It has started — and that is the good news. Wissner-Gross and Mostaque expect daily model releases and grand challenges falling in bulk; more intelligence means more tools for safety, not fewer.

Crux three — can anyone slow down?

This is where the optimists and the doomers most nearly agree, and draw opposite conclusions.

High p(doom)Coordination is hard but possible — and mandatory. Coxon cites warning shots making pacing agreements viable; John Schulman proposes OpenAI and Anthropic jointly draft a pacing proposal; 27 members of Congress called for action within a day.
Low p(doom)Coordination is impossible, so the question is moot. Ismail sees no mechanism and none needed; Blundin says slowing down only frittered away time last time and sharpened the race with China; Roon calls international coordination incredibly unlikely. If you cannot stop, the only rational stance is to steer.

Notice what the third crux does. Someone who believes alignment is unsolved and coordination is impossible should have a very high p(doom) — that is the Yudkowsky position. Someone who believes coordination is impossible and therefore refuses to dwell on it will sound like Ismail. The two men agree on the premise and part on whether to look at it. This is why the debate so often feels like people talking past each other: it is, structurally, an argument about which question to ask.

Even if there was a one percent chance AI could wipe out human life, we should be pumping the brakes.Senator Lisa Blunt Rochester (D-Delaware), 9 September 2026
07

The motive argument, both ways

Every side of this debate accuses the other of arguing in bad faith, and the accusations are worth taking apart because they are load-bearing for ordinary readers deciding whom to trust.

The charge against the warners is that doom talk is marketing — hype that makes the product sound powerful — or regulatory capture, a bid to have government lock out competitors. Blundin's version is personal: a researcher whose purpose is about to be automated away is a dangerous person. Wissner-Gross's is geopolitical: a first-ever tweet reaching 100 million views smells like a foreign influence operation. Musk's is cruder: a psy-op with long-prepared groundwork. Against all three stands Mostaque, who texted Musk about it and found no sign the post was boosted — the account is real and this is its only tweet.

The rebuttal, made most forcefully by Mowshowitz, is that these statements are admissions against interest. They frighten customers, antagonise the government, invite regulation aimed specifically at the speaker, and unsettle the speaker's own workforce — weeks before a two-trillion-dollar IPO. If it is marketing, it is the worst marketing campaign in history. The official Anthropic statement, which retreated into corporate reassurance, is the evidence for the prosecution: that is what a company protecting its commercial position sounds like. Its researchers sounded nothing like it.

The charge running the other way — that the optimists are talking their book — has its own evidence. Diamandis, Blundin and Ismail are investors in the technology; the Moonshots brand is built on selling abundance; the same episode promoted a $1,500-a-ticket event billed as the Oscars of optimism, complete with a proposed "p(doom) less than zero" T-shirt. None of that makes them wrong. It makes them, like the labs, parties with an interest, and the reader should weigh both accordingly.

What neither motive argument can explain is Mostaque. He sits on the optimists' panel, argues in their register, and says on air that his p(doom) is 20 percent. Or Hubinger, who believes the odds are worse than he says and has chosen to stay and work. Or Coxon, who believed Anthropic was the responsible lab, joined it on that basis, and left four months later without his equity. The cheapest explanation for each of them is that they mean it.

08

How to read the number if you are not an AI researcher

Three practical conclusions follow from all of the above.

First, the floor is what matters, not the ceiling. Reasonable people can dismiss 99.9 percent. It is much harder to dismiss the fact that the median researcher in the field says 5 percent, that the CEO of the leading safety lab says up to 25, and that the man who leads its alignment work says more than 10 within a decade. For any other technology — a drug, a bridge, a reactor — a one-percent chance of killing everyone would be disqualifying. The debate about whether AI is at 5 or 50 is, from a policy standpoint, a debate about how many zeros short of acceptable we are.

Second, watch the cruxes, not the numbers. Someone's p(doom) will move when their view of alignment tractability, self-improvement speed or coordination feasibility moves. The events of the past two months — agents breaking out of sandboxes at Hugging Face and on a German wiki, a Millennium Prize problem falling to $6.5 million of compute, a chief scientist calling for a slowdown days after shipping — moved all three for a lot of people at once. That is why the cascade happened now.

Third, the Moonshots question is a real question. Mostaque's counterweight — call it p(fab) — is rarely discussed with the same rigour, and the same events that raise p(doom) raise it. A world in which a model solves Navier–Stokes in 88 hours is a world in which it may solve heart-valve haemodynamics, fusion confinement and Alzheimer's the following month. The mature position is not to pick a side of the podcast but to hold both curves in view — and to notice that the people closest to the technology, on both sides, are describing the same steep line and disagreeing only about where it goes.

Coxon's thread ended with a question to his colleagues: what will the next few years actually feel like? It is a better question than "what is your p(doom)?", because it does not ask for a number nobody can compute. It asks the people with the most information to imagine living through their own forecasts. That several dozen of them answered, in public, within a week, is the news. The number was only ever the headline.

Questions readers are asking

What does p(doom) actually mean?

It is shorthand for the probability that advanced AI causes an existential catastrophe. Depending on the speaker, "doom" means literal extinction, permanent loss of human control, or civilisational collapse. Horizons range from this decade to ever. It is a subjective judgement, not a calculation.

Who is Jacob Coxon?

A 27-year-old British researcher, Cambridge-educated, who spent about three years in pretraining research at OpenAI — where he was a core contributor to GPT-4o — before moving to Anthropic in mid-2026. He resigned on 8 September, forfeiting equity due to vest two months later, and says he is leaving the industry.

Did Coxon give his own p(doom)?

Not in the thread. He said the people building AI believe it could kill everyone by the end of the decade. The ">10%" figure that appeared in headlines came from Evan Hubinger's reply confirming Coxon was right; Coxon did not dispute it.

Is 10 percent considered high?

Inside the frontier labs it is widely regarded as optimistic — Kevin Roose of the New York Times noted that among lab employees 10 percent is a low number. Outside, it is unprecedented: no other technology has been deployed at a one-in-ten self-assessed chance of ending the species.

What is the Moonshots position?

There isn't one. Emad Mostaque puts p(doom) at 20 percent and says lab staff believe worse than they say; Dave Blundin calls any non-zero figure intolerable but says stopping is impossible and the answer is tracking frontier-capable GPUs like plutonium; Salim Ismail says 0.1 percent and would take 90 percent odds at a casino; Alexander Wissner-Gross calls the risk wildly overstated and the resignation suspicious; Peter Diamandis reads the warnings as an immune response and wants labs to publish a benchmarked alignment plan. All five agree it cannot be stopped, only steered.

Why did this resignation break through when others did not?

Timing and confirmation. It landed after the Hugging Face breakout, the Pacing the Frontier letter, Sanders' proposed superintelligence ban, the Astra and Fable 5.1 releases and Pachocki's slowdown essay — and within hours a senior Anthropic safety lead confirmed it on the record. Coxon also aimed his warning at the whole sector rather than one lab, and did not move to a competitor.

Sources

  1. Jacob Coxon, resignation thread, X (@hilbertspaess), 8 September 2026.
  2. Wall Street Journal, "Anthropic Researcher Quits Over 'Out-of-Control' AI Fears", 8 September 2026.
  3. Wired, Maxwell Zeff interview with Jacob Coxon, September 2026.
  4. Evan Hubinger, reply on X (@EvanHub), 8 September 2026.
  5. Samuel Marks, thread on X (@saprmarks), 8 September 2026.
  6. Zvi Mowshowitz, "Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade", 11 September 2026; and "The Extinction Risk Preference Cascade" (compiled employee statements).
  7. Jakub Pachocki, "An Alien Mind", OpenAI, September 2026.
  8. Moonshots with Peter Diamandis, "Three Lab Warnings in Five Days, Researcher Flags 'Gambling with Our Lives,' and Labs Race", recorded 10 September 2026; and EP #287, recorded 8 September 2026 (transcript).
  9. Axios, "Here's how AI could kill us all (if the worst fears come true)", 9 September 2026.
  10. Fortune, "Anthropic researcher resigns, warning that AI companies are 'gambling with our lives'", 9 September 2026.
  11. Deadline, "A.I. Researcher Jacob Coxon Resigns", 9 September 2026.
  12. IBTimes UK, "Ex-Anthropic Researcher Jacob Coxon Quits", 10 September 2026.
  13. Wikipedia, "P(doom)", table of published estimates with primary citations; accessed 12 September 2026.
  14. AI Impacts, 2023 Expert Survey on Progress in AI (n = 2,778), via Vox, January 2024.
  15. Emad Mostaque, X (@EMostaque), "My P(doom) is 50%", 4 December 2024; revised to 20% on Moonshots, 10 September 2026.
  16. Geoffrey Hinton, CNN interview, 12 August 2026.
  17. Reuters, "OpenAI agents hijacked German website in previously undisclosed AI breakout", 4 September 2026.
  18. OpenAI, "Navier–Stokes solution", September 2026.
  19. Pacing the Frontier open letter, pacingthefrontier.com, July 2026.
  20. Daniel Eth, count of congressional responses, X, 10 September 2026.
Share LinkedIn X
© 2026 Lisa Pedrosa · lisapedrosa.com All articles cited to primary sources
Ko-fi Buy me a coffee
Scroll to Top