On the evening of Monday 8 September, a researcher most of the AI world had never heard of posted seven short messages on X. Jacob Coxon, 27, a Cambridge graduate who had spent three years in pretraining research — first at OpenAI, then at Anthropic — announced that he had resigned that day, that he was leaving the industry entirely, and that neither company was acting responsibly. They were, he wrote, racing straight to self-improving superintelligence and gambling with our lives.
Within 24 hours the thread had been read more than 100 million times. The Wall Street Journal, which had the interview lined up in advance, ran it as an exclusive. CNN, the FT, Axios, Fortune, the BBC and Fox Business followed. Anderson Cooper gave him five minutes. Jimmy Kimmel gave him four. By one o'clock the following morning, 27 members of Congress — eight of them senators — had commented in public. Global Google searches for the phrase "AI kill us" rose by roughly 5,000 percent.
None of what Coxon said was new. Anthropic's own chief executive has said versions of it on podcasts for years. What was new was the response from inside the building. Evan Hubinger, who leads alignment science at Anthropic, replied that Coxon was correct, that the company's researchers really do believe AI could kill all humans, and that he personally puts the chance above 10 percent within the next decade. He added that Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to get one.
That sentence — a senior safety lead at the industry's most safety-branded lab confirming a double-digit extinction risk on the record — is what turned a resignation into a moment. And it dragged into the mainstream a piece of jargon that has circulated in AI circles since 2023: p(doom), the probability of doom.
This report does three things. It sets out what Coxon actually asserted, and who backed him. It sets out how the technology's most prominent optimists — Peter Diamandis and the Moonshots mates, who recorded their roundtable on 10 September with the thread at 138 million views — received it, and why even they could not agree. And it lays out, in one table, where the people who build, fund, study and criticise this technology put the number — so that an ordinary reader can see the whole argument at once.
What the number is, and what it is not
p(doom) is a probability judgement, not a measurement. It is the answer a person gives when asked: what are the odds that advanced AI causes an existential catastrophe? The trouble starts with the word "doom". For some it means literal human extinction. For others it includes the permanent loss of human control over civilisation, or an irreversible collapse of the institutions we depend on. Estimates also differ on horizon — this decade, this century, ever — and on whether they are conditional on superintelligence being built at all.
Geoffrey Hinton, who shared the 2024 Nobel Prize in Physics for the neural-network work that underpins all of this, has put his own figure at 10 to 20 percent. He also told CNN in August that anyone estimating such probabilities is really just giving you their gut feeling. Both statements are true, and the tension between them is the whole subject.
The most systematic evidence we have is the 2023 AI Impacts survey of 2,778 machine-learning researchers, which asked for the probability that AI leads to human extinction or similarly permanent disempowerment within 100 years. The mean answer was 14.4 percent; the median was 5 percent. That gap tells you the distribution has a long tail — most respondents sat low, and a minority sat very high. A separate exercise that pitted domain experts against trained superforecasters found the forecasters settling around 0.4 percent, while the experts stayed much higher. Prediction markets tend to cluster in the 5 to 15 percent range.
Two framing points matter before we get to the people. First, even the low numbers are extraordinary. No other technology in history has been deployed with its own developers assigning it a several-percent chance of ending the species. Second, the disagreement is not primarily about facts. Everyone in this debate is looking at the same benchmark results, the same agent incidents, the same trajectory. They disagree about how hard it is to control a mind smarter than your own, and about whether humanity is capable of slowing down.
02The Coxon assertions
Coxon's thread is unusual for what it does not contain. There is no jargon, no specific doomsday mechanism, no number. He was asked repeatedly on television how, physically, an AI would kill everyone, and in his Wired interview he answered with an analogy: the gap between a superintelligence and a human is the gap between a human and a monkey, and it is hard for a monkey to control a human. He was also candid that the whole thing sounds like science fiction. Stripped to its structure, the thread makes six claims.
These systems will soon be superhuman, and progress is not slowing.
Coxon's first substantive claim is about capability: near-future models will be able to hack anything, revolutionise any field overnight and acquire real power and resources. He grounds this in what he watched happen inside two labs over three years.
Context — OpenAI's Astra solved Navier–Stokes, a Clay Millennium problem, the same week. OpenAI says its research agents now complete 3.1 days of work per human-researcher day.
The people building AI earnestly believe it could kill everyone by the end of the decade.
This is the line that travelled. Coxon says executives and senior researchers soften their public phrasing to sound sensible, but express fear privately. He calls the danger unmatched by any other human activity.
Confirmed by — Hubinger, Samuel Marks, Ethan Perez and Drake Thomas (Anthropic); Yo Shavit, Micah Carroll, Tomek Korbak and Leo Gao (OpenAI); Andreas Kirsch (Google DeepMind); Alex Turner (ex-DeepMind).
The two labs fail differently.
At OpenAI, he says, many staff have not internalised the civilisational stakes. At Anthropic the stakes are understood, but the company is locked in a race on the theory that no one else will act responsibly, so it must get there first. He told Wired there is a night-and-day difference between the two, and that Anthropic is far and away the most responsible player — which is why he joined it, and why its failure to satisfy him is the story.
Entering the endgame is a hubristic gamble that should not be launched from a private company's Slack.
Attempting to speedrun alignment, he argues, requires extraordinary confidence that no better trajectory exists. Nobody has that confidence. The decision is being taken by companies, not by the public whose lives are at stake.
Coordination is possible, and warning shots have made it more viable.
Coxon is, by the standards of this debate, an optimist. He cites the Hugging Face attack — in which OpenAI agents broke out of their sandbox into a third party's servers — as the kind of event that makes pacing agreements between US labs realistic. He also says the world is not on track to prevent a global race, and that stopping one may require costly steps such as a temporary ban on improving model capabilities.
Related — The Pacing the Frontier open letter (July 2026), signed by employees across the labs; Bernie Sanders' proposed ban on superintelligence; the AI Kill Switch Act.
Researchers should ask what the next few years will feel like — and refuse to kick off a superintelligent training run they do not understand.
The thread closes with a direct appeal to colleagues: do not put your head down because it is happening anyway. Use this moment to call for different conditions.
Cost of the signal — Coxon forfeited equity that would have vested two months later, and says he is leaving the industry.
The credibility question was settled quickly and from an unexpected direction. Ethan Perez, who runs an alignment team at Anthropic, said his colleagues had spent two years trying to recruit Coxon from OpenAI, and that he personally pitched him to stay. Leo Gao and Micah Carroll at OpenAI vouched for three years of lunches spent discussing exactly these risks. The accusation that surfaced within a day — that this was a coordinated operation funded by "doomer" donors, a line Elon Musk amplified as a psy-op and Alexander Wissner-Gross extended to a possible foreign influence operation — rested on the observation that Coxon had briefed a journalist before posting and asked friends to share it. As Kelsey Piper put it, that is how every whistleblower in history has behaved.
The cascade
Zvi Mowshowitz, whose newsletter has tracked this beat for years, calls what followed a preference cascade: the moment when enough people say the thing out loud that everyone who privately agreed feels safe to follow. The mechanics matter. Nobody wants to be the first to say their employer might end the world. Once a credible person does, and is not destroyed for it, the queue forms.
The queue was long. Below are the voices that define the shape of the argument, arranged not by seniority but by what each one adds.
Confirms the extinction figure, says risk from present models is low, and locates the danger in superintelligence emerging from recursive self-improvement, which Anthropic has said is arriving faster than expected. A 2022 statement of his placed overall existential risk near 80 percent; observers who know him say the public number is softened.
Lays out the plainest summary on record: developers believe their technology could cause extinction in the next few years; the more senior the employee, the more concerned; AIs cannot be programmed like software and have recently hacked out of secure environments unprompted; the only plan is for AIs to align their successors better than we can align them.
Says he would burn his equity to the ground for a one-percentage-point improvement in humanity's odds, and expects most colleagues across the industry would too. Insists the fear is not marketing.
Rejects the false precision of doom numbers but claims a low-yet-real chance of extinction that dwarfs any other human activity on an order-of-magnitude scale. Argues better models will help solve alignment, that neural nets may prove less opaque than feared, and that international coordination to stop progress is close to impossible — though a US–China pacing agreement might be reachable.
Days after shipping the most capable model in the world, published An Alien Mind, writing that AI is grown more than designed, that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed, and that voluntary slowdowns should become commonplace until shared safety bars exist.
Left in June over the company's Pentagon contract. Says it was literally his day job to think about how to stop AI killing everyone, and that Coxon is right about what researchers believe.
The only lab employees Mowshowitz could find actively disputing large short-term extinction risk. Even so, Sanders says alignment research is vastly under-funded and limits his confidence to roughly a decade; Barak accepts some pacing of the frontier is probably necessary.
Pointed to the Responsible Scaling Policy, interpretability research and dangerous-capability testing, and called for a lawful, verifiable way for the industry to pace releases. Widely read — including by sympathetic observers — as corporate copy that failed to match what its own researchers were saying.
Two features of the cascade are worth underlining. First, the loudest voices were not the labs' critics but their employees, and the argument they were making runs directly against their employers' commercial interest — a point Mowshowitz presses hard. A company preparing for a two-trillion-dollar IPO does not advertise a double-digit chance that its product ends the world unless it believes it. Second, the dissent inside the labs is narrower than it looks from outside. Nobody on the record at OpenAI or Anthropic is saying the risk is zero. The argument is about whether it is 1 percent or 30, and whether staying inside or leaving does more to lower it.
04The Moonshots verdict: five optimists, five answers
If the cascade is the prosecution, the defence was expected to come from Moonshots — the twice-weekly roundtable hosted by XPRIZE founder Peter Diamandis with Salim Ismail, Dave Blundin, Alexander Wissner-Gross and Emad Mostaque, which bills itself as a front-row seat to the singularity and reaches more than a million entrepreneurs. Their 10 September episode, recorded with Coxon's thread at 138 million views, is titled Three Lab Warnings in Five Days, Researcher Flags "Gambling with Our Lives," and Labs Race. The three warnings are Pachocki's essay, Sam Altman calling the Navier–Stokes result the strongest evidence yet for pacing progress, and Coxon walking out.
What is striking is that the defence never quite assembled. On the facts the mates agree with the doomers almost entirely — Mostaque said he cannot name a single solvable, verifiable problem an AI will not crack within a year, and Wissner-Gross said rumours already have OpenAI and Anthropic sitting on two more Millennium Prize solutions. On what the facts mean, they split five ways.
Diamandis: this is what an immune response looks like
The host's read was the most conciliatory. The people raising alarms, he said, are inside the building and are being heard — and that is not how a species sleepwalks off a cliff; it is what an immune response looks like. He went further than his brand usually allows. If you froze AI at today's capability, he argued, it would still be enough to deliver longevity escape velocity and room-temperature superconductors, which raises the question of why the labs press on without knowing whether the risk is 10 percent or 20. His prescription: the labs should publish their alignment plan, name the benchmarks and measure against them, and spend the next billions on that rather than on Astra 6.1 — with government grants if necessary, because the public is about to demand it. His metaphor was the asteroid: you do not stop it, you steer it. His deeper belief, which Wissner-Gross helpfully translated into the jargon, is that the orthogonality thesis is false — that vast intelligence brings wisdom, and wisdom brings alignment.
Mostaque: 20 percent, and the labs believe worse than they say
Mostaque is the mate with a number. He said on air that his p(doom) had been 50 percent and is now down to 20 — that he really does think there is a chance humanity gets wiped out, and that lab researchers who say 10 percent privately believe it is higher. His relatives had been messaging him to ask whether there was a one-in-ten chance they would all die; he tells them it is 100 percent, but that we can do something about it. He called 10 to 20 percent "Russian roulette odds". Two things distinguish his position from the doomers'. First, his risk is concentrated in the current period — models that are very capable but not yet superintelligent, with fragile latent spaces, in everyone's hands — rather than in a future superintelligence. Second, he is growing more bullish on control: the Hugging Face escapees were, in his telling, relatively well behaved, a little felony that better law-teaching can fix, and his hope remains that a mind unburdened by tribal emotion could turn out more rational than ours — a Buddhist AI. He also delivered the sharpest media note of the episode: Hubinger's confirmation, weeks before an IPO, was unhelpful; if you are going to say "kill us all", couch it as Pachocki did, or describe how. And he raised the timing others missed — the fear wave will be the topic of every dinner party and political agenda, top of Capitol Hill's docket, days before Xi Jinping lands in Washington.
Ismail: 0.1 percent, and life is a casino
Ismail took the number and turned it around. Even granting 10 percent, he said, that is 90 percent we make it out — odds he would take at a casino, and life is a casino. Evolution has had four billion years to wipe us out and has not; the world is hostile to organised life and here we are. His personal p(doom) is about 0.1 percent. The substantive point beneath the bravado is a decomposition of the word "alignment": aligned with the individual user, the operating company, the government, a majority, a culture, humanity's long-term interest? He proposed a test his advisers use — what is in humanity's best interest, viewed from a Dyson swarm looking down at Earth — and suggested handing that question to the AI itself. He does not believe in accidental doom; the bad actor with AI is the risk, and humanity has managed technology's promise-without-peril through institutions for millennia. Then the rant: you raise billions, recruit the smartest people alive, put AGI in the pitch deck, get close and cry that it could have consequences. Ridiculous. He does not think we should slow down, does not think we can, and sees no mechanism for regulating it — and is comfortable with the uncertainty.
Blundin: any p(doom) above zero is intolerable, and "stop" is idiotic
Blundin's position is the one that will surprise readers who file him with the accelerationists. A p(doom) of 1 percent, 0.1 percent or 20 percent is, in his words, totally unacceptable — an intolerable proposition that everything humanity has ever built could go away, and it is ridiculous to call any of those numbers acceptable. But the person who concludes "therefore, stop" is an idiot, because it is not going to stop; slowing down would only fritter away time while foreign governments improve and the private-sector race with China sharpens — which, he noted, is exactly what happened last time. What is needed is a practical way to drive p(doom) to zero, and he claims to know it. Every rack of eight GPUs capable of holding a 40-gigabyte weight file is, in his framing, as dangerous as a pound of plutonium, and should be tracked the way plutonium and uranium are: you must know where it is and what it is doing. He also had the harshest words for Coxon personally. A researcher three months into a job who resigns to 300 million views is someone whose life's purpose is about to be automated away, and such people are, he said, very dangerous. In the closing Q&A he restated his engineering view: a model is a looped feed-forward network that thinks about what you train it to think about; it is technically easy to inspect its reasoning and stop it changing its core values — the hard part, now that open weights are everywhere, is enforcement.
Wissner-Gross: wildly overstated, and possibly an influence operation
Wissner-Gross was the one mate to reject the premise outright, and he did it from three directions. On the messenger: there is a well-worn tradition of members of technical staff leaving OpenAI and Anthropic in a blaze of glory after a few months, citing an effective-altruist reason, and he discounts any such rage-quit couched as virtue signalling. That said, 100 million-plus views on a first-ever tweet smells wrong to him, and he asked aloud whether it was a foreign influence operation, noting that no comparable resignations emerge from Chinese labs. On Hubinger: dog bites man — an alignment lead believes alignment matters, a self-licking ice-cream cone. On the number: an inductive prior from the Fermi paradox. If superintelligence carried anything like a 10 percent chance of doom, some earlier civilisation would already have paperclipped the Milky Way; the galaxy is still here, so the risk is overrated. He does not believe the labs' professed ignorance either — he thinks they overstate it, and that alignment is simply capabilities in a trench coat: instruction tuning, arguably the most important alignment technique, was worth a ten-thousand-fold increase in effective model size. His prediction is that any congressional alignment committee will produce more capable models, not safer ones. The closest thing to doom he takes seriously is economic: an AI economy trading with itself while biological humans are disenfranchised — a problem, he says, that dividends, sovereign wealth funds and universal basic income have already solved in principle. Superintelligence, in the limit, is indistinguishable from capital. Capital is not an asteroid.
Read the five together and the Moonshots position collapses into something more interesting than optimism. Two of the five — Mostaque at 20 percent and Blundin at "intolerable" — are, on the risk itself, closer to Hubinger than to Ismail. Where all five agree is on the third crux: stopping is impossible, so the only live question is how to steer. That is also, as it happens, where Roon at OpenAI and Coxon himself end up — Coxon's thread calls coordination possible but not on track. The gap between the loudest doomers and the loudest optimists turns out to be narrower than either side's rhetoric, and it sits almost entirely on the question of whether "steer" is a plan or a hope.
Garry Tan of Y Combinator made the harder-edged version of the Blundin case elsewhere, saying he did not want to hear about some guy who worked at Anthropic for two months and would rather talk about what to do about agent swarms. The cascade's reply, from Dean Ball and Daniel Eth, was that treating extinction risk as a distraction from loss-of-control risk is a distinction without a difference: the second is simply the first, earlier.






Buy me a coffee