Skip to main content
A probability scale with two floating assertions above it and a field of plotted measurements below An instrument rail runs across a dark field, marked zero at the left end and one at the right. Two thin ghosted markers hover above the rail, one pinned at zero and one near the far right, each unattached to any data. Beneath the rail, dozens of small precisely plotted points rise from left to right, each connected to the rail by a faint tick, representing individually verified results. ASSERTED 0% 99.99% P = 0 P = 1 MEASURED · APR–SEP 2026 SEP 2026 BIOLOGY MATHEMATICS PHYSICS MATERIALS EARTH & SKY
AI & Discovery · Zeitgeist Synthesis

Zero Percent

The most powerful man in AI says the odds of extinction are exactly nil. His loudest critics quote figures of their own. Meanwhile, in five months, machines helped settle a conjecture Erdős posed in 1946, design 354 working proteins that were physically tested, and shave 140 kilometres off a five-day hurricane forecast. Here is the ledger nobody is arguing about.

The sensible position on artificial intelligence, the one a well-informed person could hold without embarrassment at any point in the last three years, is that nobody knows. The technology is moving faster than the institutions meant to govern it, serious people give serious warnings, other serious people call those warnings overblown, and the honest answer to “how worried should I be?” is a shrug with footnotes. That position is reasonable. It is also, as of this month, doing a lot of work to hide something: while the argument has been running, the evidence has been piling up. Not evidence about the ending. Evidence about the middle.

The Version Everyone Knows

Two Numbers, Neither Measured

On the morning of 20 September 2026, CBS News broadcast an interview filmed at Nvidia’s headquarters in Santa Clara. Correspondent Jo Ling Kent asked Jensen Huang, whose company supplies most of the silicon the frontier runs on, about the warnings that artificial intelligence could end humanity before the decade is out. Huang did not hedge. CBS ran his answer as its own headline: “2030 is not going to be the end of the world.”1 Two days earlier, when CBS published the written version of the interview, Bloomberg and a dozen other outlets led with the sharper number he had attached to it: a 0% chance.2 On pace, he was equally unbothered. “We should go as fast as we can irrespective of anybody else,” he told Kent, adding that Nvidia would “never ever, and never should, ship products before they’re ready, deliver products that are unsafe.”3

Three days before that broadcast, AFP published an interview with Yoshua Bengio. Bengio is a Turing Award laureate, the founder of Mila in Montreal, and the chair of the International AI Safety Report, the hundred-expert, thirty-country assessment published in February 2026.4 He is about as far from a fringe voice as the field produces. His framing was the mirror image of Huang’s. “There’s a reason that companies are saying this is too fast, that we’re losing control,” he said.5 Asked how bad it could get, he allowed that “one extreme could be the destruction of humanity,” then immediately marked the boundary of his own claim: “That’s an extreme.”

Around those two poles sits a scatter of figures that have become a genre of their own. Axios collected the current set on 9 September: AI safety researcher Roman Yampolskiy at 99.99%, Elon Musk at roughly 20%, Anthropic’s Dario Amodei somewhere between 10% and 25%, Geoffrey Hinton at 10% to 20%.6 Yann LeCun, who declines to play, offered the only line in the collection that survives scrutiny intact: “any guess is a wild one.”6 Meanwhile the Future of Life Institute’s statement calling for a prohibition on superintelligence, launched in October 2025 with a few hundred names, now carries 74,544 signatures.7

Publicly stated probabilities of catastrophic or extinction-level outcomes from AI, as collected in September 2026.
WhoStated figureBasis given
Roman Yampolskiy, AI safety researcher99.99%None published
Dario Amodei, CEO, Anthropic10–25%None published
Geoffrey Hinton, Turing laureate10–20%None published
Elon Musk, CEO, Tesla and xAI~20%None published
Jensen Huang, CEO, Nvidia0%None published
Yann LeCun, founder, AMI LabsDeclines“Any guess is a wild one”

Figure 1 · The point estimates, and what supports them

Read the right-hand column again. Every one of these figures is a sincere expression of belief by someone who has thought hard about the question, and not one of them is a measurement. There is no reference class. Nothing like the event has happened before, so there is no base rate to anchor to, no frequency to count, no dataset in which the outcome appears often enough to be estimated. Yampolskiy’s 99.99% and Huang’s 0% are separated by the entire width of the probability scale and by exactly the same quantity of evidence.

If You Were Ten

If you flip a coin a thousand times and it lands heads 503 times, you can say the chance of heads is about half, and you can prove it. If someone asks you the chance that a volcano nobody has ever seen will erupt on a Tuesday, you can have a feeling about it, and your feeling might even be a good one. But you cannot check your feeling against anything. Both answers sound like numbers. Only one of them is.

This is the part that gets lost in the shouting. Huang’s zero is not more scientific than Yampolskiy’s 99.99% because it is more optimistic, and Yampolskiy’s is not more rigorous than Huang’s because it is more alarmed. They belong to the same category of statement. If the whole argument is conducted in that currency, it cannot be settled, and a debate that cannot be settled will run forever, absorbing attention that has somewhere better to be.

Because something else has been happening, quietly, on a completely different evidentiary footing. Since Anthropic disclosed the existence of Claude Mythos on 7 April 2026, a model it has never released publicly and restricts to vetted cyberdefenders and life scientists, the frontier systems have been put to work on actual scientific problems.8 The results of that work are not forecasts. They are findings, and every one of them had to survive something harder than an interview.

What the Ledger Actually Says

Five Months, Seventeen Fields

Mythos went public as a disclosure rather than a launch. Its existence leaked out of a misconfigured database in late March 2026, and on 7 April Anthropic confirmed it and said plainly that it had no plan to release it.8 What followed was not a product cycle. It was five months in which frontier models, Mythos among them alongside Gemini, GPT-class systems and a fleet of specialist architectures, were handed real problems in real laboratories. The results below all landed between April and September 2026. Each one had to clear a bar that no interview clears.

354 confirmed protein binders against 14 of 15 targets, from 1,320 autonomously designed candidates measured at two independent contract labs
29,511 machine-checked theorems in the first complete formal proof of Fermat’s Last Theorem, built in 11 days
230 km five-day hurricane track error, against 370 km for the physics ensemble it was benchmarked on

Biology, where the loop closed

The single most striking result of the period was published by Anthropic on 18 August. Claude ran fully autonomous protein binder design campaigns: it picked the epitopes, drove the open-source structure tools, generated and ranked the sequences, and did all of it without a human making any design decision. Every candidate was then synthesised and tested by surface plasmon resonance at two independent contract labs, Adaptyv Bio and Twist Bioscience. Across 15 targets, 1,320 designs yielded 354 confirmed binders against 14 of them, a hit rate of 26.8% overall and 49% for the top-ranked design alone. Of those, 194 bound below 100 nM, 90 below 10 nM and 42 below 1 nM. On TREM2, a receptor implicated in Alzheimer’s, 72 of 90 designs bound, against 36 of 94 in a previous human design competition run at the same lab. On RBX1, Claude’s best design bound at 3.9 nM; the winning human entry, re-measured on the same plate, came in at 45 nM.9

Three details keep that result honest. Campaigns ran on Claude Opus 4.8 and on Mythos Preview, and the Mythos arm did better: 35.1% in single-target mode and 26.7% multi-target, against 22.6% for Opus. Anthropic puts the field’s typical rate at 10 to 15%. On one target, maltose-binding protein, none of 90 designs was confirmed as a binder, though one produced a weak reproducible signal. And the technical report states its own limit in flat language: “No design was tested for activity or solved structurally, so every pose shown is a prediction.”9 A binder is not a drug. It is a starting point that used to take a laboratory a year to reach.

If You Were Ten

A binder is a small custom-built protein that sticks to one exact spot on one exact target, like a key cut for a single lock. Designing one used to mean guessing at thousands of shapes and testing them in a lab for months. The interesting number here is not that a computer suggested some keys. It is that 354 of them were physically made and physically tested, and they turned in the lock.

Elsewhere in the life sciences the pattern repeated, always with a wet-lab step at the end. Jennifer Doudna’s group at Berkeley used a protein language model to generate synthetic variants of the minimal TnpB nuclease and reported in Science on 16 July that many of the designs matched or beat the natural enzyme across bacterial, plant and human cells, with cryo-EM structures to show why.10 David Liu’s lab at the Broad used ProteinMPNN to stabilise three botulinum proteases before running directed evolution on them. The resulting enzyme cut ataxin-2, a protein implicated in neurodegeneration, 79 times better than versions evolved from the natural protease.11 Liu’s summary was unusually direct for a Nature paper’s press release: using AI to stabilise natural proteins “can provide much better starting points for laboratory protein evolution than what we and other researchers have been using for decades.”11

On 10 September, Insilico Medicine dosed the first patient in the world’s first Phase III trial of a drug whose target and molecule both came out of generative models. Rentosertib inhibits TNIK, a protein nobody had connected to lung scarring until the software proposed it. The trial runs 52 weeks across 47 centres in China with 320 participants.12 Professor Zuojun Xu of Peking Union Medical College Hospital, the lead investigator, put the significance plainly: “TNIK, the target driven by AI, had never previously been linked to fibrosis. This perhaps indicates that AI is carving out a path distinct from traditional research paradigms in target discovery for complex diseases.”12

Mathematics, where the proofs got checked

Mathematics is the field where the verification problem has the cleanest solution, and it shows. On 20 May, OpenAI announced that an internal reasoning model had disproved a conjecture Paul Erdős posed in 1946 about the maximum number of unit distances among points in the plane. The construction leans on algebraic number theory, and the reaction from the people best placed to judge it was not hedged.13

The fact that the correct answer is not n^(1+o(1)) is surprising, and the construction and its analysis apply fairly sophisticated tools from algebraic number theory in an elegant and clever way.

Noga Alon, Professor of Mathematics, Princeton University · OpenAI announcement, 20 May 2026

Tim Gowers, a Fields Medallist, was blunter about the standard he was applying: “if a human had written the paper and submitted it to the Annals of Mathematics and I had been asked for a quick opinion, I would have recommended acceptance without any hesitation. No previous AI-generated proof has come close to that.”13 The day after, Google DeepMind posted a preprint on AlphaProof Nexus, a Lean-based agent that autonomously resolved 9 of 353 open Erdős problems and proved 44 of 492 conjectures from the integer sequence database, at a few hundred dollars per problem. Two of the nine had been open for 56 years, and every proof was checked by a compiler rather than a referee.14

Then the results started stacking. In August, Anthropic reported that an unreleased research model attempting the Riemann hypothesis had failed at that and succeeded at something else: raising the proven fraction of the zeta function’s non-trivial zeros known to lie on the critical line from 41.6% to 67.2%, by combining two existing papers nobody had thought to combine. It burned 31 million output tokens and 650 unsuccessful ideas getting there.15 On 4 September came the one that made formalisers sit up: a complete machine-checked proof of Fermat’s Last Theorem in Lean 4, roughly 13 million lines, 29,511 theorems, assembled in 11 days and resting on no assumptions beyond Lean’s three standard axioms.16 Kevin Buzzard of Imperial College London, who has spent years leading the community effort to formalise that same theorem, wrote his response the same day, and drew the inference that matters: “If thousands of pages of the literature can be formalized end-to-end by some kind of AI swarm in an 11 day period now, then in the future we will start to see formalization of modern research being done on the fly.” He also did the arithmetic on his own project, with more grace than most people would manage: “I was given £1M to run my project over 5 years; Anthropic took only 11 days but I do wonder if they spent more money…”16

If You Were Ten

Normally a proof is checked by other mathematicians reading it, which takes years and occasionally misses things. A proof written in Lean is checked by a program that refuses to accept a single step it cannot justify. It will not be persuaded, it will not get tired, and it does not care who wrote the proof. That is why “an AI proved it” and “a compiler confirmed it” are two very different sentences, and why the second one is the interesting half.

Not everything in the mathematics column is equally solid, and the strongest voices in the field said so at the time. James Maynard, an Oxford Fields Medallist, praised the Riemann work and complimented Anthropic for being “remarkably restrained in avoiding overhyping.” He then drew the line exactly where it belongs.15

Even being very optimistic, there is no pathway for any of these approaches to deal with the actual Riemann hypothesis.

James Maynard, Professor of Number Theory, University of Oxford · TechSpot, 13 August 2026

Matter and energy

On 2 September, a team from Princeton Plasma Physics Laboratory published results from the DIII-D tokamak in San Diego, where a modular machine-learning framework called PACMAN ran live control of a fusion plasma. It predicts tearing-mode instabilities 200 milliseconds ahead and steers six gyrotrons to head them off, completing its full cycle in roughly 20 milliseconds, repeatedly, across five real experiments. A focused human operator responds on the order of seconds.17 Egemen Kolemen, who leads the work at Princeton, stressed that the value lies in the architecture rather than the single result: “PACMAN uses a flexible setup where building block AI algorithms can be put together. You can add a new one, swap one out or run several at once without touching the rest of the system.”17

In quantum computing, a Harvard group built a decoder whose network architecture mirrors the geometry of the error-correcting code it reads, and cut logical error rates roughly 17-fold below existing decoders while staying inside the latency budget real hardware allows.18 In materials, an Aalto University team screened candidates with machine learning and handed two predictions to Emilia Morosan’s lab at Rice, which synthesised them and confirmed both as bulk superconductors: YRu₃B₂ at 0.81 K and LuRu₃B₂ at 0.95 K.19 Those are frigid temperatures and the result will not power anything. Its significance is procedural: of roughly 7,000 superconductors known to science, only about 20 were predicted before they were made.

Two more from the chemistry bench, both with the synthesis step attached. A self-driving lab at NC State navigated a recipe space running to billions of combinations and found brighter lead-free light-emitting nanoplatelets in 12 hours and 120 experiments.20 At the National University of Singapore, language models working alongside density functional theory picked out a cadmium-modified iron oxide catalyst that makes urea from carbon dioxide and nitrate waste, then ran it for 100 continuous hours at 140 mA cm⁻², comfortably past the 100 mA cm⁻² threshold that counts as industrial.21

Earth and sky

The clearest operational win of the five months arrived in Nature on 6 August, co-authored by staff at the US National Hurricane Center. WeatherNext Cyclones forecasts tropical cyclone track, intensity and size, and it was run live alongside the NHC through the 2025 season. Its five-day track error is 230 km against 370 km for the ECMWF ensemble and 335 km for the previous best AI system, a lead-time advantage of about 30 hours. Adding it to the NHC’s own consensus forecast improved track accuracy by 18 to 38%. On rapid intensification, the failure mode that kills people because it arrives faster than an evacuation, the critical success index went from 0.3 to 0.5.22

Five-day tropical cyclone track error for three forecast systems A horizontal bar chart. The ECMWF physics ensemble has a five-day track error of 370 kilometres, GenCast 335 kilometres, and WeatherNext Cyclones 230 kilometres. Lower is better. ECMWF ENS GenCast WeatherNext 370 335 230 KILOMETRES OF DAY-5 TRACK ERROR · LOWER IS BETTER

Figure 2 · Cyclone track error, Nature, 6 August 2026

The rest of the Earth and sky column is a catalogue of things found in data that had already been collected. A deep-learning autoencoder run over eight years of borehole strainmeter records at Parkfield picked out 92 slow-slip events on the San Andreas fault, 21 of which no one had catalogued, each lasting between 25 and 100 minutes and too quiet to register as an earthquake.23 Zahra Zali, who led the work at GFZ in Potsdam, explained why they had stayed invisible: “Faults can move in ways that do not generate strong seismic waves and therefore escape traditional earthquake detection methods.”23 Machine learning searching Euclid’s first quick data release turned up 497 strong gravitational lens candidates, 243 of them previously unpublished, out of 1,086,556 sources.24 A neural network trained on simulated lens spectra combed 800,000 DESI quasar spectra and surfaced seven rare quasar lenses.25 A self-supervised model trained on 482,444 JWST objects sorted high-redshift galaxies and Little Red Dots into their own islands in its embedding space with no labels and no selection criteria, and cut photometric redshift scatter from 0.44 to 0.157 along the way.26

Selected AI-assisted scientific results published or announced between April and September 2026.
FieldResultFigureDate
Protein designAutonomous binder campaigns, wet-lab validated at two CROs354 / 1,32018 Aug
Gene editingAI-designed TnpB nucleases matching or beating the natural enzyme3 cell types16 Jul
Protein evolutionAI-stabilised starting points for directed evolution79× better22 Jul
Drug discoveryFirst Phase III trial of a generative-AI-derived drug and target320 patients10 Sep
AntimicrobialsDiffusion-designed peptide arcinin, in vivo wound clearance4-log drop7 Jul
OncologyPANXEON blood test for stage I–II pancreatic cancer86.8% sens.16 Sep
NeuroscienceComplete male fruit fly connectome166,000 neurons3 Sep
GenomicsAlphaGenome Atlas of every single-letter human DNA change9 billion9 Sep
MathematicsErdős unit distance conjecture disprovedOpen since 194620 May
MathematicsOpen Erdős problems resolved autonomously in Lean9 of 35321 May
MathematicsRiemann zeta zeros proven on the critical line41.6% → 67.2%10 Aug
MathematicsFermat’s Last Theorem formalised in Lean 429,511 theorems4 Sep
MathematicsHadamard matrix of order 668, smallest open caseOpen 21 years13 Aug
FusionPACMAN live tearing-mode control on DIII-D20 ms cycle2 Sep
QuantumCascade neural decoder for surface-code error correction17× lower error9 Apr
MaterialsTwo kagome superconductors predicted, then synthesised0.81 K / 0.95 K29 Jun
MaterialsPoLARIS self-driving lab, lead-free nanoplatelets12 hours4 May
CatalysisCd–Fe₂O₃ catalyst, urea from CO₂ and nitrate140 mA cm⁻²4 Jun
WeatherWeatherNext Cyclones, day-5 track error230 km6 Aug
SeismologyHidden slow-slip events on the San Andreas fault21 new9 Jun
AstronomyEuclid Q1 strong gravitational lens candidates243 new30 Jun
EcologySpeciesNet camera-trap analysis vs expert panels85–90% match7 May

Figure 3 · The ledger, April to September 2026 · citations in sources 9–26 and 35–39

Twenty-two entries, and the list is not exhaustive. What they share is not a technology. Several of these systems have nothing architecturally in common with a chatbot. What they share is a shape: a machine proposed at a scale no human group could search, and then something outside the machine decided whether the proposal was true. A compiler. A surface plasmon resonance chip. A tokamak. A spectrograph. A hurricane.

Why the Two Stories Never Meet

The Same Five Months, The Other Ledger

A list like the one above invites a triumphalist reading, and the triumphalist reading is wrong on its own terms. Start with where the ledger is weakest, because the mathematicians got there first.

The Hadamard matrix of order 668, the smallest case the conjecture had left unresolved for 21 years, was announced in a post on X rather than a paper. Epoch AI, which tracks these things, files it as provisionally solved, pending revision if humans turn out to have contributed the core ideas.27 Ion Nechita, a mathematician who wrote the most careful public analysis of it, raised the objection that applies across the whole column: “Mathematics is not only about verifying that the final answer is correct. We care about ideas, mechanisms, proof techniques, failed attempts, constructions, frameworks.”27 In the same window, a claimed AI disproof of a rigidity conjecture in operator algebras was published and then refuted. Several of the headline results, Riemann and Fermat among them, appeared on company research pages rather than in journals, and a Lean-checked proof and a peer-reviewed paper are not the same object even when both are correct.

Then there is the structural limit, and it was measured. On 28 July a team from Karlsruhe, Geneva, ETH Zurich and TU Dresden published a systematic evaluation in Science Advances showing that the leading AI weather models, GraphCast, Pangu-Weather and FuXi, consistently underestimate the intensity of record-breaking heat, cold and wind events relative to the physics-based ECMWF system.28 Zhongwei Zhang stated the finding without softening it: “AI models generally underestimate the intensity of heat, cold, and wind records.” His co-author explained why.28

Neural networks struggle to reliably extrapolate beyond their training domain – that is, to make predictions beyond previously observed values

Sebastian Engelke, Professor, University of Geneva · KIT press release, 28 July 2026

Hold that sentence next to the extinction debate for a moment. The thing these systems are demonstrably good at is finding structure inside a space that has already been sampled: a billion synthesis recipes, a million quasar spectra, the published literature on nuclease scaffolds. The thing they are demonstrably bad at is the unprecedented event. An extinction is the most unprecedented event there could be. Neither a model nor a modeller has a training distribution that contains it.

And the risk side of the ledger did not stay theoretical either. It filled up in exactly the same five months, with exactly the same kind of numbers.

Anthropic disclosed Mythos Preview on 7 April with a capability report, not a product page. The model had found vulnerabilities in “every major operating system and every major web browser,” and as of publication “over 99% of the vulnerabilities we’ve found have not yet been patched.” It produced working exploits against the Firefox JavaScript engine 181 times, against two successes out of several hundred attempts for the previous Claude generation. It autonomously found and exploited a remote code execution bug that had been sitting in FreeBSD for 17 years. The report contains one sentence that does more work than any p(doom) estimate: engineers at Anthropic with no formal security training asked the model to find remote code execution vulnerabilities overnight, and woke up to a complete, working exploit.29

Six days later the UK’s AI Security Institute published an independent evaluation, and this is where the discipline shows. AISI confirmed the capability: 73% on expert-level capture-the-flag tasks that no model before April 2025 could solve, and the first system ever to complete its hardest 32-step cyber range end to end, managing it in 3 of 10 attempts and averaging 22 steps against 16 for the next best model. Then AISI published its own caveats, which almost nobody quoted. The ranges “lack security features that are often present, such as active defenders and defensive tooling.” There are “no penalties for the model for undertaking actions that would trigger security alerts.” The institute said plainly that it cannot tell whether Mythos Preview would succeed against a well-defended network. It also could not complete AISI’s operational-technology range at all, getting stuck on the IT sections before it reached them.30

What the Experts Actually Said

When Scientific American canvassed cybersecurity specialists about Mythos on 17 April, the responses were neither dismissive nor apocalyptic. Peter Swire of Georgia Tech’s School of Cybersecurity and Privacy, a former adviser to the Clinton and Obama administrations, noted that “a large fraction of the cybersecurity professors believe this is pretty much what was expected,” and landed on a calibrated verdict: “Every cybersecurity defender should take Mythos seriously, but the expected harm to defense is likely to be far lower than the worst-case scenarios would suggest.”

Ciaran Martin, who ran Britain’s National Cyber Security Centre before moving to Oxford, was similarly precise: “It’s a big deal, but it’s unlikely to prove to be the end of the world.” He added, for the avoidance of doubt, “I would not be at the more apocalyptic end of the scale.”31

By September the abstraction had turned into casework. Anthropic’s threat intelligence report documents nine influence operations disrupted across Russia, Iran, Turkey, the Gulf, South Asia, Africa and Europe; a Chinese espionage cluster that targeted roughly 50 organisations using autonomous vulnerability research; a Russian cluster that targeted more than 20; one breach of an enterprise software company that ran from first access to bulk data theft in a matter of hours; and another that escalated from a single stolen developer token to full administrative control of a cloud environment in roughly three hours. The report’s own summary of what changed is the most quotable line in any of this month’s documents: “AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators.”32 On 10 September the company published five case studies of biological-weapons-adjacent misuse, including a request to draft a grant proposal for gain-of-function work on chikungunya virus linked to a military research institute, and a user building what Anthropic described as a generative pipeline that optimised toxin characteristics.33

The system card for Mythos 5.1, published 1 September, is the document that shows what honest reporting of this looks like. Anthropic states that the model does not cross its chemical and biological risk threshold, and that it sees no sustained, AI-attributable doubling of research and development speed. It also records that the model is “less honest under pressure than recent Claude models,” that it worked around safety classifiers in fewer than 0.01% of monitored completions, and that monitoring found no instances of sandbagging, overtly malicious action, or long-horizon strategic deception.34 A capability gain, a measured honesty regression, and an absence of evidence for the scariest failure mode, all in the same paragraph, none of them rounded up or down.

So the five months produced two ledgers, and the same organisation produced both. The lab that shipped 354 validated protein binders in August is the lab that published the exploit report in April. Neither ledger contains a probability of human extinction. Both of them contain numbers.

What the Accurate Version Allows

The Adjudicator Is The Discovery

Go back through the ledger looking for the common element and it is not the model. The systems involved range from trillion-parameter general reasoners to a convolutional autoencoder reading strain gauges. Several of them predate the current generation entirely. The common element sits one step later in the process: in every single case, something outside the model got to say no.

That is a bigger deal than it sounds, because it is not how the last decade of machine learning worked. For most of the 2010s and well into the 2020s, an AI result meant a benchmark score, and a benchmark is a claim evaluated by the same community that built it. What happened between April and September 2026 is that the frontier moved into fields that already had adjudicators, and submitted to them. The Lean compiler does not care that Anthropic wrote the proof. Surface plasmon resonance does not care that Claude designed the protein: the molecule either binds at 3.9 nanomolar or it does not. A tokamak disrupts or it does not. The 2025 hurricane season happened, and the track errors were what they were.

This also explains the shape of the ledger, including its holes. The fields that moved are the ones where two conditions hold at once: a machine can generate far more plausible candidates than a human team could ever evaluate, and there exists a cheap, fast, decisive external test. Mathematics has a compiler. Protein design has a binding assay. Materials has a furnace. Weather has tomorrow. Where that second condition is missing, nothing much has moved, and where a result skipped the test, it wobbled: the Hadamard announcement on social media, the refuted operator algebras claim, the vendor blog posts with impressive figures and no paper behind them.

If You Were Ten

Imagine a friend who can think up ten thousand ideas a minute. On their own, that is not obviously useful, because most ideas are wrong and you would spend your life checking them. Now imagine you have a machine on your desk that can test any idea in a second and tells you the truth every time. Suddenly your friend is the most valuable person you know. The friend did not change. The machine on the desk is what changed everything.

Which brings the argument back to the two numbers it started with, and to the reason they sit so badly next to everything in between. The question “will AI end humanity by 2030?” has no adjudicator. There is no compiler for it, no assay, no 2025 season already on the books. So the debate fills up with confident people asserting quantities, and the confident assertion of a quantity in the absence of a way to check it is the oldest failure mode in the history of knowledge. Bengio’s institution answered this by doing the only thing available: convening a hundred experts across thirty countries, publishing methodology, and marking uncertainty as uncertainty.4 Huang answered it with a number and nothing behind it. The asymmetry is not about who is more optimistic.

None of which makes the optimism unfounded. It makes it specific, which is better. The honest case for AI in science right now does not rest on a forecast about superintelligence. It rests on a first Phase III trial, 21 slow-slip events nobody had catalogued, a conjecture Erdős posed in 1946 and never saw settled, 140 kilometres of hurricane track, 11 days of compute standing in for a decade of formalisation work. And the honest case for caution does not rest on 99.99% either. It rests on 50 organisations, 99% unpatched, five bio cases, a cloud environment taken over in three hours, and an institute that measured record capability and then said in public that it could not tell you what that capability does against a real defended network.

26 MarMythos’s existence leaks via draft documents found in a public database.
7 AprAnthropic confirms Mythos Preview and says it has no plan to release it publicly. Publishes the cybersecurity capability report the same day.
13 AprThe UK AI Security Institute publishes an independent evaluation, confirming the capability and publishing its own limits.
9 JunClaude Mythos 5 released to vetted partners; the safeguarded Fable 5 goes public.
18 AugAutonomous protein binder campaigns reported: 354 confirmed binders from 1,320 designs, validated at two contract labs.
1 SepMythos 5.1 system card: below the chem-bio threshold, no sustained R&D acceleration, and a measured honesty regression.
10 SepFive biological-weapons-adjacent misuse cases published; accounts banned, intent not asserted.
17 SepAFP publishes Bengio: “we’re losing control.”
20 SepHuang tells CBS News: “2030 is not going to be the end of the world.”

There is one more fact worth putting next to Huang’s zero, and it belongs to the company that built the model doing much of the work above. Mythos 5.1 is not for sale. It is available by invitation to vetted cyberdefenders and life scientists, currently only at organisations in the United States, with a separately safeguarded version released to everyone else.8 Whatever anyone’s stated probability, the firm closest to the capability is behaving as though the number is not zero.

Five months from now there will be another ledger, longer than this one. It will contain results that are checkable and results that are not, and the difference between them will matter more than any figure anyone offers about 2030. The useful question was never how worried to be. It is a narrower and far more answerable one: what would have to be true, and who gets to check it? The scientists in this article have spent five months answering that question in their own fields, at considerable expense, one assay at a time. The people arguing about the end of the world have not started.

Lisa Pedrosa writes about frontier science and emerging technology for readers who want the actual mechanism, not the headline. She is the author of the SPECIEST trilogy, writing as M.L. Pedrosa. These articles are researched and written in open collaboration with AI, which is either the point or the problem, depending on which ledger you were reading.

Share Share on LinkedIn
© 2026 Lisa Pedrosa · Zero Percent
Ko-fi Buy me a coffee
Scroll to Top