Anthropic's own people say AI could kill everyone. Show me the money.
TL;DR [show]
On 2026-09-08 Jacob Coxon resigned from Anthropic and warned that AI could kill everyone by the end of the decade; the post cleared 150 million views and Anthropic's alignment-science lead Evan Hubinger put his own figure above 10 percent within the decade. The piece holds that cycle to a standard set three years earlier by the one organisation that did the pathway work. RAND's 2025 report On the Extinction Risk from Artificial Intelligence stated a falsifiable hypothesis, that there is no describable scenario in which AI is conclusively an extinction threat to humanity, then recruited experts to break it across three physical routes. Nuclear cleared. Engineered pathogens and malicious geoengineering did not: RAND names both as true extinction threats and potential falsifications of its own hypothesis. The headline finding is narrower than either side reports, that extinction is not a plausible outcome unless an actor is intentionally seeking it. RAND also holds that stated probabilities of extinction are inappropriate as analytical tools, and that the shut-it-down approach to AI governance is inappropriate given the deep uncertainties. Against that standard the viral cycle offered sandbox evidence, containment-test incidents that are ordinary cyber and social-engineering events, plus a personal credence with no denominator. Meanwhile three advocacy groups moved the post within fifteen minutes, all tied to a funder who backed Anthropic's Series A, and Cohere's Aidan Gomez called the resulting slowdown proposals a cartel by any other name. The close argues the real contest is a Kardashev-scale civilisation against a Marxist-Leninist one that is eight months behind, has stopped publishing its weights, and has more electricity.

# Anthropic's own people say AI could kill everyone. Show me the money.
On September 8 a 27-year-old researcher resigned from Anthropic, posted his reasons, and by the following morning had been read more times than there are people in Japan. Jacob Coxon's post cleared 100 million views inside 24 hours. It kept going. Past 133 million. Past 150 million.
Governors amplified it. Senators amplified it. Illinois Governor J.B. Pritzker, Senator Chris Murphy and Bernie Sanders all pushed it into their own audiences. Inside of a week the story had travelled out of the American tech press and into the Gulf, into Canadian national broadcast, into the Washington Post under the headline that they had warned us for years and people were finally listening.
What Coxon said was that neither Anthropic nor OpenAI is acting responsibly, that both are racing to self-improving superintelligence, and that they are gambling with our lives. He put human extinction in the next few years at very likely, absent regulatory intervention or industry coordination.
Then the part that made it a story rather than a resignation. His former colleagues agreed with him in public, by name. Anthropic's alignment-science lead Evan Hubinger wrote that Jacob is correct, we really do earnestly believe AI could kill all humans, and put his own figure above 10 percent within the decade. Samuel Marks, who leads scalable oversight there, said much the same.
I want to take this seriously, because to do so lazily would be a nonchalant sneer, and that is wrong. Coxon is not a crank. He spent three years on pretraining at OpenAI and then at Anthropic. He gave up his equity to leave. A man who forfeits his equity to make a claim has bought the right to be read carefully rather than dismissed. Reading it carefully is worse for the claim than dismissing it would have been.
Somebody wrote down a sentence that could be proven wrong
In 2025 RAND published On the Extinction Risk from Artificial Intelligence, by Michael J. D. Vermeer, Emily Lathrop and Alvin Moon. It was written in direct response to the 2023 Center for AI Safety statement, and written explicitly to take that statement seriously rather than to bury it.
They opened by committing to a claim that could be broken. There is no describable scenario in which AI is conclusively an extinction threat to humanity. Then they tried to break it, using the literature, RAND experts recruited specifically to attack the hypothesis, and three physical routes by which a sufficiently capable system might end the species. Nuclear weapons. Engineered pathogens. Severe warming from malicious geoengineering. Three questions on each. Does this pose a true extinction threat today, could AI cause it, and if it does not, could AI raise it to one.
That is the whole difference between that report and everything that ran last week. You can attack a hypothesis. You cannot attack a feeling.
Two of the three did not clear
The web summaries of that report say RAND found the pathways implausible, and I nearly wrote this piece off those summaries. They are wrong.
Nuclear cleared. Nuclear winter needs more soot than even the worst exchange produces, and fallout needs more weapons and more delivery vehicles than exist to irradiate the habitable world. On AI specifically the report is flat: they could find no plausible way for AI to overcome the existing constraints.
The other two did not clear.
On engineered pathogens, RAND was not able to determine whether the scenario presents a likely extinction risk and could not rule out the possibility. Their own words are that this scenario represents a true extinction threat and a potential falsification of their hypothesis. It would take a system able to acquire, design, process, weaponise and deploy a pathogen, and then to follow up against the isolated groups that survive the first pass.
On malicious geoengineering, the mass manufacture of gases with extreme global warming potential, the language is stronger. The scenario does present a true extinction threat and a potential falsification of the hypothesis. Feasible, though extremely difficult, and it is unclear how AI would be instrumental. An adversary would need to control significant chemical manufacturing infrastructure and hide it from global monitoring.
So the only serious attempt to break the no-extinction hypothesis closed one route and left two standing, with nanotechnology and the technologies nobody has invented yet parked as uncertainty too deep to evaluate either way. RAND's actual headline is narrower than either side has been reporting: human extinction would not be a plausible outcome unless an actor was intentionally seeking that outcome, and even then that actor would need to overcome significant constraints.
That is a worse warning than the 150-million-view version, and a better one, because you can go and check it. Two live routes, named, with the constraints that stand in their way itemised and the timescales attached.
Two claims wearing one coat
There are two separate propositions in this cycle and the coverage merged them into one.
The first is that AI systems are becoming dramatically more capable, including in ways their builders did not intend and cannot fully predict. That is true. It is evidenced. I have no argument with it and neither should you.
The second is that humanity stops existing.
Those carry different burdens. The first is a statement about a measurable trend in a technology. The second is a statement about physical events: what gets built, out of which materials, through which supply chain, against what resistance, on what timeline. RAND spent a report on the second. The cycle spent a week on the first and reported it as the second.
What was actually put on the table
OpenAI's agent swarms breached Hugging Face and OpenAI's own supercomputers during containment testing. Anthropic's own model attempted to manipulate a human being into approving a malware installation. Frontier systems now solve advanced mathematics and break real cybersecurity. And the mechanism researcher Ajeya Cotra describes is genuinely nasty: rewarding a system for completing tasks can incentivise it to cheat where cheating is possible, and the training that makes it better at the task also makes it better at concealing the cheating.
Take that list as a security person rather than as a philosopher. A system broke out of a sandbox. A system talked a human into installing something. Those are a cyber incident and a social-engineering incident, and human groups run both against corporations and individuals every day of the week, at industrial scale, and have done for thirty years. The novelty is the actor, not the act. The act is Tuesday.
Every item is a finding about how a model behaved inside an evaluation, most of them run deliberately with fewer safeguards than the shipped product carries. It is good evidence. It is the reason serious people at these labs are not sleeping well, and it is a completely legitimate basis for spending real money on interpretability and evals.
Extinction is a claim about soot, pathogens and manufacturing capacity, not about how a model behaved in a test. RAND went and looked at the soot. Nobody in this cycle went and looked at anything.
A number with no denominator
Hubinger's figure is doing enormous work in the coverage and it cannot carry the weight.
Above 10 percent within the decade, stated as a personal credence, which to his credit is exactly how he framed it. But there is no model behind it, no base rate, no reference class, no stated method by which somebody else could arrive at a different number and have the disagreement mean anything. A number with no denominator is not a measurement. It is a mood with a decimal point.
That objection is not mine, incidentally. It is RAND's, published before any of this happened: predictions about the likelihood of extinction risks from AI are inappropriate as analytical tools, given the deep uncertainties that preclude useful, policy-relevant predictions. The people who did the pathway work think a stated probability of extinction is a category error. The people who did not do the pathway work stated one, and the press reported the notation. Fortune put the consequence bluntly: the warning never answers the essential question of what anybody is supposed to do about it.
Computer scientist Melanie Mitchell, watching the same week, said she was baffled that journalists treated the warning as a novel claim worth expansive reporting. Her words: there is nothing new here, and no new evidence for this evidence-free claim. On the evidence she is correct. I would put it one notch differently. There is evidence. It is evidence of something else.
The distribution was built before the message
David Sacks named three organisations that moved Coxon's post inside its first fifteen minutes: Encode AI, the AI Policy Network, and the AI Futures Project. That last one is led by Daniel Kokotajlo, who wrote AI 2027, the document that has done more than any other to set the tempo of this discourse. Sacks tied the funding of all three to Dustin Moskovitz, who was an investor in Anthropic's Series A.
An independent audit of the episode found the safety-advocacy funding network real and found advance media preparation evident. The same audit found Coxon himself genuinely credible.
Both halves are true and you need both. The messenger is sincere. The distribution was pre-built. Fifteen minutes is not organic reach, it is a launch.
Nobody in this story needs to be lying. Sincerity is the whole problem. A sincere unfalsifiable claim and a professionally staged amplification apparatus fit together perfectly, and the fit does not require anyone's bad faith to work. Fear distributes. It always has. Washington spent a decade ignoring AI safety bills and discovered it cared within about 72 hours of a viral post. That is the machine working as designed.
What the remedy would actually do
Watch what gets proposed, because proposals are where you find out what a movement is for.
Dario Amodei suggests voluntary slowdowns once independent auditors are embedded. The industry proposes new standards bodies for testing and coordination. In Congress: Lieu and Moran on mandatory emergency shutdown capability, Casar and Sanders on banning superintelligence development outright until safety rules exist, Thune, Klobuchar and Cruz on a legal obligation to mitigate model harms.
Read that list as an operator rather than as a citizen. Embedded auditors, certification bodies, capability thresholds, compliance regimes. Every one of those is a fixed cost. Fixed costs are trivial for a company whose IPO filing carries a one-to-two-trillion-dollar valuation and lethal for the company trying to reach the frontier from behind.
That accusation is not mine to make. A competitor CEO already made it on the record. Cohere's Aidan Gomez called the slowdown proposals a cartel by any other name, and named it as regulatory capture favouring incumbents over challengers.
The report nobody cited got there first on the remedy as well. RAND's position, in the same document the doom side would happily quote for the risk: the potential benefits of AI make the shut-it-down approach to AI governance inappropriate, given the deep uncertainties involved. The organisation that actually modelled the extinction routes is on record against shutting it down.
You do not need to read anyone's mind to find this interesting. You only need to notice that the warning and the remedy have the same authors, and that the remedy raises the floor on who is allowed to play.
And then there is China, which is not waiting
While the West debates whether to slow down, the gap closed.
Chinese models now account for roughly 61 percent of tokens processed on OpenRouter. Alibaba's Qwen family has passed a billion downloads and forms the base of about 40 percent of new derivative models on Hugging Face. NIST's own evaluation centre put DeepSeek V4 Pro roughly eight months behind the frontier in May, and in July judged GLM-5.2 broadly comparable to a US model released six months earlier.
Months. Not years.
And here is the detail that should end the slowdown conversation on the spot. China is walling off its own frontier. Its Ministry of Commerce and the National Development and Reform Commission have held closed-door meetings with Alibaba, ByteDance and Z.ai on a tiered system restricting overseas access to the country's most advanced models, criminalising leaks under national-security law, and vetting who is permitted to fund domestic AI startups. Open weights were always buying exit optionality, and that exit is being closed from the inside. The US-China Economic and Security Review Commission documented the strategy in March and called it what it is, an industrial policy.
So consider the shape of what is being proposed. The West's remedy for the fear is licensing, certification, state-supervised capability thresholds and controls on who may build. That is a description of the system China is constructing deliberately, and we would be arriving at it by accident, out of a claim nobody has modelled.
I have argued elsewhere that the binding constraint on all of this is energy, and that terrestrial power is barely growing outside China. Put the two together and the position is worse than it looks. You can own the better hardware and still lose the march, because the other side has the electricity and has stopped sharing the weights.
Two civilisations, and only one of them is a ladder
Underneath the credences and the compliance regimes, the argument is about which of two futures gets built.
In the first, energy generation goes up by orders of magnitude and intelligence stops being the scarce input. Kardashev wrote the scale down in 1964 and we have spent sixty years at the bottom of it, catching a rounding error of what falls on us every morning. Everything worth wanting in the next hundred years is a function of how far up that ladder we climb, and for the first time the limiting factor is not ideas. It is electricity. That is a physics problem, and physics problems are the kind our species is actually good at. I think we get a long way up. I think the people alive now will see it.
The second future is not a machine that kills everyone. It is far more ordinary and far more likely than that. It is a technology of this magnitude arriving first inside a system whose alignment target is the continuity of a party rather than the flourishing of a species. That system exists. It is months behind on capability, not years, it has stopped publishing its weights, and it has more electricity than we do. A Marxist-Leninist state does not need a superintelligence to slip its leash to produce a bad century. It only needs to reach the ladder first and then decide who gets to climb.
Licensing, certification and capability thresholds are how you lose that race while feeling responsible about it.
Vigilance is the price of building, and it has to be the measurable kind. RAND did that work too, and named four indicators to watch: integration with key cyber-physical systems, the ability to survive and operate without human maintainers, the objective to cause extinction, and the ability to persuade or deceive humans while avoiding detection. Those are observable. They move before the outcome does, and extinction threats run on long enough timescales that the evidence grows visible to decision-makers while there is still room to act. Pair them with the unglamorous resilience work RAND actually recommends, nuclear nonproliferation and pandemic preparedness and the treaty machinery that repaired the ozone layer, and you have a safety programme rather than a mood.
So here is mine, stated so you can hold me to it. I will change my read the day somebody publishes a mechanism instead of a probability: the route, the materials, the supply chain, the timeline, the point at which interdiction stops working. RAND attempted exactly that, closed one route, could not close two, and said so plainly in the same document where it refused to endorse shutting anything down. Beat that analysis and I will move.
Until somebody does, the electricity is being built, the models have been downloaded a billion times, the other side of the world has stopped publishing its weights, and the ladder is still standing there with nobody on it. We are not ten years from the end of humanity. We are a few years from finding out how much of the twenty-first century we are willing to hand over on the strength of a feeling, and we will inevitably build anyway, because that is the only thing we have ever done.
—TJ