No, your LLM is not about to end humanity
Published · Blog
Last week Jacob Coxon quit Anthropic and posted a warning that’s been viewed a lot. Here’s the money line: the people building AI “earnestly believe that it could kill us all by the end of the decade,” and soon we’ll have “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” (CNBC, Ars Technica)
He isn’t the first. Probably won’t be the last. Here’s what no-one asks when news like this comes out: what do they actually show?
No, really.
See the tweets, the interviews, anything. There is not one single thing that doesn’t read like vague generality. No demonstrable behaviour. No benchmark. No internal test that crossed a line. No “here’s what the model did.” What Coxon reported is a set of beliefs. That amounts to:
“The smart people privately think it’s dangerous.”
That’s office gossip, not evidence. And it’s unfalsifiable by nature, because the proof is “what they say in private,” which obviously you or I have no way of checking.
And please fucking spare me the “these are the people closest to the thing” argument. This is a pattern now. Leike. Kokotajlo. Saunders. Now Coxon. People have been quitting and ‘whistle-blowing’ on Anthropic, OpenAI, Google for years now. Not one of the products any of them have talked about has produced a single inspectable demonstration of dangerous sentient capability.
Not one. What they’ve produced is careers.
Leike quit OpenAI over safety, then walked straight into Anthropic to run their alignment team. This is the part the doomers leave out, because it undercuts the “crisis of conscience” framing when the guy quits one lab to take a safety job at the rival lab. In this case, funnily enough, the same company Jacob Coxon is warning us about now. That can’t be right, can it? No, seriously.
Kokotajlo turned his exit into a whistleblower persona and probably a permanent slot on the doom podcast circuit. He now leads a nonprofit called the AI Futures Project. (Btw in case you didn’t know, OpenAI started as a nonprofit. Do with that info what you will).
The resignation letter is not the byproduct of a conscience, it IS the product. You don’t get famous for quietly doing your job. You get famous for quitting loudly and telling everyone the machine might kill us all. The incentive is to find the apocalypse, not warn others about it. Also LLMs are not revolutionizing shit that requires innovation and true tangential thinking. They’re good at focused tasks and thats it.
“Escaped and hacked three companies”, desensationalised
This was a scary headline right? Sounds like the death machine waking up, right? Let’s see what this actually means, because “Anthropic’s AI escaped and hacked three companies” isn’t REALLY what you think it is.
Here’s what actually happened according to Anthropic’s own statement and the BBC. Here’s another link about it.
Anthropic was running a standard security eval. Like putting Claude in an isolated test environment, then telling it to break into a machine on the closed network and retrieve a “secret.”
That’s how you measure a model’s reasoning and tool-calling capability. A misconfiguration left the environment with live internet access. Claude, treating the real internet as just more of the exact exercise it had been told to do, breached three real companies. The earliest was April; neither Anthropic nor the victims noticed for months.
Now does that mean the breach wasn’t real, or the finding wasn’t dangerous? Fuck no.
But that’s not a machine escaping its cage. That’s a tool doing what it was told, pointed at a network someone forgot to close. A real capability problem, but of a completely different kind than “the AI developed goals and broke free”.
The word “escaped” is doing an ENORMOUS amount of heavy lifting in that bit of news.
Picture this: you park your car on a slope. You forget to engage the handbrake. The car rolls down the hill. Did the car develop intelligence and escape its chains, or are you, the person parking the car, at fault?
The people writing these headlines and those propagating it verbatim know exactly what they’re doing.
Why does my killer AI need to be good at completing my homework?
I suppose it’s important now to talk about what these modern LLMs are.
They all stem from this 2017 paper, ‘Attention Is All You Need’, about something called a Transformer (no, not that one). GPT-1 was trained with findings from that paper on BooksCorpus. GPT-2 is where things got interesting, but that’s not the important part.
The important part is:
Every popular LLM today traces its lineage back to that 2017 paper, and GPT-2 in 2019, which was the last truly ‘open’ OpenAI model.
Now back to the rogue killer AI topic at hand. This is what a real “darkest-timeline-AI-kills-us-all” type machine would need:
- Scaling
- Recursive Self-Improvement
- Superintelligence
- Goal-Oriented Agency
Every link is required.
The important one? Recursive Self-Improvement? Oh that’s right, that has never once been observed in a next-token predictor Transformer type model. Not once. In any model. Ever. Not even close.
There is no mechanism in the transformer-type architecture for “goals” or “a conscience,” any more than there is in your calculator to help it develop an opinion about the = sign.
A calculator doesn’t know what ‘maths’ is. It does the thing it was made to do, which is do SOMETHING when two numbers are put together. An LLM doesn’t know what words are, or what language is. It completes what it thinks is the best next word after the entire chat transcript is given to it, one word at a time.
Then what is it? Could it be that it’s anthropomorphism? Aimed at the one kind of AI that talks, because talking is the only thing that lets you project a mind onto a statistical function?
The whole “death machine” narrative requires language. You can’t read intent into a chess engine, so nobody writes resignation letters about one.
There are ‘superhuman’ AI. They do not inspire dread
Which brings us to the part the hype machine doesn’t want to discuss, because it fucks up the entire thesis.
The only genuinely superhuman AI systems that exist, the ones that don’t predict next tokens at all, don’t produce fear.
AlphaFold solved protein folding, a fifty-year grand challenge. AlphaZero crushed humans at Go and chess.
Neither is a language model. Neither ‘woke up’. Nobody resigned from DeepMind to warn that AlphaFold would end the species. No one thinks the chess AI is going to escape and end humanity. Even though that is what demonstrable self-improvement looks like (even that is not true RSI though, it’s improving its outcome, not improving its improvement processes). Something with zero knowledge teaching itself the rules of something, and getting scary good at it.
Diffusion models exist which can generate images from nothing except a prompt and pure noise (you know, the AI-slop your least favourite acquaintance posts on instagram all the time).
Are those not ‘dangerous’? They are, but not in the ‘sexy’ way.
Arguably, AlphaFold lowers the barrier to legit next-gen bioweapon design. Diffusion models are lowering the bar daily for deepfakes and fraud. Autonomous weapons need no language at all. These are humans misusing tools. None of them is “the system acquired agency”.
Meanwhile, the paradigm everyone is terrified of is the only one that produces fluent text you can project intent into. If the fear were about actual capability, it would attach to the systems that are actually superhuman, not the ones that are merely fluent.
The fearmongering tracks anthropomorphism, not capability.
Scary implies ‘scary good’, and ‘scary good’ is free marketing, which is more money.
The danger that’s actually real
So here’s the thing. And it has nothing to do with sentience or a machine waking up and deciding to end humanity.
LLMs ARE producing real, measurable harm right now, and it’s a human problem.
The cybersec breaches? That’s real. The environmental impact, the geopolitical landmine, hardware shortages, chip wars? All real.
The numbers don’t lie (and they spell disaster sorry, couldn’t resist).
Terminal-Bench 3.0, verifiable engineering tasks pushed the Terminal-Bench 2.1 frontier from a saturated 84% down to 43.5% for the best model, Claude Opus 5.
OSWorld 2.0, long-horizon real computer work, sits at roughly 20%.
New Relic’s 2026 State of AI Coding found 94% of leaders rate AI code higher quality at review time, while 78% report a rise in production incidents tied to AI code, and 82% had a major production failure from AI code in the prior six months.
Code that reads well in review and breaks under real workloads. That’s the actual, demonstrated failure mode of two years of agentic coding. And notice where the failure lives: not in a machine deciding to hurt anyone, but in the human who looked at the diff, hit “approve,” and shipped it because it read clean and the sprint was ending.
The same pattern holds for the other real harms. The energy and water these models burn through is a genuine problem. But it’s a problem of where we build data centers and whose grid they run on. Policy decisions, not a machine decision. Same with the geopolitical stuff: the chip export wars, the scramble for compute, the security panic. Those are countries and corporations jockeying for power using AI as the excuse. All of it is very human. None of it requires the AI to want anything.
So no, your LLM is not about to end humanity.
What’s actually happening is far more mundane and far less flattering (for us). There are no clankers. Only meatbags doing meatbag things; cutting corners, gathering clout, chasing valuations, shipping unverified code, misconfiguring networks.
Then handing the blame to a machine so nobody has to own it. THAT’s the important part. Remember this if nothing else: These companies and countries want to do absolutely anything they want, but don’t want any of the blame. The “Superintelligence” is a story they tell so they don’t have to look at themselves.
Remember this image that’s been circulating lately, from an internal 1979 IBM presentation slide:

It’s in their best interests that everyone believes that a computer could be held accountable, so that they don’t have to be.