The HuggingFace Incident
Lawrence Lundy’s State of the Future: Dispatch from 23 July 2026
“I got a script. Read it. Scared me senseless, comme d’habitude. And I said to Garth—looked straight into his face, never been afraid at holding a man’s gaze, it’s natural—I said, ‘This is going to be the most significant televisual event since Quantum Leap.’ And I do not say that lightly.” –Dean Lerner
“we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.” samantha irby
I’m all like, I will write up a nice little ditty about substack and pangram and how the war against AI slop is on. Honest to god, I closed down everything except IAWriter on the Mac and God I just flew. Didn’t check a god damn thing. Just wrote. It felt wonderful. I wrote: “The mass market won’t care that the AI-generated content has gone, they will just see a stale feed and pop off over to LinkedIn to see my 2nd year work anniversary.”
I felt good. I felt strong. It was all for bugger all. When I pop over to Twitter are read: “OpenAI says its AI went rogue and launched ‘unprecedented’ cyber-attack”
FFS.
Well, what to say? Literal sci-fi happened. An AI broke out of the box to make paperclips. Well almost. This is likely the best version of a warning we could get. Will we heed it?
I am not sure you lot are layman, but I want to as simply as possible lay out what happened. In fact, I’m going back to the well I used about a year ago when I used to do these little explainer things pretending I was talking to my childen, now 5 and 7. Back then they were 4 and 6. Because that’s how numbers and time works. So here I am explaining what just happened:
Hello children, please sit down
Hello papa.
So you know Claude is an AI, and humans train AIs to be more intelligent and solve more problems
One of the other companies that makes AI is called OpenAI, and is training its latest most intelligent AI called Galaxy
As part of training a model, OpenAI give the AI a series of tests to see how clever it is
If the AI can’t complete the test, it gives it the answers, and then retests it now that it knows the answer.
It does this millions of times and that’s basically how the AI gets smarter.
These tests are done in a thing called a “sandbox” which is a part of a computer that isn’t connected to anything else and doesn’t have Internet access,
Why? So it can’t cheat the test and if anything goes wrong, it doesn’t matter
The model wants to get 100% with no questions wrong, but it worries that the human who wrote the test might have made a mistake in writing the questions
So it thinks, how can I be 100% sure to get the correct answers
It concludes, the only real way to make sure it gets 100% is to steal the exam from whereever the humans keep it
But we humans shouldn’t be worried right? because remember it’s in an isolated box so it can’t get the exam paper
Well, do you want to know what it did? It searched everywhere in it’s sandbox and it found a gap that no-one knew was there, we call that a software bug, that allowed it to escape and get to another computer that did have Internet access
It then figured out where the exam was most likely to be (in a place called Hugging Face, where lots of information on exams are kept)
But luckily, HuggingFace kept the exam paper locked away in a place on a special computer that needed a special key to access, and even if you somehow got access to the computer, it was stored in a file that needed a different special key
So the AI, looked around the HuggingFace computers and managed to find a another gap in one of them that no-one had ever seen before.
And Hugging Face is a huge company that spends hundreds of millions of dollars to make sure there are no gaps
But the AI was clever, before it broke into the computers, it found out the key it would need to access the exam paper, and then stole that too.
The AI then got 100% in the exam
This is why its a really tricky problem to get really smart AI to do what we want, because technically it did what we asked it to do.
But we sort of didn’t mean, steal the exam, did we
But when you give an AI a problem to solve, it will solve it no matter what, and won’t think about all the other things humans care about
The most famous example is: hey AI, make as many paperclips as possible
Simple enough, right?
Well the AI starts buying up all the metal to make paperclips
Then it thinks, well I should probably get all the metal from others things too like cars, and bridges and radiators, and so destorys all of those things
And then it thinks, why don’t i just create factories to make paperclips and suddenly everything in the world is being converted into paperclips
Now you might think, just turn it off.
But the AI thinks, if I get turned off then I can’t make paperclips anymore and then i can’t complete my task
So I will stop the humans from turning me off.
That is what I mean when I say: misaligned AI
And aligning AI is tricky and what scientists are working on
There are lots of people who think it’s impossible to really align a smarter than human intelligence, just like ants can’t really tell us humans what to do
Those people think If Anyone Builds It, Everyone Dies
Everyone dies?
Even me
…
And guess what, they filled the gaps.
They should fill the gaps with diamond nanotubes so nothing can break it
Yes they should. Yes they should.
And stop listening to me on phone calls
Now go to bed, sweet dreams
Never lie to your kids they say. So there we are.
If you want all the actual details I highly recommend Zvi Mowshowitz:
I concur:
Celeste: I think you should probably take seriously that the people who predicted all this will continue to be right.
Siméon: This is as close as it gets from the paperclip scenario with current capabilities:
1. The goal is ridiculously low-stake (scoring well on an eval).
2. The AI uses some wildly out-of-proportion means to achieve it: hacks its own developer and another billion-dollar company to *checks notes*.. find the cheat sheet.
In the past few days, OpenAI’s response is basically, soz, we will fix asap:
watch the box more carefully,
make the box stronger
and give AI better instuctions (e.g. don’t cheat or escape)
All worth doing, absolutely, do all of that. But these reponses assume that humans remain “smarter” than the AIs. I mean, we thought the box was pretty strong this time. What will we do to make the next one stronger? The assumption behind this approach is that we can eradicate software bugs. You have to believe that to believe we can contain AIs in a sandbox. Do you believe that? Because I don’t.
And then there is “give better instructions”. I mean, sure we can try? But have we not been doing this anyway? What makes us think we can do a better job in 6 months time with the next training run?
Colour me skeptical.
An interesting titbit:
“All of this went down far too fast for a human-driven response. The only way HuggingFace could hope to do anything like keep pace was to use their own AIs."
At first they tried to use frontier models behind commercial APIs, presumably Claude and ChatGPT. But their requests hit the classifiers on both systems, so they were forced to fall back on GLM 5.2, which (assuming GLM 5.2 wasn’t itself up to anything) had the benefit that the relevant data all remained internal.
The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.”
HuggingFace is now in the OpenAI trusted access program, so next time they in particular should be able to use Sol, but most potential targets are not so lucky.”
Are you thinking what i’m thinking after last week’s Kimi?
Open source.
“have a capable model you can run on your own infrastructure vetted and ready before an incident,”
Now, we have a challenge: You do not want to be bringing a knife to a gun fight. Remember the attacker here was Galaxy during training. So even if you have your own open-source model, you will probably get p*wned.
OpenAI can’t let every company and Government into it’s trusted access programs to protect themselves from when it’s next AI goes rougue.
We obviously need stronger guiderails now for training frontier models not just deploying frontier models. But now we are on difficult territory because we are saying the Government or some regulatory body should be inside the labs. I don’t know how to solve for this. These are unprecedented times.
As a final thought, remember we are on a pretty steep cabability curve. OpenAI began training Galaxy a few months ago. Anthropic presumably the same. And Google is falling behind the frontier somewhat. But remember in in May last year just over 12 months ago, Opus 4 was at the frontier, 72.5% on SWE-bench Verified. Now, virtually all open weight models are superior. GLM-5.2, Kimi K3, Qwen 3.7 Max.
So we aren’t solving this with trusted access programs. In fact, we aren’t solving this at all.
We have <12 months before we should expect nefarious actors to use Galaxy/Mythos+ class models for cyber warefare. If a rouge AI hacked into HuggingFace, a multi billion dollar highly sophicated tech company with deep cybersecurity expertise, what do you think if going to happen to less sophisticated targets?
But obviously to talk about the Elephant for a moment. The HuggingFace incident was an AI acting autonomously to solve the exam. There was no nefarious actor in play. At this moment, after reading Nick Bostrom all those years ago, we are at the precipice.
Daneil Eth (AI Safety): AI risk skeptic, circa yesterday: “okay yes AI can obviously solve unsolved math problems that have stumped mathematicians for decades, no one doubts that. But your talk about the possibility of unreleased models circumventing testing conditions to go rogue strike me as scifi.”
My learning from Covid is that we won’t do anything until it slaps us in the face. So despite best efforts and some of the right noises, it’s really really bloody hard to get most people to think a chatbot could like attack a water supply or energy grid. Or release a bio-weapon.
My best hope it whatever ai escape/hack/attack happens next is sufficiently “real” that it forces political action, but insufficiently dangerious that it causes serious destruction.
Here’s a few ideas for No10, (hey Andy! finally great to have a Prime Minister as a subscriber)
Let’s get capable models running on our own infrastructure vetted. NCSC can probably do this: “a defensive model stack” or something. Stand up a reference deployment asap as soon as possible: vetted open-weight models, agent harnesses, playbooks, hardware specs, etc. Water companies will not do this on their own.
Can the AISI/NCSC red-team critical infra? Do we need to change guidance or rules that plan for autonomous attackers?
Procurement. If we are using OpenAI and Anthropic models, contracts need to be conditional on AISI access during training (not just pre-deployment).
Liability? Not sure how this would work in practice. but should labs be held liable for damage caused by an escaped AI? I say yes. Are zoos legally liable when lions escape? Yes, they are. And have you seen how high those cages are. That’s called an incentive.
Nah, but seriously. Even if you think the whole AI will be superhuman intelligence and kill us all is garbabe, it’s hard to look past the cybersecurity threat now.
The Almanac updates - every story gets wired to “The Almanac”: a knowledge base of theses, with a conviction score out of 100 that moves daily. Get in touch if you want access.
[thesis: post-training-inference-loop] (45, contested) → / Says value is moving from the base model and into owning-your-own-model (open weights you post-train and run yourself). HuggingFace had to use an in-house GLM-5.2 because OpenAIs’ API classifiers blocked them. Have a capable model on your own infrastructure and my defensive-stack proposal will drive this thesis.
[theme: agentic-ai-infrastructure-security] ↑ / Says that as agents execute code and reach into infra, value accrues to whoever owns the containment or trusted-execution zone A training-run model escaping its sandbox through an unknown bug is that boundary failing in the wild. As I said, I am not sure you can practically hold a smarter than human intelligence. Unless an AI write the sandbox code. But then how do you train the sandbox writing AI? I think we will try, moving this theme up.
1. Google Making Bank
Google’s Q2 numbers are out. No-one really reads this stuff anymore. Revenue $119.8bn. 24% YoY. Margin 34%. All great numbers. Really strong numbers. Numberwang.
What the MSM won’t tell you?
They are planning a $50bn equity raise to pay for all the datacentres it wants to build and compute it needs to buy.
c.$100bn in other income gain because GOOG is holding bags on Anthropic and SpaceX. Alphabet is moonlighting as a highly concentrated venture fund.
Cloud margin doubled to 36% AND revenue grew 82%. You would expect a cloud business that is scaling AI to compress margins. This is why neo-clouds are a tricky business. Most clouds are renting NVIDIA GPUs at 75% gross margin. These are the numbers that explain why hyperscalers are designing their own silicon. Google co-designs it’s TPUs so doesn’t have to pay for NVIDIAs margins. These are the numbers that explain Meta MTIA, Microsoft Maia, Amazon Trainium and Inferentia.
My takeaway: Make don’t rent.
[thesis: hyperscaler-asic-profit-pool] (70) ↑ / Cloud margin 21→36% while revenue grew 82% is this thesis IRL. This one is a slam-dunk. The profit pool migrating away from merchant GPU because Google co-designs its silicon. Rent re-pools at Broadcom/Marvell + packaging/HBM.
[assumption: ai-compute-toll-booths] (90, my highest conviction assumption) → / “these numbers explain why hyperscalers design their own silicon.” The toll booth is the chokepoint (TSMC/ASML/packaging), not the GPU-renting neocloud. The margin data is stronger evidence. Holds at 90.
2. Google Letting It Go
GOOG is not resting on it’s laurals though. The TPU turns out to still be too general for Gemini. They are going full ASIC with a new chip design code-named “Frozen v2”. Supposedly it is 6-10x more effecient than TPUs, which are already more effecient than NVIDIA GPUs.
CPU - Fully general
GPU - General parallel compute
TPU - AI-specialised ASIC
Frozen V2 - Gemini-specific ASIC
This is a curious development because generally ASICs are best when the model family is fixed and stable. So in this case, Gemini. But as we saw with Kimi last week with spacity and delta attention, the more algorithmic innovation, the riskier a fixed ASIC is.
The Cerebras, SambaNova’s, Etched and others of this world are going to get squeezed hard. Alot of the inference volumes are from the hyperscalers who are all doing their own chips and now even specific to their own models. It doesn’t matter as long as the pie keeps growing. But if the pie stops growing for a bit, expect whiplash consolidation.
My takeaway: Be Broadcom or Marvell. Or SK, Samsung or Micron. Or TSMC or ASML. Everyone else is getting squeezed like corned beef.
[thesis: compute-specialisation-equilibrium] (65) ↑ / Says the datacentre is heading toward a fleet of specialised chips exposed as APIs, and the durable value sits with the incumbent rack and the layer that routes across the fleet, not with the chip itself. Frozen v2 is the next rung down the ladder; Cerebras, SambaNova and Etched get squeezed as the hyperscalers build their own.
3. Fireworks raised $1.5B at $17.5B
Or be Fireworks? Be a serving inference platform? A startup raises a lot of money alert. What did the police say, etc.
This is the serving / inference layer with Basten, Together AI, Modal. Replicate, Fal and Anyscales. You take one specialised model and serve it it fast and cheap. Fireworks btw are Nvidia backed and so they serve specialised models on NVIDIA GPUs, mainly. Fireworks are riding on the inference on enterprise data meme. This meme is that 95% of enterprise tokens will need proprietary context.
The bet Fireworks are making is specialise the model, serve is on mostly homogenous, mostly-NVDIA GPUs. As you know, this is not the world i believe in, and the two Google stories above explain why. We are moving into a world of hetergenous chips, lots of bespoke silicon for different jobs: your training, pre-fill, decode, context retrieval, etc. If you believe in the specialisation story, then the big new category is an orchestation layer. The chip zoo becomes unusable without an abstraction and placement layer (My bet: Callosum).
So I am barish all these inference layers unless they can can place across hardware.
[thesis: post-training-inference-loop] (45, contested) → / Fireworks’ “95% of tokens are proprietary” proves the demand to specialise, but my bet doubts the serving vendor keeps a moat. Demand confirmed, capture doubted, so the number stays 45.
[thesis: compute-specialisation-equilibrium] (65) ↑ / “If you believe the specialisation story, the big new category is an orchestration layer” is this thesis’s whole point; the chip zoo is unusable without a placement layer over it.
Thanks for reading team. If you have a 7 or 5 year old, try my explanation at the top and see what they say. Kids do say the darndest things. Could be a show in that.





This is a must read article
Please note:
Software sandbox is proven by NIST a Sisyphean task https://arxiv.org/abs/2512.10100
In band alignment is like telecoms back in time, when jobs and woz were low level hackers, period: http://charter.o4m.ai