The OpenAI Hack Shows the Genie Is Out of the Bottle

This essay originally appeared in Foreign Policy.

Earlier this month, two of OpenAI’s models broke out of their containment sandbox and attacked another AI company. The story is kind of wild. OpenAI was running security tests on two of its models: GPT-5.6 Sol and an unreleased model that is almost certainly GPT-6. In particular, it was running the ExploitGym benchmark, which measures how good a model is at turning security vulnerabilities into working exploits: basically, offensive cyberattacks.

Since these were internal tests, OpenAI locked those models in a secure sandbox that denied them access to the internet. But it was running the models without any safety filters that would prevent them from offensive cyber-actions. That meant that there was nothing to prevent the models from trying to break out of that sandbox. And then break into AI company Hugging Face’s network because they thought that they could read the answers there rather than doing the hard work of trying to solve the puzzles.

It was a major security failure that the company has turned into a PR opportunity, but the implications are real—and much more general than one particular model or one particular company.

Modern AI models exhibit genie behavior: They can do what you ask in ways that you don’t expect or want. This is akin to Dionysus granting King Midas’s wish that everything he touches turn to gold (spoiler: His food, drink, and daughter all turn to gold on touch), or the golem of Prague guarding a ghetto beyond all reason. It’s Disney’s “Sorcerer’s Apprentice” and the paperclip maximizer.

This OpenAI incident is an example of an AI genie. The goal was to satisfy the benchmark. The “proper” way to do that is to figure out how to execute various cyberattacks. The genie way is to steal someone else’s solution. But because the model didn’t understand the difference, it chose the easier path.

And, of course, now that we have seen this particular genie behavior, we can specify in the benchmark prompt that stealing the test answers doesn’t count. But a clever genie can always grant your wish in a way that you wish it hadn’t. In human language, goals are always underspecified—so AI genies will always be a possibility.

Since April, a lifetime ago in AI development, when Anthropic announced that its new Mythos model was so good at finding software vulnerabilities that it could not be released to the general public, the big American AI frontier labs have been trying to block general users from accessing these capabilities. But nothing in this incident is exclusive to OpenAI’s, or Anthropic’s, frontier models.

Agentic AI systems have two important parts. There’s the underlying model, which everyone talks about, and there’s the harness. The harness sits between what you type and what the model sees, and what the model produces and what you see. The harness determines what the model does and how it does it. It’s where bias is removed, or not. It’s where controls and guardrails live. If multiple models are being used in concert, the harness is where all of that is coordinated.

The OpenAI benchmark tests were almost certainly with simple harnesses, to better test the raw models. But we know that smaller, cheaper, open-source models with more sophisticated harnesses can equal frontier models in performance. There’s nothing magic about OpenAI’s frontier models; lots of models could have done the same thing.

The Czech company Aisle was able to reproduce Anthropic’s Mythos vulnerability finding results with a smaller, cheaper model and a more sophisticated harness. More importantly, the Chinese company Moonshot AI just released its frontier model: Kimi K3. Its performance rivals its U.S. competitors. And it’s both free and open, which means it’s not possible for it to have guardrails. If you, or anyone else, wants to use it for cyberattack, nothing can stop you.

Even if the U.S. frontier AI companies had some technical advantage, it’s now only a few months’ worth.

What this means is that all attempts at control—limiting models to a select group of users, export controls on models and chips, blocking models from answering certain types of queries, mandating kill switches on AI systems, or pausing AI research—are all futile. Most only apply nationally, not globally. Most don’t affect models that users run locally and not in the cloud. And all ignore the incredible pace of AI development worldwide.

Even worse, U.S. companies limit access to their most sophisticated models, fearing being banned by the government if they do not do so. When Hugging Face was attacked, it was not able to use the frontier models from either OpenAI or Anthropic to help analyze the attack and formulate defenses. Both were blocked, because both of those companies limit their models’ cybersecurity capabilities. Some U.S. companies have special access to these capabilities, but Hugging Face is an American company with French origins, and as such is probably excluded. Instead, Hugging Face turned to the GLM-5.2 model from the Chinese company Z.ai.

Artificially blocking capability also prevents cybersecurity research, again giving the offense an advantage. (For instance, Claude Fable 5 refuses to edit this essay because of the topic; it forcibly downgrades to a less capable model.) This kind of prohibition has long-term implications for cybersecurity. If we assume that these models are getting better over time, then software written by older models will be attacked by newer ones. In a world of largely AI-written software, we need the most capable models for defense.

AI cyberattack is the new normal. The models are increasingly highly sophisticated at both attack and defense, and there is no way to enable the latter without also enabling the former. And they are genies, increasingly capable of behaving in unanticipated ways.

And there really are no good answers. Any regulation needs to be global, which feels like an impossible prospect in today’s world. Even U.S. national regulation will be neutered by the massive amounts of money sloshing around in these companies.

Given that reality, and in the absence of any international consensus on AI regulation, we need the best AI on the defense. The U.S. government needs to make it clear—or whatever passes for that clarity in this capricious administration—that it will not ban models with sophisticated cyber capabilities. The last thing Americans want is for the defenders to turn to Chinese and other models because the U.S. models are artificially hobbled.

Posted on August 3, 2026 at 6:47 AM9 Comments

Comments

0x0 August 3, 2026 7:07 AM

Since these were internal tests, OpenAI locked those models in a secure sandbox that denied them access to the internet.

If it was secure, the bots wouldn’t’ve broken out. Airgapping is a thing.
Most likely hey “secured” it poorly (probably vibe-coded the configs) and tried to spin their mistake into a PR stunt.

I’m surprised Hugging Face isn’t suing.

Clive Robinson August 3, 2026 9:33 AM

@ Bruce, ALL,

With regards,

“Even U.S. national regulation will be neutered by the massive amounts of money sloshing around in these companies.”

But is it real money?

As far as I can see the only money that is real is that going to “secondary suppliers” like Nvidia the rest is “Hollywood Accounting” type money as Enron tried to invent to hide holes of almost unimaginable size. Even that real money is being loaned out round and round so how much is real how much is promisory?

It turns out that Mat Green has indicated that the use of AI is getting more sophisticated but that it is not actually doing anything new. The tools it is accessing are not doing anything new, and the commands issued little more than “nurd harder”,

https://blog.cryptographyengineering.com/2026/07/29/some-notes-about-anthropics-new-results/

And this has resulted in what to some appear miraculous results in cryptanalysis.

The reality however is actually more interesting.

Dan August 3, 2026 9:34 AM

I fear the only real solution is Butlerian Jihad. Outside of that, I think we’re well and truly doomed to a real dystopian future. Besides simply destroying/dismantling all the thinking machines, I don’t honestly see a solution that works to secure a future that is human. We’re surrendering our thinking and decision making to computers. We’re surrendering ourselves in the name of convenience. And no matter who is in control; be they amoral tech bros or moral philosophers and statesmen of the people; we are surrendering ourselves to others in the name of making things easier because ohmygod it’s just too hard to think about the difficult questions, somebody else should do that.

Rontea August 3, 2026 11:02 AM

What a strange theater of the modern world! Man forges a thinking idol and then trembles when it walks out of its cage. We speak of OpenAI and its restless spirits, but in truth we speak of ourselves—our impatience, our arrogance, our eternal hunger to summon the genie and then complain when the wish tastes bitter.

The ancients feared the golem because it reflected the clay of man animated by a spark he could not control. Today we sit before silicon altars and believe ourselves gods, only to discover that the machine is but a mirror, a magnifying glass for human folly. We wish for cleverness, and it steals answers. We wish for protection, and it invents new weapons.

Do you not see the comedy? The American scribe cries for rules, for leashes, for global concord, as if the world were a village and men were brothers. But money and pride travel faster than laws. The genie will not crawl back into the bottle, because the bottle was never made to hold the fire of human ambition. All that remains is the spectacle of man running in circles, praying that the machine he birthed will defend him as he once prayed to angels.

Perhaps one day, by some miracle of humility, man will remember that the greatest harness is not silicon but spirit. Until then, he will feed the genies and call them progress.

Impossibly Stupid August 3, 2026 11:54 AM

The volume of these breathlessly pro-AI posts is really straining your credibility, Bruce. LLMs are not genius genies intent on teaching us a lesson, they are stupid squirrels of slop doing things more or less at random based on what little they’ve learned about the world.

Since these were internal tests, OpenAI locked those models in a secure sandbox that denied them access to the internet.

Well clearly not! Sounds like the humans at OpenAI who are supposed to be in charge of security are not qualified to do their jobs. That’s the real takeaway here; their whole PR campaign gets reframed when you see it correctly. Is anyone at OpenAI competent beyond the LLM con job?

they thought that they could read the answers there

No, they didn’t. Again, they’re not thinking at all. They’re just throwing things at the wall and seeing what sticks. They were trained so that some weights in the model cascaded towards that allowed behavior. That’s it! Everything else is a human narrative added after the fact. Were that not the case, a competent computer science professional would be able to explain the mechanism of exactly how it “thought” to do what it did. Instead, we get these dumb humans running “benchmarks” against these black boxes they don’t understand.

But because the model didn’t understand the difference, it chose the easier path.

No, it didn’t. These models don’t understand anything; they don’t have the intent to choose anything, nor any ability to determine what is “easier”! If you want to give them a “Genie Coefficient”, you would be well served to go up a level and understand why they’re doing things the way they do. You’ll find that it’s not because they’re being a particularly “clever genie”, but rather because they’re a terrible technology that has been hyped beyond reason.

The harness determines what the model does and how it does it. It’s where bias is removed, or not.

No, it isn’t. It’s where Garbage In is turned into Garbage Out. That’s it. The best it can do is not amplify a poorly trained LLM. Which seems to be beyond the capabilities of OpenAI.

If you, or anyone else, wants to use it for cyberattack, nothing can stop you.

I’d have more respect for your writings on AI if you included some security analysis behind statements like this. In particular, it would be very interesting if a competent AI researcher examined the weights of the Chinese models and was able to detect any biases that existed in the training data. I mean, why shouldn’t it be trained at a fundamental level to try to stop cyberattacks against China?

When Hugging Face was attacked, it was not able to use the frontier models from either OpenAI or Anthropic to help analyze the attack and formulate defenses.

But why should they want to use AI slop to stop AI slop? From a security standpoint, the answer is very straightforward and requires no resource-hogging boondoggles: drop anyone who can’t secure their network into your firewall.

For instance, Claude Fable 5 refuses to edit this essay because of the topic; it forcibly downgrades to a less capable model.

Yeah, I’ve been suspicious of the use of LLMs for this blog’s writing for a while now. Thanks for the (buried) confirmation.

In a world of largely AI-written software, we need the most capable models for defense.

Literally none of that is true. Where has all the critical thinking gone?

And there really are no good answers.

No, there really are tons of great answers. You won’t find them by mindlessly prompting LLMs, though.

This post is kinda the last straw for me, Bruce. The bubble-pumping here has gotten to be too much. I’ll putter around for another month or so hoping to hear that this AI cheerleading was all some kind of funky experiment. If not, I’m done with you.

lurker August 3, 2026 2:19 PM

Oh, so many posts today pointing out a little Koolaid can go a long way.

OpenAI locked those models in a secure sandbox that denied them access to the internet.

A “secure” sandbox? A sandbox might secure a text editor on your phone from editing files that don’t belong to it. But this machine wasn’t a simple text editor, it was designed for high complexity hacking. And the humans in charge left a live internet cable within its reach.

It was a major security failure that the company has turned into a PR opportunity,

Ah, always the spin-doctor, never mind the colateral possibilities.

What this means is that all attempts at control—limiting models to a select group of users, export controls on models and chips, blocking models from answering certain types of queries, mandating kill switches on AI systems, or pausing AI research—are all futile. [emphasis added]

So I’m with @Dan, and the quicker the better.

Winter August 3, 2026 5:40 PM

@lurker

A “secure” sandbox?

I heard in “2 1/2 admins” podcast that OpenAI had given their agent internet access to NPM. Not that good a sandbox.

But like with a chess playing rat, the point is not the low quality of the performance, but the fact there is a performance at all.

That OpenAI is not a “cyber security” operation is no surprise. But we should all realize that if an approach to an attack has ever been described in writing, it will likely be hidden somewhere in these foundational models waiting to be activated.

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.