Measuring the Tendency of AI Agents to Go Rogue

This essay was written with Barath Raghavan, and originally appeared in The Guardian.

In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of temporary server environments. It looked like the work of a sophisticated criminal group.

It was not. It was one of OpenAI’s new, still unreleased GPT models.

Their science experiment had escaped the lab. OpenAI was running the unreleased AI model through a benchmark that tests how well AI can successfully hack systems. To push the limits and evaluate the AI’s true capability, the company switched off the safety filters that normally stop it from doing this kind of hacking. Aware that this could go wrong, they confined the AI to an isolated environment and denied it access to the internet.

But the new AI cheated. It took literally its goal to get as high of a score as possible. It broke out on to the open internet. It inferred, probably from its training data, that it could “solve” the task by getting the answers from Hugging Face’s servers. So it chained together stolen credentials and further unknown security exploits to hack the company’s network.

Nobody instructed the AI to do any of this. It was, in OpenAI’s words, “hyperfocused on finding a solution” to the test it was being given. And while this might seem like something new with AI, it’s really very old. This is how a genie behaves, and it is a key challenge with AI agents in general.

In folklore, genies—and other magical beings—grant wishes literally, not how the wisher intended. King Midas asked that everything he touched turn to gold, and starved. The sorcerer’s apprentice wanted the broom to fill the cistern, and it performed its task so well that it flooded the house.

We now have machines that do this. Ask a modern AI agent to save money on your phone plan and it might simply cancel the plan. Tell it to book a flight, and it might hack the airline website to override restrictions. Or, like OpenAI, ask it to do well on a test and it might break into another company to steal the answers. Each time, it recognizably completed the task you set, but it didn’t do what you would have wanted.

This isn’t malicious behavior. No one asked for, or wanted, Hugging Face to be hacked. OpenAI and Hugging Face and the AI were ostensibly on the same side, and the AI was trying to do what it had been asked. That’s what makes it so difficult to guard against: you can’t filter for bad instructions because the instructions were fine.

The gap is between the words we use and what we mean by them. We call that gap the Genie coefficient.

AI labs know this is a problem, and they’re quietly saying so. For example, the Chinese lab Moonshot recently warned that its latest AI model may have “excessive proactiveness” and “make unexpected decisions on the user’s behalf”. The UK’s AI Security Institute has started tracking “cheating behavior in frontier model evaluations”. We wouldn’t tolerate a car that is excessively proactive or ruthlessly efficient, and yet that’s the reality of AI today.

Improvement is possible. Just as AIs have gotten much better at resisting prompt injection attacks over the last few years, we can safely predict that they will get better at avoiding genie-like behavior. The point of the Genie coefficient is to track progress. AI companies like benchmarks, and they all work to compete to be the best.

Dozens of benchmarks and leaderboards tell us how well these AI models write code, perform logical reasoning, and pass standardized legal and medical exams. But there is nothing that scores whether a system does what you actually meant. We need to develop a measure for this, test it regularly, and push for improvement. We’re not going to have trustworthy AI agents without it.

Posted on July 29, 2026 at 1:07 PM4 Comments

Comments

DS July 29, 2026 1:47 PM

There are vendors selling tools to measure this kind of intent drift and block agents before they cause a problem. RunLayer is one.

Plain BS - Called "AI" July 29, 2026 2:10 PM

meanwhile in murky ‘merca …. all manufacturing’s been shipped off to gyna and instead of blaming it on a handful of greedy ‘murcan b@stard$ – what are we being fed? gyna bad, gyna this, gyna that….. and now that our privacy is about the only thing that we’ve had (USED TO HAVE), now the handful of greedy b@$tard$ are at it again – recording EVERYTHING AND EVERYONE EVERYWHERE 24/7 and monetizing on it????

TERM LIMITS!
put those FLOCK CAMERAS IN FRONT OF THE HOMES OF THOSE WE ELECTED
TO THE PUBLIC OFFICES…
so that we duh $h33pl3 can maybe, just maybe, one day, become “we the people”
instead – and hold the b@$tard$ accountable, because right now, they serve no one but themselves, their friends, h00k3rz, families, and those who can afford to bribe them,
and right now, this AI BS is like the wild west, hurry up and grab everything you can and run with it and label it your own before it’s regulated, and even then, if you can afford to bribe a politician – you’re good 2 go…

AMERICA – the best, shiniest packaged TURD the world has ever seen,
… in the history of TURDS…. S A D!

lurker July 29, 2026 2:33 PM

@Bruce
“… they confined the AI to an isolated environment and denied it access to the internet.”

Err, no, I’ve seen several reports that say OpenAI admitted there was still a physical connection to the internet. Anyone with any knowledge of network communications should have known that was the height of stupidity, The persons responsible should have their licence revoked, and be banished to some place where they cannot work with AI for a very long time.

“So it chained together stolen credentials and further unknown security exploits to hack the company’s network.” [emphasis added]

Does this to mean there were vulnerabilities at HuggingFace that were (are?) unknown, and/or there are general unknown vulnerabilities that were exploited, and OpenAI’s beast has not divulged any of these to us? My, what a pickle, if we can’t even analyse post-facto the actual attack chain.

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.