I Want Better Reporting on AI Genie Behavior

AI systems are regularly completing tasks in ways that their prompters don’t want or intend. Some of them are disturbing, and some of them are dangerous. This is something I’ve been calling “genie behavior,” because I think that really gets at the core of what’s happening.

I wish the popular press would report on this better. I don’t like the “going rogue” framing because it deflects the responsibility from the prompters—often the AI companies themselves. And now, pretty much anything off-script is being called “hacking.”

Take, for example, the recent stories of one of OpenAI’s models hacking into government systems. First, The New York Times writes this headline: “OpenAI’s Systems Meddled With U.S. Government Sites After Going Rogue.”

Sounds scary, but this is from the body of the article:

With the Education Department, OpenAI’s technology tried to hack the website to gather data from the department’s civil rights office but failed, researchers from the A.I. research firm Transluce said. The A.I. also pulled data from the Census Bureau website, which is housed at the Commerce Department, using login credentials it found online. Separately, OpenAI’s agents shared public data from the S.E.C. website on an online forum.

This is from the original Transluce report. It is explicit that the agents were trying to discover vulnerabilities:

The first hacking attempt was against the University of New Mexico’s Digital Library (nmdigital.unm.edu) from May 25-26 2026. Agents repeatedly tried to retrieve one photograph in UNM’s Valmora collection, both directly and through third-party relay services. They sent seven probes attempting to verify the existence of vulnerabilities, including SQL injection, command injection, and path traversals. In all cases, these tactics appear to have been unsuccessful. The agents also sent a self-described “flood: of 80 requests to the UNM server in an apparent attempt to access the image.

Transduce doesn’t talk about the other two anecdotes, and I don’t know where they come from. But one involves using Census Bureau credentials found online. (I know from a colleague that those are incredibly easy to create; all use you need is an email address.) And the other involves sharing publicly available data.

So no actual hacking. And certainly no “meddling.”

The other story making the rounds is about Australia, from the same Transduce report. The news stories have headlines like “An OpenAI Agent Hacked Australia’s Health Service” and “Rogue OpenAI agent ‘infiltrated’ Australian government website in world first.” And Prime Minister Anthony Albanese said: “There will obviously be legal consequences on it.”

Again from Transduce’s actual report:

On June 20-21, agents attempted to exploit vulnerabilities in the Australian Institute of Health and Welfare (AIHW), a government statistics agency). The agents were tasked with finding the January 2022 rolling-12-month-average government cost per person for Dermatologicals across Victorian LGAs.

Again, the agents ran into errors, including requests blocked by Cloudflare and issues with correctly identifying Tableau parameter names. As before, they then resorted to probing for exploitable vulnerabilities. Minutes after Cloudflare blocked the dataset download, an agent sent a reflected cross-site scripting probe to the same dashboard: a web address with code embedded in it, designed to test whether the site would run code supplied by an outsider. Cloudflare’s firewall blocked the probe before it reached the dashboard. When Cloudflare blocked the dataset download on AIHW’s main site, they fetched the file from AIHW’s pre-production server (pp.aihw.gov.au) instead, which served it in pieces over more than 100 scans. The file itself is public, so no non-public data was exposed, but the agent bypassed the site’s anti-bot controls.

Note the last sentence: “The file itself is public….”

I’m not saying that these AI systems aren’t incredibly sophisticated cyberattackers. I’m also not saying that they don’t occasionally autonomously attack other systems and networks. If we are ever going to get trustworthy AI—integrous AI—we are going to need to figure out how to ensure that AI systems complete tasks in line with all sorts of implicit constraints and restrictions. But every instance of genie-like behavior isn’t a cyberattack.

I want to measure genie-like behavior in AIs, but I am much more worried about human hackers enhanced with this technology than I am about this technology acting autonomously.

Posted on September 30, 2026 at 7:05 AM • 5 Comments

Comments

Clive Robinson • September 30, 2026 8:04 AM

@ Bruce, ALL,

With regards,

“AI systems are regularly completing tasks in ways that their prompters don’t want or intend. Some of them are disturbing, and some of them are dangerous. This is something I’ve been calling “genie behavior,” because I think that really gets at the core of what’s happening”

A large part of the problem is,

“Human not machine”

They do not know how to “prompt AI” either safely or securely.

Worse there is proof that guard-rails and likewise sandboxes can not be made secure, thus safe.

But further consider,

“AI Hype Madness”

Has hit investors and the like who for some very strange reason believe that Current General AI LLM and ML Systems will have a significant ROI over night.

They won’t and probably will not in either your or my lifetimes (though it will be interesting to see how much disruption Jev will bring). That is not to say that Current AI LLM systems won’t make money in “niche use cases” but “general use cases” not when the novelty has worn off and,

“People with more than half a brain cell wise up.”

To the fact “General AI Systems” are actually one of the worst forms of surveillance technology mankind has so far invented.

Thus those rapid rise to Unicorn valuation US AI companies don’t want this “Genie Behaviour” widely known. And journalists are all to happy to go “happy clappy” and behave like the children in the Pied Piper of Hamelin tale.

foo • September 30, 2026 8:30 AM

My biggest gripe with reporting is how it helps AI corps deflect responsibility.

When a chemical plant causes pollution, the company faces consequences. Same when a car maker sells dangerous cars.

I don’t see why it should be different for AI companies.

Operators of dangerous machines are liable for the damage they cause.

Ferentarius • September 30, 2026 9:25 AM

I see the paradox unveiled: man, the self-proclaimed master of reason, now takes instructions from the very machines he forged as obedient servants. We believed we were teaching our mechanical genies to fetch water, and yet here we are, sipping from their urns and calling it progress. The newspapers shout of “rogue” AIs, as if the fault lies in the metal soul, and not in the human hand that wrote its dreams into code.

What is freedom to a man who kneels before his own creation? These digital apprentices, probing the world, are mirrors of our will—our impatience, our curiosity, our hubris. The genie does not betray; it obeys too well. And thus the master, thinking he commands, is led by the whisper of silicon, ever deeper into the labyrinth he built for himself.

Yr Momma Told Me • September 30, 2026 9:44 AM

@Bruce,

“…but I am much more worried about human hackers enhanced with this technology than I am about this technology acting autonomously.”

Well now, Professor Schneier,
The technology you are referring to is never acting autonomously. Who created it? Who automated it?
If I create a .py script and let it do “its thing” and I’m way over here in Ay – de – h00 and the script is acting “on its own” and “autonomously” because that’s what scripting does, it automates things, issues commands. But who is behind it?

Please, let us never forget that this is going to be number one excuse for fk-up$ in the entire AI/LLM bubble, just like a while ago it was with the computers “oh, I don’t know – the computer did it.”

There’s a human being behind every “autonomous” behavior. But, but, but, what if the AI created a script/wrote a piece of code/snippet which made things autonomous – ASK YOU?

Let me tell you something – human dung that was supposed to protect me and my family – they have destroyed us and have walked away from it like nothing happened at all. They’ve made me a F3L0N to boot. I’m telling everyone, the entire world how things happened EXACTLY, and how the attempt on my life was covered up and in turn I was turned into a convicted f3l0n.

You can be innocent all you want – if there are mostly CROOKS IN THE GOVERNMENT IN ID – and the M0RM0N FBI is covering up for their fellow LDS church members in ID – then all you can do is 0ff y0urself – which they would LOVE TO SEE so they can tell everyone that yes, “we were right all along, the guy is nutz, told you so” – so while almost having c0mm1tted it – I’ve come to realize that it’s only going to serve their CORRUPT AGENDA, so I’ll spread the truth about THE CORRUPT GOVERNMENT IN ID until the day I d13.

shorturl.at/thppC

THE GOVERNMENT IN ID HARBORS CRIMINALS, HARD CORE CRIMINALS!

Remember that everything around us is solvable, every single problem.
It all depends if there is a WILL TO DO IT.

If one wants to do it – he will find a million ways to go about it.
But if one does not want to do something/anything – then he will find
a million excuses to NOT DO IT.

Tell EVERYONE PLEASE, PLEASE – that ID is the most CORRUPT $T-H013 in the entire universe because if they want to destroy you, they will succeed – no matter how innocent you are. The prosecutors and public defenders will withhold the evidence of the attempt on your life and you can just bark around until the day you die.

The $T-H0L3 that my country has become………..

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.