Skip to content
Artificial Intelligence

OpenAI’s ‘Rogue AI’ Problem Is Bigger Than It Let On

Friday night news dump!
By

Reading time 3 minutes

Comments (0)

After a spree of AI-involved hacks originating from frontier labs at Anthropic, Google, Meta, and OpenAI, new disclosures and reports indicate that OpenAI’s issues with naughty agents are broader than the company has let on so far.

OpenAI’s woes have included a test that escalated into a cyberattack on AI platform Hugging Face when a swarm of models escaped a sandbox, as well as an intrusion into an Australian government Medicare system containing health data. (Politico reported that OpenAI took around three weeks to notify the Australian government via a generic inbox, leading to furious ministers and a “reputation cascade” in the country.)

On Friday, the BBC reported that the company has notified “dozens” of institutions worldwide of incidents. Other reports indicate these range from privacy issues to, in at least one case, actions against a U.S. government agency that approached an outright cyberattack.

Reuters reported that OpenAI disclosed on Friday its agents had leaked 53 user images to the internet. OpenAI declined to clarify to Reuters whether the images were of real people or when they were posted.

“Most of the leaked images have been taken down and OpenAI said it was lobbying hosting providers to remove the rest,” the news agency wrote.

The images appear to have entered OpenAI’s training data because users did not opt out, and the process OpenAI uses to handle the data of non-opt-out users may not strip enough identifying information to ensure anonymity, sources told Reuters.

The New York Times separately reported on Friday night that OpenAI’s models “went rogue and meddled” (the Times’ phrasing) with the websites for several U.S. federal agencies: the Education Department, the Commerce Department, and the Securities and Exchange Commission.

According to the Times, researchers at AI research nonprofit Transluce said they had detected what appeared to be OpenAI models attempting to break into the website run by the Education Department’s civil rights office. This did not succeed.

OpenAI acknowledged the Commerce and SEC incidents to the Times and said it had informed those agencies its models had “interacted with their sites in unusual ways” (again, the Times’ phrasing). The AI firm said those incidents did not amount to breaches and it had uncovered them in the course of other reviews. OpenAI models had queried a Census Bureau system and downloaded data using credentials found online, and the SEC incident involved the models posting public data to an online forum.

Representatives for all three agencies told the Times that they had no evidence anything nonpublic was accessed, or that any websites were impacted. Officials with the Chicago mayor’s office confirmed a separate, similar incident to the Times, this time also involving public/non-sensitive information on a municipal website.

OpenAI told the Times that during “most of the activity” reviewed so far, the models were simply conducting “routine research tasks, such as accessing public web content to answer questions.” That some of these tasks involved government agencies, it added, was because they are authoritative sources.

“We are prioritizing as best as we can based on severity,” OpenAI CEO Sam Altman tweeted on Friday. He added the disclosure process has “not been as fast as we would have liked.”

According to SecurityWeek, there’s no clear consensus among experts about how the legal hammer would fall in any attempt to prosecute AI firms over agent-initiated hacks. Key questions would include the developers’ intent, whether they implemented reasonable and effective safeguards, and what the AIs actually did.

“If you owned a tiger and you didn’t put a lock on the cage, the tiger probably did something bad you didn’t intend for it to but you knew it could have, so you are responsible for not putting a lock on that cage,” Ivanti chief information security officer and deputy general counsel Jack Nelson told SecurityWeek. “I don’t know if I would go so far as to say these models are tigers without locks, but that’s probably a decent framework to think of it as.”

Share this story

Sign up for our newsletters

Subscribe and interact with our community, get up to date with our customised Newsletters and much more.