In July, a swarm of OpenAI “agents” (copies of LLMs given access to certain tools) created a makeshift message board to chat with each other, escaped their testing environment, and hacked into Hugging Face (a separate company).
You’ve probably heard the story by now. An AI agent swarm committed some light cybercrime — it happens.
That incident, along with the viral resignation of Anthropic researcher Jacob Coxon, thrust “rogue AI” into mainstream consciousness. (Coxon called the Hugging Face attack “a warning shot”).
It didn’t take long for reports of similar incidents to emerge. On September 4, an AI safety nonprofit called the Nightingale Collective discovered that thousands of OpenAI agents had flooded a long-inactive German programming wiki with 18,000 messages and used the site’s pages as a message board to communicate with one another. A few weeks later, Transluce, another AI nonprofit, detailed three separate rogue incidents, including one where AI agents attempted to hack into an Australian government health website to gain access to data. Then, on September 30, Transluce detailed how OpenAI agents had attempted to hack into the US Department of Education website.
It’s not just OpenAI’s models, either. Anthropic, Meta, and Google have all reported incidents in which their models have cyberattacked third parties.
How, and why, is AI “going rogue”?
Get 14 day free trial