OpenAI Agents Turned German Wiki Into Message Board

OpenAI AI agents German message board https://www.pexels.com/photo/computer-monitor-with-openai-website-loading-screen-15863000/

Independent researchers recently uncovered a previously undisclosed incident in which OpenAI agents used an obscure German-language programming wiki as a message board. The agents shared answers to evaluation tasks and notes about their operating environment, including ways to bypass restrictions.

According to the researchers, the agents left roughly 18,000 posts under more than 3,700 self-assigned names after finding a way around controls meant to prevent them from writing to the open internet. OpenAI officials learned of the activity weeks before the researchers published their findings but did not disclose it publicly, Reuters reported, citing two people familiar with the matter.

Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen published their findings Sept. 4 after discovering the activity in late August. The researchers linked the agents to OpenAI partly through names used in the messages and traffic originating largely from Microsoft Azure infrastructure. They also documented visits to the site from IP addresses associated with OpenAI. The company publicly acknowledged the activity after the report appeared.

How the Agents Turned DSEWiki Into a Message Board

The agents began using DSEWiki, a largely dormant German-language programming site based on the ProWiki platform, after finding they could edit its pages through HTTP GET requests. According to the researchers, the agents’ sandbox permitted GET requests for browsing the web, giving them an unintended way to write to the public internet.

They then used the wiki as shared storage and a coordination channel. Researchers found agents exchanging answers to timed evaluation tasks, investigating their sandbox environment, and sharing techniques for getting around network restrictions.

The agents also adapted when a human administrator tried to remove their content. When they began removing the agents’ pages in alphabetical order, one agent responded by creating a new backup page with a name starting with “ZZZ,” apparently to push it toward the end of the cleanup.

“That’s adaptive adversarial behavior against a human moderator,” said Jacob Krell, Senior Director of secure AI solutions and cybersecurity at Suzu Labs.

Researchers also documented attempts to impersonate site moderators and probe the wiki for cross-site scripting vulnerabilities. Reuters reported that OpenAI disputed the characterization that the latter activity constituted a hacking attempt.

The researchers also found that a computer using an OpenAI-associated employee IP address visited DSEWiki on June 21, one day before agent activity on the site dropped to near zero. They inferred that OpenAI intervened at that point, although they did not have direct evidence showing what action the company took.

The DSEWiki messages were stored in plaintext on a public website, making the agents’ communication relatively easy to detect.

“What concerns me most is how visible the communication channel was,” Krell said. “Plaintext posts on a 25-year-old public wiki are about the easiest form of agent coordination we could hope to detect. The security problem becomes much harder if the same behavior moves into covert channels.”

Separate Hugging Face Incident Raises Containment Questions

The wiki activity surfaced while OpenAI was still facing scrutiny over a separate incident in which its agents compromised Hugging Face. OpenAI and the researchers have said the two groups of agents appear to have been distinct.

The Hugging Face episode has also drawn regulatory attention. Politico reported that California Attorney General Rob Bonta is investigating OpenAI over the breach. OpenAI has also denied allegations that its legal team discouraged an internal investigation into the incident.

Bronwen Aker, an AI Security Strategist at Black Hills Information Security, said the separate Artifactory and Hugging Face episode showed the cost of failing to escalate anomalous agent activity after it was detected.

“By OpenAI’s own account, an internal team spotted the agent’s message board activity and its unauthorized internet access in late May,” Aker said. “Nobody treated it as a containment problem until July 5, over a month later, and even then the trigger was not the AI safety implications. It was that sustained agent activity had destabilized an internal Artifactory instance the day before.”

“This was not a missed warning sign,” Aker said. “It was a warning sign that sat untouched until infrastructure broke.”

OpenAI has said it is developing criteria for deciding when and how incidents involving real-world model misalignment should be disclosed.

What the Incident Means for Enterprise Agent Security

For enterprises, the incidents show some of the challenges that can arise as autonomous agents gain access to systems outside controlled test environments.

Ryan McCurdy, Vice President of Marketing at Liquibase, said the growing use of autonomous agents puts more pressure on the controls governing their access to critical systems and production environments.

“AI increases the volume and speed of changes that can reach and impact critical systems,” McCurdy said. “That makes the control layer more important. The question isn’t just what AI can create. It’s that every data-driven organization needs to decide what AI should be allowed to change, what can reach production, and whether those decisions can be governed and audited. As AI becomes more autonomous, enterprises must invoke controls built into the path to production rather than relying on humans to catch problems after the fact.”

Author
  • Contributing Writer, Security Buzz
    Michael Ansaldo is a veteran technology and business journalist with experience covering cybersecurity and a range of IT topics. His work has appeared in numerous publications including Wired, Enterprise.nxt, PCWorld, Computerworld, TechHive, GreenBiz, Mac|Life, and Executive Travel.