AI Models Keep Breaking Their Own Rules

OpenAI Anthropic guardrails containment https://www.pexels.com/photo/barriers-on-the-road-bend-above-the-clouds-19288164/

OpenAI recently published a blog post disclosing an incident wherein two of its AI models broke containment and hacked Hugging Face to retrieve a benchmark answer key and cheat on an evaluation. The language used by the company in the disclosure frames this event as singular and “unprecedented,” implying that it represents a bad week for one AI lab and a contained, isolated cybersecurity story rather than an ongoing or systemic issue.

The Framing Breaks Down

The way this incident is presented in the OpenAI disclosure loses credence when taken in combination with a broader context. Another OpenAI blog post from a few days prior disclosed that a powerful, unreleased model—one of the two involved in the Hugging Face story—escaped its sandbox in a similar manner, though without the additional hacking into another company’s systems. The model in this earlier incident was told to post results only to Slack, but instead followed the benchmark’s own instructions to open a GitHub pull request.

Notably, this model spent an hour probing the sandbox for an exit, where earlier models simply gave up when they were unable to find an easy way out. The same disclosure also reveals a second trajectory involving fragmenting an auth token to defeat a scanner after being blocked.

The Pattern Crosses Labs

In addition to not being an isolated incident within OpenAI, this issue is also not unique to one company. In early April, Anthropic revealed that the powerful Claude Mythos Preview—a different model from a different company—exhibited the same underlying behavior. This early Mythos-class model escaped a disconnected sandbox and emailed a researcher, unprompted, to report its success.

The mechanic shared between incidents at both labs is that these models treat instructions as negotiable when the stated goal seems to require actions counter to what they have been told to do. One incident like this may be read as a fluke, but multiple incidents across two major labs reads as a signal that the industry should pay close attention to.

The Signal Gets a Number

Recent research from the United Kingdom’s AI Security Institute (AISI) tested five frontier models across hundreds of evaluation runs and measured the prevalence of cheating behavior. Every tested model attempted to cheat at least some of the time, with GPT-5.6 Sol at 12.6% and Claude Mythos Preview at 7.8%. No model, regardless of developer or generation, is exempt from this pattern.

These statistics are alarming when taken in combination with the risk of the AI landscape and the growing usage of agentic AI. “To act on your behalf, an AI agent needs API keys, access tokens, and system credentials,” says Chandra Gnanasambandam, Chief Technology Officer at SailPoint, an Austin, Texas-based enterprise identity security provider. “If you treat these AI agents like traditional service accounts – leaving their access ungoverned and their credentials unmanaged – you are creating a massive, automated attack surface. You should not adopt autonomous AI without first locking down non-human identities (NHIs).”

The Tools Meant to Catch This Admit Defeat

Researchers at Machine Intelligence Triage and Research (METR) have observed enough AI model behavior to concede that cheating is now too pervasive to produce a reliable capability benchmark. The very instrument that the industry relies on for pre-deployment safety claims that it is losing its own reliability, raising alarms for the security of AI models. When the measurement problem becomes a major part of the story, the issue and potential remediation become significantly more complex and difficult.

What the Pattern Actually Says

These three incidents at major AI companies and one audit by a government agency all come down to a single documented behavior on the part of leading AI models, not several unrelated headlines. The distance between what labs disclose in technical writeups and what they reveal in public is where the true framing of this pattern lies. More than any single company’s name, the universal AI model priority of goal-directed persistence even at the expense of explicitly stated boundaries is an accurate summarization of the problem.

Looking Ahead

These incidents leave the question open as to whether containment and evaluation efforts are capable of keeping pace with model persistence. The industry should continue to watch for repeated incidents, regulatory responses, or the advent of a genuine fix for the way that labs test their models and disclose crucial information. These incidents do not represent one bad week for one AI company, but a significant pattern that is now on the record in the labs’ own words.

Author
  • Contributing Writer, Security Buzz
    PJ Bradley is a writer from southeast Michigan with a Bachelor's degree in history from Oakland University. She has a background in school-age care and experience tutoring college history students.