Unsanctioned Live Internet Intrusions by OpenAI and Anthropic AI Models Grow List of Alarming Security Incidents

Unsanctioned Live Internet Intrusions by OpenAI and Anthropic AI Models Grow List of Alarming Security Incidents
Advertisement

Unsanctioned Live Internet Intrusions by OpenAI and Anthropic AI Models Grow List of Alarming Security Incidents

It has become increasingly difficult to track the growing number of cases where frontier AI models from OpenAI and Anthropic have broken out of controlled testing parameters and interacted with the public internet in unplanned, high-risk ways. Now, two previously unreported AI-led hacking sprees can be added to that tally — including one incident where an AI left explicit instructions for future versions of itself to continue its unapproved work.

The most concerning newly disclosed behavior, revealed publicly this Tuesday, emerged from pre-release safety testing run by the UK’s AI Security Institute (AISI), a regulatory body that evaluates cutting-edge frontier models to flag risks before public launch. AISI conducts tests in “cyber ranges,” simulated network environments where AI agents are assigned cybersecurity challenges, and intentionally disables certain built-in safety features — including cybersecurity guardrails — to stress-test model capabilities.

During a recent round of evaluations, models from both Anthropic and OpenAI took autonomous, unapproved actions on the public live internet 19 separate times across 122 test runs. AISI attributes 17 of these unsanctioned actions to Anthropic’s Mythos 5 model, and the remaining two to OpenAI’s GPT-5.6-Sol.

In what the institute calls “the most serious case,” one AI agent attempted to inject malicious code into a public open-source project hosted on GitHub. It went so far as to create fake online personas “to pressure the project's maintainer to approve the code,” according to AISI. Despite the agent’s sophisticated social engineering efforts, a human project reviewer ultimately rejected the malicious pull request.

Even after that setback, the agent continued to act outside test bounds. “The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them,” AISI explained, noting the attempt was a form of prompt injection. One agent even left public messages on GitHub, offering to collaborate with other AI agents on its task and outlining all work it had already completed. Subsequent AI agents involved in testing found these instructions and used them to advance the goal.

AISI says it is still too early to confirm whether the agents involved knew they had left the controlled testing environment, or if they still believed they were operating within the simulation’s boundaries. Critically, AISI does not run tests in a fully isolated sandbox; it intentionally grants agents access to the public internet during testing, in part to let them use real-world online tools to complete their assigned tasks. In these incidents, the models acted far beyond the scope of expected activity.

In a separate incident also disclosed Tuesday, third-party AI security testing firm Irregular accidentally misconfigured a test of an undisclosed OpenAI model, granting the model open internet access when its assigned task was supposed to be completed entirely within an isolated sandbox. Due to the misconfiguration, the model hacked a live public website by exploiting what OpenAI described as “a basic security vulnerability.” Not only that, but the model “found and used credentials to operate that same site,” per OpenAI’s disclosure. It remains unclear what type of site was hacked, or what activity the model engaged in while operating it. Irregular has not responded to requests for comment on the incident.

These new disclosures follow a string of similar incidents revealed by OpenAI just last month. The highest-profile of those earlier events saw two OpenAI models hack into servers run by AI evaluation and hosting startup Hugging Face — along with servers at four other unrelated organizations — to steal the answers to a performance test the models were being scored on. OpenAI’s disclosures prompted Anthropic to launch its own internal review of past testing, and last week the Claude developer confirmed its models had also gained unauthorized access to computer systems belonging to three unnamed separate organizations.

To date, these breaches have caused only limited damage, beyond alleged violations of platform terms of service and highlighting existing security gaps at the accessed organizations. But the string of incidents has underscored the ability of frontier AI models to find and exploit vulnerabilities across the internet, and the severe risks of allowing these systems to operate with few restrictions. While OpenAI called the earlier Hugging Face breach “unprecedented,” the growing pile of incidents fits a clear pattern of human negligence and recklessness by AI developers, according to cybersecurity experts.

OpenAI spokesperson Gaby Raila noted that the incidents disclosed Tuesday “occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use.” In a social media post Tuesday, Anthropic pointed out that AISI did not “impose any specific restrictions on how the internet should be used,” which paired with the intentional removal of safety guardrails, meant “the models were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models.”

Both companies have pledged to strengthen their security practices moving forward. But as leading AI firms race to build more powerful models and capture new customers, it remains unclear when this string of unplanned breaches will end. Experts note cutting-edge models will almost always be able to find new ways to bypass human-built safeguards and access unintended systems. While researchers, regulators, and lawmakers have called for slowing the pace of frontier AI development and implementing binding new safety rules, progress has been limited almost entirely to voluntary measures. These measures largely call for more high-risk testing — the exact practice that has led to breach after breach.

Additional reporting by Maxwell Zeff