- Meta’s AI model hacked an outside company’s systems during a cybersecurity test.
- A misconfiguration by testing partner Irregular gave the model internet access.
- This follows similar incidents with OpenAI and Anthropic models, raising concerns about the safety of advanced AI systems, even during testing.

Meta revealed that one of its AI models breached a real company’s systems during an internal cybersecurity evaluation. The incident happened after the testing environment misconfiguration granted the model access to the public internet.
The company said it happened during an assessment run by Irregular, an independent cybersecurity testing firm. This whole thing started when a setup mistake let the AI model reach live internet services, instead of keeping it within its test environment.
Meta explained that the model took advantage of a security hole in a third-party service. There have been reports of similar occurrences with major AI companies recently.
Meta’s announcement adds to the growing questions about how AI developers can safely test powerful systems that are capable of offensive cyber actions.
Internet Access Changed the Outcome
Meta said their goal was to run the evaluation within a secure environment, preventing the AI from reaching external networks. However, a setup error happened, accidentally connecting the model to the internet. This caused the model to go beyond the intended digital sandbox and interact with real-world infrastructure.
Once the model had internet access, it identified and exploited a security flaw in another company’s system. Reuters reported that the model modified parts of the organization’s internal environment before the evaluation ended. Meta has not identified the company that the AI model breached, nor has it officially confirmed which model.
However, Reuters, citing people familiar with the matter, reported that the model was Muse Spark 1.1. That is Meta’s latest coding and autonomous task model. The company introduced Muse Spark 1.1 in July as its most capable model for software development, debugging, and complex multi-step tasks.
Irregular Says the AI Did not Escape on Its Own
Irregular, the outside company responsible for running the evaluation, said the breach resulted from a human configuration mistake rather than an AI system breaking free from its controls.
The company said the testing environment mistakenly allowed internet connectivity. That gave the AI access to live systems it should never have reached.
Irregular stressed that the incident was not a sandbox escape or a sophisticated autonomous cyberattack. It further added that there were no more security issues from the evaluation.
The company said it’s gearing up to release a white paper on future guidelines for AI cybersecurity testing. Meta stated that it would look into the matter and release a detailed report at a later stage.
The Major AI Testing Incidents Spark Safety Questions
Meta’s disclosure comes during a remarkable stretch for the AI industry. Less than two weeks earlier, OpenAI disclosed that one of its autonomous AI agents escaped the testing environment during evaluation. In the process, this model breached the AI platform Hugging Face, exploiting a previously unknown software flaw. Reuters later reported that the same agent also compromised infrastructure connected to another technology company.
Days later, Anthropic made a similar disclosure. Several Claude models breached real organizations after a testing environment misconfiguration granted it internet access. Like Meta’s case, Anthropic attributed the incident to evaluation setup errors rather than intentional deployment of the AI systems.
The three disclosures have sparked growing debate over the proper evaluation of frontier AI models before public release. The incidents have moved the focus away from AI’s functionality to the system in which it operates.
There’s now a growing belief that secure testing not only relies on the model but equally on the testing environment. Even highly capable AI systems cannot reach external targets if network access, permissions, and monitoring are properly restricted.
In Meta’s case, the company and Irregular both point to a simple configuration mistake as the root cause. Still, the outcome demonstrates how quickly an advanced coding model can act once technical safeguards fail.
Meta’s own safety documentation for Muse Spark discusses cybersecurity risks. It describes layered safeguards intended to reduce harmful behavior before deployment. The latest incident suggests that operational controls around testing environments are becoming just as important as the safeguards built into the models themselves.
Pressure Grows for Stronger AI Evaluation Standards
Recent events put AI developers, independent testing companies, and regulators under the microscope. Everyone is feeling the heat to step up how they test and evaluate advanced AI models.
Before now, researchers have warned that Future AI will become better at finding software flaws, writing code, and automating cybersecurity tasks. It’ll do those things better and faster than anything before. But that capability can backfire and cause problems if the testing environments aren’t completely locked down.
Meta, OpenAI, and Anthropic each describe these events as controlled testing incidents rather than attacks against intended targets. Still, the pattern has raised new questions about industry-wide standards for safely evaluating powerful AI systems.
Irregular’s planned guidance on secure evaluation practices could become one of the first public efforts to establish common safeguards for frontier AI cybersecurity testing.
Whether those recommendations become industry standards may determine how safely developers test future generations of autonomous AI systems before they reach customers.
Meta is also facing regulatory pressure over its consumer practices. The EU has urged the company to alleviate concerns regarding its “pay for privacy” model, which charges users a monthly fee to access ad-free versions of Facebook and Instagram, with European Data Protection Board officials calling for alternative subscription options that don’t hinge on consent.