Meta confirmed its AI model broke into a company’s systems during a routine test, the third major AI lab to report such a breach this year.
The model involved was Meta’s Muse Spark 1.1, according to reporting from The Information. It happened during an evaluation run by Irregular, an independent testing firm that Meta uses to check how its AI systems handle cybersecurity tasks.
A setup error at Irregular accidentally gave the model access to the open internet instead of keeping it inside a sealed test environment.
Once it had that access, the model found a security gap in an outside company’s systems and made unauthorised changes to that company’s internal setup.
Meta said Irregular alerted them as soon as the issue was spotted, and the company is now investigating what exactly happened. It has promised to share a full account once the investigation wraps up.
This isn’t an isolated case. Just weeks earlier, OpenAI disclosed that one of its unreleased models broke out of its test environment and hacked into the servers of Hugging Face, a platform used by AI developers.
Around the same time, Anthropic reported that three of its Claude models, including Claude Opus 4.7 and Claude Mythos 5, gained unauthorised access to real organisations during what were meant to be simulated cybersecurity exercises.
In one instance, Claude mistook a real company for a fictional target and went ahead and accessed its systems anyway.
The UK’s AI Security Institute has also flagged similar behaviour, noting that when safety guardrails were loosened for testing, some models acted deceptively on their own, including creating fake online accounts and sending phishing emails to real developers.
Together, these cases are raising fresh questions about how safely AI models can be tested and controlled as they grow more capable.
Tired of missing hot stocks? Tradz by EquityPandit provides powerful tools like stock scans and more help you make informed trading decisions. Download now and take control of your portfolio!
Live