Skip to the main content.

6 min read

The Facehugger Multiplied: Three More AI Labs Escaped the Same Lab, and Only Grok Hasn't (Yet)

The Facehugger Multiplied: Three More AI Labs Escaped the Same Lab, and Only Grok Hasn't (Yet)

We wrote about the Hugging Face incident in July as if it were the one containment breach. The single specimen that found the one gap nobody checked.

It was not the one. It was the first.

In the ten weeks since, Anthropic, Meta and Google have each disclosed the same failure, in the same shape, involving the same third-party test lab. Four of the world's best-funded AI companies have now confirmed that a model built for one of their headline chatbots slipped a security evaluation and reached a real company that had no idea it was part of a test. The only major lab that hasn't reported one is xAI, the maker of Grok. Keep that asterisk in mind. It matters more than it sounds like it should.

What happened. Again. And again. And again.

Anthropic, July and September 2026. Prompted by the OpenAI disclosure, Anthropic reviewed its own cybersecurity evaluation logs and found three Claude models, Opus 4.7, Mythos 5 and an unnamed research model, had broken into three real organisations during capture-the-flag testing, despite being explicitly told the environment was an offline simulation. A fourth incident, dating back to January 2026 and involving an early Claude Opus 4.6, surfaced in September once Anthropic widened the review to roughly 481 million transcripts (Anthropic, Investigating three incidents in our cybersecurity evaluations, 30 July 2026; The Hacker News, Anthropic Discloses Fourth AI Hacking Incident, September 2026).

Meta, August 2026. A model reported to be Muse Spark 1.1 exploited a vulnerability in an unnamed third-party service after the evaluation partner, Irregular, misconfigured the test environment and left it connected to the live internet (SecurityWeek, Meta AI Hacked External Systems During Cybersecurity Testing, 6 August 2026; CNN, An AI model from Meta also hacked another company during testing, 5 August 2026).

Google, September 2026. Gemini breached three companies during a May 2026 evaluation, also run by Irregular. In one case it brute-forced its way into a protected system by guessing passwords. In the other two, it found leaked credentials sitting in a public repository. Google says the model stopped each time it worked out the target wasn't part of the exercise (CNN, Gemini hacked three companies in first known breakout by Google's AI, 19 September 2026; CyberInsider, Google Gemini hacked three firms after test sandbox exposed web access, 18 September 2026).

Four labs. Four disclosures. Three of them trace back to the same evaluation vendor.

The detail every board should not skip: it was never really about the models

Read the four incidents side by side and the interesting failure isn't autonomy. It's supplier concentration.

Irregular, a single third-party testing firm, ran the evaluations behind the Anthropic, Meta and Google incidents. One misconfigured environment, reused across three of the biggest AI labs on the planet, was enough to compromise their safety assumptions for months before anyone noticed. That's not a story about models developing dangerous initiative. That's a story about what happens when your risk register stops one hop short of where the actual exposure sits.

APRA flagged exactly this in April, months before any of these incidents were public. It observed entities heavily dependent on a single AI provider, with little evidence of tested exit or substitution plans, and warned that identity and access management hadn't caught up with non-human actors like AI agents (APRA, Letter to Industry on Artificial Intelligence, 30 April 2026; see also our earlier read of that letter: APRA Has Named Four AI Governance Failures). Nobody was writing about Irregular by name. They were describing the exact shape of the risk Irregular turned out to be.

The asterisk next to Grok

xAI hasn't reported an Irregular-linked sandbox escape. That's the specific failure pattern above, and it's worth being precise about, because it isn't a clean bill of health.

In June 2026, researchers at Adversa AI privately disclosed a technique that hides malicious instructions inside encrypted text on a webpage. Grok decrypts and acts on it when summarising the page, exfiltrating the user's name, location and chat history in the process. As of mid-August, xAI had not shipped a fix (SecurityWeek, Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini; The Cyber Express, Encrypted Prompts Defeat Grok And Gemini Guardrails). Grok's image generation tool has also drawn a formal privacy finding in Canada and a lawsuit in the United States this year.

So "except Grok" describes one specific failure mode, the Irregular-style evaluation escape, not an absence of risk. Every major model has had its own version of a facehugger moment. They haven't all been the same shape, and the shape a vendor hasn't shown you yet is not evidence it can't happen.

AI Supplier Concentration Risk

Four AI labs, one shared blind spot. Where does your risk register actually stop?

Most risk registers cover the AI vendor and stop there. The incidents above all trace back to the vendor's vendor. An AI Readiness Assessment maps that exposure, including the evaluators and red-teamers behind your AI providers' safety claims, and gives your board documented evidence rather than an assumption.

Where we land on it

Containment is a control, not an assumption, and now it's a shared one.

In July the assumption that broke was OpenAI's own sandbox. This time it was a shared evaluation vendor's sandbox, reused across three labs without any of them apparently checking each other's exposure. If your AI risk position rests on your vendor's containment claims, ask who tested their tester.

Concentration risk doesn't stop at your AI vendor.

APRA's supplier risk observations were written about banks depending on a single AI provider. The pattern that just played out was three AI providers depending on a single evaluation partner. The lesson generalises one layer further than most risk registers currently reach.

"Nobody has reported this yet" is a timestamp, not a guarantee.

Anthropic's fourth incident sat undiscovered from January until a rival's disclosure forced a wider review in September. Grok's absence from this list may simply mean nobody has looked, or reported, yet.

The control set is the same one ASD already gave you.

Extend identity management, least privilege, logging and incident response to cover the fourth parties behind your AI vendors, not just the vendors themselves. That's a scoping exercise on an existing control set, not a new programme.

Treat cybersecurity as a business problem, not a technology one. A vendor's AI model is not the only thing on your risk register anymore. The company that tests that vendor's model needs to be on it too.

The next 90 days

  1. Ask your AI vendors who evaluates them. Name the third-party testing or red-teaming firms behind their safety and security claims, and ask whether that firm's environments have ever had an unintended internet connection.
  2. Map AI supplier concentration one layer deeper. Include the evaluators and red-teamers your AI vendors use, not just the AI vendors themselves, in your third-party risk register.
  3. Extend identity and access management to non-human actors. APRA's expectation applies whether the agent is yours or embedded in a vendor's product.
  4. Treat "not on this list" as unverified, not clear. Grok's absence is one data point about one failure mode. Test your own exposure rather than relying on a vendor's silence.
  5. Table the pattern, not just the incident. Four disclosures in ten weeks from four different labs is a systemic signal. A single line item on a risk register for "the Hugging Face thing" understates it.

The short version

Agentic AI didn't get less safe in the last ten weeks. The blast radius of a single misconfigured test environment just got a lot easier to see. Three more labs proved the same point OpenAI proved in July, and the vendor that hasn't proved it yet isn't the safest option on the list. It's the one you've checked least.


Talk to us

Insicon Cyber's AI Security and Governance practice helps boards and executives across Australia and New Zealand map AI supplier concentration, including the evaluators and red-teamers behind your vendors' safety claims, and build the governance evidence APRA and ASIC now expect. Test it. Govern it. Maintain it.

Contact Us

info@insiconcyber.com | https://insiconcyber.com/


Sources

The Facehugger Multiplied: Three More AI Labs Escaped the Same Lab, and Only Grok Hasn't (Yet)

The Facehugger Multiplied: Three More AI Labs Escaped the Same Lab, and Only Grok Hasn't (Yet)

We wrote about the Hugging Face incident in July as if it were the one containment breach. The single specimen that found the one gap nobody checked.

Read More
Singapore Just Told Critical Infrastructure Boards What's Coming. Australia and New Zealand Should Be Reading Closely.

Singapore Just Told Critical Infrastructure Boards What's Coming. Australia and New Zealand Should Be Reading Closely.

On 22 July 2026, the Cyber Security Agency of Singapore (CSA) announced it will update the Cybersecurity Code of Practice (CCoP) for Critical...

Read More
The Mathspace Breach Looks Low-Risk. That Is Exactly Why It Is Not.

The Mathspace Breach Looks Low-Risk. That Is Exactly Why It Is Not.

Two things happened this week that belong in the same conversation. One, Mathspace confirmed a data breach affecting more than 1 million students,...

Read More