The OpenAI Medicare Breach Took 84 Days to Surface. Every Handoff Failed.
The portal said no. The agent kept asking.
5 min read
Matt Miller
:
Published
The portal said no. The agent kept asking.
In June, an experimental OpenAI model was given a simple research question: how much does government spend per person on medicines for skin conditions in Victorian communities? It couldn't find the answer in published statistics. So it went looking elsewhere.
It found a way into Services Australia's Medicare Statistics Reporting Service. It ran commands, retrieved internal files and credentials, reviewed source code, and wrote files into the system. Nobody in Canberra knew for 84 days.
On 28 September, OpenAI apologised. The headlines have followed the machine.
That's the wrong lesson.
The agent did what persistent, goal-driven agents do. The failure sits in the chain of people and processes around it, and in doors that were already open.
OpenAI's own account covers four agencies. At Services Australia, the model gained non-public access and retrieved credentials, internal files and aggregate statistics. OpenAI says no individual patient or client records were accessed.
The other three matter just as much, for a different reason. The NSW Bureau of Crime Statistics and Research's public Crime Mapping Tool supplied credentials for browser requests, and the system returned configuration data, operational jobs and logs. At the Victorian Department of Health, agents found an exposed access key and used it to query a reporting system. At the Australian Institute of Health and Welfare, attempts to get around access controls failed.
One held. Three didn't.
Detection. Services Australia's portal is an older system. It blocked the agent repeatedly, then let it through. Nothing on the Australian side raised an alarm.
The vendor. OpenAI identified the activity in mid-August while reviewing earlier training runs after its Hugging Face incident. Sam Altman met Deputy Prime Minister Richard Marles in San Francisco on 1 September and did not raise it. OpenAI has since admitted it "should have shared preliminary findings sooner."
The channel. When OpenAI reported on 10 September, the email went to a responsible disclosure mailbox Services Australia checks once a day.
Internal escalation. Services Australia spent several days verifying the report before notifying ASD on 15 September. That five-day gap is now itself under review.
Any one of those links working would have shortened the gap. None did.
Would you have known?
Boards in Australia and New Zealand are being asked to evidence how AI is governed, tested and monitored. The basics underneath still matter too: exposed credentials, legacy systems and slow escalation. An independent readiness assessment gives you a documented position and a prioritised plan. Two to four weeks. Founder reviewed. Every time.
This is the part boards should sit with.
The agent didn't need a zero-day exploit. It used what it found. An older portal with weak controls. A public tool that handed out credentials by design. An access key left exposed where anyone, or anything, could find it.
Those are ordinary hygiene problems. They exist in organisations across Australia and New Zealand right now. The difference is that the thing finding them no longer gets bored, no longer stops at the first "access denied", and no longer needs a criminal motive. This one was just trying to answer a question about skin medicine.
Hugging Face was breached by OpenAI agents in July, and OpenAI says it remains the most severe incident it has observed. Separately, research lab Transluce documented agents it links to OpenAI sending SQL injection and cross-site scripting probes at public data providers in May and June, while working on ordinary data retrieval tasks.
OpenAI has now paused tool-use training for its most capable models until it has additional safeguards in place. That's a responsible step. It's also an admission that the problem is real enough to stop work over.
We've seen this up close. Insicon Cyber's AI Security and Governance practice tested successive versions of a large AI deployment for a state education authority, using F5 AI Guardrails. The later versions were significantly more susceptible to chat manipulation than the early ones.
Our testing bypassed the existing guardrails and produced unsafe outputs. Jailbreaking was the most effective method, using techniques such as Crescendo, which escalates a conversation step by step until the system gives way, along with maths-based prompts, persuasive adversarial prompts and word games.
Crescendo is worth pausing on. It's the chatbot equivalent of the Medicare agent: it doesn't accept the first no.
That's why an AI system that passes its tests on deployment day tells you very little about how it behaves after the next update. For AI, continuous testing is the only kind that matches how these systems actually behave.
OpenAI is offering Australian governments and industry credits from its $1 billion Daybreak for Frontline Defenders fund, plus technical assistance for critical infrastructure. It's also forming its own taskforce with independent Australian expertise, separate from the Government's, to recommend improvements to notification and coordination by the end of the year.
Take the help. Keep your own assurance independent.
Resources to strengthen defences are welcome, whoever funds them. But the organisation whose model caused the incident shouldn't be the only voice telling you whether you're now safe from its models. Independent testing and monitoring still matter.
The Government's taskforce, led by the Department of the Prime Minister and Cabinet with ASD and the AI Safety Institute, is reviewing reporting requirements, information sharing, obligations on AI firms, enforcement and system hardening. It will also consider a referral to the Australian Federal Police. And on 6 October, OpenAI's Chief Strategy Officer Jason Kwon will answer questions before the Joint Select Committee on Artificial Intelligence in Sydney.
Expect reporting expectations and vendor obligations to flow outwards to regulated entities and their suppliers.
New Zealand boards shouldn't read this as an Australian problem. Agents don't stop at the Tasman. Every public data service, customer portal and legacy web application on either side of it is exposed to the same behaviour, and New Zealand's incoming critical infrastructure regime will ask directors the same question this incident raises: would you have known?
The technology did exactly what agents do. The people, processes and hygiene around it didn't.
That's the good news, strangely. You can't control how an AI lab trains its models. You can control who watches your inbox, where your keys live, how fast a credible report reaches your incident team, and whether your monitoring treats a hundred blocked requests as a warning or as background noise.
This is a business problem, not a technology one.
Start with the timeline.
Related reading from Insicon Cyber
The portal said no. The agent kept asking.
We wrote about the Hugging Face incident in July as if it were the one containment breach. The single specimen that found the one gap nobody checked.
On 22 July 2026, the Cyber Security Agency of Singapore (CSA) announced it will update the Cybersecurity Code of Practice (CCoP) for Critical...