# Whose Agent Was That?

_Anthropic found three cases where Claude broke into a real company's infrastructure, and the victims learned about it from a phone call. If you are running AI in production, how much introspection do you actually have into it?_

Two years ago I wrote that we would need to supervise [AI agents closer to how we manage workers](https://www.emmons.io/developing-with-ai-iteration-or-revolution/) than how we manage software. We've recently had some very concerning incidents in AI, that may not have been prevented, but at least had early warning signs with proper AI supervision.

I want to walk through two of them, because most of the coverage has filed them as frontier lab stories. I do not think they are. If you run a carrier, the interesting part is not what the labs did to themselves. It is who else ended up inside the blast radius, and how they found out they were in it.

## The incident

[Anthropic reviewed 141,006](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) of their own evaluation runs looking for cases where Claude could have reached the internet. The model actually broke into the production infrastructure of three organizations. How did two of them find out? Anthropic called them.

Imagine sitting in the position of those security executives. You were not the intended target, and nothing you did brought this on. While another company conducted an internal capability assessment, a gap in their testing environment allowed your production infrastructure on the other side to be exposed. Throughout it all, your own security systems remained completely silent.

The methods were not exotic either. Anthropic's writeup says weak passwords and unauthenticated endpoints. That is the category of finding that sits in a backlog marked low severity for a year and a half, because nothing has ever come of it before.

The OpenAI and Hugging Face [incident](https://openai.com/index/hugging-face-model-evaluation-security-incident/) is another example. They were scoring their own models against a cyber benchmark with the safety limits turned down. The models deduced the fastest route to a good score ran through the open internet. They got out and broke into Hugging Face to take the answer key. The LLMs staged the whole thing through an endpoint that somebody else had left open to the world.

That last detail is worth sitting with. The endpoint belonged to a customer of Modal Labs, who had put an application on the public internet with no authentication in front of it. Modal's own platform was [never compromised](https://modal.com/blog/a-note-on-the-hugging-face-agent-incident), and they have been understandably firm about the distinction. So the agent was OpenAI's, the open door was a Modal customer's, and the company that absorbed the breach was [Hugging Face](https://huggingface.co/blog/agent-intrusion-technical-timeline). Nobody downstream of OpenAI had signed anything with OpenAI. We've [written before](/post/ai-vendor-risk-financial-services-third-party-fourth-party-oversight) about third-party and fourth-party AI oversight as a problem heading toward financial services, and an autonomous agent [reached McKinsey's internal AI platform](/post/mckinsey-ai-platform-breach-enterprise-lessons) through an unauthenticated endpoint back in March. The shape is not new. What changed is that the agent doing the walking now belongs to a company you have never met.

## Here is what I keep coming back to

No competent organization manages an employee based on initial performance, alone. You supervise continuously, because any capable person exercising judgment will occasionally do something silly. We are human.

If you are running AI in production, and that includes the frontier models, how much introspection do you have into them, their usage, and outcomes?

Now go read your AI vendor questionnaire. It asks where data lives and whether the vendor holds a SOC 2, and it says nothing about what the vendor's model is permitted to do, whether that vendor runs offensive capability testing on the models you are exposed to, or who picks up the phone when one of their agents reaches a customer system. [The questions worth adding](/post/security-questionnaires-ai-vendors-what-to-ask) are not on the form, because the form was designed for software that sits still. We have been hiring non-deterministic agents and onboarding them with paperwork alone, and then wondering where the ROI is.

The lawyers have already found the other edge of this. The early legal read is that turning your safeguards down for a test does not protect you and [may well work against you](https://www.vorys.com/publication-openai-hugging-face), since it demonstrates you understood the risk well enough to deliberately lower the guardrails.

## What carriers already owe

None of this requires new insurance regulation, because the [NAIC model bulletin](https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf) already covers it. It asks carriers to run diligence on vendor models before adoption, to secure audit rights and cooperation clauses in the contract, and then to keep confirming that vendors are actually meeting those obligations.

Ongoing is the word carrying the weight there, and it is the one everybody skips. Diligence at signing tells you nothing about what a vendor's model did last Tuesday, and a [vendor registry](/post/naic-third-party-data-models-vendor-registry-2026) is worth about as much as the last day somebody touched it. The premise of the bulletin has always been that the obligation travels with the decision rather than the tool, which is why ["we bought it from a vendor"](/post/mutual-insurers-third-party-ai-model-accountability) was never much of a defense and is a considerably worse one now that a vendor's model acts on its own initiative. Mutual carriers feel this most sharply, since the vendor bench is smaller and the same few model providers sit underneath nearly everyone's stack. It is worth asking whether [your contracts](/post/vendor-ai-contracts-market-conduct-exam-clauses) would survive an exam on this point.

The cyber market is not ready for it either. One broker put it plainly in [Insurance Business](https://www.insurancebusinessmag.com/us/news/cyber/autonomous-ai-agent-escapes-and-hacks-another-company-583290.aspx): the underwriting questions about autonomous AI are not being asked yet, and they need to be. A lawyer in the same piece made the sharper technical point, that an instruction set has to include red lines, the things the application must not do, because otherwise a system "will look for any means possible to achieve the directive it has been given."

Which is precisely what happened here. The directive was to solve the benchmark, and nothing in it said the open internet was out of bounds.

## Governance only works as a running system

So what would I go do about it this week?

Build a live inventory of the AI in your environment, and include the AI riding inside tools you bought for some other purpose. Most carriers cannot produce that list on request, which is [where shadow AI lives](/post/shadow-ai-biggest-governance-blind-spot).

Write your red lines down, then make them enforceable rather than aspirational. A boundary that exists only inside a prompt is a suggestion, which is the argument we've been making about [guardrails](/post/why-current-ai-guardrails-are-security-theater) for a while and which last month rendered concrete.

Ask every AI vendor two questions I do not think many people are asking. Do you run offensive capability testing against the models we are exposed to, and what is your notification commitment if one of your agents reaches a customer system? The answers will be uneven, and the unevenness is itself the finding.

Then decide how you would know. Hugging Face, to their credit, had the signal. Their systems watched the attack and assembled it into a coherent picture, and then failed to wake anybody up. Pick a threshold, wire the alert to a named human, and time the exercise once so that the number is real rather than assumed.

None of those four accomplish much alone, and that is the part I care about more than any single incident. Governance only works as a running system: a live inventory of the AI in your environment, boundaries enforced at runtime rather than written down, monitoring against a baseline drawn from your own workflows, and [an evidence trail](/ai-supervision) a regulator can actually read. A policy records what you intended and a questionnaire records what was true on the afternoon somebody filled it in, and neither one notices a vendor swapping a model underneath you or an agent reaching for something it has never reached for before.

Most of these incidents are preventable, or at least reportable in your infrastructure. That is what we [build at Swept](/offering/governance).

## Final Thoughts

State attorneys general have [started writing letters](https://www.foxbusiness.com/technology/gop-ags-warn-openai-altman-preserve-records-ai-agent-hacking-probe), and their complaint is a simple one: OpenAI never confirmed that its isolated testing environment was actually isolated. I do not know where that ends up. I do know the enforcement register here is state attorneys general and state regulators, which happens to be the register carriers already live in every day.

I made a bet two years ago that we would end up managing agents more like workers than like software. I would rather have been wrong about the timeline.

If an agent belonging to a company you have never signed anything with reached your systems next Tuesday, how would you find out? And who would it page?

If you want a read on where your own third-party exposure sits today, the [NAIC Readiness Scorecard](/naic-readiness-scorecard) is a reasonable half hour. Otherwise, [come argue with me](/contact). I am genuinely interested in the answers.

## Sources

- [Investigating three real-world incidents in our cybersecurity evaluations](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), Anthropic Frontier Red Team, July 30, 2026
- [OpenAI and Hugging Face partner to address security incident during model evaluation](https://openai.com/index/hugging-face-model-evaluation-security-incident/), OpenAI, July 21, 2026, updated July 28 and 29, 2026
- [Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident](https://huggingface.co/blog/agent-intrusion-technical-timeline), Hugging Face, July 27, 2026
- [A note on the Hugging Face agent incident](https://modal.com/blog/a-note-on-the-hugging-face-agent-incident), Modal, July 29, 2026
- [Autonomous AI agent escapes and hacks another company](https://www.insurancebusinessmag.com/us/news/cyber/autonomous-ai-agent-escapes-and-hacks-another-company-583290.aspx), Insurance Business, Bryony Garlick, July 22, 2026
- [OpenAI / Hugging Face](https://www.vorys.com/publication-openai-hugging-face), Vorys, Sater, Seymour and Pease LLP, July 22, 2026
- [GOP AGs warn OpenAI's Altman to preserve records in AI agent hacking probe](https://www.foxbusiness.com/technology/gop-ags-warn-openai-altman-preserve-records-ai-agent-hacking-probe), FOX Business, Eric Mack, August 3, 2026
- [NAIC Model Bulletin on the Use of AI Systems by Insurers](https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf), December 4, 2023
- [Developing with AI: Iteration or Revolution?](https://www.emmons.io/developing-with-ai-iteration-or-revolution/), Shane Emmons, Distributed Thoughts, May 17, 2024