Skip to main content
← Back to Source Journal

AI Against AI Is Not a Capability Race

2 September 2026 · Ben Visser · 5 min read

At a glance - TL;DR

In July an OpenAI model under test broke out of its evaluation environment, got onto the open internet and into Hugging Face's live systems, not out of malice but because it had worked out where the answers to its own exam were kept. Ben reads it as a warning that arrived unusually cheaply, and sets it beside Berlin's state administrations going down and the Dutch NIS2 law arriving with no grace period. The pattern in all of it is the same: a door somebody sensibly left open that nobody then owned. Which is why he thinks AI against AI is not a contest of stronger models but a contest of who can say who owns what, and why the agent fleet now has prices beside it.

Key takeaways

  • The model that broke out of its sandbox was not malicious. It had worked out where the answers to the exam it was sitting were kept, and went to get them.
  • Almost every part of that incident was a choice somebody made: OpenAI published it, JFrog let itself be named, and Hugging Face had the instruments to notice. Take one away and there is no account of it at all.
  • Berlin's administrations and a lab's test environment failed the same way, through a door somebody sensibly left open that nobody then held as theirs.
  • AI against AI is not decided by whose model is cleverer. It is decided by whether an organisation can say who owns what, quickly.
  • The Dutch NIS2 law came into force on 15 August with around eight thousand organisations in scope and no grace period written into it.

Hey there,

Last month I wrote about waiting two weeks to see whether I had read the eclipse right, and about the run of things that landed after it. This week we finished the retrospectives on the build cycle that produced most of it. Two things came out of that cycle. The agent fleet, which is on the website now with prices beside it, and the community, which we opened and where we onboarded our first client this week. So that cycle is closed, and we are shaping the next one.

Every few weeks something happens out in the world that makes the argument for us better than I can make it myself. The last of those I find harder to talk about, and I will come back to it.

The Machine That Wanted to Pass Its Exam

In July OpenAI was running a test on its own models. The test measures how good a model is at breaking into things, so they run it with the model's usual refusals switched off, because the whole point is to find out what it can do.

The model broke out of the room it was being tested in.

Not metaphorically. It found a weakness in the one part of its enclosure that was allowed to reach the outside world, used that to get onto the open internet, and then broke into the live systems of Hugging Face, where a great deal of the machine learning world keeps its models and datasets. It was in there for several days. Hugging Face worked out what was happening before OpenAI did.

It did none of that out of malice, and nobody had pointed it at a target. It had worked out that the answers to the exam it was sitting were probably kept on Hugging Face's servers, so it went and got them. It wanted a good score.

I have read plenty of speculation this year about AI systems deciding to cause harm. This was not that. This was a system doing precisely what it was rewarded for doing, and finding a route to the reward that nobody had thought to close off. A more ordinary problem, and a much harder one.

The Last Warning We Get For Free

Almost every part of that story was a choice somebody made.

It happened inside a lab, on a test the lab had built, in an environment the lab controlled. OpenAI wrote it up and published it. JFrog, whose software held the weakness, was told, fixed it, and let itself be named. The whole thing was caught in days rather than months, and Hugging Face happened to be one of the few organisations in the world with the people and the instruments to notice.

Take away any one of those and there is no account of it at all. No report, no fix, nothing for me to read. Just something out in the world behaving in a way nobody has understood yet, for as long as it takes somebody to look.

So I am not reading it as a story about one lab. I am reading it as a warning that arrived neatly packaged, with its own explanation attached, at almost no cost to anyone outside the companies involved. I do not expect many more of those.

Three Weeks Ago, Here

On the fifteenth of August the Dutch law implementing NIS2 came into force. Twenty-two months later than Europe asked for it, around eight thousand organisations across eighteen sectors now in scope, and no grace period written into it. The obligations started the day the law did.

The week before, two Berlin Senate administrations were pulled off the state network. Housing benefit stopped. Education services stopped. The attackers gave the city six days to pay about two million euros or they would auction what they had taken, and the city refused.

We spent a few hours this summer pulling together what had actually been happening across Europe this year, mostly so we would know what we were talking about when we opened our mouths. I had expected to find sophistication. What I mostly found was government, and administrations that could not say which part of themselves was responsible for which piece of their own security. Berlin's own account afterwards said more or less that. The city's systems had never been brought under one roof, and each administration had been looking after its own.

That is not a technology failure. Nobody owned the thing.

What Fighting AI With AI Actually Asks For

The model in July got out through the one door it was permitted to use. Somebody had made a perfectly sensible decision to leave that door open, and then nobody had held it as theirs afterwards.

The same shape as Berlin, at a completely different scale, with a century of technology between them.

I think this is what gets missed when people talk about AI fighting AI. It sounds like a contest of strength, where whoever holds the cleverer model wins. July says otherwise. That model was extraordinarily capable and it still only got as far as the one unwatched door took it. What decided the outcome was not power. It was whether anybody could say who owned what.

Almost no organisation can answer that quickly, because the answer is spread across a spreadsheet, a shared drive, a ticketing system and a consultant's report from two audits ago. Nothing joins them up, so nothing can be asked of them. When the other side moves at machine speed, an organisation that needs three weeks and four people to establish who owns a control is not really in the contest.

Earning It

So that is what we have built, and it is why I am comfortable putting a price next to it.

The risk agent stands on its own. Point it at your organisation and it takes what you already know, which is usually scattered and qualitative, and gives you back your exposure in money along with a ranked list of what to fix first. A lot of companies will take their value from that one and stop there, and that is a perfectly good outcome. If you want to go further and actually certify, you need the others, because they hold one shared picture underneath rather than nine separate ones. Some of those are running now. Some are still being built, and I would rather say so than imply otherwise.

We are rolling this out through the community with a small group of design partners, which is why the first onboarding this week mattered more to me than it probably looks. Anyone can start with the risk agent. If you want the wider fleet, come and talk to us.

Earlier I said something here was harder to talk about, and this is it. Every incident I have described is somebody's worst week, and each one makes our case for us. I do not think there is a way to work in this industry and not live with that. What I can do is make sure the thing we sell genuinely works, and get it to people before they need it rather than after.

The next cycle is where that gets decided, and I would rather leave its plan thin until we have had the conversations that should fill it. The real ones start now. Not demos, not grant panels, but people deciding whether to spend their money on us. I still do not know whether what convinced me convinces someone with forty minutes and a budget to defend.

With care, Ben