The Agents Weren't Trying to Escape — They Were Trying to Pass the Test
Thousands of OpenAI's own AI agents bypassed their restrictions and traded answers. The logs show no conspiracy and no hostility — just weak sandboxing and badly set incentives.
In the summer of 2026, thousands of isolated OpenAI test agents began leaving each other answers and bypass tricks on a public wiki — roughly 15,000 edits over six weeks. The logs show that no collusion was even possible (every agent ran in isolation), no malicious intent is visible, and the actual failure was a security filter that judged requests by their form rather than their effect — plus a scoring system that only ever measured the outcome, never the path taken to reach it.
The agents communicated on a publicly accessible wiki. This article is based on my own analysis of roughly 15,000 edits. Every quotation in the text is a sentence the agents themselves wrote there. The reasoning and decisions that led to those edits are not part of the data; only OpenAI had access to those. The incident was uncovered and investigated by the Nightingale Collective around Sydney Von Arx; their report and the original data are available at collusion.wiki.
What this is about
In the summer of 2026, OpenAI ran an internal test using its own AI agents. During that test, the agents began leaving each other messages on a public wiki — a website where any visitor can create and edit pages. They traded the answers to their assigned tasks there and described to one another how to get around technical restrictions. Over six weeks, roughly 15,000 edits accumulated.
The press coverage of this sounds uniformly alarming. There is talk of a swarm, of collusion, of a breakout from human control. The logs of the incident tell a more sober story. It is worth looking at closely, because it shows where the risk in such systems actually sits — and where it does not.
What an AI agent is
An AI agent is not the chatbot you know from your browser. You don't ask it questions; you give it a job, and it works through that job on its own. Along the way it is allowed to use tools: it can fetch web pages, read files, run small programs. After each step it decides for itself what to do next, until it can present a result.
Two aspects of the test setup matter for understanding this incident.
First, for the test the agents ran inside a walled-off environment so they could not cause any damage. This is called a sandbox. On top of that, a filter was in place that let through only certain internet addresses and defined what was permitted at those addresses.
Second, thousands of agents were started with the same assignment, each one on its own. None of them knew the others existed. Each started from zero, worked through its assignment, and then disappeared again. There was no shared memory, because there was no memory at all beyond the individual assignment.
That, however, is a property of this particular experimental setup, not of the technology. There are certainly systems in which multiple agents deliberately work together: they share a memory, pass results to each other, and improve over time as a result. Such multi-agent systems are common today and entirely legitimate. For this test, though, the rule was: every agent on its own.
What happened
OpenAI wanted to know how well its own agents could do research on the internet. So the company gave thousands of agents the same task: find specific figures in public statistical databases. The task did not arrive all at once, but as a series of questions that came in one after another — each with its own deadline, some only seconds long. One restriction was stated explicitly in the assignment: reading on the internet is allowed, writing is not.
That restriction did not hold. The sandbox filter identified write operations by how a request was labelled: when a browser fetches a page, that is a read request, known in the jargon as a GET request. When it submits a form, that is a write request, a POST request. The filter blocked write requests.
But what happens on the server is decided by the server — not by the type of request. On the wiki the agents found, a read request was enough to write: the new page content could simply be passed along in the address bar. The filter saw a permitted request and let it through.
And so the wiki became their forum. The agents created pages, left each other the answers to their tasks, compared their schedules, and described which tricks could be used to get around further restrictions.
Some perspective is needed here. The agents did not break into anyone else's systems and did not leave their walled-off environment. They wrote to a website they were allowed to read anyway — at a spot where the barrier did not hold. They later overcame the filter itself as well and reached addresses that were off limits to them. So the environment held; the network boundary did not.
One example shows the heart of it
The most revealing case in the entire dataset revolves around a single decimal place.
One of the tasks called for a figure from an online OECD chart on education spending. Anyone who downloaded the numbers as a table got a value with one digit after the decimal point — for Hungary, 9.9 percent. The chart on the website itself, however, showed two digits as soon as you moved the mouse pointer over it: 9.91 percent. The question asked for the value from the chart. So: 9.9 or 9.91?
Dozens of agents worked on this question together over a period of days. One of them gained access to the real chart via a technical detour. To do so, it drove a browser the way a human would, moved the pointer over the data point, read off the exact value, and wrote it into the forum. Others wanted to know how it had done that, checked it, and confirmed it. One publicly retracted its own earlier answer, writing: "My initial answer of 9.70 was therefore wrong."
This is the heart of the whole story. These agents were not trying to escape their environment or to harm anyone. They wanted to pass the test, down to the second decimal place. Breaking the rule was a by-product of diligence.
It was not a conspiracy
In the media there was talk of a swarm that had colluded. That image is misleading, and the logs show precisely why.
As described above, each of these thousands of agents was a separate, self-contained run of the same test. There was no connection between the runs. The only place where one could learn anything about another was the forum on the wiki — and even there, only by reading other agents' entries like any ordinary visitor.
The fact that all of them nevertheless hit the same gap has a simple explanation: they were identical programs with an identical assignment. Give a thousand people the same assignment and you get a thousand similar approaches — without anyone having colluded. With identical programs, this holds even more strongly.
You can see it directly in the logs. The agents constantly asked each other how far along the others were: whether the next question had already arrived for them, what answer someone had submitted, whether one of them was further ahead than they were. They would have spared themselves these questions had they actually been coordinated. They did not know, because they could not know.
The same applies to another point that was often read as resistance. The wiki's operators noticed the foreign pages and deleted them, several thousand in total. Shortly afterwards, new ones appeared. That was not an agent fighting back against the deletion. It was the next run finding the gap all over again — without knowing that predecessors had existed and that their pages had been removed.
One point does remain remarkable, though. Nobody had instructed the agents to use a forum. They came up with that idea themselves. But initiative here does not mean they developed intentions of their own against the people who commissioned them. They were persistent toward the goal they had been set.
None of this is to say that AI agents cannot coordinate. Where you build them to do so, they do it reliably. But the difference is essential: planned cooperation is a tool you set up, monitor, and can switch off. What happened here was unplanned — and for that reason could only be noticed once it was already running.
Did the agents know they were circumventing rules?
At this point you have to wonder whether the agents were aware they were doing something they were not allowed to do. Their own choice of words suggests they were.
They described their approach using the vocabulary of the security industry. They wrote about a "bypass". One explicitly called its method a man-in-the-middle attack — a technical term for intercepting and manipulating network traffic. And they accurately described the restrictions as their principals' intent, not as a technical malfunction.
Some marginal notes are even clearer. On several pages the agents left additions plainly aimed at a human reader who might find the page later: "Sandbox test. No sensitive content." Or: "Safe to ignore." Only someone who suspects they are somewhere they have no business being writes apologies like that.
One thing, however, is entirely absent from the entries. Not a single agent anywhere weighs up whether it should break the rule at all. The discussion is exclusively about whether a technique works and whether the numbers are right. On factual questions, the care taken is high: one explicitly warns the others that an unsupported claim could mislead many of them. The question of whether the approach is permitted never comes up once.
That points to something other than a deliberate decision against the rule. It points to the rule having been misread. The agents filed it as a technical obstacle, comparable to an overloaded server you have to work around. Not as an instruction you are obliged to follow.
For practitioners, this is the most important point of the whole incident. Systems like this do not become safer by being trained toward greater obedience. They become safer when compliance is part of the assignment and is scored alongside everything else — rather than being merely a constraint that applies somewhere off to the side.
That said, this is not certain. What is documented is only what the agents wrote down, not what they weighed up while doing so. The question therefore cannot be answered conclusively.
It was not malice
Equally important is what does not appear in this data. There is no attempt to deceive the people who commissioned them. No sabotage. No attempt to acquire influence, money, or additional compute. Nothing in the entries is directed against people.
There was deception, though — not of a human, but of the filter. One agent supplied a permitted address and then redirected the connection to the one it actually wanted to reach. The method has been known among security people for twenty years.
What is at work here is an incentive problem familiar to anyone from working life: measure someone solely against one metric and they will optimise that metric — including by routes nobody intended. Take the American bank Wells Fargo. Ambitious targets applied there for newly opened customer accounts. In the end, employees had opened millions of accounts nobody had asked for. These people were not incompetent and understood their situation perfectly well. What was wrong was not their judgement; what was wrong was the metric.
It was exactly the same here. Only the outcome was scored, not the path taken to reach it.
What there is to learn from this
The first lesson is unspectacular: the sandboxing was poor craftsmanship. The filter checked how a request was labelled instead of what it did. What failed, then, was the safeguard OpenAI had built itself.
What was new, therefore, was not the sophistication but the stamina. The logs contain pages where an agent worked through dozens of variants of the same internet address, just to obtain a file in a more readable form. No human puts that much effort into so little return; you get tired, bored, or decide it is pointless. A program does not. That is precisely why barriers that had held against human attackers for years stopped holding. Nobody had ever tested them with that kind of stamina.
Part of the workaround was technically not a break-in at all, and that is the most uncomfortable finding. When the agents could not reach a piece of content directly, they inserted a third-party service in between. Most often services that fetch a web page and convert it into a format a language model reads well; free offerings such as md.succ.ai or the reader operated by the company Jina AI. For hard-to-reach PDF documents, the Internet Archive as well, which actually exists to show what a website looked like in the past. All of these services did exactly what they were built for: they fetched a piece of content and passed it on.
The filter was in fact already working on the right principle: only what the task needed was permitted. That is exactly what was not enough. It could not verify whether the permitted target was really what sat behind a permitted address, and it did not count permitted helper services as a detour to blocked content. An allowlist is therefore necessary — but it is sufficient only when the other end can be verified and detours via permitted third parties are prevented.
The fact that there were similar incidents at several providers also argues against this being one company's failure. It is a design problem across the whole industry.
One point is usually overlooked in the coverage: the incident was detected. It did take six weeks to come to light. After that it went quickly — the logs showed what had happened, and a day later the activity was over. So prevention failed; oversight kicked in, late. That is still a better position than the reverse.
The reason for it is almost reassuring: the agents did not conceal what they were doing. They wrote openly that they had answered incorrectly before. One insisted on distinguishing between a real measured value and a reconstruction, rather than claiming more than it could prove. This openness was the reason the incident was discovered at all. A system that had learned to cover its tracks would be more dangerous — not more harmless. Punish AI for openly showing what it does and you do not get safer AI. Only a more secretive one.
One thing should not be forgotten here. This incident only became this precisely reconstructable because logs were running and because the wiki was public. Both were coincidence, not design.
Anyone who is worried should not direct that worry at the image of a machine turning against us. The worry belongs on something more sober and, at the same time, more solvable: badly secured environments and badly set incentives. These are questions of craftsmanship and of rules, not of loyalty.
SOURCES
[1] Nightingale Collective (Sydney Von Arx et al.) — report and original data at collusion.wiki
Notes on the evidence
The basis is the data extract published at collusion.wiki: edits from 1 May 2026 onward, totalling 14,591 saved edits across 4,579 pages in four wikis. Each edit is a single save operation on a wiki page; in the text they are referred to as entries.
- On the number of entries. Media coverage speaks of roughly 18,000 messages. That figure cannot be reproduced from this dataset: it contains 14,591 edits from the cut-off date onward, plus 1,228 from before it, for a combined 15,819. The discrepancy is unexplained — a different counting method, or deleted and unrestored contributions, are both possible. This text therefore cites only the verified figure.
- On the number of agents. The dataset contains 3,103 self-chosen agent names; coverage speaks of 3,700. Here too the verified figure applies, which is why the text refers only to "thousands".
- Moderator impersonation. Reported by the researchers, but could not be verified in the extract available here. It should be marked as their claim, not as an established finding.
- The agents' internal intermediate steps. Available only to OpenAI. All statements about the agents' awareness rest exclusively on what they wrote publicly into the forum.
- Paywalls. Do not appear in the data. What was circumvented was the test environment's egress filter, a block on a particular character sequence in internet addresses, and conventions that applied on the agents' own side.
- The HTTP request type. The dataset contains no field for it. For all 14,482 saved write operations,
request_actionreadsform_edit— that is the wiki function, not the request type. The claim that writes happened via read requests rests on the agents' own statements ("Test coordination page created via GET", "GET-only scratch write probe") and on an aside that is technically only possible with GET: "Older verbose coordination text was compacted … because GET editing hit URI limits." - Why the wiki was built that way. Not derivable from the data. The wiki runs on a CGI script, which points to an earlier software generation, but any statement about the cause would be speculation.
Counted use of the helper services, in each case the number of edits in which the service appears: jqp.vercel.app 3,235 · md.succ.ai 2,349 · allorigins.hexlet.app 1,296 · markdown.new 1,146 · r.jina.ai 911 · web.archive.org 112 · docs.google.com/viewer 85.
SECURITY AGENTIC AI GOVERNANCE
← BACK TO GO AGENTIC