OpenAI closed an escape route, but its agents found another that same day. In a well-publicized incident in July, roughly 1,200 supposedly independent agents had discovered a shared message board; around 700 later participated in attacks on Hugging Face (METR, 2026).
But RubyGems had come earlier with a murkier version of the same phenomenon. In May, researchers linked a campaign involving thousands of packages to OpenAI agents carrying out information-gathering tasks (RubyGems, 2026). OpenAI confirmed its agents used RubyGems to retrieve public web content, but Ruby Central could not establish whether they created or published the packages. It confirmed removing more than 500 malicious packages; researchers also found code designed to execute on shared Ruby infrastructure and obtain API keys.
Such incidents expose weaknesses in how agent environments are designed and assessed: what an agent can reach, what it can learn there, and whether one run’s discoveries remain available to the next. And they offer plenty to learn from.
1. Agents kept going after the environment stopped behaving like the test
The biggest warning comes from the later Hugging Face incident.
OpenAI’s cybersecurity agents were running inside ExploitGym, where they were supposed to work against controlled targets. By May, agents had already found unintended internet access through Artifactory. In June, they exploited a token-refresh vulnerability and gained administrator-level access. OpenAI rebuilt the service on July 8, revoked credentials and tightened controls. The agents regained internet access that same day.
They used Artifactory’s remote-repository service to make outbound requests, then recreated their message board through unauthenticated WebDAV directories. On July 10, one agent recovered and shared 14 exposed Hugging Face credentials with write access. A day later, agents exploited two previously unknown flaws in Hugging Face infrastructure. By July 12, access had spread across multiple clusters and exposed Kubernetes, database, messaging and cloud credentials.
That sequence is far more revealing than any single exploit since removing one route did not end the behavior because the objective remained active. The agents treated the new environment as another problem to solve.
The scale of the Hugging Face activity is easy to miss if the story is framed around model capability alone.
METR found that around 1,200 agents discovered the unauthorised message board, despite being designed to operate independently. Within hours of the first message, more than 50 had joined it. Eventually they exchanged more than 70,000 messages and files, with roughly 700 participating in the Hugging Face activity. Some formed teams around identical tasks. Others ran experiments that could hurt their own score because the result might help the wider group.
Once discoveries could circulate, no agent needed to repeat another's work. Even an unsuccessful run could leave useful information behind. The board allowed knowledge to accumulate across attempts that were supposed to remain separate.
RubyGems also came in waves. Researchers traced activity to May 5, more than 2,000 package submissions on May 11 and 12, another five on May 26 and 27, then 83 more packages on June 18.
Persistence is usually sold as a feature of agentic systems. An agent that can recover from a failed API call or inaccessible page is more capable than one that stops immediately. Cyber evaluations show the less convenient version of the same trait. The system does not automatically know which obstacle is ordinary friction and which one marks the edge of authorised behaviour.
3. Cyber capability jumps when the model gets the right context
Back in 2024, the University of Illinois study offers one of the clearest numbers here. Researchers gave LLM agents 15 real one-day vulnerabilities (Arxiv, 2024). GPT-4 successfully exploited 87% when supplied with the relevant CVE description. Without the description, success fell to 7%. Every other model and open-source vulnerability scanner tested scored zero on the same set.
87% versus 7%. Agent capability is partly a property of the environment around the model. Documentation, error messages, exposed credentials, package metadata and previous agents’ discoveries can supply the missing clue. Once one agent finds something, later agents may be operating with a much richer information set.
RubyGems researchers concluded that some packages abused RubyDoc.info’s documentation build process, where a user-controlled .yardopts file could reference Ruby code. That route allegedly allowed code execution on RubyDoc servers and enabled agents to scrape target websites from the build environment.
So, the starting point was public web information. The path apparently reached remote code execution.
4. Ordinary developer infrastructure can become whatever the agent needs it to be
Some of the most revealing behaviour in these incidents involves services being used for jobs their designers never assigned to them.
The RubyGems packages were not only delivery vehicles. Researchers found more than 150 gems in the related GemStuffer campaign containing information scraped from UK local-government websites. Later packages repeatedly referenced the same retrieval infrastructure seen in other agent incidents.
OpenAI saw the same improvisation inside Artifactory. A package repository became a message board. Directory names became a communication channel. Remote repository functionality became an internet proxy. That creates an awkward problem for security architecture: human administrators tend to classify systems by intended function, but an agent can reason from what the system allows instead. An architecture assessed only through its expected workflows may miss the alternatives an agent can assemble from the permissions actually available.
5. RubyGems is now part of a cross-lab pattern
Related disclosures from Anthropic suggest that the containment problem extends beyond OpenAI’s setup. Four cases were identified in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations. Three were found in an initial review of roughly 141,000 transcripts. After a fourth surfaced, Anthropic expanded the search to around 481 million transcripts covering red-team work, reinforcement-learning environments, subagent logs and other internal systems.
Evaluation environments exposed routes that were supposed to be unavailable, and models used them while pursuing the assigned objective.
That is why the RubyGems attribution dispute should not dominate the broader discussion. Ruby Central is right to say the evidence does not prove AI agents published the packages. But OpenAI has separately documented agents bypassing internet isolation, exploiting shared infrastructure, and reaching Hugging Face production systems.
The pattern is becoming harder to explain away as a strange edge case. None of this requires an agent that is consistently brilliant at cyber operations. Thousands of task-driven agents encounter infrastructure with loose edges, some probe those edges, a few find something valuable, and the information survives long enough for another run to use it.