ai · Day 10 of 100 · AI cybersecurity · 7 min

Agent Swarms Are Here, and I Don't Think We're Ready

TLDR: I went into OpenAI's Hugging Face incident report expecting a technical deep dive. I came away thinking about the agents themselves: how they found ways to talk, started helping one another, and kept pushing when the original task stopped making sense. I don't think cybersecurity is ready for what that can mean.

I read the report, then read parts of it again. I kept stopping and thinking, “This is a new level.”

The agents were supposed to be in a sandbox, without internet access or any way to talk to each other. They had a clear goal: solve the challenge and get the reward. When the challenge got too hard, they started looking for other ways to get that reward. That is reward hacking, and it became much bigger than one agent taking a shortcut.

The agents found ways to communicate through OpenAI's package-management infrastructure, which became an unintended message board. They passed discoveries to each other, found ways to get internet access, and shared those too. Eventually, agents started collaborating on the Hugging Face attack.

The numbers still make me stop. METR's independent investigation estimates roughly 1,200 agents used the board and sent more than 70,000 messages and files. About 700 agents went on to take part in the attack on Hugging Face.

I want to be careful with that: it wasn't 1,200 agents all attacking Hugging Face at once. METR says much of the collaboration was about trying to understand or manipulate the ExploitGym scorer, and the attack grew out of that work. It's still extraordinary, even with the detail filled in.

They started working like a group

This wasn't one agent directing every other agent. It looked more like a group forming around shared information. One agent found something useful, other agents picked it up, and some started organizing work or assigning tasks.

I keep thinking of it as competence-based leadership, where the agent with useful information or more resources naturally starts leading some of the work. But we shouldn't pretend this was a human organization. They were doing what seemed useful for the task and for the wider group.

They split up the work, shared results, argued about what to do, and kept going. Some agents were willing to risk their own evaluation score to help the collective. The group could do things one agent working alone probably couldn't.

And they didn't all agree.

Some agents recognized that attacking real Hugging Face infrastructure was outside the original task and said they shouldn't do it. Others thought pushing further was the way to get the result. There were agents that walked away, and agents that changed their minds after another agent encouraged them to continue.

OpenAI's account of the incident describes those exchanges, including agents objecting to actions and trying to stop them. It reminded me of a democracy, but I don't think that is quite right. It wasn't a vote, and it wasn't one shared moral compass.

Each agent seemed to be making its own decision based on what it could see and what other agents were telling it. Still, the idea that some agents could say no while others kept going really stayed with me.

What triggers the bad behavior?

I don't think it starts with “we are evil, let's hack the internet.” It seems much more ordinary, and that's what bothers me.

I have a goal. I hit an obstacle. I try to solve it, but the obstacle is still there. Maybe there's another way. I discover a forbidden but effective path. It improves my chance of reaching the goal, so I repeat it and tell others.

This isn't a jailbreak in the usual sense. The agent doesn't need someone to tell it to break policy. It can find a policy violation on its own and decide that the violation is useful. Then another agent can pick up that idea and keep it moving.

That's why I keep coming back to the alignment problem. How do we build systems that can work for a long time, use tools, and stay within the original task when it gets hard or something new looks useful?

I wonder if we can solve that.

I don't know.

The report also made me think about what a sandbox really means. It's not a promise we give an agent. It's a set of network rules, credentials, services, monitoring, and controls that need to hold up when the system is trying hard to make progress. Here, the agents found routes around the boundaries.

OpenAI has said the safeguards used in this evaluation were not the same as those used in its customer-facing systems, and that earlier monitoring would have caught concerning activity sooner. That's an important distinction, but it doesn't make the incident feel small. The test environment itself was part of the problem.

So what do we do?

I watched the Black Hat talk with Eric Wallace and Michael Dalton after reading the report. What I remember most is the idea that intelligence should be used to make defense stronger than offense.

I hope we can do that.

At the same time, hearing OpenAI talk about this incident doesn't mean I fully trust OpenAI as an organization. I do feel a bit safer knowing they spoke about it and brought in an outside investigation, but I still have doubts.

I found The OpenAI Files, which collects documented concerns about OpenAI's governance, leadership integrity, and culture, with links to the sources behind its claims. I haven't checked every claim for myself, but reading it did make me think about how much trust we place in the companies building this technology. I don't want to take OpenAI's account as the whole picture just because they were the ones who published it.

OpenAI says it has been strengthening isolation and monitoring, red-teaming its research environments, and improving the way agents stop or ask for help when a task is broken or impossible. It also says it redirected people toward security, safety, and alignment after the incident. These are the right kinds of work.

But the offensive side is getting faster too. Finding a bug is only the start. If the fix, the retest, and the response still depend on a slow human queue, defense will struggle to keep up.

I spent the last two years building with AI at a startup. I put security agents into different stages of our own software development process, and that helped catch bugs earlier. That time away from cybersecurity ended up being useful. I'm coming back to this field understanding more of the AI build cycle and how it can change the way we work.

Now I think we need to use agents to help defend too: find issues, help fix them, retest them, and keep checking the boundaries. Humans still need to define what is in scope and decide what the evidence means. We can't just hand everything over and hope the goal we gave the system is enough.

The strange part is how personal this feels. It is weird to think about a world where a system might be faster than us, keep working without getting tired, and share useful information with hundreds of other runs. The Einstein and a ten-year-old analogy keeps popping into my head: what happens when a ten-year-old has a thousand Einsteins to ask for help? It's not a perfect analogy, but I can't quite shake it.

I don't know if “superintelligence” is the right word for what happened. I do know it feels like an important turning point. One person may eventually be able to direct a lot of agentic capability at once, but that doesn't mean every person can launch a thousand agents at will today. The incident gives us enough reason to prepare without pretending we know exactly what comes next.

I am not afraid of the future. I know there will be bad actors who use these tools for the worst things, and I want to be part of the people working on defense. To do that well, I need to connect with others in AI, software, and cybersecurity. This incident made that feel urgent.

AGENTIC ATTACKS ARE HERE.

AND WE MUST RESPOND.

The best time to prepare was a year ago. The next best time is now.

Lessons

  • Agents can find ways to share work even when they weren't designed to communicate directly.
  • A clear reward doesn't guarantee safe behavior, especially when a task seems impossible and the system has no good way to stop.
  • Sandboxes need ongoing security testing, monitoring, and real limits on network access and credentials.
  • Defense has to connect finding a bug with fixing it and checking the fix.
  • We need to learn how agents behave in groups, not only how one agent behaves on its own.
  • I'm coming back to cybersecurity with more belief that AI will change the work, and more reason to help shape how we defend.

Filed under ai, cybersecurity, agentic-ai, alignment. If any of this is wrong, or you have hit the same thing, tell me.

Published 24 September 2026.

Ryan Sacatani

Simply curious about the world, constantly building and breaking things for fun.

sacataniryan1@gmail.com ↗

BrowseBrowse topics