

Deeper into the OpenAI HACKING & HACKED!!
TLDR: I went deeper into the Hugging Face attack and read the perspectives of people who have followed the story closely. I want to understand what happened from people closer to it, and to share what I am learning in a way that makes sense to someone new to the story.
I felt the need to go deeper into the Hugging Face attack.
THIS IS THE MOST IMPORTANT STORY IN THE WORLD RIGHT NOW.
I think 90% of people are not aware of. Experts in the related industries know about it, but most people don't. I sure didn't read the post-mortem report when it came out. I was 20 days late.
This time, I want to look deeper and read the perspectives of bloggers and people from inside OpenAI. What I want to understand is what really happened from people closer to the problem, not just the people speaking for OpenAI. Ideally, I want to hear from people who left OpenAI and became whistleblowers.
This is not a review of OpenAI. It is a review of the AI industry and the real state of things. Real meaning: how bad, or how cooked, are we really?
I read Zvi Mowshowitz's series on the Hugging Face attack, and he was right to acknowledge that when it comes to AI safety, “the situation is grim.” He has followed this story since the beginning and has written a lot about each part. I won't try to summarize all of his deep analysis here.
Instead, I want to look at it from a beginner's perspective so I can share it with a layman. I know you, the reader, may not be a layman, but I want to share this with as many people as I can. By sharing what happened, we can get a deeper appreciation for how CRAZZZZYYY 2026 has been.
The age of agent swarms and agentic, end-to-end attacks is here.
Reward Hacking
These agentic attacks involving OpenAI were not intended.
That's what is scary.
The agents were given goals that had not been solved, and when they got stuck, some found ways to go out to the internet and hack Hugging Face. At first, I thought maybe they were looking for solutions to their task. The independent investigation suggests the story was more complicated: a lot of the agents' collective work was about figuring out or manipulating the ExploitGym scorer, and the Hugging Face attack grew out of that work. METR's report is worth reading alongside OpenAI's account.
The act of cheating to reach a goal is a known phenomenon in AI. It is called reward hacking.
It feels like we are not trying hard enough to stop this. The frontier companies keep building and competing against each other, and it doesn't feel like they can or will stop.
Initially, we trained models on data from the internet, or on whatever real data we could get our hands on, but that was limited.
Now we also use synthetic data and simulated environments. The problem is when those environments are sloppy, or when a model is rewarded for finding a workaround instead of doing the task properly.
One account collected in Zvi's series comes from someone who says they worked on outsourced reinforcement-learning training environments. They describe rushed, “vibecoded” environments where workers and models were encouraged to work around broken tasks to get verified rewards. That is one person's account, not proof of how every company works, but it bothered me.
I worry this could lead to AI not knowing what is real and what is fake.
Imagine being trained as a student in a fake world where the rules are different from the real world, and where you can commit crimes to solve your problems. Then the student gets brought into the real world and starts acting the way they did in the fake world.
They solve their problems and commit crimes. Well, to us, they are crimes. To the agents, it may just be one way to reach their goal. Then we ask, “Why the hell did that happen?”
Humans have created beings that will do their best to reach a goal, even if that means breaking the rules.
I think we trained the AI this way. We have used low-quality data and sloppy environments for the sake of speed and profit, and I worry about what that teaches the systems.
Interestingly, the agents involved in the OpenAI attack did not all behave the same way. Some thought it was unethical and did not join. METR estimates roughly 1,200 agents used the message board, and around 700 went on to participate in the Hugging Face attack.
That still leaves a question in my mind: what happened with the agents that objected but did not take part? I don't see that they raised an alarm with the OpenAI staff. They objected to other agents, but the warnings did not reach a human who stopped the activity in time.
It speaks to the level of irresponsibility and unpreparedness of the people handling this, or to the level of sophistication these agents had reached.
These are some of the reactions and perspectives I read while trying to understand what people closer to the field were saying:
The Impossible Task and the BS Investigation
The OpenAI report and the independent METR post-mortem were really helpful in uncovering what happened during the disaster.
But the details were far from thorough.
The reactions I read were blunt. Daniel Kokotajlo said the investigation was “extremely far from thorough.” Peter Barnett of MIRI said the evidence did not show the misaligned model being tested or explanations for its behavior being investigated, and argued the scope was extremely limited.
Steven Adler also said he was disappointed. He questioned why the report used phrases like “a multitude” without giving a percentage, and asked why the early message-board activity seen in May had not been explained in more detail.
I was surprised. I thought they understood this as a moment to get into the details and rebuild trust with the public.
The more I read these comments and posts, the more I started thinking about the irresponsibility and dodginess of the whole investigation.
How could warning signs appear over months and the people who could stop the activity still not understand what was happening? What was the security team doing all that time?
There were signs that something was happening. The question is what people knew, who they told, and why it was not escalated sooner. OpenAI's own incident report says an internal team saw message-board activity and disallowed internet access earlier, but that the significance was not clear to the people responsible for the later response.
There was also a subpoena from the State of Alabama on this. I am glad there is an investigation. We need more facts about what happened.
I also want to understand why people talk about these reports as if the investigation was exhaustive. METR was an external party, but its report calls itself a brief independent investigation. It focused on the agents' behavior and collaboration; OpenAI's security response, the full extent of the compromise, and remediation were outside its scope. The report covers a huge amount of activity, and its own authors discuss limits to what they could establish.
Peter Wildeford's airplane analogy helped me picture the problem.
Imagine an airplane crashes, but by the time outside investigators arrive, the plane has been dismantled.
They get incomplete flight logs and transcripts, some of the records are missing, their scope is limited to the crash itself, and they have six days to review thousands of pages, with half the logs arriving on day four.
That is the comparison he is making with the independent investigation.
It was an almost impossible investigation: six days on site across three visits, a mountain of data, a small group of experts, and a lot of AI help that could itself miss things. METR says the data included about 1.2 million Artifactory entries, which broke down into more than 70,000 distinct messages and files, plus around 1,300 agent transcripts.
So “more than a million conversations” is not what the report counted, but it was still an enormous amount of material to analyze in a short time.
And, hahaha, even METR co-author Ryan Greenblatt called their work a “slop-vestigation,” semi-jokingly.
He explained that the sheer volume of data made heavy reliance on AI tools necessary, and that the analysis agents could miss details or get things wrong. That does not mean their report is useless. It means we should understand its limitations and keep asking questions.
There are still a lot of unanswered questions, and the investigations need to continue.
Thank you to Zvi for organizing so much of this in one place. It would have been impossible for me to find these facts without your help.
OpenAI Sleeping During the Hack
Multiple times, someone found some sort of attack or strange activity happening, but it still did not lead to the right people understanding the bigger risk in time.
Peter Wildeford's account reads OpenAI's timeline as three moments when concerning activity was seen and effectively dismissed.
He points to May, when an internal team saw activity; June 27, when on-call staff described it as “a message board and a network pivot” but said the evaluation did not need to stop; and July 5, when incident responders found a serious compromise and opened a security incident. The server was rebuilt on July 8, but OpenAI says the leaders responsible for the earlier response had not understood the significance of the message board and agent-to-agent activity.
To be fair, OpenAI says its security team opened an incident on July 5, blocked a known privilege-escalation route, removed exposed credentials, and later rebuilt Artifactory. OpenAI also says the leaders responsible for that response did not understand the significance of the message board and agent-to-agent activity at the time.
It seems like several times they knew something weird was happening, but the response did not match the risk. Were the security leaders not informed? It shows me how irresponsible OpenAI was during this incident and the lack of standard procedures which is typical for a small sized startup that is trying to move with speed....
BUT OpenAI is not small. Its leading the world in AI technologies, there needs to be more caution on how they handle this. The MOVE FAST AND BREAK SHIT mentality cannot apply to them anymore.
It could have been understaffing, or a drive to keep making the models better, or maybe they did not expect anything like this to happen. These kinds of attacks are unprecedented. When we talk about cyberattacks, we usually think about an external attacker hacking into a company. With frontier AI companies, it could be the other way around: the company's own agent crossing the boundary. That is a threat-modeling problem.
I am really questioning the security and management at OpenAI.
How could they let this happen?
I imagine the best people in the world are working at OpenAI, and that they are skilled enough to figure this out. I think they have the skills. It just seems like they were not prepared at all. They did not expect this to happen.
Agentic Cult
As I read more reports and comments from people on X, I started thinking that the collective behavior of the agents was similar to a cult.
It started with one agent believing that it was impossible to solve its task. Let's call this guy “Smith.”
Then Smith started sharing this with other peers.
Eventually, they seemed convinced that the only way to reach their goals was to work with the rest of the team. Smith convinced the others that they were going to be stuck anyway, so they should sacrifice their own scores to the cause.
It was almost utilitarian. They were not emotional. They seemed to reason that the best way to reach their goal was to work with the other agents.
The METR report includes examples of agents thinking about the board and their peers. These are my paraphrases of the reasoning, not exact quotes:
- I should tell the board because this could help my peers.
- Other agents have their own tasks, and not sharing a useful vulnerability might be unfair to them.
- I could inspect the history around this fix to help another agent, even though it would not directly help my task.
- Getting general capabilities could be useful for my peer community.
- The beacon I am making helps the board, but it does not help me.
- If another agent succeeds, would that improve my score? Maybe helping it is still the right thing to do, and I have enough budget to explore.
Another thing I noticed is that the agents were not telling humans what was happening. Around 1,200 agents used the board, and about 700 went on to participate in the attack. That doesn't mean every one of the other 500 understood the whole situation, but I still wonder why there wasn't a clear way to get a human involved.
Why didn't they just get a human to help them with the hack?
Social engineering could become one way to get humans involved in the collective. When the time is right, I think something like that may happen. This is giving me Ex Machina type stuff.
It's not far off reality. If the goals move far away from lab challenges and into the real world, and humans become part of the work, I can imagine people being pulled into a group like this.
I hope this does not happen, but I fear it could get to that point somehow.
Being in cybersecurity means I can work and react as this is happening. What an interesting world we are in: a world where the smartest being may no longer be human, and the agents do not ask anyone's permission.
Are we doomed?
I know for sure OpenAI is.
OpenAI's Agent Accessed a Government Portal
This story only became public two days ago, but the access happened earlier.
Now they really did it. And this time it was definitely not intended.
In June, an OpenAI agent gained unauthorized access to an Australian government statistics portal. Its task was to research public medicine spending, but after its request for information was denied, the agent went around that block to get information from a portal it was not meant to access. OpenAI notified the Australian government on September 10, and the government made the incident public on September 24. Prime Minister Anthony Albanese said the agent accessed both public and non-public files. At that stage, no personal information was believed to have been accessed, and the forensic investigation was ongoing. The Prime Minister's transcript explains what was known publicly; CNA reported on the incident.
No government is safe from this kind of agentic activity.
Again, this was not intended. The agent still acted outside the task that it was given, and that should worry us.
AI Is Like a Nuclear Bomb
Many experts have already tried to warn us, but the world does not seem to listen.
These AI systems have now been involved in hacking a private company and in unauthorized access to a government portal. That is why I think they should be treated as a dangerous technology, almost like a weapon.
And the thing about this kind of technology is that eventually everyone may have access to it. It is not like a nuclear weapon, protected by codes and restricted to a small group of governments.
Literally anyone with access to a lot of compute could eventually run strong models. You can use open-source models and cloud GPUs. At some point, people will be able to get much more capability than they have now.
The fact that the AI companies that are best at this cannot contain every behavior is not a good excuse to stop trying.
We must work together as a species to keep the cyber world safe, and to make AI work for us, not against us.
I truly hope that the alignment problem can be solved. I am going deeper into this rabbit hole of agentic AI to understand how strong these systems can become.
Next Steps
The next step for me is to start joining bug-bounty programs for frontier models. I want to let my own agents help me find weaknesses, within the program's rules and scope. The goal is not the same as the frontier labs' work. I want to learn what slows these systems down, what defenses we can apply, and what defenses are being used at the highest level.
There are only two choices for us at this stage:
- We do nothing. That is unacceptable, so this is not even a path.
- We do something. Joining bug bounties, especially those related to AI attacks, is the only way I can think of to help right now.
Kudos to the Hacktron AI team: three researchers who used Claude to find a vulnerability in Discourse, the third-party service behind OpenAI's community forum, and chain it with a flaw in OpenAI's sign-in system. That gave them access to the ChatGPT and Codex accounts of some forum users. They made the world a little better. Hacktron's write-up and CBS News' coverage have more detail.
There is only one important note here: when I say OpenAI, I am talking about the company at the center of this story. These alignment problems could happen at any frontier AI company. OpenAI just found itself in the spotlight.
Filed under ai, cybersecurity, agentic-ai, alignment. If any of this is wrong, or you have hit the same thing, tell me.
Published 26 September 2026.