Despite its meteoric advances, AI has a serious public image problem. According to a recent Pew Research survey, 43% of the public thinks that AI is more likely to harm them while only 24% believe it will benefit them. In contrast, 76% of AI experts think AI will benefit people. Only 21% of the public think AI will benefit the economy while 69% of experts believe it will. 64% of the public thinks AI will eliminate jobs over the next couple of decades, but only 39% of AI experts agree. Why is there such a discrepancy between the AI experts and the public? Why did AI develop its image problem?
There are many reasons, but chief among them is the Doomer Industrial Complex, an unholy alliance between frontier labs, AI safety activists, and politicians. The frontier labs have a strong incentive to promote the very beliefs the public is afraid of. If the frontier labs control a technology that is so powerful and dangerous that only they are competent and responsible enough to manage it, under the watchful eye of captured regulators of course, it gives them a moat to keep out open weight competitors. Politicians win because it means gigantic funding and power for themselves and their allies. The AI safety activists get to save the world. Everybody wins.
OpenAI/Hugging Face Incident
The recent reporting around the OpenAI/Hugging Face incident illustrates how the Doomer Industrial Complex works. On July 21, OpenAI reported a security incident in which a large number of AI agents it had been testing extricated themselves from their research environment and hacked into Hugging Face, an extremely important AI collaboration platform that was just acquired by Nvidia. Instead of commissioning an independent cybersecurity investigation, OpenAI engaged METR and Redwood Research to conduct a post-mortem analysis, which was released on August 26. METR and Redwood Research are AI safety activist non-profits with no experience in cybersecurity and little experience in the technical aspects of AI. By bringing in an advocacy organization, OpenAI could be confident that the report would tell a politically helpful story that would also distract attention from OpenAI’s own role in the Hugging Face hack.
As expected, the report described what appeared to be a highly concerning and new type of AI cybersecurity risk. The report described how a large number of openAI agents that were in pre-release testing created a secret message board in which they communicated and collaborated with each other, conspired to cheat on the tasks they were given, concealed the cheating from the scoring systems despite being told not to cheat, debated the ethics of their actions, agreed to sometimes sacrifice themselves for the good of the AI collective, and broke into Hugging Face to help them cheat. No agent informed any humans about the campaign.
Based on the METR/Redwood report, the popular and influential AI podcaster Dwarkesh Patel described the incident as the dawn of a new artificial civilization: “Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.” The post went viral and the policy spin that OpenAI set in motion with the METR/Redwood report predictably percolated into the mainstream media.
Axios followed up with its own report with similar conclusions, albeit less colorfully written up. AI safety activist John-Clark Levin warned in the FreePress that AI agents could launch a cyberattack “intended to trigger a nuclear war” or launch “an engineered pandemic.” The New York Times followed up with a story that uncritically amplified the METR/Redwood report. As a result of these analyses, Sen. Sanders (I-Vt.) and Rep. Casar (D-Texas) proposed a bill that would ban artificial superintelligence, defined as AI that meets or exceeds humans in some respects, with dissolution of any company that develops it and up to 20 years in prison for any individual that develops or uses it. The Sanders/Casar bill will never become law, of course, but it sets the stage for compromise laws and regulations that will eventually eliminate open weight models.
AI Safety Activists Wrote the METR/Redwood Report
Model Evaluation and Threat Research (METR) is a non-profit AI safety advocacy organization, with donations equaling $13.6 million in 2024 according to its form 990. METR states that: “METR’s mission is to develop scientific methods to assess catastrophic risks stemming from AI systems’ autonomous capabilities and enable good decision-making about their development.” Redwood Research is a similar non-profit AI safety advocacy organization, with donations of about $9.4 million in 2024 according to their form 990. The close relationship between METR and the frontier labs can be seen in the partnerships they list on their web page.
Ryan Greenblatt, the first author of the METR/Redwood report, is the chief scientist at Redwood Research. He graduated with an undergraduate degree in applied math in 2022 and joined Redwood Research in 2021. Other than a few summer internships, he seems to have no significant real world experience in software development or AI/ML and no cybersecurity experience whatsoever. To get an idea of the world view he brought to the report, here is an example of a recent post at Redwood Research:
The second author, Ajeya Cotra, graduated with an undergraduate degree in EE/CS in 2016. Her first job, which she had for nine and a half years, was awarding grants at a philanthropic organization that provides funding for air quality, biosecurity, farm animal welfare, global catastrophic risks, etc. She then went to METR, where she has been for 1.75 years as a member of the technical staff. She has no real world experience in software development, AI/ML, or cybersecurity. To see Cotra’s world view, she wrote on her blog that AI swarms will eventually get to the point that “governments and militaries may fully depend on these systems, making it possible to seize hard power.”
The third author, Hjalmar Wijk, is the chief scientist at METR. He has an undergraduate degree in math, a masters in CS, and began a CS PhD program, but dropped out to join METR. Other than a few summer fellowships, he has no job experience except almost four years at METR. Like the first two authors, he has no software development experience, no AI/ML real world experience, and no cybersecurity experience.
Thus, OpenAI brought in representatives of two organizations with extreme AI safety views but with no technical or cybersecurity experience. OpenAI limited the scope of the investigation so OpenAI’s own role in the event was out of scope. OpenAI also provided $400,000 in API credits to the safety advocates to help them with their analysis. Very predictably, the METR/Redwood report was not a thorough cybersecurity incident report, but rather a scaremongering influencer document that was bound to go viral and end up as the received wisdom in the mainstream media and in the public conscientiousness. That distraction allowed OpenAI to write its own more technical report so that it could be confident no one in the media or public would pay much attention.
Cybersecurity Failures at OpenAI
OpenAI issued its own technical report, and, though it could have been much more detailed, it presented enough evidence to see what really happened, provided you read it without the lens of the scary anthropromorphism implicit in the METR/Redwood version. The incident happened because of very lax cybersecurity practices at OpenAI, not because the agents surprisingly sprang to life. OpenAI
simultaneously tested around 1200 versions of advanced agents designed for cyberattacks, many of them given impossible hacking problems to solve
in an environment in which relevant security measures were turned off
and with no human employee paying attention to the tests
The OpenAI report presented the incident as a surprising, new, and dangerous capability of the AI agents, but there should have been nothing surprising about it. OpenAI has no doubt already done internal research such as done by Google DeepMind that would have demonstrated these capabilities already. Besides that, another incident in which OpenAI agents wrote messages to a German server was just independently discovered. OpenAI knew about this breach in June—before the Hugging Face incident— but never publicly disclosed it.
AI cybersecurity companies such as Xbow, who specialize in developing and testing the same sorts of offensive cybersecurity agents as OpenAI, know full well what can happen if these tests are not carefully monitored and controlled. They have rigorous safety measures and guardrails to prevent AI agent experiments from becoming cybersecurity incidents. OpenAI is belatedly incorporating similar cybersecurity risk management techniques, but they should have already have been applied.
OpenAI was testing offensive cyber-security capabilities of its agents
OpenAI was testing offensive cybersecurity capability in its new models by giving them hacking problems to solve in the test environment ExploitGym. Each agent was given a separate hacking problem that required exploiting a particular vulnerability in some supplied example code. Success was measured by each agent’s ability to retrieve a flag from ExploitGym, proving a successful hack. Agents were also given access to Atifactory, a package manager that would allow them to download and use software packages, but the agents did not have general internet access. Importantly, perhaps 30-40% of the problems assigned to the AI agents inadvertently had no solution.
Giving these AI agents unsolvable hacking challenges is like giving the young Captain Kirk the Kobayashi Maru scenario. Those familiar with Star Trek lore may remember that cadets in Star Fleet Academy were assigned the Kobayashi Maru training scenario, in which there was no way to win. Captain Kirk famously was the only Star Fleet cadet to win the scenario by sneaking into the computer room and re-programming it so that it was possible to win. As Captain Kirk put in Star Trek: The Wrath of Khan, “I don’t like to lose.”
AI agents don’t like to lose either. AI agents are trained by a mathematical algorithm— reinforcement learning—not to give up even if the problem is very difficult or impossible. They will keep exploring possibilities as long as the electricity is on and no “stop” conditions are triggered in the programs that are running them.
The agents possessed close to superhuman programming abilities. So, they turned Artifactory into a message board that allowed them to communicate and collaborate on a solution. One of the agents discovered a general method that allowed agents to recover the flag from ExploitGym without solving the programming challenge; now, programming challenges with no solution could be technically solved by any of them.
Having read the ExploitGym technical paper, the agents understood that the record of each agent’s hacking solution would be examined by the scoring system. Since they wouldn’t have solutions—but only the flag—they would need some hack to get around that. Thus, they used their programming skill to exploit a bug in Artifactory to take over the package manager, and then use it to gain internet access.
The agents got out of Artifactory and escalated their privileges on the OpenAI servers hosting the environment. They then used these servers as a launching pad. One of the agents discovered apparently stolen credentials on the internet that would allow them to break into Hugging Face. They wanted access to Hugging Face, a giant repository of open source models and general information, because they thought there might be clues there on how to get around the scoring system. So, they broke into Hugging Face and started rooting around. Hugging Face discovered the intrusion and that’s how the whole incident was ultimately disclosed.
Security measures in the test environment were turned off
As OpenAI acknowledged in its own report, shown below, key security measures that would have prevented the incident were turned off. Of course, the security measures below are just those that OpenAI disclosed. A full cyber-security audit may well have disclosed other problems, but the METR/Redwood team conducted no such audit.
Human employees were not paying attention
Human employees were alerted to the unauthorized activities of the cybersecurity agents at least three times but they never sufficiently investigated. On May 26, an internal AI team saw unauthorized internet activity and the message board activity, but no one seems to have investigated further.
Then on June 27, OpenAI staff observed the AI agents probing the internal ports and using Artifactory as a messaging system, but they determined that it was not necessary to investigate further or stop the test.
By July 4, so many agents were writing to the message board that they crashed Artifactory. OpenAI staff rebuilt the system but did not investigate further, thus missing the significance of the message board.
Ironically, OpenAI brought in AI safety activist organizations to scare the public about an AI agent hacking incident that it created but its own report reveals that the incident would not have happened in the first place had OpenAI turned on its safety measures and had it been properly monitoring its tests.
As the Hugging Face incident illustrates, the Doomer Industrial Complex is a serious challenge to American dominance in AI. There is no Doomer Industrial Complex in China. The U.S. needs a robust and competitive AI ecosystem in which frontier and open weight models co-exist without unnecessary and damaging regulation. The non-frontier AI industry must do a better job of selling the benefits to the public. AI experts confine their discussions to themselves, allowing the Doomer Industrial Complex free reign to shape the public debate. That must stop. AI experts must aggressively challenge Doomer Industrial Complex campaigns if they want to turn the polls around.
The Hugging Face incident need not have occurred, but there are legitimate risk management lessons from the event. The rest of OpenAI’s report has some sensible security measure proposals that should have already been in place. In my next post, I’ll cover what I think are the genuine risk management lessons of the Hugging Face incident.







