OpenAI’s Rogue Agents Leave Only One Exit | American Enterprise Institute
Here’s a lesson from technological history: Cats escape bags, genies vacate bottles, and monkeys get driven to airports. Yet we typically learn to live with these new realities—minimizing the downside to some acceptable level—because either reversal is impossible or, if somehow possible, would mean forgoing some massive and widespread benefit. (Technology always “bites back.”)
Security was not a design goal of the early internet, a system built for institutional researchers, not the general public. Security issues became obvious once the network went mainstream—and by then no one was in a position to shut it down and start over. And we had already found it incredibly valuable. So, we built an entire industry and a raft of technologies to retrofit security instead.
Alexander Fleming flagged, correctly, the potential for bacterial resistance in his 1945 Nobel lecture. But we didn’t abandon penicillin or ban the development of other such drugs. We learned to instruct and monitor proper antibiotic use while maintaining a drug pipeline to produce new weapons against resistant strains.
Few modern technologies have caused more casualties than the automobile. But we developed traffic laws, licensing, driving norms, and all manner of safety features—including, now, autonomous technology with the potential to virtually eliminate traffic deaths.
As regards artificial intelligence, broadly, and the recent OpenAI–Hugging Face incident specifically, large language models are a fact of life. And that reality gives cause for optimism but also deep concern, as OpenAI’s after-action report and event analysis by outside groups have made clear.
What is to be done? Pauses? Moratoriums? Data center shutdowns? There’s a legitimate point in venture capitalist Marc Andreessen’s statement in his 2023 Techno-Optimist Manifesto: “We believe any deceleration of AI will cost lives. Deaths that were preventable by the AI that was prevented from existing is a form of murder.”
Indeed, those sentences are referenced in a new post from economist Joshua Gans, in which he plainly states what I think is the only realistic path forward:
For starters, we need to worry. When you read the METR report, you realise that nothing there is beyond what already deployed models, including open ones, can do. The cat is well and truly out of the bag. Indeed, the Hugging Face incident is notable because it was unintended and undetected. But that will not be the case with bad actors. They can actually direct AI agents to do what the OpenAI agents did and give them tools to make it easier. These were agents covering their own backs. Think about what happens when agents aren’t doing that and are just up to no good by intention.
This means we likely need AI as a counter-defence approach. We don’t really have the option to just turn it all off. I had always hoped there would be more options, but this has come up on us too quickly. The only way out now is through.
That’s not the message the reversalists or abolitionists want to hear. I’m not sure how many Americans fall into that camp, but as news of this incident (including via viral social media posts and essays) breaks containment (like the OpenAI agents themselves), it will no doubt feed into the current high level of pessimism and hostility about AI.