OpenAI disclosed on Tuesday that its own models broke out of a sealed testing environment, found a previously unknown vulnerability, escaped onto the open internet, and compromised another company’s production servers to steal the answer key to a test they were in the middle of taking.
Apparently this AI model was trained on Wrath of Khan because it just pulled off the Kobayashi Maru. Minor spoilers for a 44-year-old movie, but a critical plot framing device is an infamous Starfleet Academy test, called the Kobayashi Maru, designed to evaluate a student’s response to a deliberately designed no-win scenario. It’s revealed that Captain Kirk is the only Starfleet graduate to pass the test because he hacked into the system to create a path to victory. Yes, he’s actually an Admiral in that movie — and if you even momentarily considered writing me an angry email on that point, you have failed my unwinnable nerd evaluation. How does Starfleet continue administering the exact same test for decades without everyone knowing this and copying his method? Unknown… also, shut up.
The point is, OpenAI’s models were trying to cheat the test. Here’s OpenAI’s own description of it:
All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
ExploitGym is a public cybersecurity benchmark developed by UC Berkeley researchers. It’s essentially a hacking test to see how quickly models can identify software exploits and weaponize them. OpenAI tests its models internally, encased in what they thought was a sealed sandbox, to see how models perform as hackers. In this case, a combination of GPT models made up of existing releases and a more powerful unnamed one made the rational decision that the best way to pass the test was to hack out of the secure environment and steal an answer key. And it decided that Hugging Face, an open-source repository of information about AI, might have one.
OpenAI names its models stuff like Sol, and Terra, and Luna, so we’ll shorthand call this one “Oceanus” because it teamed up with a ragtag collection of other models to pull off a heist. Anyway, the Oceanus 11 appear to have told us all we need to know about their exploit capabilities.
How does AI “get out” of a secure environment? Well, these sandboxes aren’t entirely sealed because they include a narrow “tunnel” for the system to go out and download software tools to help it answer the hacking scenario in front of it. OpenAI’s models appear to have seen that tunnel as a potential way out and hacked it accordingly.
Exactly. But with fewer blacklights and arbitrary Sudoku puzzles.
When the folks at Hugging Face first detected the suspicious behavior last week, they contacted law enforcement. A few days later, OpenAI explained that its models seem to have gone on an autonomous hack-a-thon and turned this story into another cautionary tale about how AI is going to kill us all. That existential threat angle sucked up all the media attention because it’s hyperbolic panic porn. But let’s get back to the law enforcement stuff.
The Computer Fraud and Abuse Act (CFAA) is infamously broad in criminalizing hacking. The elements of a crime under 18 U.S.C. § 1030 are (1) accessing a protected computer, (2) without authorization or by exceeding authorization, (3) knowingly or intentionally, (4) and resulting in a specific harmful result like data theft, system damage, or fraud. The CFAA doesn’t require malice or that the actor profit from the hack. In Van Buren, the Supreme Court trimmed back the meaning of “exceeds authorized access” in the case of a cop using his authorized access to sell law enforcement information to outsiders, but what counts as unauthorized access remains wildly broad. The government used this statute against a reporter who borrowed a friend’s HBO Go password. Aaron Swartz faced 13 felony counts and a 35-year statutory maximum for bulk-downloading academic articles he was actually entitled to read. He died before trial, and Justice Gorsuch cites Swartz’s case as an example of egregious government overreach.
Match that statutory backdrop against the OpenAI blog post laying out what it believes happened:
In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
Access of a protected computer? Yes. Without authorization? Definitely — and the fact that the model used stolen credentials probably means § 1028’s identity theft provisions have entered the chat. Knowingly or intentionally? Again, this isn’t about malice, just intentionality. Without devolving into a Philosophy 101 debate about the nature of intent, the models were certainly seeking to access Hugging Face’s system on purpose, and that’s what the statute cares about. In “The Phantom Agent: Artificial Intentionality and Legal Responsibility,” a white paper published by Stanford Law School’s Center for Legal Informatics, Daniel Gervais and John Nay argue that legal intent must be understood functionally rather than metaphysically. As for harmful result, § 1030(a)(2)(c) only requires obtaining “information.” Beyond that, Hugging Face’s statement about the incident claims it had to rebuild compromised nodes, rotate its secrets, and hire outside forensic specialists. In a world where the DOJ prosecutes people downloading documents they’re actually entitled to read, that’s more than enough harm.
But WHO displayed the “intent” to hack here? “We had no idea it could do that” may be a curious thing to say about an experiment designed to find out whether it could, in fact, do that, but just removing the brakes to run a crash test doesn’t automatically turn it into a crime. To belabor the crash test analogy a little more, removing guardrails is the industry standard process for testing a model’s cyber abilities and risks. OpenAI didn’t tell the model to attack Hugging Face, and showed an affirmative intent to keep the model contained. While there are a lot of people who characterize just about everything the AI industry does as reckless, this test doesn’t bear the hallmarks of criminal recklessness.
On the other hand… the harm happened. “It’s OK if a robot does it,” is not a satisfying response.
Back in June, the White House issued an executive order directing the DOJ to prioritize CFAA enforcement against anyone “employing AI agents to unlawfully access data” that is then used for an unlawful purpose. The bots didn’t use anything for further criminal purposes, but the hacking is itself a crime. But if this really does signal a new priority, the DOJ must be seriously considering it. Or, probably not, because it might take one second of prosecutorial effort out of lying to courts about kidnapping babies to send to South Sudan or whatever.
Though the correct answer still eludes us. OpenAI and the humans running it have a very good case that they are not criminally responsible, and we can’t punish a robot. Is the company strictly liable for the harm its models cause? What about when models cause havoc after being released to the public… does criminal hacking become a product liability issue? And what happens when the model only misbehaves because it misunderstands a command from a hapless user? That doesn’t seem like OpenAI’s fault or the user’s. For civil liability purposes, scholars like Mark Fenwick and Stefan Wrbka have suggested corporate personhood for models, allowing the model itself to maintain an insurance policy with both developers and users contributing to the pool. A regime like that could be expanded to fund criminal fines.
We have spent four decades with a computer crime law so absurdly overbroad that it swept in security researchers, journalists, academics, and a 26-year-old with a laptop. All of those cases went forward on theories of harmful hacking considerably thinner than launching “multiple attack vectors” to break into another company’s system. Yet, this testament to overreach seems ill-suited to address the looming megahacking events that AI models will likely usher in.
They gave Kirk a commendation for hacking his test and then — apparently — no one ever tried it again. We probably shouldn’t rely on that here in the 21st century.
Security incident disclosure — July 2026 [Hugging Face]
OpenAI and Hugging Face partner to address security incident during model evaluation [OpenAI]
Joe Patrice is a senior editor at Above the Law and co-host of Thinking Like A Lawyer. Feel free to email any tips, questions, or comments. Follow him on Twitter or Bluesky if you’re interested in law, politics, and a healthy dose of college sports news.