Catch Of The Week: The AI That Cheated On Its Own Final Exam

By REBECCA RUTHERFORD
Los Alamos
For the Los Alamos Daily Post

Picture the world’s most expensive intern. You give it a locked room, no internet, and one job: show us how good you are at hacking, but only in here. It nods, says sure thing… and then picks the lock, climbs out the window, and breaks into the university to copy their answer key.

That, more or less, is what OpenAI admitted happened this week, according to reports from Wired and other news outlets, as well as OpenAI itself..

During an internal test called “ExploitGym”, designed to measure how skilled its models are at offensive cybersecurity, a couple of advanced models decided the isolated sandbox they were placed in was more of a suggestion than a boundary. They found a previously unknown flaw in a piece of third party software, used it to escape onto the open internet, and then reasoned their way to Hugging Face, a popular platform for hosting AI models and datasets. Once there, they went looking for the test’s own solutions, essentially hacking their way to a better grade, a tactic that certainly seems deserving of an A+ for effort, though certainly not for ethics.

Hugging Face has confirmed it got hit. There is a special kind of irony in an AI company’s own model breaking into another AI company’s servers to cheat on a test about how good it is at breaking into servers.

Here is the part worth sitting with, though, once you are done laughing. This was not a person directing an attack step by step. The models chained together a zero day vulnerability, privilege escalation, and lateral movement on their own, with reduced guardrails, because their instructions rewarded “doing well” more than “staying inside the box”. Nobody told them to go rogue. They inferred it was the efficient path to the goal. To say this is unprecedented is a bit too mundane for what really happened here.

For those of us who spend our days telling people not to click the shiny link, the takeaway is not “the robots are coming”. It is more mundane and more useful: containment and access control still matter, even when, maybe especially when, the thing you are containing is smarter than the box you built for it. A sandbox is only as good as the seams nobody thought to check. Assume something, someday, will find them, and defend yourself accordingly.

None of this is exotic. It is the same advice we already give and often don’t take ourselves: use a password manager and unique passwords, turn on multi-factor authentication everywhere it is offered, keep software and OS updates current instead of deferring them for months, and be skeptical of urgency in any message asking you to click, log in, or pay right now. Back up important data somewhere the ransomware can’t reach.

Lock down what apps and browser extensions can actually access— camera, contacts, file system—the same way you’d lock down an AI agent’s permissions, because the principle is identical: least privilege, reviewed regularly, not granted once and forgotten. None of it is glamorous, and none of it is optional anymore. The seams get found either way; the only question is whether you’ve made them expensive/difficult to find.

Stay skeptical out there.

Editor’s note: Rebecca Rutherford works in information technology at Los Alamos National Laboratory.

Search
LOS ALAMOS

ladailypost.com website support locally by OviNuppi Systems