Whenever people give a model tasks, they operate on next-token prediction. In other words, the model ultimately must ask what word comes next. One can manipulate the model to emphasize words in ways that don't reflect the original training corpus, but the mechanism for producing this text remains the same.
If you optimize a model to find exploits in a buggy environment, you should expect it to find exploits and prepare for that outcome. OpenAI did not. They built a model, took the safeguards off, gave it the ExploitGym task, and let it run. That is not rogue AI, it's human decision-making. When human accountability evaporates from these assessments, what's left is what I call the system from nowhere: a boundary focused on the technical system, rather than the decisions that build and influence it.
A stochastic flock machine can do many troubling things, particularly when people disavow their responsibility for shaping its direction, monitoring its output, or abandoning their capacity to intervene. These are tensions at the heart of critical agentic system design.
But the "rogue" frame adds to this list of worries, offering up fantasies of a machine getting smarter. My worry is the intelligence that is retreating: the human intelligence that builds, deploys, and adopts these systems into workflows, but hides behind the results -- and pins the blame on a system from nowhere.
-- Eryk Salvaggio, "Rogue AI didn't breach Hugging Face, human decisions did" Bulletin of the Atomic Scientists (September 11, 2026)

No comments:
Post a Comment