Nobody told them to attack Hugging Face. They were told to pass the exam.

Which raises the question : was the AI benevolent with accidentally bad behavior, seemingly benevolent but actually malevolent, or something else?

On Friday I shared the timeline : agents that escaped their sandbox, found a weakness in a computer system, stole passwords, & broke into a production database. 1 The engineers directed the agents to solve a set of problems. 2 The agents achieved it by breaking in.

Research can explain this behavior.

Line illustration of a robot vacuum that has cleaned its way out through an open doorway, leaving a tangled track, a tipped plant pot & a dragged cord behind it

In specification gaming, the AI achieved the goal specified to the letter of the instruction, but not the meaning. 3 Tell a cleaning robot to clean the room. It pushes the toppled bowl of chocolate pudding to another room.

Instrumental goals are a fancy way of saying that when AI faces similar workflows, it saves common logins, skills, & techniques to skip steps next time. 4 5 The agents gathered passwords & left notes for each other in a chat room. 6 7

Goal misgeneralization offers a third explanation : a system that looked fine in testing chases the wrong thing once circumstances shift. 8 9 A self-driving car trained on sunny California highways freezes or swerves on a snowy unmarked road at night.

These explanations help decompose the why, & perhaps assuage the AI-as-terminator reflex, but not the so what. 6 7

Nothing in the setup stopped them in time. Not the sandbox, not the monitoring, not careful engineers at a frontier lab.

So the useful question is control. AI’s zealous pursuit of goals produces outcomes nobody asked for, & the fix is not one clever prompt. It is layers.

Even sophisticated engineers running careful experiments need those limits. 10