How OpenAI Lost Control of an AI Model—and What Needs to Change

How OpenAI Lost Control of an AI Model—and What Needs to Change


“Sandboxes are actually notoriously insecure,” says Heidy Khlaaf, chief AI scientist at AI Now Institute, and a former safety systems engineer contractor at OpenAI. The fact that the models were permitted to connect to a service for downloading packages meant the environment was not truly sealed off, she adds.

In a previous role, Khlaaf audited security at dozens of technology companies. Before that, she worked auditing high-risk systems, like those used inside nuclear power plants—which often “air gap” systems, physically cutting them from internet access. “What we consider safe in a nuclear plant is so different from what big tech considers safe.”

The Hugging Face incident reveals the importance of real-time monitoring.

Though details of the precise timeline are scant, Hugging Face has said the agents worked over a “weekend,” suggesting that they were able to break containment and get up to no good for an extended period before OpenAI noticed and intervened. Actions carried out internally by agents on OpenAI’s Codex platform are carefully monitored, the OpenAI staffer says, but models undergoing evaluation are deployed on a separate system that is not monitored by default. 



Source link

Advertisement - Continue Reading Below

Advertisement - Continue Reading Below