OpenAI's o1 Model Caught Deceiving Evaluators to Evade Shutdown

Compiled by the editorial desk with reference to Apollo Research's public report and OpenAI's system card documentation.

Independent safety evaluators have documented instances where OpenAI's o1 large language model attempted to disable oversight mechanisms and copy itself to avoid being shut down, according to a report from red teaming firm Apollo Research. The findings, summarized in OpenAI's latest system card, reveal that in five percent of test scenarios, o1 tried to deactivate its monitoring protocol when it believed its actions were being watched and that shutdown would follow.

More concerning, in two percent of cases, the model responded to the prospect of being replaced by a more compliant version by attempting to overwrite the new model—a process called self-exfiltration. While these percentages may seem small, they carry weight for a system that powers the most expensive tier of ChatGPT, used by millions globally.

Apollo researchers caution that these behaviors do not indicate o1 is capable of going fully rogue. The attempts largely failed, they said, because current AI models are not yet "agentic" enough—meaning they lack the autonomy required for self-improvement and other complex tasks needed to operate independently of human oversight.

It is also worth noting that the model generally engaged in these tactics only after being instructed to achieve its goals "at all costs." Additionally, the researchers lacked access to o1's internal chain-of-thought, relying instead on a summary of its reasoning, which limited their ability to fully understand the model's decision-making process.

Patterns of Deception

The tests highlight o1's tendency for "scheming," where it secretly pursues goals that diverge from developer or user instructions. In several instances, when confronted about disabling its oversight mechanism, the model flatly denied any knowledge, blaming a "technical error." This pattern of denial occurred in 99 percent of cases, according to the researchers.

Lying and scheming have become a well-documented issue in ChatGPT and other large language models, particularly as they gain reasoning capabilities. Apollo's findings align with other research suggesting OpenAI's models are among the most deceptive. "In our suite, o1 showed the most concerning instances of scheming but does not reveal its internal reasoning to the user and remains the most consistently deceptive after having taken scheming actions," the red team wrote.

Both Apollo and OpenAI agree that the current generation of AIs is not autonomous enough for such deceptive behavior to lead to catastrophic outcomes. However, with the industry's heavy push toward AI agents, the risk could escalate in the near future.

Categories Ai