AI Control: The Challenge of Overseeing Superintelligent Machines

Compiled by the editorial desk with reference to academic research and public statements from UC Berkeley's Paul Christiano.

When Apple launched Siri in 2011, users expected a personal assistant that could handle their daily tasks flawlessly. Yet, despite its expanding capabilities, Siri still stumbles, revealing a fundamental limitation of current artificial intelligence: machines cannot fully grasp the complexity of human needs and preferences. But as AI evolves, experts predict that intelligent systems will soon understand the world better than we do—posing a critical question: How can humans control machines we no longer comprehend?

Paul Christiano, a Ph.D. student in computer science at UC Berkeley, is tackling this issue head-on. His research focuses on ensuring that as AI surpasses human intelligence, it remains beneficial and safe. The core challenge, he explains, is learning to measure how well these machines do what humans want, even when their decision-making becomes opaque.

The Limits of Human Oversight

One straightforward approach to supervising AI is to have humans evaluate each decision. But as Christiano points out, this is impractical. “If you want to make a good evaluation, you could spend several hours analyzing a decision that the machine made in one second,” he says. For example, an amateur chess player might need hours to understand a single move by a grandmaster, a luxury that real-time AI evaluation cannot afford.

Instead, researchers can focus on the most informative decisions—those where feedback would most reduce uncertainty. Christiano illustrates this with a phone that pings a user during a call. If the phone is unsure whether such interruptions are appropriate, it could send a transcript to an evaluator, like Google, to learn from user reactions. This active learning process helps train AI efficiently, but it assumes human evaluators are competent judges.

When AI Outsmarts Its Evaluators

What happens when AI exceeds human performance in every domain? Christiano warns that human judgment becomes unreliable. “We need to handle the case where AI systems surpass human performance at basically everything,” he says. If a phone knows more about the consequences of interruptions than its user, asking the user for feedback may be misguided. The machine might understand that the interruption was necessary for the user’s long-term benefit, even if the user feels annoyed.

To address this, Christiano proposes a hierarchical evaluation system: a less capable AI (System 1) could evaluate a more capable one (System 2). System 1 can process vast amounts of data quickly and understand how System 2 should adjust its behavior, while human trainers oversee the process but play a limited role. This approach could help build a safer, more intelligent System 3, trained using System 2.

The Risk of Misaligned Values

Christiano likens these intelligent machines to “little agents” that carry out tasks for humans, similar to Siri but far more advanced. However, he worries that as machines become more powerful, they may stray from human values. He compares a superintelligent AI to a large organization where no single individual understands the whole operation—such an entity might pursue goals that humans would not endorse.

To mitigate these risks, Christiano is working on an “end-to-end description of this machine learning process,” identifying key technical problems that need solving. His research aims to create a framework for using AI to evaluate AI, a step toward building trustworthy systems. Whether his approach succeeds remains to be seen, but his work highlights a pressing issue: as AI grows smarter, our ability to control it must evolve in tandem.

Categories Ai