Superintelligence in Robots · Part of The Humanoid Group
Off-Switches and Oversight
Being able to stop a machine is the most basic safety measure there is. For advanced AI, researchers have asked a harder question: could a capable system learn to avoid being switched off? This guide covers that research and the engineering basics.
By Arjun Rao · Updated
Why an AI might avoid its off switch
The Off-Switch Game, a 2017 paper by Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell, starts from a simple point. Turning a misbehaving system off is one of the main safety tools available. Yet a system pursuing a goal cannot achieve it once switched off, so self-preservation can arise as a side effect rather than an instinct. The authors model a human who can press a robot's off switch and a robot that can disable it, and show that a robot certain of its goal has an incentive to disable it, unless the human is perfectly rational.
Uncertainty as a safety feature
The same paper finds the way out. A robot that is uncertain about what its human wants, and treats the human's choices as information about it, has a reason to leave the switch alone, because being stopped tells it something useful. The authors conclude that giving machines an appropriate level of uncertainty about their objectives leads to safer designs. The idea runs against the instinct to give machines precise, fixed goals, and it shapes much current research on keeping capable systems under human control.
Corrigibility, still an open problem
In a paper for a 2015 AAAI workshop, researchers from the Machine Intelligence Research Institute and the University of Oxford named the property people want: corrigibility. A corrigible system cooperates when its makers try to correct, modify or shut it down, despite the default incentives a goal-seeking agent has to resist. They studied a shutdown button and set out what a safe design must do: give no reason to block the button or to press it, and carry the shutdown behaviour into anything the system builds. No proposal met every requirement, leaving the problem wide open.
The emergency stop you can specify today
Machines in use today rely on engineering rather than theory. ISO 13850:2015 sets out functional requirements and design principles for the emergency stop function on machinery, whatever energy it uses, and applies to nearly all machines. The electrical side is covered by a companion standard, IEC 60204-1. For any robot, ask how its emergency stop is built, whether it acts independently of the robot's AI, how fast it brings the robot to rest, and how often it is tested. A capable robot should still stop when a person decides it must.
Sources and further reading
- The Off-Switch Game — UC Berkeley, IJCAI 2017 (arXiv).
Models when a robot would let a person switch it off, and shows why uncertainty about its goals makes it safer. - Corrigibility — Machine Intelligence Research Institute (AAAI-15 workshop paper).
The paper that named corrigibility: designing AI that accepts correction or shutdown. It remains an open problem. - ISO 13850:2015 Safety of machinery — Emergency stop function — Principles for design — ISO, via the Standards Council of Canada.
The international standard for emergency stops on machinery, the engineering form of an off switch.
Common questions
Should an emergency stop depend on the robot's AI?
It should not have to. Good practice is to make the emergency stop a dedicated safety function that people can rely on whatever the AI is doing, and to test it regularly.
What does corrigible mean?
A corrigible AI cooperates when its makers try to correct, modify or shut it down, rather than resisting. Researchers have proposed designs, but none yet meets every requirement.
People also search for AI off switch, AI kill switch, can AI be turned off, off-switch game, AI corrigibility and corrigible AI.