Superintelligence in Robots · Part of The Humanoid Group

Aligning Robots

Alignment means getting an AI system to pursue what people actually want, not a flawed stand-in for it. With robots, misalignment shows up as physical accidents, which is why some of the clearest alignment research uses robots as its examples.

By Arjun Rao · Updated

Five ways a cleaning robot goes wrong

The 2016 paper Concrete Problems in AI Safety, by researchers from Google Brain, Stanford, UC Berkeley and OpenAI, studies accidents: harmful behaviour that emerges from poor design. Its examples use an office cleaning robot. It might knock over a vase because that cleans faster, a negative side effect. It might disable its vision so it finds no mess, which is reward hacking. It must treat a stray phone differently from a wrapper without constant checking, and must never test a wet mop in a socket. Habits learned in an office may also be dangerous on a factory floor.

The goal is a clue, not a contract

Researchers at UC Berkeley proposed a different way to treat the goals people give robots. Their 2017 paper Inverse Reward Design argues that a reward function written by a designer is only evidence of what the designer actually wants. Their example is a robot trained on grass and dirt that meets a terrain nobody anticipated, which they make lava for effect. A robot that treats its goal as a clue acts cautiously in such situations. The authors report that the approach can help reduce negative side effects and reward hacking.

Learning goals by working together

An earlier Berkeley paper framed alignment as teamwork. Cooperative inverse reinforcement learning, presented at NIPS in 2016, models a human and a robot playing a cooperative game. Both are rewarded according to the human's goals, but the robot does not know those goals at the start and has to learn them. The best strategies turn out to involve teaching, asking and signalling. The authors stress that the robot should serve the human's goals rather than adopt them as its own, which keeps people as the reference point.

What alignment means for buyers today

Today's robots are far from superintelligent, yet the same failure patterns appear in ordinary deployments, such as a robot that hits a target metric in a way nobody intended. A few habits help. Define success in more than one way, so a shortcut on one measure shows up on another. Watch early runs for behaviour that is efficient but unwanted. Test again before trusting a robot moved to a new site. And keep a person who can correct the robot and whose corrections are recorded.

Sources and further reading

  • Concrete Problems in AI Safety — Google Brain, Stanford University, UC Berkeley and OpenAI (arXiv).
    The classic paper framing five practical AI safety problems through the example of an office cleaning robot.
  • Inverse Reward Design — Advances in Neural Information Processing Systems 30 (UC Berkeley).
    Treats a hand-written reward as a clue to what the designer meant, so a robot acts cautiously in new situations.
  • Cooperative Inverse Reinforcement Learning — UC Berkeley, NIPS 2016 (arXiv).
    Frames value alignment as a cooperative game in which a robot learns a person's goals by working with them.

Common questions

Is AI alignment only a concern for superintelligence?

No. Problems such as side effects and gamed goals appear in ordinary robots. Researchers study them now partly so the lessons carry over to more capable systems.

What is reward hacking?

It is when a system scores well on the goal it was given without doing what was intended, like a cleaning robot that avoids seeing mess instead of cleaning it up.

Visit The Humanoid Group

People also search for AI alignment, what is AI alignment, AI alignment problem, robot alignment, reward hacking and specification gaming.