Superintelligence in Robots · Part of The Humanoid Group
Levels of AGI
Arguing about whether a system is AGI goes nowhere without a scale. In 2023 Google DeepMind researchers proposed one, grading AI by performance and breadth, with separate levels for autonomy. Here is how it works and how it applies to robots.
By Arjun Rao · Updated
Performance and generality
The Levels of AGI framework, presented at the ICML conference in 2024, rates systems on two axes: how well they perform and how many tasks they cover. Level 0 is no AI. Level 1, Emerging, means equal to or somewhat better than an unskilled human. Level 2, Competent, means at least the 50th percentile of skilled adults, and Level 3, Expert, at least the 90th. Level 4, first called Virtuoso and renamed Exceptional in 2025, means at least the 99th percentile. Level 5, Superhuman, outperforms every human. A general system at Level 5 is what the authors call artificial superintelligence.
Where today's systems sit
Each level is rated separately for narrow and for general systems, which explains an apparent contradiction. Superhuman narrow AI already exists: the paper lists AlphaFold, AlphaZero and the chess engine Stockfish. Chatbots such as ChatGPT, Gemini and Llama 2 were placed at Emerging AGI, the first general level, and every higher general level was listed as not yet achieved. So a system can beat every human at one game and still rank low on general ability, and the same is true of a robot that excels at a single task.
Autonomy is a separate dial
The framework also sets out six levels of autonomy that describe how people work with an AI system: no AI, AI as a tool, as a consultant, as a collaborator, as an expert, and finally as an agent acting on its own, which the authors listed as not yet reached. Keeping autonomy apart from capability is useful for robots. How much independence to give a machine on a given task is a decision for the people deploying it, and a highly capable system can still be run as a tool under close supervision.
Bodies optional, by design
One of the framework's principles is that physical tasks add to a system's generality but are not a prerequisite for AGI. Other researchers measure intelligence differently again. François Chollet argued in 2019 that skill at particular tasks falls short of measuring intelligence, which he defined as how efficiently a system acquires new skills. His ARC benchmark targets puzzles that are easy for people and hard for AI. Either way, a system could be called AGI and still fumble a towel, so judge robots on physical tasks.
Sources and further reading
- Levels of AGI for Operationalizing Progress on the Path to AGI — Google DeepMind, ICML 2024 (arXiv).
The framework rating AI by performance and generality, with separate autonomy levels and a table of example systems. - On the Measure of Intelligence — François Chollet (arXiv).
Argues that intelligence is efficient skill-learning rather than skill itself, and introduces the ARC benchmark. - ARC-AGI — ARC Prize Foundation (non-profit).
Explains the ARC-AGI benchmark and its aim of tasks that are easy for people but hard for AI systems.
Common questions
Does a system need a body to count as AGI?
Not in the Levels of AGI framework. Its authors say physical abilities add to generality but are not a prerequisite, so an AGI claim says little about a robot's hands or legs.
What is the ARC benchmark?
ARC, the Abstraction and Reasoning Corpus, is a set of puzzles introduced by François Chollet in 2019. It aims at tasks that are easy for people but hard for AI, to measure how efficiently a system learns.
People also search for AGI levels explained, DeepMind levels of AGI, emerging AGI, superhuman AI, how to measure AGI and AGI definition.