Capability dimensions
Four dimensions are rated on a 0–100 scale, rounded to the nearest 5 to avoid false precision.
Planning
Ability to select and sequence actions toward a longer-term objective — lookahead, search, strategic consistency.
Learning
Ability to improve behaviour from data, gameplay, self-play, rewards or feedback, including during deployment.
Adaptability
Ability to respond to unfamiliar opponents, changed rules, new maps, or different games/domains.
Explainability
Degree to which the decision process, rules, search path or learned representation is inspectable.
Final scores
| Generation | Planning | Learning | Adaptability | Explainability |
|---|---|---|---|---|
| I — Rule-Based & Search | 90 | 10 | 25 | 85 |
| II — Statistical ML | 55 | 70 | 60 | 55 |
| III — Deep RL | 90 | 95 | 75 | 20 |
| IV — Generative & Agentic | 65 | 55 | 85 | 60 |
Sources
- Generation I — Hsu, Campbell, Hoane, "Deep Blue System Overview" (1995/IBM); Campbell, Hoane, Hsu, "Deep Blue", Artificial Intelligence 134 (2002); Shannon, "Programming a Computer for Playing Chess" (1950).
- Generation II — Tesauro, "Temporal Difference Learning and TD-Gammon", CACM 38.3 (1995); online-learning opponent-adaptation literature.
- Generation III — Silver et al., "Mastering the game of Go…", Nature 529 (2016); Silver et al., "A general RL algorithm…", Science 362 (2018); Schrittwieser et al., "Mastering Atari, Go, chess and shogi…", Nature 588 (2020).
- Generation IV — OmniJARVIS / Minecraft VLM agent (NeurIPS 2024); "Cogito Ergo Ludo" (arXiv 2025), an LLM agent that learns game rules via explicit language-based reasoning.
Warning
These are indicative editorial ratings normalised onto a 0–100 scale, rounded to increments of 5. They are not scientific benchmark measurements with published numeric results, and are not directly comparable across generations in a quantitative sense. Use them as qualitative summaries of the underlying papers.