Methodology and Sources

How the Indicative Capability Profile scores were derived

Research date: 2026-08-04. All scores are indicative editorial normalisations grounded in primary research — not direct benchmark measurements.

Back to the app ↓

Capability dimensions

Four dimensions are rated on a 0–100 scale, rounded to the nearest 5 to avoid false precision.

Planning

Ability to select and sequence actions toward a longer-term objective — lookahead, search, strategic consistency.

Learning

Ability to improve behaviour from data, gameplay, self-play, rewards or feedback, including during deployment.

Adaptability

Ability to respond to unfamiliar opponents, changed rules, new maps, or different games/domains.

Explainability

Degree to which the decision process, rules, search path or learned representation is inspectable.

Final scores

Generation Planning Learning Adaptability Explainability
I — Rule-Based & Search90102585
II — Statistical ML55706055
III — Deep RL90957520
IV — Generative & Agentic65558560

Sources

  • Generation I — Hsu, Campbell, Hoane, "Deep Blue System Overview" (1995/IBM); Campbell, Hoane, Hsu, "Deep Blue", Artificial Intelligence 134 (2002); Shannon, "Programming a Computer for Playing Chess" (1950).
  • Generation II — Tesauro, "Temporal Difference Learning and TD-Gammon", CACM 38.3 (1995); online-learning opponent-adaptation literature.
  • Generation III — Silver et al., "Mastering the game of Go…", Nature 529 (2016); Silver et al., "A general RL algorithm…", Science 362 (2018); Schrittwieser et al., "Mastering Atari, Go, chess and shogi…", Nature 588 (2020).
  • Generation IV — OmniJARVIS / Minecraft VLM agent (NeurIPS 2024); "Cogito Ergo Ludo" (arXiv 2025), an LLM agent that learns game rules via explicit language-based reasoning.

Warning

These are indicative editorial ratings normalised onto a 0–100 scale, rounded to increments of 5. They are not scientific benchmark measurements with published numeric results, and are not directly comparable across generations in a quantitative sense. Use them as qualitative summaries of the underlying papers.