We Are Online Since 1998

Gotta Catch ‘Em All, Slowly: What AI Playing Pokémon Teaches Us About Machine Intelligence

admin
By admin
7 Min Read

Pokémon Red and Blue first came out on the Game Boy in Japan in 1996. Since February 2025 they have served as an informal public benchmark for large language models (LLMs). Models from Anthropic, Google and OpenAI have played the same 8-bit adventure on livestreams, and their logged steps, hours and failures give measurable data on what current AI systems can and cannot do.

From Crowd Play to Model Play

  • February 2014: the “Twitch Plays Pokémon” stream let viewers control Pokémon Red through chat commands. More than 1 million participants finished the game in about 16 days.
  • February 2025: an Anthropic researcher launched “Claude Plays Pokémon” on Twitch with Claude 3.7 Sonnet. The stream drew about 2,000 concurrent viewers in its early days.
  • May 2025: the independent “Gemini Plays Pokémon” project, run by developer Joel Zhang, completed Pokémon Blue with Gemini 2.5 Pro after several hundred hours of play.
  • August 2025: GPT-5 completed Pokémon Red in 6,470 steps on the GPT_Plays_Pokemon Twitch channel. OpenAI’s o3 had needed 18,184 steps.
  • May 2026: Claude Opus 4.7 beat Pokémon Red, about one year after the first Claude stream.

The Numbers Against a Human Baseline

  • Children typically finish Pokémon Red in 20–40 hours.
  • By January 2026, Claude Opus 4.5 had played for more than 500 hours and taken 170,000 steps without finishing.
  • Opus 4.5 later reached Victory Road at roughly 230,000 steps with all eight gym badges.
  • GPT-5’s 6,470-step run took about 7 days of real time, and o3’s run took more than 15 days.
  • Gemini 3 Pro completed the sequel, Pokémon Crystal, without losing a single battle.

Where the Models Got Stuck

The logs from these runs point to specific, repeated failure points:

  • Visual recognition: Opus 4.5 was stuck in Silph Co. for weeks because it did not recognize the Card Key lying on the floor. Opus 4.6 recognized the Card Key but could not see the switches in the Victory Road boulder puzzles.
  • Spatial navigation: Opus 4.5 spent 4 days circling a gym without finding the entrance.
  • Multistep dependencies: Opus 4.5 skipped the Gold Teeth in the Safari Zone. That item is required to get HM Strength, and the model went back for it only after clearing Pokémon Mansion.
  • Efficiency between versions: Cinnabar Mansion took 112,000 steps with Opus 4.5 and 3,000 steps with Opus 4.6.
  • Battle strategy: GPT-5 leveled up one Pokémon almost exclusively, which reviewers described as a brute-force approach rather than balanced team building.
  • Endgame resources: Opus 4.7 needed two attempts at the Elite Four and managed healing items better on the second try.

The Role of the Harness

Every model plays through a “harness,” a layer of software that feeds it screen data and turns its decisions into button presses. Harness design directly affects results:

  • The Gemini 2.5 Pro setup used the mGBA emulator, a grid overlay on the screen, data read from the game’s RAM, a text-based map tracker and a memory summary roughly every 100 actions.
  • That setup also used two specialized sub-agents, “Pathfinder” and “Boulder Puzzle Strategist.”
  • The Gemini developer stepped in on limits for escape-item use and once to work around a known game bug.
  • The Opus 4.7 harness added tools for saving screenshots and zooming into parts of the screen.
  • Because the harnesses differ, comparisons across models are not controlled experiments.

Independent developers often host the logs, dashboards and harness code for such projects on tech-focused domains. An io domain registration is a frequent choice: .io is the country-code domain of the British Indian Ocean Territory and is widely used by developer tools and AI startups.

What the Experiments Reveal

  • Knowledge is not the same as execution: the models can state type matchups, gym order and item locations, yet still fail to walk through a doorway they cannot see clearly.
  • Perception is the main bottleneck: the Opus 4.7 release notes cite “substantially better vision,” including spotting switches and trees that can be cut down, as the change that got the model past Victory Road.
  • Progress came in increments: each Claude version from 3.7 Sonnet to Opus 4.7 cleared obstacles that had stopped its predecessor, without a single breakthrough release.
  • Persistence partly makes up for weak planning: long runtimes let models recover from errors that would have permanently blocked earlier versions.
  • Hybrid systems are emerging: Jev, a decision model that is not an LLM, reportedly beat Pokémon Red in under a week, with Claude Opus 5 coaching it past dead ends.

The Cost of Slow Play

Hundreds of hours of continuous inference on frontier models use substantial computing power and electricity. Energy use is one of the main issues raised in discussions of the environmental trade-offs of generative AI for sustainability-focused businesses. Step counts are therefore a practical metric as well as a speed measure. GPT-5 needed about one-third of o3’s steps, so it also needed roughly one-third of the model calls.

Key Takeaways

  • Pokémon Red tests perception, navigation, memory and long-horizon planning in one environment.
  • A child’s 20–40 hours compares with hundreds of hours and 100,000+ steps for most LLM runs.
  • Harness design, vision quality and memory tools explain most of the gap between models.
  • Each new model generation has cut the steps needed for specific sections of the game, sometimes by more than 90%.
Share This Article
Leave a comment
Need Help?
How can I help you?