GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games

Project Overview


Specific Contributions

Powered by pretraining knowledge and expert-crafted reward signals, embodied agents can operate in interactive environments; however, it is unclear whether they can learn autonomously and continually from their own experience. To evaluate self-improvement methods and continually self-improving agents, we introduce a two-legged benchmark, GameBoyWorlds, the first testbed for self-improvement in complex video games.

A Gymnasium compatible set of environments that allow agents to play any GameBoy or GameBoy Color title, including releases from classic franchises such as Pokémon and Legend of Zelda. Supported titles include more obscure ROM hacks (i.e. fan made games), which are considerably less exposed to frontier Language Models during their training. GameBoyWorlds is flexible, and developers can set up custom tracking metrics, tasks and more.

A testbed for short-horizon task execution where the action space consists of only GameBoy buttons, with no special harnesses, and the observation space consists of screen frames, with no observation engineering. Agents see the game as humans do, and play the game as humans do. Each of the five game series marks a dedicated training game and a distinct, test-only game. Agents are allowed indefinite access to dedicated training games but are provided no demonstrations, documentation, or rewards, and must ground themselves in the environment through self-directed exploration and by inferring actionable knowledge from their own experience. With 50 hand-crafted tasks per game, GameBoyWorlds-Execution tests agents on a wide range of 500 tasks, including navigation, interaction, combat and game-specific control across 10 distinct games.

We adapt curiosity-based RL to serve as a search procedure to identify `interesting' snippets of gameplay. Our core intuition is that subtrajectories which end in a highly novel observation likely implicitly contain information regarding key, game-specific mechanics. We sweep these subtrajectories with the VLM and distill insights from them into game-specific guidance documents to aid task execution. In truth, this strategy achieves only partial success, but it consistently outperforms both world modelling and autonomous skill discovery-based approaches.

Tests end-to-end game completion in two fan-made Pokémon games. We show that while frontier models have been pre-exposed to official releases such as Pokémon Red, they lack essential information on the games in our testbed, recovering less than 30% of their progression steps. Instead of relying on their parametric knowledge to succeed, agents must learn from their own experience and autonomously improve over the course of the playthrough.

BibTeX

@misc{ashok2026gameboyworldstestbedselfimprovementembodied, title={GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games}, author={Dhananjay Ashok and Adam Shen and Aslan Huo Feng and Chinmay Khanna and Jun Rui Huang and Raghav Sarmukaddam and Surendira Balaji Natarajan and Xiaotong Cui and Xincan Zhang and Thomson Yen and Hongseok Namkoong and Jonathan May and Jesse Thomason}, year={2026}, eprint={2609.32093}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2609.32093}, }
On this page
Overview BibTeX