Execution Track

Can embodied agents self-improve their short-horizon task execution capabilities?

Leaderboard

GameBoyWorlds-Execution is a benchmark for self-improvement, not simply task performance, so our metric of interest is task completion gain, i.e., the increase in task completion after self-learning in the training environments. We first show that there is room for self-improvement to push model performance, even at the frontier.

Task success rates (%) for frontier models on the GameBoyWorlds benchmark. Select a game series below, and hover any bar for the exact model, game title, and success rate.

    Success and Failure

    One task Gemini 3.6 Flash completed and one it did not, drawn at random from each game's benchmark run. Select a game below.

       Success

       Failure

      Self-Improvement

      Task completion gain (percentage points) over the Gemma 4 31B base agent, per game. Everything above the line helped; everything below it hurt. Hover a marker for the underlying rates, and click a method in the legend to show or hide it.

        No approach consistently improves performance across all game environments.

        World Modelling −3.4% average task completion

        Qualitatively, we observe that the world model is unable to predict the next observation with sufficient accuracy to serve as a useful tool.

        Autonomous Skill Discovery −44.3%  (40.7% → 22.6%)

        The skill discovery approach performs even worse, with catastrophic collapses on PokémonCrystal (36.0% → 4.0%), PokémonRed (38.8% → 8.2%) and ZeldaOracleOfSeasons (40.0% → 14.0%). The performance degradation is always greater for test-only games, suggesting an overfitting to the training distribution. We hypothesize the extent of degradation is caused by an inability to accurately judge task completion during practice, a known failure scenario for skill acquisition methods.

        Curiosity Guides +2.0%  (40.7% → 41.5%)

        The only method that achieves partial success is our novel utilization of curiosity-based exploration to write game-specific documentation. Here, task completion increases by as much as 22% on select games like BombermanPocket (72.0% → 88.0%), and the boost applies to both games seen during training time and unseen test games like SwordOfHope2 (66.0% → 72.0%). However, even this approach leads to considerable performance degradation on some games (DejaVu1, −29%; DejaVu2, −25%; PokémonCrystal, −17%).

        The failure of these standard implementations establishes GameBoyWorlds-Execution as a challenging benchmark and calls for further research on acquiring short-horizon task execution capabilities via self-improvement.

        Submissions

        Run your agent with scripts/benchmark/run_benchmark.sh, which writes one CSV per game to results/benchmark/<game>/. Upload it here and we will add your model to the leaderboard.

        Must be a .csv file, 5MB max.

        Benchmark Tasks

        Browse the benchmark by game, or see the full test suite on GitHub.

        On this page
        Leaderboard Success & Failure Self-Improvement Submissions Benchmark Tasks