Can embodied agents self-improve their short-horizon task execution capabilities?
GameBoyWorlds-Execution is a benchmark for self-improvement, not simply task performance, so our metric of interest is task completion gain, i.e., the increase in task completion after self-learning in the training environments. We first show that there is room for self-improvement to push model performance, even at the frontier.
Task success rates (%) for frontier models on the GameBoyWorlds benchmark. Select a game series below, and hover any bar for the exact model, game title, and success rate.
One task Gemini 3.6 Flash completed and one it did not, drawn at random from each game's benchmark run. Select a game below.
Task completion gain (percentage points) over the Gemma 4 31B base agent, per game. Everything above the line helped; everything below it hurt. Hover a marker for the underlying rates, and click a method in the legend to show or hide it.
No approach consistently improves performance across all game environments.

Qualitatively, we observe that the world model is unable to predict the next observation with sufficient accuracy to serve as a useful tool.
The skill discovery approach performs even worse, with catastrophic collapses on PokémonCrystal (36.0% → 4.0%), PokémonRed (38.8% → 8.2%) and ZeldaOracleOfSeasons (40.0% → 14.0%). The performance degradation is always greater for test-only games, suggesting an overfitting to the training distribution. We hypothesize the extent of degradation is caused by an inability to accurately judge task completion during practice, a known failure scenario for skill acquisition methods.

The only method that achieves partial success is our novel utilization of curiosity-based exploration to write game-specific documentation. Here, task completion increases by as much as 22% on select games like BombermanPocket (72.0% → 88.0%), and the boost applies to both games seen during training time and unseen test games like SwordOfHope2 (66.0% → 72.0%). However, even this approach leads to considerable performance degradation on some games (DejaVu1, −29%; DejaVu2, −25%; PokémonCrystal, −17%).
The failure of these standard implementations establishes GameBoyWorlds-Execution as a challenging benchmark and calls for further research on acquiring short-horizon task execution capabilities via self-improvement.
Run your agent with scripts/benchmark/run_benchmark.sh, which writes one CSV per game
to results/benchmark/<game>/. Upload it here and we will add your model to the
leaderboard.