Minigames Benchmark Code

Seventy short games that computer‑use agents must finish from pictures alone

Games
70
Paces
4
Limit per game
180 s
Categories
21

News

Figure 1. All seventy games, each shown at seed 0. Every game is a 640×480 picture and a handful of keys. Select one to see what it asks of a player.

Abstract

Computer-use agents are usually measured on office software and web forms, where the screen waits patiently for the next click. Minigames Benchmark measures them where it does not: seventy small games that a person can finish in about a minute, covering aiming, timing, memory, navigation, physics and puzzles. An agent is given a short written brief and then only screenshots of the game, and acts only through the mouse and keyboard of a Linux desktop, with no access to game state.

Every game is scored the same way: one point if it is finished inside three minutes of game time, nothing otherwise. Each is played at four paces, from a game that freezes until the agent asks for time to pass, to one that runs at full speed while the agent thinks. The four are designed to separate how well an agent plays from how quickly it decides. Games are deterministic, and every episode is replayed from its recorded inputs before it is scored. As a baseline, a 27-billion-parameter open model running the reference harness on one consumer GPU finishes 16 of the 70 games when thinking is free, and 10 at full speed. An expert human finishes all 70.

Why It Is Hard, and Why It Lasts

The Screen Does Not Wait

Most computer-use tasks sit still until the next click. Games move. An agent has to see, decide and act inside a loop that the game, not the agent, sets the pace of.

No Single Trick Wins

Seventy games ask for aiming, timing, memory, navigation, physics, logic and language, each with its own controls. One agent has to play all of them with nothing built for any one game.

Pixels, and Nothing Else

There is no page structure, accessibility tree or game state to read. The agent gets a short brief and the picture, the same as a person sitting down at the game.

Room to Grow

Tick mode makes thinking free, so today’s slower models can already make progress. Each clock pace then asks for more speed, up to full speed, where the game runs as fast as a person plays it. As models, harnesses and hardware get faster, scores should climb from left to right across the paces, and the full-speed column has the most headroom of all.

Results

There is no leaderboard yet. A score depends on three things at once: the harness, the model and the hardware they run on, and in clock mode it also depends on how fast they are together. Ranking systems that differ in all three would say little about any one of them, so for now this page reports one baseline, run end to end on one machine and described exactly, beside an expert human.

Figure 2. Games finished inside the limit, out of seventy, at each pace. The expert human finished every game.

What the Baseline Shows

  1. Speed costs games. The same system finishes 16 games when thinking is free and 10 at full speed.
  2. It only wins where waiting is safe. On a clock, it finished no game in which standing still can lose. Only in tick mode did it finish any of those, and only two.
  3. Puzzles and words carry the score. Seven of the ten games it finished at full speed are puzzles, quizzes or word games.
  4. Free thinking is slow. Each press of F1 took a median of 8.5 seconds, so one game’s three-minute limit would take about eight and a half hours.

How It Works

Game time is counted in ticks, sixty to a second of play, and every game is written for that rate. The pace decides when a tick passes, and there are two ways to decide it.

--clock RATE

Clock Mode

The game runs on its own at RATE ticks per second of the agent’s time. It never waits: while the agent thinks, the picture it is thinking about goes out of date. Any agent can play it unchanged.

--ticks N

Tick Mode

The game stands still until the agent presses F1, which lets N ticks pass. Thinking is free, but the agent must press F1 to move time on, and hold a key down across the press for it to count.

Pace Flag Game speed Wall time per game, at most A full run takes, at most
Ticks 3 --ticks 3 Waits for the agent 3,600 presses of F1 As long as the agent takes
Clock 6 --clock 6 Ten times slower 30 minutes 35 hours
Clock 20 --clock 20 Three times slower 9 minutes 10.5 hours
Clock 60 --clock 60 Full speed 3 minutes 3.5 hours

A full run is all seventy games played one at a time, each to its limit. Games that end sooner, and --jobs, which plays several at once, both shorten it.

Configurable Beyond the Four Paces

The four paces are what the results here use, not a limit on what you can run. Either mode takes any rate from 1 to 60. A run can be narrowed to chosen games or tags, played on other seeds, run several episodes in parallel, recorded to video, stopped and resumed, and in tick mode given idle and total timeouts. Every option is in docs/BENCH.md.

An Episode

  1. Two containers start. The game’s container holds an X server and has no network. The agent’s container is yours.
  2. The agent is given the task. It is told the screen size, the pace and the game’s brief, but never the seed.
  3. It plays from screenshots. It watches the shared display and works the mouse and keyboard, as on any Linux desktop.
  4. The game runs to its end. It is won, lost or out of time. In tick mode it also ends when the agent’s program exits, or when an optional timeout passes.
  5. The inputs are replayed. The game is played again from the record of what landed on which tick, and an episode that comes out differently is refused.
The agent's container and the game's container share only an X display socket. Screenshots go to the agent and mouse and keyboard input goes to the game. Every input is recorded by tick, the game is played again from that record, and the episode is scored only if the replay reaches the same outcome. Agent container yours, any Linux image Game container no network Model Harness Mouse, keyboard X server, 640 × 480 screenshots shared socket input inputs.txt, by tick Played again Same outcome An episode is scored only if its replay reaches the same outcome.
Figure 3. One episode. The two containers share nothing but the display.

Scoring

  • A game finished inside its limit scores one. Anything else scores zero.
  • The limit is three minutes of game time, the same for every game. Time taken is reported beside the score and never folded into it.
  • There are no lives. Where a wrong move is possible, it ends the round, except in one typing game that allows five slips.

Try a Game at an Agent’s Pace

These are the benchmark’s own games, running in this page. Pick one, then switch between clock and tick mode to feel what an agent works against. Your input is shown as the benchmark would record it. The whole collection, from its menu, is at play.minigames.sodhi.tech.

Most games need a keyboard. On a phone, try one played with the pointer, such as Clicking Circles.

Pace

Full speed. The limit takes three minutes.

Game time
0:00.00
Wall time
0:00.00
Ticks
0
Pointer
–
Input log

    What the Games Test

    The games were designed to cover as wide a range of logic and gameplay as seventy short rounds allow: reflexes and planning, reading and counting, steering and aiming, remembering and exploring. Every game carries one tag for how its round is settled, one for how it can be lost, and any that fit of what it asks for. Select a tag to see its games in the table below.

    All Seventy Games

    Filter by tag, or select a game to watch it and read what it asks of a player.

    Settled by
    Lost by
    Asks for
    Game Inspired by Skill under test Controls Tags

    Other games and products are named only to say what inspired a game. Their names belong to their owners, and Minigames Benchmark is not affiliated with or endorsed by any of them.

    How the Games Were Made

    Every game was written from scratch for the benchmark, in Rust, and drawn in software a pixel at a time. Many are new takes on familiar kinds of game, so a person knows roughly what to do at a glance, but none borrows the name, characters or art of anything published. Every game has been playtested and beaten by multiple human players.

    They all keep the same rules, so that an agent learns the benchmark once rather than seventy times:

    • Everything a player needs is in the picture. There is no score, timer or progress counter beside it.
    • One key does one thing everywhere: rising is Up or Space, and one step is one press of an arrow.
    • What is at stake is shown as a word, a color and a shape, never by one of them alone.
    • A seed and the inputs decide every picture exactly, on every machine, and the tests check it.

    Run Your Agent

    Bring your own harness. Anything that runs in a Linux container, looks at an X display and drives its mouse and keyboard can play: a frontier model behind a computer-use API, a local model in a harness of your own, or a scripted baseline. The display, the clock, the recording, the replay and the scoring stay on the benchmark’s side.

    Your Own Agent

    1. Write a Containerfile. Start from minigames-agent, which has Python, pyautogui and xdotool, or from any Linux image with sh and sleep.
    2. Name the command that starts it. The benchmark puts the task in place of {task}.
    3. Run the four paces. Give the benchmark the image and the command, once per pace.

    The Reference Harness

    The agent image ships with harness, a deliberately minimal starting point. Each turn it sends one screenshot to any OpenAI-compatible model and runs the pyautogui code that comes back, remembering nothing between turns. In tick mode it presses F1 after every answer.

    Shell
    # get the benchmark and build the images
    git clone https://github.com/nsodhi-cmu/minigames && cd minigames
    uv sync --extra bench
    podman build -t minigames-game .
    podman build -t minigames-agent agent
    podman build -f my-agent.Containerfile -t my-agent my-agent-source
    
    # run your agent at the four paces, then compare them
    uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --ticks 3  --out runs/ticks3
    uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --clock 6  --out runs/clock6
    uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --clock 20 --out runs/clock20
    uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --clock 60 --out runs/clock60
    uv run python -m minigames.bench --report runs/ticks3 runs/clock6 runs/clock20 runs/clock60

    Docker works too: build with docker build -f agent/Containerfile and run with --engine docker. The full guide, including local models and the reference harness’s settings, is in the README, and the harness itself is in agent/.

    The same clone also builds the games as a Python module, for rendering frames or a training loop without containers. It is not on PyPI; uv sync builds it into the project’s environment, and the README shows how to use it.

    Inspired by
    Skill under test
    Controls
    Round ends
    Layout
    Tags