The Screen Does Not Wait
Most computer-use tasks sit still until the next click. Games move. An agent has to see, decide and act inside a loop that the game, not the agent, sets the pace of.
Seventy short games that computer‑use agents must finish from pictures alone
Computer-use agents are usually measured on office software and web forms, where the screen waits patiently for the next click. Minigames Benchmark measures them where it does not: seventy small games that a person can finish in about a minute, covering aiming, timing, memory, navigation, physics and puzzles. An agent is given a short written brief and then only screenshots of the game, and acts only through the mouse and keyboard of a Linux desktop, with no access to game state.
Every game is scored the same way: one point if it is finished inside three minutes of game time, nothing otherwise. Each is played at four paces, from a game that freezes until the agent asks for time to pass, to one that runs at full speed while the agent thinks. The four are designed to separate how well an agent plays from how quickly it decides. Games are deterministic, and every episode is replayed from its recorded inputs before it is scored. As a baseline, a 27-billion-parameter open model running the reference harness on one consumer GPU finishes 16 of the 70 games when thinking is free, and 10 at full speed. An expert human finishes all 70.
Most computer-use tasks sit still until the next click. Games move. An agent has to see, decide and act inside a loop that the game, not the agent, sets the pace of.
Seventy games ask for aiming, timing, memory, navigation, physics, logic and language, each with its own controls. One agent has to play all of them with nothing built for any one game.
There is no page structure, accessibility tree or game state to read. The agent gets a short brief and the picture, the same as a person sitting down at the game.
Tick mode makes thinking free, so today’s slower models can already make progress. Each clock pace then asks for more speed, up to full speed, where the game runs as fast as a person plays it. As models, harnesses and hardware get faster, scores should climb from left to right across the paces, and the full-speed column has the most headroom of all.
There is no leaderboard yet. A score depends on three things at once: the harness, the model and the hardware they run on, and in clock mode it also depends on how fast they are together. Ranking systems that differ in all three would say little about any one of them, so for now this page reports one baseline, run end to end on one machine and described exactly, beside an expert human.
Game time is counted in ticks, sixty to a second of play, and every game is written for that rate. The pace decides when a tick passes, and there are two ways to decide it.
--clock RATE
The game runs on its own at RATE ticks per second of the agent’s time.
It never waits: while the agent thinks, the picture it is thinking about goes out of date.
Any agent can play it unchanged.
--ticks N
The game stands still until the agent presses F1, which lets N ticks pass.
Thinking is free, but the agent must press F1 to move time on,
and hold a key down across the press for it to count.
| Pace | Flag | Game speed | Wall time per game, at most | A full run takes, at most |
|---|---|---|---|---|
| Ticks 3 | --ticks 3 |
Waits for the agent | 3,600 presses of F1 | As long as the agent takes |
| Clock 6 | --clock 6 |
Ten times slower | 30 minutes | 35 hours |
| Clock 20 | --clock 20 |
Three times slower | 9 minutes | 10.5 hours |
| Clock 60 | --clock 60 |
Full speed | 3 minutes | 3.5 hours |
A full run is all seventy games played one at a time, each to its limit. Games that end sooner, and --jobs, which plays several at once, both shorten it.
The four paces are what the results here use, not a limit on what you can run. Either mode takes any rate from 1 to 60. A run can be narrowed to chosen games or tags, played on other seeds, run several episodes in parallel, recorded to video, stopped and resumed, and in tick mode given idle and total timeouts. Every option is in docs/BENCH.md.
These are the benchmark’s own games, running in this page. Pick one, then switch between clock and tick mode to feel what an agent works against. Your input is shown as the benchmark would record it. The whole collection, from its menu, is at play.minigames.sodhi.tech.
Most games need a keyboard. On a phone, try one played with the pointer, such as Clicking Circles.
Full speed. The limit takes three minutes.
The games were designed to cover as wide a range of logic and gameplay as seventy short rounds allow: reflexes and planning, reading and counting, steering and aiming, remembering and exploring. Every game carries one tag for how its round is settled, one for how it can be lost, and any that fit of what it asks for. Select a tag to see its games in the table below.
Filter by tag, or select a game to watch it and read what it asks of a player.
| Game | Inspired by | Skill under test | Controls | Tags |
|---|
No game matches every filter. Clear one to see more.
Other games and products are named only to say what inspired a game. Their names belong to their owners, and Minigames Benchmark is not affiliated with or endorsed by any of them.
Every game was written from scratch for the benchmark, in Rust, and drawn in software a pixel at a time. Many are new takes on familiar kinds of game, so a person knows roughly what to do at a glance, but none borrows the name, characters or art of anything published. Every game has been playtested and beaten by multiple human players.
They all keep the same rules, so that an agent learns the benchmark once rather than seventy times:
Bring your own harness. Anything that runs in a Linux container, looks at an X display and drives its mouse and keyboard can play: a frontier model behind a computer-use API, a local model in a harness of your own, or a scripted baseline. The display, the clock, the recording, the replay and the scoring stay on the benchmark’s side.
minigames-agent, which has Python, pyautogui and xdotool, or from any Linux image with sh and sleep.{task}.
The agent image ships with harness, a deliberately minimal starting point.
Each turn it sends one screenshot to any OpenAI-compatible model and runs the pyautogui code that comes back,
remembering nothing between turns. In tick mode it presses F1 after every answer.
# get the benchmark and build the images
git clone https://github.com/nsodhi-cmu/minigames && cd minigames
uv sync --extra bench
podman build -t minigames-game .
podman build -t minigames-agent agent
podman build -f my-agent.Containerfile -t my-agent my-agent-source
# run your agent at the four paces, then compare them
uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --ticks 3 --out runs/ticks3
uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --clock 6 --out runs/clock6
uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --clock 20 --out runs/clock20
uv run python -m minigames.bench --run-agent my-agent "my-agent {task}" --engine podman --clock 60 --out runs/clock60
uv run python -m minigames.bench --report runs/ticks3 runs/clock6 runs/clock20 runs/clock60
Docker works too: build with docker build -f agent/Containerfile and run with --engine docker. The full guide, including local models and the reference harness’s settings,
is in the README,
and the harness itself is in agent/.
The same clone also builds the games as a Python module, for rendering frames or a training loop without containers.
It is not on PyPI; uv sync builds it into the project’s environment, and the
README shows how to use it.