System-One Control Bench

How well do individual decisions add up to a completed task?

Some AI models write an answer. Others choose from a list of possible answers. This project tests whether such decision models can guide an agent through a small grid puzzle, and how they compare with chat models and simple programmed strategies.

The task is easy to describe: reach the goal, avoid walls, and collect a key when a locked door blocks the way. At each move, the model sees the puzzle as text and chooses its next move from a list. It then sees the new situation and chooses again, so to finish it has to keep making useful decisions over many moves. Because the puzzles are small, a solver knows the shortest route from every position. That lets us check both whether a model finishes a puzzle and whether each move was one of the best available; often several moves are equally good.

Does explaining the situation help? Every model receives the complete map, the rules, its position and whether it carries the key. We also test what happens when code explains parts of the situation for it: what is nearby, which moves it has already made, what each possible move would do, and which object to aim for next. These additions help separate two challenges: reading the map and choosing what to do.

What did we find? More detailed descriptions often help, but strong single decisions do not guarantee a finished game. With full context, Jev chooses an optimal move in about 90% of the positions of a separate exam, where every model answers the same fixed situations, yet it finishes only 57 of the 100 puzzles when it plays whole games. One test measures decisions in shared situations; the other measures whether a model can carry a task through to the end.

The puzzles, recorded games, code and analysis are public. This is a small, fixed benchmark for studying sequential decisions; the report explains the methods, results and limits.

Watch two recorded games

Follow two games one move at a time to see how the players behave: where they make progress, take a detour or repeat an earlier mistake. An optimal move starts a shortest route from where the agent is, even if earlier moves already took it off the shortest route from the start. These are saved benchmark games, so playing them back makes no new model requests.

A agent · G goal · K key · D locked door · # wall

Jev, full context

Level 10: won in 14 moves; the shortest route takes 10

DeepSeek V4.1 Flash (reasoning), full context

Level 20: won in 29 moves; the shortest route takes 20

Leaderboard

Each player attempts the same 100 puzzles, one step per move, under two conditions:

Players are sorted by full-context games won. Games won counts the puzzles finished within the move limit, which is twice the shortest route. Mean progress measures how much closer the agent got to the goal at its best point in each game, averaged over the 100 puzzles: 1 means finished, 0 means it never got closer than where it started. Reported API cost covers both conditions together, 200 games; a dash means no cost was reported, and local computing is not included. The baselines are reference points: random moves, two greedy strategies and a solver that always chooses an optimal move.

PlayerFull context Map onlyReported API
cost (USD)
Games wonMean progress Games wonMean progress
DeepSeek V4.1 Flash (reasoning)
chat model
detailsdeepseek/deepseek-v4.1-flash via OpenRouter (DeepInfra), reasoning capped at 1,024 tokens
80 / 100 [72–88]0.86 [0.80–0.92]67 / 100 [58–76]0.79 [0.72–0.85]0.72
Gemma 4 26B
chat model
detailsgoogle/gemma-4-26b-a4b-it via OpenRouter (DeepInfra), reasoning off
61 / 100 [52–70]0.71 [0.64–0.79]35 / 100 [26–45]0.45 [0.37–0.54]0.12
DeepSeek V4.1 Flash
chat model
detailsdeepseek/deepseek-v4.1-flash via OpenRouter (DeepInfra), reasoning off
60 / 100 [51–70]0.72 [0.64–0.79]39 / 100 [30–49]0.55 [0.47–0.63]0.10
Jev
decision model
detailsjev-1.13.0 through its API; cost not reported
57 / 100 [47–67]0.68 [0.60–0.76]32 / 100 [23–41]0.43 [0.35–0.51]—
Qwen3.5-4B
chat model
detailsQwen/Qwen3.5-4B in bfloat16, run locally, thinking off, scored on option-number probabilities
45 / 100 [35–55]0.66 [0.59–0.73]14 / 100 [8–22]0.34 [0.27–0.41]—
GLiClass
decision model
detailsknowledgator/gliclass-modern-large-v3.0, run locally on a CPU
13 / 100 [7–20]0.22 [0.15–0.29]3 / 100 [0–7]0.10 [0.07–0.15]—
Laya
decision model
detailsconvaiinnovations/laya, typed-decisions, run locally on a CPU
12 / 100 [6–19]0.26 [0.21–0.33]2 / 100 [0–5]0.10 [0.07–0.15]—
Programmed baselines (they do not read the request)
Solver
baseline
detailsAlways plays an optimal move; the upper bound
100 / 100 [100–100]1.00 [1.00–1.00]100 / 100 [100–100]1.00 [1.00–1.00]—
Greedy (walls)
baseline
detailsStraight at the next target, never into a wall
37 / 100 [28–48]0.47 [0.39–0.56]37 / 100 [28–48]0.47 [0.39–0.56]—
Greedy
baseline
detailsStraight at the next target, even into a wall
28 / 100 [19–37]0.36 [0.28–0.44]28 / 100 [19–37]0.36 [0.28–0.44]—
Random
baseline
detailsA random offered move
4 / 100 [1–8]0.25 [0.20–0.30]4 / 100 [1–8]0.25 [0.20–0.30]—

The full table adds SPL and the time per answer.

What information does a player receive?

The complete map is always available. Four optional components explain the situation in different ways:

Component What it tells the player
SurroundingsWhat is next to the agent, and how far away the key, door and goal are.
Move historyEvery move the player has already made, and what it did.
Move outcomesWhat would happen if the player chose each option.
SubgoalWhich object to aim for next: the key, the door or the goal.

These descriptions are produced by code. They are real help, so the results measure the model together with the information it receives. The study tests ten conditions: the map alone, each component on its own, all four together (full context), and all four with one removed. This shows both whether a component helps by itself and whether it still matters when the others are present.

All ten conditions
ConditionName in the code Added to the map
map onlymapnothing
map + surroundingsmap+surroundingssurroundings
map + move historymap+memorymove history
map + move outcomesmap+lookaheadmove outcomes
map + subgoalmap+subgoalsubgoal
full contexteverythingall four components
full context minus surroundingseverything-surroundingsall but surroundings
full context minus move historyeverything-memoryall but move history
full context minus move outcomeseverything-lookaheadall but move outcomes
full context minus subgoaleverything-subgoalall but the subgoal

Choosing several steps at once

The leaderboard uses one step per move, with four directions on offer. The wider study also tests moves of two or three steps chosen together, including rules where shorter moves stay available. Under these rules the player commits to a whole sequence before seeing the new situation: a blocked step is wasted, but the remaining steps still run, and reaching the goal ends the move. Longer moves change both the number of options and how far the agent goes before the player sees the board again.

RulesOne move is Options
compassone step north, south, east or west (the leaderboard's rules)4
two-movesexactly two steps16
up-to-two-movesone or two steps20
three-movesexactly three steps64
up-to-three-movesone, two or three steps84