2048 as a decision bench
TypeSafe Jev ranks a few swipes and says how sure it is. It does not chat. One bad swipe ends the run, so you can see whether the prompt helped or not.
We already spent a September search on this same mid-climb board. How far to look ahead, whether past moves help, whether forcing a plan helps, whether 16 or 64 sequence options beat four directions. A lot of that felt obvious to a human and still lost. Those arms are not scrubbable in the matrix yet. What you can play above is the leftover question: how much board evidence to send, and whether a short lecture on a bare board helps Laya the way captions help Jev.
Jev, planner arm
6704
Chain, wall, and next-tile lines
Single run · peak 512 · replay 223 steps
Laya, corner arm
3496
Board alone, lock a corner and side
Single run · peak 256 · replay 67 steps
Planner arm, recorded scores
6704 vs 2864
Chain and wall lines
Single run vs Single run. Steps describe each replay.
This matrix, origin seed 101
Multi-seed cells report median peak tile and final score. The grid shows the tile IQR and run count. Its animation and step count belong to the median-matching replay. Legacy cells are single runs. Bold scores lead within each model. The highest AI score is green. Board saturation scales with final score, using medians for multi-seed cells and excluding the controls. Steps include the opening's approximate four-move offset.
| Prompt | TypeSafe Jev | Laya | Decisions | Decisions PNG |
|---|---|---|---|---|
| Board alone, no captions | Single run: 256 · 2876 14 steps | Single run: 256 · 3480 64 steps | Median, IQR 256–512, n=20: 256 · 3196 22 steps | Median, IQR 256–256, n=20: 256 · 3112 31 steps |
| Pictures of each swipe option | Single run: 256 · 2876 14 steps | Single run: 256 · 2844 5 steps | Median, IQR 256–512, n=20: 256 · 4236 114 steps | Median, IQR 256–512, n=20: 512 · 5008 117 steps |
| Chain, wall, and next-tile lines | Single run: 512 · 6704 223 steps | Single run: 256 · 2864 10 steps | Median, IQR 512–512, n=20: 512 · 6040 197 steps | Median, IQR 512–1024, n=20: 512 · 7532 297 steps |
| Plan lines plus recent swipes | Single run: 256 · 3192 29 steps | Single run: 256 · 2872 10 steps | Median, IQR 512–512, n=20: 512 · 6068 180 steps | Median, IQR 512–512, n=20: 512 · 6936 248 steps |
| Board alone, "play like a master" | Single run: 256 · 2876 14 steps | Single run: 256 · 2904 17 steps | Median, IQR 256–256, n=20: 256 · 3056 27 steps | Median, IQR 256–256, n=20: 256 · 3152 36 steps |
| Board alone, lock a corner and side | Single run: 256 · 4316 131 steps | Single run: 256 · 3496 67 steps | Median, IQR 256–512, n=20: 256 · 3152 35 steps | Median, IQR 256–256, n=20: 256 · 3156 42 steps |
In the original single-seed recording, TypeSafe Jev only left 256 with chain, wall, and next-tile captions. Laya's captions, planner, and history runs stopped while legal swipes remained. Those are failures to choose a swipe, not blocked boards. Its corner rule scored 3496, compared with 3480 for board alone. One game per arm cannot establish which prompt works best for Jev or Laya across seeds.
In the 20-run Decisions recording, planner captions raised the median tile from 256 to 512 for both text and PNG. Adding history kept the median at 512. Text median scores were 6040 for planner and 6068 for history; PNG scores were 7532 and 6936. History did not consistently improve the result. The master and corner lectures kept median tiles at 256 in both columns.
All 240 Decisions games ended before the safety cap, including two refusals counted as failures. None reached 2048. PNG planner reached 1024 in six of 20 games, compared with three for text, but both had median tile 512. That suggests a difference worth testing on other opening boards; it does not establish PNG superiority. This matrix helps compare prompts on one mid-board position. Jev and Laya need multi-seed recordings before these columns support a model ranking.
Click a cell for the evaluate prompt at that step. That prompt is the thing that changes between rows. Expectimax and the human column are controls that ignore the row. They provide a ceiling for comparison; their original recordings reached 2048.
Decisions and Decisions PNG use these same rows and the same opening board. Text sends the ASCII grid. PNG sends a picture of the current board and keeps the row's option captions. The public Decisions call is POST /v1/decisions with gpt-6-luna. The body is model, input, and questions. There is no reasoning or effort field. Responses has reasoning.effort. The recorder can send that as low. That call was not made.
Each Decisions cell states its run count. Multi-seed cells show median outcomes; legacy peaks stay labelled as single runs. Greedy and random on seeds 101, 202, and 303 still match the September control rows at a 40-move cap: greedy +256, +32, +60, all dead; random +1248 alive, +80 dead, +1332 alive. Medians +60 and +1248. Rows are in openai-decisions-results.jsonl.
What we already tried
The search kept surprising me in the same way: ideas that feel responsible to a human player often shortened the run. I wish those games sat in the matrix too. They do not yet. Here is the short ledger.
| Idea | Human hope | What happened |
|---|---|---|
| Look ahead 0, 1, or 2 swipes | A picture of the board after this swipe, then the next, should beat a bare grid. | Depth 2 was the jump. Depth 0 and 1 died early on the same opening board. Most of the payload is those follow-up pictures. |
| Past swipes in the state | Four recent moves stop it repeating a swipe that just failed. | Cheap to send. Dropping them was uneven, not a clear win. The pictures mattered more than the history line. |
| Short ask vs "plan the next two" | If the pictures already show the plan, a one-line ask should match a paragraph. | At a 120-move cap the short line looked fine. Played out, it died earlier. "Choose a swipe" died on every seed. The sentence is doing work the pictures do not replace. |
| Strategy lectures | "Play like a master," lock a corner, keep large tiles together. The advice a human would give. | On TypeSafe Jev those lectures shortened runs. Two aims in one paragraph fought. A concrete goal beat a policy speech. |
| Pick among 16 pairs or 64 triples | Force a short plan by ranking every ordered sequence, not just four directions. | Every 16-way and 64-way game died with 256 still the largest tile. The first swipe itself was worse. One 4-way choice won. |
| Commit the plan without re-asking | Once it names a route, play the next one or two swipes before asking again. | Worse than asking every swipe. Follow-ups often did not move after the real spawn. Replan after what actually landed. |
| Chain, wall, and next-tile captions | Put the distinguishing fact on each option's board, not in a glossary above. | This is what carried TypeSafe Jev past 256 and, later, to 2048 on some boards. Labels on the picture beat a paragraph at the top. |
Captions on each option beat lectures. One 4-way choice beat forcing a multi-swipe plan. Looking two swipes ahead beat looking none. A corner lecture hurt Jev and barely helped Laya once the board was already bare. The live lab kept the caption style that survived. This page puts that style next to thinner ones, and next to Laya, on one shared seed you can scrub.
Method for the matrix
One origin seed (101), one opening board (score 2844, largest tile already 256). Multi-seed recordings vary the spawn stream and use the same seed list across models. Each cell plays until death, 2048, or a move safety cap. Models choose a single swipe per call. Rows vary only the English payload. Expectimax and Human reuse one replay across every row. The earlier sequence, commit, and depth arms used the same opening board and seeds 101, 202, and 303. They are summarised above rather than replayed in the grid. A keyed Decisions run merges into this file with pnpm exec tsx scripts/jev-vs-laya-play-matrix.mts --openai from apps/web. That command also writes the three-seed summary. The page records 20 runs per cell by default. Use --runs=1 for a shared spawn stream and head-to-head replay. With multiple runs, each cell shows the replay nearest its median tile, then median score, with ties going to the earlier run. The summary file uses the prompt-search spawn, the seed itself. This fixed opening measures play from a mid-board stress position. A single peak, including PNG reaching 1024, is not evidence of model superiority.