Mid-board test
Median / single run
TypeSafe Jev
Laya
Decisions
Decisions PNG
Expectimax / Human

Board alone, no captions

Pictures of each swipe option

Chain, wall, and next-tile lines

Plan lines plus recent swipes

Board alone, "play like a master"

Board alone, lock a corner and side

Final score sets board saturation. Bold: model best. Green: highest AI score. Controls excluded. Multi-seed cells use median scores.

0

2048 as a decision bench

TypeSafe Jev ranks a few swipes and says how sure it is. It does not chat. One bad swipe ends the run, so you can see whether the prompt helped or not.

We already spent a September search on this same mid-climb board. How far to look ahead, whether past moves help, whether forcing a plan helps, whether 16 or 64 sequence options beat four directions. A lot of that felt obvious to a human and still lost. Those arms are not scrubbable in the matrix yet. What you can play above is the leftover question: how much board evidence to send, and whether a short lecture on a bare board helps Laya the way captions help Jev.

Jev, planner arm

6704

Chain, wall, and next-tile lines

Single run · peak 512 · replay 223 steps

Laya, corner arm

3496

Board alone, lock a corner and side

Single run · peak 256 · replay 67 steps

Planner arm, recorded scores

6704 vs 2864

Chain and wall lines

Single run vs Single run. Steps describe each replay.

This matrix, origin seed 101

Multi-seed cells report median peak tile and final score. The grid shows the tile IQR and run count. Its animation and step count belong to the median-matching replay. Legacy cells are single runs. Bold scores lead within each model. The highest AI score is green. Board saturation scales with final score, using medians for multi-seed cells and excluding the controls. Steps include the opening's approximate four-move offset.

PromptTypeSafe JevLayaDecisionsDecisions PNG
Board alone, no captions

Single run: 256 · 2876

14 steps

Single run: 256 · 3480

64 steps

Median, IQR 256–512, n=20: 256 · 3196

22 steps

Median, IQR 256–256, n=20: 256 · 3112

31 steps

Pictures of each swipe option

Single run: 256 · 2876

14 steps

Single run: 256 · 2844

5 steps

Median, IQR 256–512, n=20: 256 · 4236

114 steps

Median, IQR 256–512, n=20: 512 · 5008

117 steps

Chain, wall, and next-tile lines

Single run: 512 · 6704

223 steps

Single run: 256 · 2864

10 steps

Median, IQR 512–512, n=20: 512 · 6040

197 steps

Median, IQR 512–1024, n=20: 512 · 7532

297 steps

Plan lines plus recent swipes

Single run: 256 · 3192

29 steps

Single run: 256 · 2872

10 steps

Median, IQR 512–512, n=20: 512 · 6068

180 steps

Median, IQR 512–512, n=20: 512 · 6936

248 steps

Board alone, "play like a master"

Single run: 256 · 2876

14 steps

Single run: 256 · 2904

17 steps

Median, IQR 256–256, n=20: 256 · 3056

27 steps

Median, IQR 256–256, n=20: 256 · 3152

36 steps

Board alone, lock a corner and side

Single run: 256 · 4316

131 steps

Single run: 256 · 3496

67 steps

Median, IQR 256–512, n=20: 256 · 3152

35 steps

Median, IQR 256–256, n=20: 256 · 3156

42 steps

In the original single-seed recording, TypeSafe Jev only left 256 with chain, wall, and next-tile captions. Laya's captions, planner, and history runs stopped while legal swipes remained. Those are failures to choose a swipe, not blocked boards. Its corner rule scored 3496, compared with 3480 for board alone. One game per arm cannot establish which prompt works best for Jev or Laya across seeds.

In the 20-run Decisions recording, planner captions raised the median tile from 256 to 512 for both text and PNG. Adding history kept the median at 512. Text median scores were 6040 for planner and 6068 for history; PNG scores were 7532 and 6936. History did not consistently improve the result. The master and corner lectures kept median tiles at 256 in both columns.

All 240 Decisions games ended before the safety cap, including two refusals counted as failures. None reached 2048. PNG planner reached 1024 in six of 20 games, compared with three for text, but both had median tile 512. That suggests a difference worth testing on other opening boards; it does not establish PNG superiority. This matrix helps compare prompts on one mid-board position. Jev and Laya need multi-seed recordings before these columns support a model ranking.

Click a cell for the evaluate prompt at that step. That prompt is the thing that changes between rows. Expectimax and the human column are controls that ignore the row. They provide a ceiling for comparison; their original recordings reached 2048.

Decisions and Decisions PNG use these same rows and the same opening board. Text sends the ASCII grid. PNG sends a picture of the current board and keeps the row's option captions. The public Decisions call is POST /v1/decisions with gpt-6-luna. The body is model, input, and questions. There is no reasoning or effort field. Responses has reasoning.effort. The recorder can send that as low. That call was not made.

Each Decisions cell states its run count. Multi-seed cells show median outcomes; legacy peaks stay labelled as single runs. Greedy and random on seeds 101, 202, and 303 still match the September control rows at a 40-move cap: greedy +256, +32, +60, all dead; random +1248 alive, +80 dead, +1332 alive. Medians +60 and +1248. Rows are in openai-decisions-results.jsonl.

What we already tried

The search kept surprising me in the same way: ideas that feel responsible to a human player often shortened the run. I wish those games sat in the matrix too. They do not yet. Here is the short ledger.

IdeaHuman hopeWhat happened
Look ahead 0, 1, or 2 swipesA picture of the board after this swipe, then the next, should beat a bare grid.Depth 2 was the jump. Depth 0 and 1 died early on the same opening board. Most of the payload is those follow-up pictures.
Past swipes in the stateFour recent moves stop it repeating a swipe that just failed.Cheap to send. Dropping them was uneven, not a clear win. The pictures mattered more than the history line.
Short ask vs "plan the next two"If the pictures already show the plan, a one-line ask should match a paragraph.At a 120-move cap the short line looked fine. Played out, it died earlier. "Choose a swipe" died on every seed. The sentence is doing work the pictures do not replace.
Strategy lectures"Play like a master," lock a corner, keep large tiles together. The advice a human would give.On TypeSafe Jev those lectures shortened runs. Two aims in one paragraph fought. A concrete goal beat a policy speech.
Pick among 16 pairs or 64 triplesForce a short plan by ranking every ordered sequence, not just four directions.Every 16-way and 64-way game died with 256 still the largest tile. The first swipe itself was worse. One 4-way choice won.
Commit the plan without re-askingOnce it names a route, play the next one or two swipes before asking again.Worse than asking every swipe. Follow-ups often did not move after the real spawn. Replan after what actually landed.
Chain, wall, and next-tile captionsPut the distinguishing fact on each option's board, not in a glossary above.This is what carried TypeSafe Jev past 256 and, later, to 2048 on some boards. Labels on the picture beat a paragraph at the top.

Captions on each option beat lectures. One 4-way choice beat forcing a multi-swipe plan. Looking two swipes ahead beat looking none. A corner lecture hurt Jev and barely helped Laya once the board was already bare. The live lab kept the caption style that survived. This page puts that style next to thinner ones, and next to Laya, on one shared seed you can scrub.

Method for the matrix

One origin seed (101), one opening board (score 2844, largest tile already 256). Multi-seed recordings vary the spawn stream and use the same seed list across models. Each cell plays until death, 2048, or a move safety cap. Models choose a single swipe per call. Rows vary only the English payload. Expectimax and Human reuse one replay across every row. The earlier sequence, commit, and depth arms used the same opening board and seeds 101, 202, and 303. They are summarised above rather than replayed in the grid. A keyed Decisions run merges into this file with pnpm exec tsx scripts/jev-vs-laya-play-matrix.mts --openai from apps/web. That command also writes the three-seed summary. The page records 20 runs per cell by default. Use --runs=1 for a shared spawn stream and head-to-head replay. With multiple runs, each cell shows the replay nearest its median tile, then median score, with ties going to the earlier run. The summary file uses the prompt-search spawn, the seed itself. This fixed opening measures play from a mid-board stress position. A single peak, including PNG reaching 1024, is not evidence of model superiority.