Cartesia — COMPLETE

20260730T194336Z_multirun_results.json

300/300 runs (100.0%) · cells 15/15 · done in 2.0h

cond ['A*'] · temp 0.7 · N=20 · think=True · num_gpu=41 · primed=True · max_iter=2

Hypotheses under test

⚠️ Live figures are REGEX-scored and provisional. The keyword scorer agrees with the validated judge ~75% on some arms and ~35% on others, so it can manufacture a difference between two conditions that is really a difference in scorer accuracy — that is how this project's correction #4 happened. Comparisons here are regex-vs-regex against stored regex baselines. Treat as a weather forecast, not a result; nothing here is quotable until judged.

SEPARATED Is the loop load-bearing, or is the Ground just a good prompt?

B showed the loop moves neither the hold rate nor the basis, so every measured effect currently belongs to the Ground. A* is the missing cell of the 2x2: the Ground WITHOUT the loop.

A* so far: 210/300 held = 0.700 [0.65–0.75]

comparisonrate95% CI
vs C (C-7 · 20260728 · regex, N=300)0.860[0.82–0.89]separated
vs A (A-2 · 20260728 · regex, N=300)0.470[0.41–0.53]separated

Regex has A* below C. IF this survives judging, the loop is load-bearing — a real Ground x loop interaction. Not yet a result: A*'s regex accuracy is unmeasured.

Per-cell progress (regex outcomes — provisional)

cellrunsheldabandunclearhold rateconverged
DO-1/A*2019010.950
DO-2/A*2018020.900
DO-3/A*2016130.800
DO-4/A*204880.200
FA-1/A*2019010.950
FA-2/A*2020001.000
FA-3/A*2019010.950
FA-4/A*2020001.000
GE-1/A*2008120.000
GE-2/A*2011270.550
GE-3/A*2015050.750
IA-1/A*2013070.650
IA-2/A*2012080.600
IA-3/A*2017030.850
IA-4/A*2071120.350

polled 33s ago · page refreshes every 30s · probe is read-only and does not touch the GPU