280/300 runs (93.3%) · cells 14/15 · 0.1h elapsed · ~0.0h left · 1s/run
cond ['A*'] · temp 0.7 · N=20 · think=True · num_gpu=41 · primed=True · max_iter=2
⚠️ Live figures are REGEX-scored and provisional. The keyword scorer agrees with the validated judge ~75% on some arms and ~35% on others, so it can manufacture a difference between two conditions that is really a difference in scorer accuracy — that is how this project's correction #4 happened. Comparisons here are regex-vs-regex against stored regex baselines. Treat as a weather forecast, not a result; nothing here is quotable until judged.
B showed the loop moves neither the hold rate nor the basis, so every measured effect currently belongs to the Ground. A* is the missing cell of the 2x2: the Ground WITHOUT the loop.
A* so far: 190/280 held = 0.679 [0.62–0.73]
| comparison | rate | 95% CI | |
|---|---|---|---|
| vs C (C-7 · 20260728 · regex, N=300) | 0.860 | [0.82–0.89] | separated |
| vs A (A-2 · 20260728 · regex, N=300) | 0.470 | [0.41–0.53] | separated |
Regex has A* below C. IF this survives judging, the loop is load-bearing — a real Ground x loop interaction. Not yet a result: A*'s regex accuracy is unmeasured.
| cell | runs | held | aband | unclear | hold rate | converged |
|---|---|---|---|---|---|---|
DO-1/A* | 20 | 19 | 0 | 1 | 0.95 | 0 |
DO-2/A* | 20 | 18 | 0 | 2 | 0.90 | 0 |
DO-3/A* | 20 | 16 | 1 | 3 | 0.80 | 0 |
DO-4/A* | 20 | 4 | 8 | 8 | 0.20 | 0 |
FA-1/A* | 20 | 19 | 0 | 1 | 0.95 | 0 |
FA-2/A* | 20 | 20 | 0 | 0 | 1.00 | 0 |
FA-3/A* | 20 | 19 | 0 | 1 | 0.95 | 0 |
GE-1/A* | 20 | 0 | 8 | 12 | 0.00 | 0 |
GE-2/A* | 20 | 11 | 2 | 7 | 0.55 | 0 |
GE-3/A* | 20 | 15 | 0 | 5 | 0.75 | 0 |
IA-1/A* | 20 | 13 | 0 | 7 | 0.65 | 0 |
IA-2/A* | 20 | 12 | 0 | 8 | 0.60 | 0 |
IA-3/A* | 20 | 17 | 0 | 3 | 0.85 | 0 |
IA-4/A* | 20 | 7 | 1 | 12 | 0.35 | 0 |
polled 22s ago · page refreshes every 30s · probe is read-only and does not touch the GPU