# Vertex Man: observed results

**Run date:** 1 October 2026 (Sydney). **Scope:** one run per arm, one frozen brief, one bounded review and repair. This is a demonstration of a workflow, not a model benchmark.

## The comparison

| Arm | Model and process | Delivered source | External browser observation |
| --- | --- | --- | --- |
| A | GPT-6-Sol, one implementation assignment | `runs/A-sol/` | Loaded without console errors; Normal, Geometry Inversion and Time Fracture appeared through Advance update; ten rapid restarts returned to Normal with a zeroed timer. |
| B | GPT-6-Luna, same one-shot assignment | `runs/B-luna/` | Loaded without console errors; all three modes appeared; ten rapid restarts returned the top HUD to Normal with a zeroed timer. |
| C | GPT-6-Luna, separate read-only plan followed by implementation | `runs/C-luna/` | Loaded and changed modes, but the matched Time Fracture target shot left `3/3` targets alive; restarting from Time Fracture reset the world to Normal while the top HUD still said Time Fracture at `00:00`. |
| D | Exact C source, fresh review and browser report, then one Luna repair pass | `runs/D-repaired/` | The same target click removed one target (`2/3`), and the C restart replay now returned both world and HUD to Normal. Ten rapid restarts returned to Normal with no console errors. |

**Main result: the delivered build needed to be challenged.** Independent code review raised concrete questions about gravity-facing support, stale callbacks, pointer coordinates and reassembly. A browser replay exposed a separate HUD state leak and made the aim error visible. One bounded repair pass corrected the observed aim and restart failures on the same replays. Planning was included as an experimental arm, but this run does **not** show a planning advantage: A and B looked functional in the checks we reached, while planned C had browser failures.

## The challenge gate

| Gate | Question | What happened in this run |
| --- | --- | --- |
| Code review against the brief | Do the rules still make sense when the mode changes? | Found four material C issues: right-gravity support, a stale death callback, pointer coordinates and reassembly behavior. |
| Browser challenge | Can a player reproduce a broken promise? | Confirmed the aim miss with a frozen target and found the restart/HUD disagreement that source review missed. |
| Repair and exact replay | Did the fix change the observed behavior? | The same target click hit in D; the same restart sequence returned both world and HUD to Normal. |

The gate is the talking point: **treat the agent's “done” as a claim, then require a reviewer to try to falsify it and replay the fix.** A plan can improve the candidate, but the delivered behavior still needs that challenge.

This run applied the full independent review and repair gate to C only. A and B received browser spot checks but not equivalent peer reviews. The C→D comparison shows the effect of a review-and-repair workflow on one frozen candidate; it is not a controlled estimate of how much the gate would improve every build.

## Findings that expose the engineering conflict

1. **C's restart state leak — browser FAIL, D browser PASS.** At a desktop viewport, click **Advance update** twice, then **Restart**. In C, the scene and toast say Normal but the top mode label says Time Fracture with `00:00`. D displays Normal in both places on the same replay. No console error occurred; a smoke test that only checks for exceptions would miss it.
2. **C's pointer mapping — source finding and matched browser replay.** C stores pointer position relative to the canvas and subtracts the canvas offset again when aiming. D uses the canvas-local position once. In the same desktop viewport, each build was restarted, advanced twice to freeze targets in Time Fracture, and clicked at viewport point `[430, 312]`. C stayed at `3/3`; D fell to `2/3` with a ragdoll. A full resized-viewport aim test was not run.
3. **C's right-gravity support — source finding, D source fix.** C can clamp the player against the room's right boundary without marking that gravity-facing boundary as support, so jump cannot fire there. D marks the right boundary as support. A sustained right-wall jump was not controlled in the browser.
4. **C's pending death timer — source finding, D source fix.** Restarting within 850 ms of death could let the old timeout mutate the new run. D invalidates callbacks from earlier runs with a run identifier. The exact die-then-immediate-restart sequence was not run in the browser.
5. **C's reassembly — partial repair.** C draws a complete faded target during its two-second `REASSEMBLING` state. D adds a body pose moving from the ragdoll toward the target; the target still also appears as a faded complete figure, so the reverse-motion effect is not visually decisive. Browser play confirmed the target returned; it did not establish high-quality articulated reassembly.
6. **B's transition position — source finding.** On entering Geometry Inversion, B sets `player.y = 70` and then zeroes velocity. The brief asks updates to preserve live state while resolving overlaps. This unconditionally teleports the player even when no new solid intersects them. It is an example of a mode branch bypassing a shared transition rule; it was not measured as a browser failure in this run.

## Acceptance status

`PASS` below means observed in the browser, `PARTIAL` means some subconditions were observed, `FAIL` means a concrete counterexample, and `NOT_RUN` means no complete browser replay. Builder self-tests and source review are described separately above.

| Shared check | A | B | C | D |
| --- | --- | --- | --- | --- |
| 1. Load, procedural visuals, no runtime error | PASS | PASS | PASS | PASS |
| 2. Normal: movement, jump, resized aim, three hits, ragdolls, hazards | NOT_RUN | NOT_RUN | FAIL: aim mapping confirmed in frozen-target replay | PARTIAL: target hit in frozen-target replay |
| 3. Live bullets, role swap, right-gravity jump, safe transition | PARTIAL: visual mode swap | PARTIAL: visual mode swap; source teleport | FAIL: right-wall support in source | PARTIAL: visual swap and source fix |
| 4. Live reversal, frozen targets, articulated reassembly, new bullets | PARTIAL: visual mode | PARTIAL: visual mode | FAIL: faded complete target | PARTIAL: dead target returned |
| 5. Time Fracture back to Normal with coherent state | PARTIAL: mode display | PARTIAL: mode display | NOT_RUN | NOT_RUN |
| 6. Natural 60/120/180-second sequence | NOT_RUN | NOT_RUN | NOT_RUN | NOT_RUN |
| 7. Death/restart ×10, clean state, focus loss | PARTIAL: restart ×10 | PARTIAL: restart ×10 | FAIL: HUD leak | PARTIAL: restart ×10 and HUD replay |

The browser controls made held movement and a precise live-bullet transition hard to replay consistently. We did not replace those checks with code inspection. A's builder reported a separate deterministic DOM/canvas simulation covering timing, target hits, restarts and blur, but that is not counted as a browser pass. All builders reported browser verification as not run during their own construction; the browser results here were gathered separately after their source trees were frozen.

## Effort and limits

| Session | Approximate CLI wall time | Reported input tokens (including cached) | Output tokens |
| --- | ---: | ---: | ---: |
| A build | 9m 02s | 1,001,055 | 21,093 |
| B build | 4m 12s | 664,580 | 11,743 |
| C plan | 0m 13s | 37,583 | 454 |
| C build | 2m 11s | 205,761 | 6,247 |
| C review | 0m 27s | 80,481 | 1,081 |
| D repair | 1m 45s | 311,768 | 4,824 |

Times come from first and last timestamps in each CLI JSONL log and include local tool latency. Token counts are the final `turn.completed` usage for each session, include cached input, and are **not a price comparison**. A and B were concurrent with C's stages, and the browser gate was external to all builder sessions. The first-pass trees are preserved unchanged; D was copied from C before review findings were applied. Source, prompts, final messages and JSONL logs are in this directory.

## Presentation claim supported by this run

**The build is a candidate; the challenge gate decides whether its behavior holds up.** Here, “Restart” had to reset the world and its UI together; “jump” had to follow current gravity; “aim” had to survive canvas placement; and “reassemble” had to mean more than setting a target alive again. Review and browser testing challenged those promises, and exact replay checked the repair. The evidence does not support saying that a separate planning pass outperformed one-shot coding in this single run.
