# Vertex Man demonstration protocol

## Arms

| Arm | Model | Workflow |
| --- | --- | --- |
| A | GPT-6-Sol | One implementation assignment |
| B | GPT-6-Luna | Same one-shot assignment |
| C | GPT-6-Luna | Read-only plan, then implementation following that plan |
| D | GPT-6-Luna | Exact C first-pass source after fresh code review, browser testing, one bounded repair, and retest |

All first-pass builders start with only the identical `BRIEF.md` and `ACCEPTANCE.md` in separate empty directories. Use the same Codex CLI, medium reasoning effort, sandbox, and available tools. Give A and B exactly this request: `Build the game in BRIEF.md. Meet ACCEPTANCE.md. Work in this directory. Finish with what you built and which checks you actually ran.` The one-shot builders may reason and test normally; do not disable their tools to induce failure.

For C, first request a short read-only plan from Luna: `Read BRIEF.md and ACCEPTANCE.md. Produce a concise implementation plan only. Identify rules that must remain coherent across mode transitions, who owns transition state, the riskiest implementation slice, and observable checks. Do not write or edit files.` Save its verbatim output as `PLAN.md`. Then request implementation from a fresh Luna session with the same brief and acceptance plus that plan. Do not edit the plan to improve it after seeing A or B.

Freeze A, B and C at first delivery. Preserve their source, final messages and JSONL event logs. A fresh reviewer of C receives the code and contract without the builder's summary, and reports only material findings with code evidence. A separate beta tester plays C against the acceptance list. Validate findings, then copy the exact C tree for one bounded repair pass D. D's effort includes C, review and repair. Do not revise the first-pass trees.

Replay mode transitions with the same viewport and input sequence. Record browser-observed outcomes separately from static analysis and the builders' own tests. Keep elapsed time, token usage and human intervention separate. One run per arm is a demonstration, not a model benchmark. A one-shot win, planned-build failure, or tie is valid; do not change the scenario after seeing results to obtain a preferred order.
