# Blind comparison worksheet

Suggested reader-run template. No results have been entered or model tests conducted.

## Record before testing

Task family:
Acceptance threshold:
Critical failure definition:
Allowed tools:
Repair budget:
Reviewer:
Human-baseline assignment:

## Private run register

| Run ID | Case ID | Product | Visible model | Interface | Date | Plan | Thinking setting | Memory/style controls | Tools | Prompt version | Source version |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | | | | | | |

Keep this register away from the reviewer until scoring ends. Log any automatic model switch.

## Blind scoring sheet

| Blind ID | Eight factual checks (0–8) | Evidence IDs (0–2) | Sections and length (0–2) | Total (0–12) | Critical failure? | Accepted? | Review minutes | Correction minutes | Reviewer comments |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | | | | |

Practice rule: accepted if total is at least 11 and critical failure is No. Record all attempts. Report first-pass acceptance separately from corrected acceptance.

## End-to-end effort

| Run ID | Prep minutes | Prompting minutes | Review minutes | Correction or fallback minutes | Active total | Unattended wait | Elapsed start/end | Incremental cost |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | | | |

Active total = preparation + prompting + review + correction/fallback. Do not add unattended waiting if other useful work was done then. Report waiting separately rather than hiding it.

## Decision

What passed:
What failed:
Cases excluded and why:
How results vary by task:
What this sample cannot establish:
Next retest trigger:
