# Two-week pilot worksheet

Proposed reader-run protocol. All fields below are blank for the reader's pilot.

## Register before the pilot

Task and output:
Allowed account and source classification:
Source version:
Participants and coded IDs:
Reviewer:
Case selection and complexity groups:
Within-person assignment method:
Model/interface/settings/date:
Prompt version:
Repair limit and manual fallback:
Quality checklist:
Critical failure rule:
Continue/change/stop thresholds:
What happens if the evidence is inconclusive:

## Log every assigned case

| Case ID | Participant code | Complexity | Assigned condition | Prompt version | Prep minutes | Draft/prompt minutes | Review minutes | Correction/fallback minutes | Active total | Unattended waiting | Completed | First-pass accepted | Final accepted | Critical error | Notes |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | | | | | | | | | | |

Active total sums the four active-labour columns. Include all reviewers' effort. Keep abandoned attempts; report counts and reasons. Do not silently divide only by accepted outputs if failed attempts consumed time.

Setup/training minutes and who contributed:
Incremental licence/usage costs:
Missing data and exclusions (before looking at results where possible):

## Decision note

Conditions compared and dates:
Assigned/completed/first-pass/final-accepted counts:
Critical-error counts:
Total active minutes including failures:
Per-completed-case minutes and range:
Task or experience differences (no employee rankings):
Total pilot effort with setup:
Predeclared gates passed/failed:
Remaining uncertainty:
Continue / change and retest / stop:
Owner and retest trigger:
