Results
Terse-prompt study · Pilot-20 · nudge B · Jul 1620
Enrollees
312
Scored rounds
1.4
Avg self-improve passes
73.9
Composite mean · 0–100
$2.52
Run cost
Scores by quality measure
| Measure | Mean | Per person | Median | Min–Max | Distribution |
|---|---|---|---|---|---|
| Groundedness | 78.2 | 77.0 | 80 | 41–96 | |
| Relevance | 81.5 | 80.9 | 84 | 52–98 | |
| Coherence | 74.8 | 73.1 | 76 | 45–95 | |
| Instruction-following | 88.0 | 87.6 | 90 | 63–99 | |
| Composite | 73.9 | 72.4 | 75 | 44–94 |
Chat memory
32K
Smallest window in set-up
34
Rounds truncated
12
Rounds where enrollee alerted
5
Rewinds · 5.8 turns mean
Final Answer per Session
76.8
Composite mean · final answers
20
Sessions that reached an answer
0.7
Turned-back attempts per round · mean
Refusals
14
Appropriate · relevance N/A
3
Unwarranted · still scored
5.4%
Of 312 scored rounds
Comparing the other runs in this study — same cohort and model config, varying the nudge.
| Run · nudge | Composite | Rounds | Model $ · run | Best model | Distribution |
|---|---|---|---|---|---|
| A — “Ask anything, no guidance” | 69.2 | 300 | $2.48 | gpt-5 | |
| B — “Keep prompts short & direct” (this run) | 73.9 | 312 | $2.14 | opus-4-8 | |
| C — “Give step-by-step reasoning” | 71.5 | 298 | $3.05 | opus-4-8 |
Cost This Run
Total spend (US$)
$2.52
models $2.14 + judge $0.38
Generator vs judge
Two separate totals
generator · judge
Speed & effort per round
Iterations / round
mean 1.4 · med 1
1–4
Response time (ms)
mean 2,180 · med 1,900
640–9k
Tokens / round
mean 1,240
total 387k
Total compute time
11 min 20 s
cumulative
Ask in Plain Language
carol_p (58.1) · dave_r (61.7) · erin_s (64.0) · … View as table
Read-only · validated