Ask
Global · all contexts Read-only
Conversation
No scope
PN
Which experiments' runs cost the most per composite point, and how many of each used a local judge?
Three experiments stand out. Politeness · SLM is the most expensive at $0.041 per composite
point — its runs lean on a hosted judge. Terse-prompt is the most efficient ($0.028), and
2 of its 4 runs used a local judge. Local-judge runs averaged 31% cheaper at similar scores.
Cost per composite point, by experiment
TableChart
| Experiment | Runs | $ / composite | Local-judge runs |
|---|---|---|---|
| Politeness · SLM | 5 | $0.041 | 1 of 5 |
| Nudge-B sweep | 6 | $0.034 | 3 of 6 |
| Terse-prompt | 4 | $0.028 | 2 of 4 |
Local vs hosted judge — avg run cost
TableChart
read-only · validated
Show query · joins 4 tables
Save prompt