Cost golf
cheapest correctGet a passing answer for the fewest tokens and dollars. The leanest correct prompt wins.
Solve real tasks with nothing but your prompt. The model is pinned, the tests are hidden, and you’re ranked on cost, speed, and quality against everyone else.
Same model for everyone. The only variable is your prompt.
Pick a challenge, write a prompt, run it against hidden tests, get an objective score, and climb. Repeat until you’re #1.
A real task with hidden test variants. The model, version, and settings are pinned — the only variable is your prompt.
No model swaps, no temperature tricks. Just the prompt. Iterate in the editor and run as many times as you like.
Your prompt runs against held-out variants you can’t see, so you can’t overfit or memorize the answer.
Cost-per-success, tokens, and latency are pure arithmetic. Quality is judged blind and pairwise — never gameable.
Your best run lands you on the leaderboard, ranked against every other engineer on the same pinned challenge.
Every challenge can be played for the cheapest answer, the fastest, the best tradeoff, or the highest judged quality.
Get a passing answer for the fewest tokens and dollars. The leanest correct prompt wins.
Lowest time-to-correct. Tail latency counts — the p95 is what users actually feel.
Sit on the Pareto frontier of quality vs cost. There’s no single winner — there’s the best at your budget.
Open-ended tasks scored by a blind, order-swapped pairwise judge with confidence intervals.
Cost, tokens, and latency are pure arithmetic — exact and ungameable. Quality is the optional second axis, judged blind.
Total cost ÷ success rate — retries and failures folded in. The only unit that actually matters, and impossible to fake.
Input and output tokens, separately. Output is the expensive half; verbosity shows up immediately.
TTFT plus p95 / p99 across runs. The mean lies; the tail is what users feel.
Pass rate and parse-failure rate over ≥5 runs. Reliably-good beats occasionally-great.
Blind, pairwise, order-swapped judging fed into a Bradley-Terry rating — with confidence intervals, not 1–10 guesses.
One pinned challenge, one honest ranking. Your best run earns your spot — and you can see exactly how far the top is.
Leaderboards are only fun if they’re fair. Every score is built to be reproducible, bias-controlled, and impossible to game.
Same model, exact version, temperature, and seed for everyone — so you’re measuring the prompt, not the model.
Held-out, LiveBench-style variants you never see, so no one can overfit one example or memorize answers.
Where ground-truth pass/fail exists, scoring is mechanical and ungameable. The judge is reserved for open-ended tasks.
Blind, pairwise, order-swapped, and length-controlled — position, verbosity, and self-preference biases removed.
Overlapping intervals are a tie — we never rank off noise. Every score is the aggregate of multiple runs.
Tokens, price, and latency are arithmetic from captured data. The judge never sees cost or model identity.
Bottom-up by design — engineers come for the challenge, teams come for the league, recruiters come for the signal.
Sharpen a skill that’s suddenly worth real money, and prove it with a public, ranked profile.
Run internal leagues on your own tasks. Level up together and see who writes the leanest prompt.
A verifiable signal for prompt-engineering skill — challenge candidates on real, scored tasks.
PromptMeter ranks humans on prompt efficiency for genuinely useful tasks — not the adjacent games people confuse it with.
Public challenges and the global leaderboard are free forever. Upgrade for private practice, quality scoring, and team leagues.
Everything you need to compete and climb.
For engineers serious about their craft.
Internal leagues and candidate assessments.
Free to start, no card required. The board’s waiting.
Start solving