PromptMeter
The competitive arena for prompt engineering

LeetCode for prompt engineering.

Solve real tasks with nothing but your prompt. The model is pinned, the tests are hidden, and you’re ranked on cost, speed, and quality against everyone else.

Same model for everyone. The only variable is your prompt.

#14 · Summarize a contract clause
Medium
// your prompt
You are a precise legal assistant.
Summarize the clause in one sentence.
Return only the sentence. No preamble.
8 / 8 tests passed$0.07/run · 412 tok
#12
rank
$0.07
cost/success
+38%
vs #1
40+real-world challenges
12k+prompts ranked
runs aggregated per score
100%hidden, rotating tests
The loop

Five steps from prompt to leaderboard.

Pick a challenge, write a prompt, run it against hidden tests, get an objective score, and climb. Repeat until you’re #1.

  1. 01

    Pick a challenge

    A real task with hidden test variants. The model, version, and settings are pinned — the only variable is your prompt.

  2. 02

    Write your prompt

    No model swaps, no temperature tricks. Just the prompt. Iterate in the editor and run as many times as you like.

  3. 03

    Run hidden tests

    Your prompt runs against held-out variants you can’t see, so you can’t overfit or memorize the answer.

  4. 04

    Get scored

    Cost-per-success, tokens, and latency are pure arithmetic. Quality is judged blind and pairwise — never gameable.

  5. 05

    Climb the board

    Your best run lands you on the leaderboard, ranked against every other engineer on the same pinned challenge.

Challenge modes

Same task. Different way to win.

Every challenge can be played for the cheapest answer, the fastest, the best tradeoff, or the highest judged quality.

Cost golf

cheapest correct

Get a passing answer for the fewest tokens and dollars. The leanest correct prompt wins.

Speed run

fastest correct

Lowest time-to-correct. Tail latency counts — the p95 is what users actually feel.

On the frontier

best tradeoff

Sit on the Pareto frontier of quality vs cost. There’s no single winner — there’s the best at your budget.

Quality weighted

judged blind

Open-ended tasks scored by a blind, order-swapped pairwise judge with confidence intervals.

How you're ranked

Objective where it can be. Fair where it can’t.

Cost, tokens, and latency are pure arithmetic — exact and ungameable. Quality is the optional second axis, judged blind.

Cost per success

Total cost ÷ success rate — retries and failures folded in. The only unit that actually matters, and impossible to fake.

Token efficiency

Input and output tokens, separately. Output is the expensive half; verbosity shows up immediately.

Latency

TTFT plus p95 / p99 across runs. The mean lies; the tail is what users feel.

Reliability

Pass rate and parse-failure rate over ≥5 runs. Reliably-good beats occasionally-great.

Quality (Elo)

Blind, pairwise, order-swapped judging fed into a Bradley-Terry rating — with confidence intervals, not 1–10 guesses.

Leaderboard

Every prompt, ranked against the world.

One pinned challenge, one honest ranking. Your best run earns your spot — and you can see exactly how far the top is.

Extract structured invoice dataMedium · Cost golf
model pinned
#EngineerScore
  • 1tokenwhisperer1842
  • 2prompted_1809
  • 3lessismore1771
  • 12youyou1604
  • 13chainofthot1598
Why it’s legit

A ranking you can actually trust.

Leaderboards are only fun if they’re fair. Every score is built to be reproducible, bias-controlled, and impossible to game.

/The model is pinned

Same model, exact version, temperature, and seed for everyone — so you’re measuring the prompt, not the model.

/Tests are hidden & rotating

Held-out, LiveBench-style variants you never see, so no one can overfit one example or memorize answers.

/Objective first

Where ground-truth pass/fail exists, scoring is mechanical and ungameable. The judge is reserved for open-ended tasks.

/Bias-controlled judging

Blind, pairwise, order-swapped, and length-controlled — position, verbosity, and self-preference biases removed.

/Confidence intervals

Overlapping intervals are a tie — we never rank off noise. Every score is the aggregate of multiple runs.

/Cost is computed, not judged

Tokens, price, and latency are arithmetic from captured data. The judge never sees cost or model identity.

Who it’s for

Practice, compete, and get hired on it.

Bottom-up by design — engineers come for the challenge, teams come for the league, recruiters come for the signal.

Engineers

Sharpen a skill that’s suddenly worth real money, and prove it with a public, ranked profile.

  • Daily & rotating challenges
  • Instant cost/latency feedback
  • A shareable skill profile

Teams

Run internal leagues on your own tasks. Level up together and see who writes the leanest prompt.

  • Private team challenges
  • Internal leaderboards
  • Shared prompt patterns

Hiring

A verifiable signal for prompt-engineering skill — challenge candidates on real, scored tasks.

  • Candidate challenge sets
  • Objective, anti-cheat scoring
  • Skill-based shortlisting

What it’s not.

PromptMeter ranks humans on prompt efficiency for genuinely useful tasks — not the adjacent games people confuse it with.

  • not jailbreakingthat’s HackAPrompt
  • not one-model token golfthat’s Prompt Golf
  • not model benchmarkingthat’s LMArena
  • not vibes-based scoringtests are pinned & hidden
Pricing

Free to compete. Pay to go deeper.

Public challenges and the global leaderboard are free forever. Upgrade for private practice, quality scoring, and team leagues.

Free

$0forever

Everything you need to compete and climb.

  • Unlimited public challenges
  • Cost, speed & pass-rate scoring
  • Global leaderboards
  • Public skill profile
Most popular

Pro

$12per month

For engineers serious about their craft.

  • Private practice challenges
  • Quality (Elo) scoring mode
  • Full run history & analytics
  • Compare your prompts side-by-side

Teams & Hiring

Customtalk to us

Internal leagues and candidate assessments.

  • Private team leaderboards
  • Custom challenge authoring
  • Candidate skill assessments
  • SSO, roles & audit logs

Pick a challenge. See where you land.

Free to start, no card required. The board’s waiting.

Start solving