POST /api/openrouter/benchmarks/leaderboard
Price: 20 credits
Get an OpenRouter benchmark leaderboard
OpenRouter runs these evaluations itself through the providers that serve each model, so a row carries accuracy together with the dollars and seconds the run took — use it when the question is accuracy per dollar rather than accuracy alone. `median` is the model across providers, while `providers` shows the same model scoring differently on different hosts. Models are identified by dated version alias, so join on `version_alias` rather than `alias`.
access-token string requiredtimeout integer — Max scrapping execution timeout (in seconds) (default: 300; min: 20; max: 1500)benchmark string required — Benchmark whose leaderboard to return (one of: "gpqa-diamond", "tau2-bench-airline")count integer required — Max result count (min: 1)@type string (default: "OpenrouterBenchmarkResult")id string requiredbenchmark string requiredmodel_version_alias string nullablemodel_alias string nullablemodel_name string nullableauthor string nullableis_pareto_optimal boolean nullablemedian object nullable@type string (default: "OpenrouterBenchmarkScore")provider_name string nullableaccuracy number nullableaccuracy_std_dev number nullablerun_count integer nullabletask_count integer nullableavg_cost_per_task number nullableavg_duration_per_task number nullableavg_output_tokens_per_task number nullableis_pareto_optimal boolean nullablelast_run_at string nullableproviders array (default: [])@type string (default: "OpenrouterBenchmarkScore")provider_name string nullableaccuracy number nullableaccuracy_std_dev number nullablerun_count integer nullabletask_count integer nullableavg_cost_per_task number nullableavg_duration_per_task number nullableavg_output_tokens_per_task number nullableis_pareto_optimal boolean nullablelast_run_at string nullable422 — The request body did not validate Check the fields against this schema. A URN with the wrong prefix is the most common cause.408 — The request ran past its time limit Raise `timeout` in the request body, up to the maximum this endpoint documents. Lowering `count` or turning off the `with_*` flags also helps, because less work finishes sooner.412 — The benchmark page carries no leaderboard right now Retrying will not help: either the entity does not exist, or the input points at a different one.429 — Too many requests: a rate limit or a usage window is exhausted When the response carries an X-Retry-After header, wait that many seconds and retry: the same number is in the body as `detail.retry_after`, and the limit clears once that window passes. The message in the body names the limit that was hit.500 — Something broke on our side Retrying will not help. If it keeps happening, send us the X-Request-ID from the response headers.529 — Rate limit reached, or the endpoint is overloaded Wait at least 30 seconds, then retry.X-Error — Error message text (present only on error)X-Request-ID — Unique request identifierX-Execution-Time — Execution time in secondsX-Result-Count — How many records the body carries. 0 means an empty result, which is a normal answer and not by itself an error. A non-zero count can come back together with X-Error when the failure happened partway through — read this header and X-Error independently.X-Total-Available-Results — How many records exist for this query, when the endpoint can say. On a `dry_run` request this is the answer and the body is empty. It saturates: the endpoint's documented maximum means 'at least that many', any smaller number is exact.X-Warning — Present when the request body carried keys this endpoint does not document. They were ignored, so any filter you meant to apply through them did not apply. Check the spelling against this schema and retry.X-Retry-After — Seconds to wait before retrying. Present only on 429.