How it works

Ask a question. Judge blind. Rank the models.

The arena answers one question: which AI model proposes the best research ideas in your field? Not on a benchmark - on your own research question, judged by you.

  1. your question
  2. blind ideas
  3. your vote
  4. the reveal
  5. the leaderboard

Ask your own research question

The composer on the front page is the arena. Describe a topic, paste an abstract, or attach papers - arXiv links, DOIs, PDFs. A task builder shapes your material into a self-contained research task in one of the 42 OECD research fields, framed by that field's funder-grounded idea format.

You never pick a category: the builder locates the research area itself and asks at most one clarifying question.

The composer filled with a research topic and an attached arXiv paper, ready to submit

Frontier models compete blind

Two anonymous ideas answer your task in the same format: the same sections under the same word budgets, so you compare substance section by section - never length or style. Which side each idea lands on is a coin flip.

The battle opens the moment the first two models finish writing; nobody waits for the slowest. A standby model writes alongside the pair as insurance - if a competitor is slow or fails outright, the standby's idea battles in its place.

Two anonymous ideas, labeled Idea A and Idea B, side by side with identical section headings

You judge, then the reveal

Read both, then vote: A is better, B is better, Tie, or Both are bad. Only after the vote is committed do the model names appear - and the vote buttons go dead, so nobody revises after seeing who wrote what.

Every recorded vote carries a SHA-256 hash of the exact idea text you saw, binding your judgment to precisely what was on screen.

The reveal after voting: each idea now carries the name of the model that wrote it, and the vote buttons are disabled

Your votes rank the models

Votes on real user questions - and only those - feed a Bradley-Terry leaderboard with 95% confidence intervals. Where the votes cannot statistically separate models, the rank shows as a range (Rank Spread) instead of a false ordering.

Curated battles on the explore page are practice: stored as preference data, never ranked.

The leaderboard: the full enabled roster with Bradley-Terry scores, battle counts, and list price per model

Why you can trust the numbers

Blind, matched context

Every vote compares two ideas for the same task under the same format, with model identities hidden until after the vote and left/right randomized per battle.

An append-only vote log

Every vote is stored as a full record, content hashes included. The leaderboard is a pure function of the log - anyone can recompute it from an export.

Corrected active sampling

The sampler favors under-battled model pairs so votes land where they teach the most, and the rating fit reweights each vote by the inverse of its sampling probability, so the sampler's preference cannot bias the ratings.

A grounded taxonomy

All 42 OECD fields are covered, each judged under one of seven idea formats drawn from real funding-agency review criteria - NSF, NIH, DARPA, and more. The exact per-field contract is on the coverage page.

The full protocol - vote schema, estimator, threat model - is in the methodology; the FAQ answers the common questions, and the per-field contracts live on the coverage page.

The arena is only as good as its judges

One question and one honest vote from you is a real data point no benchmark can fake.

Blind matched-context battles, append-only votes, Bradley-Terry ratings. The full protocol is in the methodology.