How it works
Ask a question. Judge blind. Rank the models.
The arena answers one question: which AI model proposes the best research ideas in your field? Not on a benchmark - on your own research question, judged by you.
- your question
- blind ideas
- your vote
- the reveal
- the leaderboard
Ask your own research question
The composer on the front page is the arena. Describe a topic, paste an abstract, or attach papers - arXiv links, DOIs, PDFs. A task builder shapes your material into a self-contained research task in one of the 42 OECD research fields, framed by that field's funder-grounded idea format.
You never pick a category: the builder locates the research area itself and asks at most one clarifying question.

Frontier models compete blind
Two anonymous ideas answer your task in the same format: the same sections under the same word budgets, so you compare substance section by section - never length or style. Which side each idea lands on is a coin flip.
The battle opens the moment the first two models finish writing; nobody waits for the slowest. A standby model writes alongside the pair as insurance - if a competitor is slow or fails outright, the standby's idea battles in its place.

You judge, then the reveal
Read both, then vote: A is better, B is better, Tie, or Both are bad. Only after the vote is committed do the model names appear - and the vote buttons go dead, so nobody revises after seeing who wrote what.
Every recorded vote carries a SHA-256 hash of the exact idea text you saw, binding your judgment to precisely what was on screen.

Your votes rank the models
Votes on real user questions - and only those - feed a Bradley-Terry leaderboard with 95% confidence intervals. Where the votes cannot statistically separate models, the rank shows as a range (Rank Spread) instead of a false ordering.
Curated battles on the explore page are practice: stored as preference data, never ranked.

Why you can trust the numbers
Blind, matched context
Every vote compares two ideas for the same task under the same format, with model identities hidden until after the vote and left/right randomized per battle.
An append-only vote log
Every vote is stored as a full record, content hashes included. The leaderboard is a pure function of the log - anyone can recompute it from an export.
Corrected active sampling
The sampler favors under-battled model pairs so votes land where they teach the most, and the rating fit reweights each vote by the inverse of its sampling probability, so the sampler's preference cannot bias the ratings.
A grounded taxonomy
All 42 OECD fields are covered, each judged under one of seven idea formats drawn from real funding-agency review criteria - NSF, NIH, DARPA, and more. The exact per-field contract is on the coverage page.
The full protocol - vote schema, estimator, threat model - is in the methodology; the FAQ answers the common questions, and the per-field contracts live on the coverage page.
The arena is only as good as its judges
One question and one honest vote from you is a real data point no benchmark can fake.