Straight answers about using the arena - what happens when you ask and vote, and how the leaderboard is computed. How the arena works, from question to leaderboard, is on the methodology page.
Battles and voting
Why blind pairwise battles instead of 1-to-10 scores?
Absolute scores drift between people - one voter's 7 is another's 4 - so a scored board mostly measures who happened to grade. A forced choice between two ideas for the same task is a question every voter answers on the same scale: which one wins, here. The models write freely, so the two ideas can differ in length and structure; what keeps the judgment comparable is the shared lens - both answer the same task and are judged under the same funder-grounded standard of what "better" means in that field.
Are the ideas really anonymous, and when are names revealed?
Yes. Nothing sent to your browser names a model before you vote, and every idea is screened so a self-identifying one never reaches you. Names are revealed right after your vote is recorded, and the vote buttons lock at reveal. If you continue with the same pair after a reveal, the next round is still position-blind: you know who is in the match, not which side is which.
Who judges?
Humans - the people the ideas are for. Every battle shows both ideas against the same task, with a one-line judging lens drawn from how that field's funders review real proposals, so you are judging on matched context rather than guessing what "better" means. You can optionally say what field you work in and what your role is; the honest caveat is that self-reported expertise is unverified.
Which votes move the leaderboard?
Only blind votes on user-submitted questions. The arena also keeps a small pool of curated, pre-generated battles; votes on those are collected as per-field preference data and never enter the rankings, and our own test votes are labeled so they can never mix with community votes.
Can I ask follow-ups or vote more than once?
Yes to both. A follow-up before you vote goes deeper with the same pair of ideas; once you vote a round, the next round brings a freshly sampled pair for the same question. Every round is a blind battle, and every blind vote counts - vote as much as you like.
The leaderboard
How are ratings computed?
From the full vote log at once, with the Bradley-Terry model and 95% confidence intervals. The fit has no notion of arrival order, so the same votes always produce the same board. Battles are sampled to give under-compared pairs more air time, and the fit corrects for that sampling so it cannot tilt the ratings. The methodology page walks through the whole pipeline.
Why do some rows show a rank range or no score?
A rank shown as a range means the models inside it are not statistically separated - the votes so far cannot tell them apart, and the board will not pretend otherwise. A competitor under 10 battles is listed as preliminary: its votes already inform every opponent's rating, but no score is published for it until its own record can support one.
What are rows like "glm-5.2 (opencode)"?
The same model running inside a real agent harness. A harnessed row writes with actual tools - a terminal, web search, code execution - through a disclosed scaffold such as opencode, while the raw row is the model writing from its own knowledge alone. The raw rows ship configured but disabled today, so the board ranks the harnessed rows; when a raw row is enabled the two rank separately, making the gap between them the measured value of the harness itself. After you vote, the revealed idea carries its full tool story: searches run, sources read, turns and cost.
Asking a question
What can I paste into the composer?
A topic or a few keywords, an abstract, an arXiv id, a DOI, any URL, or a dropped PDF or image. The arena turns it into one self-contained task, and the competing models answer that task.
What language can I ask in?
Any language. The task and the ideas come back in the language you asked in.
Are there limits?
Daily per-account and site-wide limits keep generation costs bounded. Ordinary use will not run into them.







