Methodology
Blind matched-context battles, append-only votes, Bradley-Terry ratings.
Your question
Blind battle
Vote, then the reveal
The thread continues
Ranking
Every pairwise vote feeds the same Bradley-Terry fit with its recorded sampling propensity; nothing else ranks.
Why this design
Blind and matched-context
Model names are hidden until the vote is cast, and both ideas answer the same task in the same format, so preference is about the ideas.
Append-only votes
Every vote is stored in full and never rewritten; the leaderboard is a pure function of the vote log that anyone can recompute from an export.
Propensity-weighted ratings
The sampler steers votes toward under-battled pairs, so each vote records its draw probability and the Bradley-Terry fit reweights by its inverse.
Harnesses are disclosed
Every ranked row is a model plus its harness, and after you vote the revealed idea carries its full tool story: searches run, sources read, turns and cost.