Benchivo is a growing database of head-to-head AI model challenges. One prompt goes to several models. Every output is published exactly as it came back. Whatever can be measured is measured, whatever cannot is argued in the open, and readers vote on the rest.

Why not another leaderboard

Public benchmarks have saturated. The top models cluster within a point of each other on tests that everyone has been optimising against for years, and the resulting ranking tells you very little about the work you actually do.

What is missing is not another score. It is the evidence underneath one: the prompt, the real output, the file that either renders or does not, and a note on what the number leaves out. That is what this site collects.

What makes it hard to copy

After fifty challenges, anyone can copy the idea. After a few thousand properly documented tests — with stored outputs, recorded model versions, historical rounds and community votes attached to each — the database itself becomes the thing that is hard to reproduce.

Independence

Benchivo is not affiliated with OpenAI, Anthropic, Google or xAI. Models are tested through ordinary consumer accounts. Where a method is weak, themethodology pagesays so rather than rounding it away.

Who runs this

Get in touch

Corrections, disputed results and challenge suggestions are all welcome athello@benchivo.com. If you think a result is wrong, say which one and why — the raw outputs are published precisely so that argument is possible.