Idle tomato tycoon: 582 lines vs 2,530
ChatGPT vs Claude

Based on 1 real-world tests
One shared test, and its verdict is still open.
1 shared test
One test has separated them on stated requirements. It is at the top of the chart below.
| Measure | ChatGPT | Gemini |
|---|---|---|
| Average score | — | — |
| Web app completion time | 7m 0s | 9m 0s |
| Community | 0% | 0% |
A blog page people finish: 46 KB vs 4.4 MBWebsites
Share of the requirements the brief actually stated, decided by reading the files each model produced. Sorted by the gap between the two, so the tests that separated them come first.
| Test | ChatGPT | Gemini |
|---|---|---|
| A blog page people finish: 46 KB vs 4.4 MB | 7/7 | 6/7 |
A blog page people finish: 46 KB vs 4.4 MB
ChatGPT by 2m 0s7m 0s vs 9m 0s
Measured in each provider’s own web app, so these include network and interface behaviour. They are not model inference latency.
ChatGPT vs Claude

Gemini vs DeepSeek

Gemini · 4 attempts

Gemini 3.6 Flash (High) vs Gemini 3.1 Pro (High)

Gemini 3.5 Flash (High) vs Gemini 3.1 Pro (High)

ChatGPT vs Gemini
