Epoch AI Benchmarks Frontier Models, Ranks OpenAI and Anthropic Highest in Composite Index
Epoch AI released its latest Automation Reports, benchmarking frontier AI models. The reports indicate that top models achieved scores up to 65% on Epoch's tasks, with Anthropic and OpenAI tying for the highest composite index score. Google, however, led in user preference rankings, according to Epoch AI's findings.
3 articles from 2 outlets covered this story. Their coverage differs on 2 points. The underlying claim is sourced from a benchmark.
What do all outlets agree on?
2 outlets covered “Epoch AI Benchmarks Frontier Models, Ranks OpenAI and Anthropic Highest in Composite Index”. All of them report the following:
- Epoch AI released Automation Reports
- Reports benchmark frontier AI models
- Anthropic and OpenAI tied in AI Composite Index
- Google led in preference ranking
- Top models scored up to 65%
Did outlets disagree about this?
Yes. Coverage of “Epoch AI Benchmarks Frontier Models, Ranks OpenAI and Anthropic Highest in Composite Index” differs on 2 points. Each account below is how a different outlet described the same event:
Frontier AI models are still unable to fully perform Epoch AI's tasks, despite their benchmark scores.
Frontier AI models achieved scores up to 65% and specific rankings in Epoch AI's benchmark.
Which outlets covered this?
All 3 articles found on this story, grouped by the stance of the piece. Every link goes to the original publisher.