Safety/benchmark/2026-09-22

OpenAI and Researchers Introduce Benchmarks for AI Safety and Vulnerability

OpenAI announced the launch of MentalHealthBench, a new benchmark designed to evaluate AI responses in mental health contexts, a development confirmed by Investing.com. Concurrently, researchers introduced MobileCybench, a benchmark focused on evaluating agent vulnerability discovery. While both initiatives contribute to AI safety and performance assessment, the explicit connection between these two distinct benchmarks is not detailed across all reports.

5 articles from 5 outlets covered this story. No difference in framing or figures was found between them. The underlying claim is sourced from a benchmark.

Did outlets disagree about this?

No. All 5 outlets covering “OpenAI and Researchers Introduce Benchmarks for AI Safety and Vulnerability” reported it the same way — no difference in framing or figures was found between them.

Which outlets covered this?

All 5 articles found on this story, grouped by the stance of the piece. Every link goes to the original publisher.

What related stories are there?

Which companies does this involve?

Get the week in AI in one email

What happened, which outlets reported it, and where their coverage differed. One issue a week.

The first issue hasn’t gone out yet. Subscribe and it’s the one you’ll get.

We’ll send the digest and nothing else. One-click unsubscribe. Privacy.