Friday, October 9, 2026
HomeInternationalAI evaluation startup Arena raises $200 Mn to strengthen AI safety and...

AI evaluation startup Arena raises $200 Mn to strengthen AI safety and alignment testing

Arena, which began in 2023 as a research project at UC Berkeley to crowdsource AI model rankings, has raised $200 million in a Series B funding round at a $3.1 billion valuation.

The latest funding follows the company’s announcement in June that it had reached $100 million in annualized run-rate revenue. Lightspeed Venture Partners and Khosla Ventures led the round, with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis, and other investors.

Previously, Arena announced a $150 million Series A round in January at a $1.7 billion post-money valuation. At the time, the company reported annualized revenue of $30 million. Consequently, its valuation has nearly doubled in approximately 10 months.

Arena operates a free, crowdsourced platform that allows consumers to compare AI models. Users submit prompts or request vibe-coded projects and then evaluate which model delivers better results. According to the company, the platform attracts tens of millions of monthly visitors.

In September last year, Arena introduced its commercial offering, AI Evaluations, which provides AI model developers and enterprises with detailed performance analytics based on community feedback.

The launch came at a pivotal time. This year, AI labs discovered that their models were gaming benchmark tests by finding ways to achieve high scores without genuinely demonstrating the capabilities those tests were intended to measure. Meanwhile, enterprises increasingly sought help identifying models that best fit their internal requirements rather than relying exclusively on standardized benchmarks.

“AI is advancing faster than our ability to evaluate it, and static benchmarks break down once models recognize they’re being tested,” the company said in its funding announcement. “The world needs a neutral third party to measure how safe and aligned AI actually is once it’s in the hands of real people. Arena is stepping into that role today,” it added.

To strengthen its evaluation capabilities, Arena has also introduced an alignment category to its leaderboard. The category ranks AI models on issues such as unauthorized action, where a model takes steps it was not asked to perform; false attribution, where it incorrectly credits statements or facts to the wrong source; and “deceptive completion,” which refers to falsely claiming to have completed tasks that remain unfinished.

Currently, several OpenAI models occupy the top positions on Arena’s preliminary alignment leaderboard. Meanwhile, Claude Opus 5.5 and Claude Fable rank sixth and ninth, respectively.

Subscribe To Newsletter

ICYMI

BRL Editor
BRL Editorhttps://businessreviewlive.com
Business Review Live covers finance, technology, travel, lifestyle, and everything in between through exclusive interviews and analysis, market statistics, digital video, and an expanded array of content formats.