Arena: Building the Evaluation Layer for the AI Economy

Benchmark leadership used to be enough in the earliest days when AI innovation was more like a research competition. The “smartest” model got the medal. Not anymore. As AI remakes enterprise infrastructure, how those models perform in real world scenarios is what the enterprise wants to know. That’s given rise to a new category of infrastructure for selecting, evaluating, and measuring models.
Few companies sit closer to the center of that trend than Anastasios Angelopoulos's Arena.
The Berkeley-born startup first became known through its public model leaderboards, where their community of millions of users compare AI models through head-to-head evaluations of genuine prompts and interactions. For Anastasios and his cofounders Wei-Lin Chian and Ion Stoica, this start as an academic effort to understand model quality has evolved. Over the past year and a half since the company incorporated, it’s evolved at an astonishing pace going from a research project to the agentic model evaluation platform.
Today, millions of users interact with models through Arena's platform, generating a constant stream of feedback about how systems perform in practical settings rather than controlled benchmark environments. That’s created one of the largest collections of real-world AI evaluation data in the industry. Arena users run multi-turn agentic workloads. Arena turns users task completion feedback into quality signals. And there’s clearly a significant appetite for this data with the company going from zero to $100M ARR in 8 months and from zero to $200M ARR in 11 months. The data needs of AI labs and enterprises continue to shift to highly specialized datasets, rubrics, and evals where model can learn from real world tasks.
That kind of revenue growth will catch any investor’s eye, including ours. For DTC, we are always looking forward to the next bottleneck, the next pain point for enterprises. Arena is already addressing several of these from model routing, fine-tuning OSS models, AI sovereignty, model upgrading etc. enabled by strong evals
An Increasingly Fragmented Market
It is highly unlikely that AI will ever be dominated by a single lab’s models. Instead, organizations will use a constantly changing mix of frontier and open-source models, each with different cost structures, strengths, weaknesses, and governance requirements.
Model selection is a significant operational challenge that’s not getting any easier. Choosing the right model for customer support may require different tradeoffs than choosing one for software engineering, document analysis, or agentic automation. Those decisions will need to account for performance, reliability, cost, latency, compliance, and security, often simultaneously.
Evaluation, in other words, starts to look less like a benchmarking exercise and more like a traffic control system.
Addressing More Complex Use Cases
The challenge is becoming even more pronounced as AI evolves beyond chat prompts to writing production-quality software, navigating tools, executing multi-step business processes, and recovering from mistakes on its own. As agents take on more responsibility, enterprises need to understand not only whether a model can complete a task, but whether it can be trusted to operate safely, reliably, and within human intent.
That's why Arena's launch of Code Arena, Agent Mode, and its new Alignment Index is significant. Code Arena extends the company's evaluation framework into one of AI's most important enterprise workloads: software development. Agent Mode reflects the shift from prompt-driven interactions to agents carrying out complex objectives across extended workflows. The Alignment Index adds a new dimension of measuring how AI systems behave in real-world scenarios where trust, oversight, and alignment are taking on greater importance.
We believe this is where evaluation is headed: beyond benchmarks to understanding how AI systems perform, behave, and can be trusted in complex environments so enterprises can make informed decisions on which model, when, and for what. Arena is connecting the dots between enterprise needs, model capabilities, and getting the job done.
Strength of the Team
Arena was founded by Anastasios, Wei-Lin, and renowned Berkeley computer scientist Ion Stoica, whose previous companies helped define major infrastructure categories in cloud computing and distributed systems. What’s most impressive about this trio is that they’ve transformed curiosity and community into a new way of looking at AI innovation. They started it as an academic effort to better understand AI model performance evolved it over 18 months into one of the most important evaluation platforms in the industry today.
Our Investment
The AI industry spent the last several years focused on building increasingly capable models. The next phase will be about understanding and getting the most out of them as efficiently, accurately, and cost-effectively as possible. We’re excited to support Arena in this investment round because we believe they’re the team to build the foundational evaluation layer for today’s and the future’s AI stack.





