Benchmarks
The models the agent runs on
We don't train our own model — we pick the strongest available ones and stand behind the choice. Here is where they sit on independent scoreboards, and how we test the agent on our own tasks.
Independent scores
The Artificial Analysis leaderboard in full, strongest model first. Highlighted rows are the models the agent runs on.
Our own testing
Public benchmarks measure the model, not the agent working on your data. For that we keep 17 domain sets and over a hundred scenarios — procurement, spreadsheets, documents, market research — graded by rubrics and a judge model.