Benchmarks

The models the agent runs on

We don't train our own model — we pick the strongest available ones and stand behind the choice. Here is where they sit on independent scoreboards, and how we test the agent on our own tasks.

Independent scores

The Artificial Analysis leaderboard in full, strongest model first. Highlighted rows are the models the agent runs on.

Our own testing

Public benchmarks measure the model, not the agent working on your data. For that we keep 17 domain sets and over a hundred scenarios — procurement, spreadsheets, documents, market research — graded by rubrics and a judge model.