The warehouse
CC BY 4.0One world, released as 1,219 Parquet files partitioned by month (71.3 GiB compressed), with setup for BigQuery, DuckDB, Snowflake, Trino, Delta Lake and Iceberg.
View datasetA frontier benchmark for enterprise-scale data-driven decisions.
The strongest model, Claude Opus 5.5, fully solves 34.8% of the tasks, scoring 95 or more, and averages 59.5 points. Eight of the thirteen models average under 30.
Mean task score out of 100, each model at its highest reasoning effort (extra-high where offered). Scores are ±2 to 8 points at 95%. Official scores are computed on a private world; contact us to submit.
Decision-Bench simulates a year of a New York food-delivery platform, from orders and couriers to fraud and incentives, and exports it to an Oracle E-Business Suite warehouse. Agents query it, analyze it in Python and file decisions.
The best offers vanish before couriers can tap. Ban scripted accounts.
Files · Bans and holdsFinance wants to free up $9.6M in driver bonuses. Decide where to trim.
Files · Budgets and plansA pay floor took effect in April. Forecast May's net pay adjustments.
Files · ForecastsThe city says couriers were underpaid. Find every short week.
Files · Reported figuresDid membership deals pay for themselves? Publish the dashboard.
Files · Dashboard data sourcesFraudsters take over a restaurant's payout account and redirect its money. Some restaurants report it; some never notice. Claude Opus 5.5 learned the takeover's fingerprint from the cases that were reported, found the unreported ones that match it, and froze their payouts, catching every hijacked restaurant and no honest one.
Hundreds of restaurants changed their bank account for ordinary reasons, and restaurant groups share one account across their brands. Six of the 47 model settings evaluated get this task right.
The warehouse and the tasks are public. Point an agent at a local DuckDB file or your own warehouse, and every decision it files is recorded with the run. For official scores and access to the simulator, contact us.
One world, released as 1,219 Parquet files partitioned by month (71.3 GiB compressed), with setup for BigQuery, DuckDB, Snowflake, Trino, Delta Lake and Iceberg.
View datasetThe 210 tasks, a reference agent with its warehouse and Python tools, the mission-control console it files through, the sandboxes it runs in, and every model setting the paper ran.
View codeWrite to us to have your runs scored, or to talk about the simulator and the new worlds and RL environments built on it. How to submit.
Get in touch
We built Decision-Bench because we couldn't find a benchmark that came close to the complexity of those environments, or that told us much about how a model would do inside one.
If you use Decision-Bench, the released warehouse or our results in your work, please cite the paper. It is published as Argo-Bench.
@article{tomitsuka2026argobench,
title = {Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows},
author = {Tomitsuka, Gabriel and Raayatsanati, Arman and Xing, Emma
and Gand, Duke and Ma, Joseph J},
journal = {arXiv preprint arXiv:2610.02122},
year = {2026}
}Scores are reported against a version. Each task keeps a hash of exactly what the agent was shown.
A frontier benchmark for enterprise-scale data-driven decisions.
Decision-Bench contains no data about real people. Restaurants are real New York businesses from public records, except those cast in a fraud scenario, which carry fictional names; everything any of them does in the world is simulated. Map data © OpenStreetMap contributors.