🏅 DABStep Leaderboard
The Data Agent Benchmark for Multi-step Reasoning (DABStep) is looking to measure and push the state-of-the-art in Data Analysis by LLMs. The benchmark is composed of ~450 data analysis questions (Dataset Link) centered around 1 or more documents that agents will have to understand and cross reference in order to answer correctly.
We have set up a notebook to quickly get an agent baseline using the free Huggingface Inference API: Colab Notebook
Check out the official technical reports here: Adyen Report Hugging Face Report
Join the discussion on the discord server!
Reproduce the baseline results with the agent code we open sourced here
Benchmark Validation Standards
All submissions are initially added to the Unvalidated Leaderboard. The Adyen/Hugging Face team will attempt, with the participation of the respective submission team, to validate any entries that rank within the top 10.
Validation confirms that a submission's results were achieved using a novel approach involving data analysis agents.
Notice Regarding Validations
Validation requests are closed. With DABStep-v2 launching in the coming weeks, we are no longer accepting new validation requests. The Validated Leaderboard is now closed to new entries; previously validated submissions remain as-is.
The Unvalidated Leaderboard remains open — you can still submit and receive a score.
Thank you to everyone who has contributed to DABStep so far. We're excited for the next phase of this benchmark and will share more details on v2 soon.
Submissions
Scores are expressed as the percentage of correct answers.
Each question calls for an answer that is either a string (one or a few words), a number, or a comma separated list of strings or floats, unless specified otherwise. There is only one correct answer. Hence, evaluation is done via quasi exact match between a model’s answer and the ground truth (up to some normalization that is tied to the “type” of the ground truth).
We expect submissions to be json-line files with the following format.
Mandatory fields are: task_id and agent_answer. However, reasoning_trace is optional:
{"task_id": "task_id_1", "agent_answer": "Answer 1 from your agent", "reasoning_trace": "The different steps by which your model reached answer 1"}
{"task_id": "task_id_2", "agent_answer": "Answer 2 from your agent", "reasoning_trace": "The different steps by which your model reached answer 2"}
Our scoring function can be found here.