Table of Contents
- Catalyst Scores 82.75% on SpreadsheetBench Verified (400): Mastering Real-World Excel at Scale
- 1. The Benchmark: From NeurIPS Spotlight to Industry Gold Standard
- 2. What is SpreadsheetBench Verified (400)?
- 3. The Verified Leaderboard: Independent, Blind Evaluation
- 4. Why 82.75% (331/400) Matters
- 5. Under the Hood: The Engineering Behind the Score
- What’s Next: Toward SpreadsheetBench V2
Catalyst Scores 82.75% on SpreadsheetBench Verified (400): Mastering Real-World Excel at Scale
Evaluating spreadsheet AI agents is notoriously hard. Clean toy tables with simple questions like “What was total revenue in cell B12?” give teams a false sense of reliability. But place that same model in front of actual business workbooks - with multi-level merged headers, currency symbols mixed with raw numbers, disjointed summary tables, and multi-sheet relational lookups - and standard frontier LLMs crumble.
Today, I’m thrilled to share my latest milestone: Catalyst has officially been evaluated and verified on the global SpreadsheetBench Verified (400) leaderboard, achieving a verified score of 82.75% by passing 331 out of 400 tasks (verified through submitted output workbooks and execution traces on Sep 04, 2026).
In this deep dive, I’ll unpack what makes the SpreadsheetBench Verified (400) subset the gold standard of spreadsheet intelligence benchmarking, how it was created, how the evaluation works, and the core architectural principles that powered Catalyst to an 82.75% completion rate on the hardest spreadsheet benchmark in AI.
1. The Benchmark: From NeurIPS Spotlight to Industry Gold Standard
SpreadsheetBench was developed by researchers at Tsinghua University and Renmin University, and was accepted as a Spotlight paper at NeurIPS 2024. It quickly became the premier and most widely adopted benchmark for evaluating AI agents on spreadsheet tasks.
When leading frontier labs need to prove their agents can handle real-world Excel workflows, this is the benchmark they turn to:
- Microsoft used SpreadsheetBench to evaluate Copilot in Excel.
- OpenAI benchmarked ChatGPT Agent against its rigorous instruction sets.
- Anthropic tested Claude across its multi-sheet transformations.
Real Problems, Not Lab Toys
Unlike synthetic benchmarks generated in a vacuum, SpreadsheetBench was built from real questions pulled directly from online Excel forums. These are not stylized toy prompts; they represent actual end-users who were genuinely stuck on actual spreadsheets in real production environments.
The benchmark spans the entire spectrum of what professionals demand from spreadsheets:
- Data Finding & Extraction: Locating elusive values conditioned on multi-attribute criteria.
- Complex Formula Writing: Synthesizing nested
INDEX/MATCH,XLOOKUP,SUMIFS, dynamic arrays, and multi-condition boolean filters. - Multi-Sheet Manipulation: Orchestrating references and cascading calculations across multiple sheets.
- Handling Irregular Layouts: Navigating merged cells, missing headers, inconsistent label offsets, and mixed datatypes.
- Cross-Table Summarization: Combining detached tables across worksheets with differing schemas into unified structures.
The Online Judge Mechanism: Zero Partial Credit
Every task comes equipped with multiple test cases with differing data across the same underlying structure.
A solution cannot simply pass on one file: it must execute against several test sheet variations successfully. Just like an online judge in competitive programming, your system either passes all test cases or it fails. There is no partial credit. This eliminates brittle agents that overfit or memorize specific coordinates, rewarding systems that truly comprehend spreadsheet schemas and logic.
2. What is SpreadsheetBench Verified (400)?
The original benchmark had 912 tasks. While comprehensive, real-world community datasets naturally accumulate edge cases: ambiguous prompt wording, non-deterministic formula outputs, or tasks where multiple conflicting valid answers existed.
To establish an unassailable benchmark standard, in late 2025 the authors collaborated with Shortcut’s Fundamental Research Labs to release SpreadsheetBench Verified (400) - an expert-annotated subset of 400 instances.
The Four-Layer Quality Filter
The curation process was exceptionally rigorous:
- Ambiguity & Non-Determinism Removal: Dropped any prompt whose requirements admitted multiple subjective interpretations or produced non-deterministic outputs.
- Filtering Out Easy Tasks: Basic tasks with trivial solutions were pruned out, concentrating the dataset on high-complexity reasoning.
- Four-Layer Expert Review:
- Layer 1: Automated Consistency Checks for data and schema integrity.
- Layer 2: External Spreadsheet Specialists reviewing real-world Excel validity.
- Layer 3: Internal Expert Review by Shortcut’s research engineers.
- Layer 4: Final Validation directly signed off by the original Tsinghua and Renmin benchmark authors.
The Verified (400) set is now the gold standard for rigorous evaluation of spreadsheet AI agents.
3. The Verified Leaderboard: Independent, Blind Evaluation
SpreadsheetBench maintains two distinct leaderboards:
- V1 - Full (912 Tasks)
- V1 - Verified (400 Tasks)
On the Verified (400) Leaderboard, results undergo strict independent verification. Participating teams execute the benchmark on their own infrastructure and submit comprehensive submission packages—including all generated output spreadsheets (.xlsx), step-by-step execution traces, and raw inference running logs. The maintainers then execute their official deterministic ground-truth evaluation suite against the submitted workbooks and audit execution logs to verify and certify the score.

Official SpreadsheetBench V1 - Verified (400) Leaderboard: Catalyst achieves 82.75% (Rank #15) on Sep 04, 2026.
Note: As at the time of writing this blog post, the table below reflects the official verified standings on the SpreadsheetBench V1 - Verified (400) benchmark track:
| Rank | Model / System | Status | Score | Date | Organization |
|---|---|---|---|---|---|
| 1 | JT AlphaData | Verified | 98.50% | Aug 14, 2026 | CMCC JIUTIAN |
| 2 | Qingqiu Agent | Verified | 98.25% | Jul 11, 2026 | Kingsoft Office |
| 3 | WPS AI | Verified | 96.75% | Jul 26, 2026 | Kingsoft Office |
| 4 | Data Analysis Agent | Verified | 96.50% | Jun 16, 2026 | ByteDance Lark Base & NovaBase Team |
| 5 | arito | Verified | 95.50% | Jul 15, 2026 | arito |
| 6 | Fundy from Emblem | Verified | 95.50% | Aug 12, 2026 | Emblem |
| 7 | Tetra-Beta-2 | Verified | 94.25% | Mar 07, 2026 | DealGlass |
| 8 | GPT for Excel | Verified | 92.50% | May 05, 2026 | Talarian |
| 9 | Leni | Verified | 91.25% | Apr 20, 2026 | Leni Inc. |
| 10 | GRID Agent | Verified | 91.25% | Jul 15, 2026 | GRID |
| 11 | Nobie Agent | Verified | 91.00% | Jan 06, 2026 | Nobie |
| 12 | Atoms Agent v0.3 | Verified | 90.50% | Aug 06, 2026 | Acephalt Inc. |
| 13 | Shortcut.ai | Verified | 86.00% | Dec 02, 2025 | Shortcut.ai |
| 14 | Kyra | Verified | 84.25% | Mar 06, 2026 | Vertex Consulting |
| 15 | Catalyst | Verified | 82.75% | Sep 04, 2026 | Catalyst |
| 16 | Decide Agent | Verified | 82.50% | Jan 29, 2026 | Decide AI |
| 17 | fabric-rlm (MiniMax M3) | Verified | 82.50% | Jul 26, 2026 | Fabric RLM |
Verified results are independently validated by the benchmark maintainers via submitted output spreadsheets, execution traces, and official ground-truth evaluators. Unverified results are evaluated by external third parties, such as OpenAI and Microsoft.
4. Why 82.75% (331/400) Matters
Securing 331 out of 400 tasks passed on the Verified subset demonstrates several crucial engineering milestones:
1. Robustness Against Filtered, High-Difficulty Tasks
The Verified 400 subset deliberately stripped out the basic, trivial tasks from the original 912 benchmark set. Passing 82.75% of these high-difficulty tasks means Catalyst reliably handles nested formulas, dynamic cell coordinate bounds, formula cross-referencing, and multi-sheet transformations.
2. Generalization Without Brittle Overfitting
Because each task tests multiple workbook variations under strict online judge criteria (all-or-nothing scoring), an 82.75% pass rate proves that Catalyst’s reasoning generalizes across structural mutations. The agent isn’t overfitting to hardcoded cell locations; it understands the relational semantics of the spreadsheet.
3. Competing Alongside Enterprise Tech Giants
The Verified leaderboard is filled with enterprise systems developed by massive research labs: Kingsoft Office (WPS AI, Qingqiu), ByteDance (Lark Base), CMCC JIUTIAN, and dedicated AI spreadsheet companies. As a solo developer building and fine-tuning Catalyst independently, placing at 82.75% right alongside well-funded engineering teams confirms that my architectural thesis and lean execution model are fundamentally sound.
5. Under the Hood: The Engineering Behind the Score
How did Catalyst achieve an 82.75% verified pass rate on these complex workflows? It comes down to three foundational design choices in my agent architecture:
Catalyst Execution Pipeline:
┌────────────────────────┐ ┌────────────────────────┐ ┌────────────────────────┐
│ Schema Extraction & │ ──> │ Deterministic Code │ ──> │ Sandboxed Verification │
│ Topological Coordinate │ │ Synthesis & Planning │ │ & In-Memory State │
└────────────────────────┘ └────────────────────────┘ └────────────────────────┘1. Schema-First Extraction vs. Token Dumping
Feeding raw spreadsheets into an LLM context window causes severe attention degradation. When an LLM reads thousands of comma-separated cells, numbers blur together, header associations get lost, and coordinate math fails.
Catalyst instead performs schema-first extraction:
- Identifies detached data islands and bounding boxes.
- Decouples table headers and hierarchical metadata from row data.
- Supplies the agent with a lean structural digest and targeted data previews rather than raw token dumping.
2. Deterministic Code Synthesis Over Direct Cell Guessing
Catalyst never asks the LLM to directly predict mathematical calculations or coordinate values. Instead, the model writes sandboxed programmatic transformations executed against an in-memory workbook representation.
If an operation requires calculating a weighted moving average across filtered dates, Catalyst generates the exact formula or transformation logic, executes it deterministically, and validates the result before finalizing the file.
3. Strict Coordinate and Boundary Guards
Many spreadsheet agent failures stem from coordinate drift - e.g., placing a formula in E15 instead of E16 due to an unseen merged title row. Catalyst implements structural boundary detection that validates target ranges against existing table schemas, preventing accidental overwrites and syntax errors.
What’s Next: Toward SpreadsheetBench V2
With SpreadsheetBench V2 now officially live (introducing SpreadsheetBench 2 on arXiv and the SpreadsheetBench-2 repository), the evaluation frontier has expanded from isolated formula manipulation to end-to-end business spreadsheet workflows—spanning financial modeling, debugging, and visualization in production-scale workbooks with cross-sheet dependencies.
I am actively benchmarking Catalyst against these V2 workflows, continuously refining its action space, tool orchestration, and sandbox execution layers.
If you’d like to explore more:
- Try the live application: catalyst.samuelolubukun.com
- Check the official benchmark repository: SpreadsheetBench on GitHub
- Inspect the dataset: SpreadsheetBench Verified 400 on HuggingFace
- Explore V2: SpreadsheetBench V2 on GitHub | V2 Dataset on HuggingFace
