AI Engineering
Spreadsheet Intelligence
Benchmarking
LLM Evals
Catalyst
SpreadsheetBench

Catalyst Scores 82.75% on SpreadsheetBench Verified (400): Mastering Real-World Excel at Scale

September 10, 2026 11 min read
Catalyst Scores 82.75% on SpreadsheetBench Verified (400): Mastering Real-World Excel at Scale

Catalyst Scores 82.75% on SpreadsheetBench Verified (400): Mastering Real-World Excel at Scale

Evaluating spreadsheet AI agents is notoriously hard. Clean toy tables with simple questions like “What was total revenue in cell B12?” give teams a false sense of reliability. But place that same model in front of actual business workbooks - with multi-level merged headers, currency symbols mixed with raw numbers, disjointed summary tables, and multi-sheet relational lookups - and standard frontier LLMs crumble.

Today, I’m thrilled to share my latest milestone: Catalyst has officially been evaluated and verified on the global SpreadsheetBench Verified (400) leaderboard, achieving a verified score of 82.75% by passing 331 out of 400 tasks (verified through submitted output workbooks and execution traces on Sep 04, 2026).

In this deep dive, I’ll unpack what makes the SpreadsheetBench Verified (400) subset the gold standard of spreadsheet intelligence benchmarking, how it was created, how the evaluation works, and the core architectural principles that powered Catalyst to an 82.75% completion rate on the hardest spreadsheet benchmark in AI.


1. The Benchmark: From NeurIPS Spotlight to Industry Gold Standard

SpreadsheetBench was developed by researchers at Tsinghua University and Renmin University, and was accepted as a Spotlight paper at NeurIPS 2024. It quickly became the premier and most widely adopted benchmark for evaluating AI agents on spreadsheet tasks.

When leading frontier labs need to prove their agents can handle real-world Excel workflows, this is the benchmark they turn to:

  • Microsoft used SpreadsheetBench to evaluate Copilot in Excel.
  • OpenAI benchmarked ChatGPT Agent against its rigorous instruction sets.
  • Anthropic tested Claude across its multi-sheet transformations.

Real Problems, Not Lab Toys

Unlike synthetic benchmarks generated in a vacuum, SpreadsheetBench was built from real questions pulled directly from online Excel forums. These are not stylized toy prompts; they represent actual end-users who were genuinely stuck on actual spreadsheets in real production environments.

The benchmark spans the entire spectrum of what professionals demand from spreadsheets:

  • Data Finding & Extraction: Locating elusive values conditioned on multi-attribute criteria.
  • Complex Formula Writing: Synthesizing nested INDEX/MATCH, XLOOKUP, SUMIFS, dynamic arrays, and multi-condition boolean filters.
  • Multi-Sheet Manipulation: Orchestrating references and cascading calculations across multiple sheets.
  • Handling Irregular Layouts: Navigating merged cells, missing headers, inconsistent label offsets, and mixed datatypes.
  • Cross-Table Summarization: Combining detached tables across worksheets with differing schemas into unified structures.

The Online Judge Mechanism: Zero Partial Credit

Every task comes equipped with multiple test cases with differing data across the same underlying structure.

A solution cannot simply pass on one file: it must execute against several test sheet variations successfully. Just like an online judge in competitive programming, your system either passes all test cases or it fails. There is no partial credit. This eliminates brittle agents that overfit or memorize specific coordinates, rewarding systems that truly comprehend spreadsheet schemas and logic.


2. What is SpreadsheetBench Verified (400)?

The original benchmark had 912 tasks. While comprehensive, real-world community datasets naturally accumulate edge cases: ambiguous prompt wording, non-deterministic formula outputs, or tasks where multiple conflicting valid answers existed.

To establish an unassailable benchmark standard, in late 2025 the authors collaborated with Shortcut’s Fundamental Research Labs to release SpreadsheetBench Verified (400) - an expert-annotated subset of 400 instances.

The Four-Layer Quality Filter

The curation process was exceptionally rigorous:

  1. Ambiguity & Non-Determinism Removal: Dropped any prompt whose requirements admitted multiple subjective interpretations or produced non-deterministic outputs.
  2. Filtering Out Easy Tasks: Basic tasks with trivial solutions were pruned out, concentrating the dataset on high-complexity reasoning.
  3. Four-Layer Expert Review:
    • Layer 1: Automated Consistency Checks for data and schema integrity.
    • Layer 2: External Spreadsheet Specialists reviewing real-world Excel validity.
    • Layer 3: Internal Expert Review by Shortcut’s research engineers.
    • Layer 4: Final Validation directly signed off by the original Tsinghua and Renmin benchmark authors.

The Verified (400) set is now the gold standard for rigorous evaluation of spreadsheet AI agents.


3. The Verified Leaderboard: Independent, Blind Evaluation

SpreadsheetBench maintains two distinct leaderboards:

  1. V1 - Full (912 Tasks)
  2. V1 - Verified (400 Tasks)

On the Verified (400) Leaderboard, results undergo strict independent verification. Participating teams execute the benchmark on their own infrastructure and submit comprehensive submission packages—including all generated output spreadsheets (.xlsx), step-by-step execution traces, and raw inference running logs. The maintainers then execute their official deterministic ground-truth evaluation suite against the submitted workbooks and audit execution logs to verify and certify the score.

SpreadsheetBench V1 Verified 400 Official Leaderboard featuring Catalyst at 82.75%

Official SpreadsheetBench V1 - Verified (400) Leaderboard: Catalyst achieves 82.75% (Rank #15) on Sep 04, 2026.

Note: As at the time of writing this blog post, the table below reflects the official verified standings on the SpreadsheetBench V1 - Verified (400) benchmark track:

RankModel / SystemStatusScoreDateOrganization
1JT AlphaDataVerified98.50%Aug 14, 2026CMCC JIUTIAN
2Qingqiu AgentVerified98.25%Jul 11, 2026Kingsoft Office
3WPS AIVerified96.75%Jul 26, 2026Kingsoft Office
4Data Analysis AgentVerified96.50%Jun 16, 2026ByteDance Lark Base & NovaBase Team
5aritoVerified95.50%Jul 15, 2026arito
6Fundy from EmblemVerified95.50%Aug 12, 2026Emblem
7Tetra-Beta-2Verified94.25%Mar 07, 2026DealGlass
8GPT for ExcelVerified92.50%May 05, 2026Talarian
9LeniVerified91.25%Apr 20, 2026Leni Inc.
10GRID AgentVerified91.25%Jul 15, 2026GRID
11Nobie AgentVerified91.00%Jan 06, 2026Nobie
12Atoms Agent v0.3Verified90.50%Aug 06, 2026Acephalt Inc.
13Shortcut.aiVerified86.00%Dec 02, 2025Shortcut.ai
14KyraVerified84.25%Mar 06, 2026Vertex Consulting
15CatalystVerified82.75%Sep 04, 2026Catalyst
16Decide AgentVerified82.50%Jan 29, 2026Decide AI
17fabric-rlm (MiniMax M3)Verified82.50%Jul 26, 2026Fabric RLM

Verified results are independently validated by the benchmark maintainers via submitted output spreadsheets, execution traces, and official ground-truth evaluators. Unverified results are evaluated by external third parties, such as OpenAI and Microsoft.


4. Why 82.75% (331/400) Matters

Securing 331 out of 400 tasks passed on the Verified subset demonstrates several crucial engineering milestones:

1. Robustness Against Filtered, High-Difficulty Tasks

The Verified 400 subset deliberately stripped out the basic, trivial tasks from the original 912 benchmark set. Passing 82.75% of these high-difficulty tasks means Catalyst reliably handles nested formulas, dynamic cell coordinate bounds, formula cross-referencing, and multi-sheet transformations.

2. Generalization Without Brittle Overfitting

Because each task tests multiple workbook variations under strict online judge criteria (all-or-nothing scoring), an 82.75% pass rate proves that Catalyst’s reasoning generalizes across structural mutations. The agent isn’t overfitting to hardcoded cell locations; it understands the relational semantics of the spreadsheet.

3. Competing Alongside Enterprise Tech Giants

The Verified leaderboard is filled with enterprise systems developed by massive research labs: Kingsoft Office (WPS AI, Qingqiu), ByteDance (Lark Base), CMCC JIUTIAN, and dedicated AI spreadsheet companies. As a solo developer building and fine-tuning Catalyst independently, placing at 82.75% right alongside well-funded engineering teams confirms that my architectural thesis and lean execution model are fundamentally sound.


5. Under the Hood: The Engineering Behind the Score

How did Catalyst achieve an 82.75% verified pass rate on these complex workflows? It comes down to three foundational design choices in my agent architecture:

Catalyst Execution Pipeline:
┌────────────────────────┐     ┌────────────────────────┐     ┌────────────────────────┐
│  Schema Extraction &   │ ──> │   Deterministic Code   │ ──> │ Sandboxed Verification │
│ Topological Coordinate │     │  Synthesis & Planning  │     │   & In-Memory State    │
└────────────────────────┘     └────────────────────────┘     └────────────────────────┘

1. Schema-First Extraction vs. Token Dumping

Feeding raw spreadsheets into an LLM context window causes severe attention degradation. When an LLM reads thousands of comma-separated cells, numbers blur together, header associations get lost, and coordinate math fails.

Catalyst instead performs schema-first extraction:

  • Identifies detached data islands and bounding boxes.
  • Decouples table headers and hierarchical metadata from row data.
  • Supplies the agent with a lean structural digest and targeted data previews rather than raw token dumping.

2. Deterministic Code Synthesis Over Direct Cell Guessing

Catalyst never asks the LLM to directly predict mathematical calculations or coordinate values. Instead, the model writes sandboxed programmatic transformations executed against an in-memory workbook representation.

If an operation requires calculating a weighted moving average across filtered dates, Catalyst generates the exact formula or transformation logic, executes it deterministically, and validates the result before finalizing the file.

3. Strict Coordinate and Boundary Guards

Many spreadsheet agent failures stem from coordinate drift - e.g., placing a formula in E15 instead of E16 due to an unseen merged title row. Catalyst implements structural boundary detection that validates target ranges against existing table schemas, preventing accidental overwrites and syntax errors.


What’s Next: Toward SpreadsheetBench V2

With SpreadsheetBench V2 now officially live (introducing SpreadsheetBench 2 on arXiv and the SpreadsheetBench-2 repository), the evaluation frontier has expanded from isolated formula manipulation to end-to-end business spreadsheet workflows—spanning financial modeling, debugging, and visualization in production-scale workbooks with cross-sheet dependencies.

I am actively benchmarking Catalyst against these V2 workflows, continuously refining its action space, tool orchestration, and sandbox execution layers.

If you’d like to explore more:

Samuel Olubukun

Samuel Olubukun

Full Stack AI Engineer

I'm a Full Stack AI Engineer focused on applied AI, autonomous agents, and production-grade web applications.

Tags:
AI Engineering
Spreadsheet Intelligence
Benchmarking
LLM Evals
Catalyst
SpreadsheetBench