Table of Contents
- Catalyst Hits #10 on SpreadsheetBench V1: Benchmarking Autonomous Spreadsheet Intelligence
- 1. What is SpreadsheetBench?
- 2. Where Catalyst Stands on the Global Leaderboard
- 3. The Benchmark Verification Journey
- 4. Why Catalyst Performs: The Engineering Behind the Score
- 5. Key Engineering Lessons from Benchmarking at Scale
- Conclusion
Catalyst Hits #10 on SpreadsheetBench V1: Benchmarking Autonomous Spreadsheet Intelligence
Building an AI agent that chats about data is relatively straightforward. Building one that accurately parses nested multi-table workbooks, calculates exact formulas without mathematical hallucinations, and executes complex sheet-level transformations deterministically across thousands of real-world test cases is a completely different engineering frontier.
Today, I’m excited to share that Catalyst has officially been evaluated and verified on the global SpreadsheetBench Leaderboard on the V1 - Full (912) track, securing the #10 spot globally with a verified score of 48.57%.
Official Benchmark Result
SpreadsheetBench V1 - Full (912) Track
In this post, I want to unpack what SpreadsheetBench actually measures, the challenges behind rigorous evaluation in spreadsheet agents, how the submission got officially verified on the board, and the architectural choices that allowed Catalyst to outperform standalone frontier LLMs.
1. What is SpreadsheetBench?
The benchmark was developed by researchers at Tsinghua University and Renmin University and accepted as a Spotlight paper at NeurIPS 2024 (ArXiv:2406.14991). It is the most difficult and most widely adopted benchmark for evaluating AI agents on spreadsheet tasks.
When frontier labs need to prove their agents can handle real Excel workflows, this is where they come:
- Microsoft used it to evaluate Copilot in Excel.
- OpenAI used it to benchmark ChatGPT Agent.
- Anthropic used it to measure Claude.
Real Problems, Not Synthetic Lab Problems
Most AI spreadsheet benchmarks in the past suffered from synthetic simplicity: clean, single-table CSVs with trivial questions like “What is the sum of Column B?”. In contrast, SpreadsheetBench consists of real tasks pulled directly from online Excel forums—not synthetic problems designed in a lab, but actual questions from actual users who were stuck on actual spreadsheets in production environments.
The tasks span the full range of what people genuinely need from Excel:
- Finding and extracting data: Isolating specific rows and cells conditioned on complex business rules.
- Writing complex formulas: Synthesizing nested lookups, matrix transformations, conditional aggregations, and array formulas.
- Manipulating cells across multiple sheets: Coordinating dynamic inter-sheet dependencies.
- Handling non-standard layouts: Merged cells, multi-level headers, missing headers, inconsistent label offsets, and mixed datatypes.
- Summarizing distributed data: Reconciling data living across different tables in varying formats into clean summaries.
Spreadsheet Manipulation Actions Tested:
├── Cell-Level Operations (Find, Extract, Sum, Count, Calculate, Fill)
└── Sheet-Level Transformations (Modify, Highlight, Delete, Format, Reshape)The Online Judge Mechanism: No Partial Credit
Comprising 912 instructions and 2,729 test cases (an average of 3 validation workbooks per task), each task comes with multiple test cases. A solution can’t just work on one spreadsheet. It has to work on several variations of the same structure with different data.
This functions exactly like an online judge in competitive programming: your code either passes all cases or it doesn’t. No partial credit.
This matters because it filters out brittle solutions: an agent that memorizes patterns or overfits to one file layout will fail. An agent that actually understands spreadsheet structure and relational semantics will generalize.
2. Where Catalyst Stands on the Global Leaderboard
The official SpreadsheetBench Leaderboard tracks systems developed by leading AI research labs, enterprise office suites, and specialized spreadsheet agent developers.

Official SpreadsheetBench V1 - Full (912) Leaderboard snapshot with Catalyst ranked at #10 (48.57%).
Note: As at the time of writing this blog post, the table below reflects the official verified and unverified standings on the SpreadsheetBench V1 - Full (912) benchmark track:
| Rank | Model / System | Status | Score | Date | Organization |
|---|---|---|---|---|---|
| 1 | Qingqiu Agent | Verified | 83.11% | June 2026 | Kingsoft Office |
| 2 | JT AlphaData | Verified | 77.85% | August 2026 | CMCC JIUTIAN |
| 3 | WPS AI (Seed 2.0) | Verified | 73.46% | June 2026 | Kingsoft Office |
| 4 | Gemini in Google Sheets | Verified | 70.48% | March 2026 | |
| 5 | Univer | Verified | 68.86% | November 2025 | Univer |
| 6 | Lingxi | Verified | 66.89% | December 2025 | Kingsoft Office |
| 7 | Bluebox | Verified | 62.90% | October 2025 | Bluebox Labs |
| 8 | Shortcut.ai | Verified | 59.25% | October 2025 | Shortcut.ai |
| 9 | Copilot in Excel (Agent Mode) | Unverified | 57.20% | September 2025 | Microsoft |
| 🔟 | Catalyst | Verified | 48.57% | August 2025 | Catalyst |
| 11 | ChatGPT Agent w/ .xlsx | Unverified | 45.50% | July 2025 | OpenAI |
| 12 | Claude Files Opus 4.1 | Unverified | 42.90% | September 2025 | Anthropic |
| 13 | ChatGPT Agent | Unverified | 35.30% | July 2025 | OpenAI |
| 14 | OpenAI o3 | Unverified | 23.30% | July 2025 | OpenAI |
Key Takeaways from the Standings
- Outperforming Raw Frontier Systems: Catalyst (48.57%) outperformed the unverified ChatGPT Agent w/ .xlsx (45.50%), Claude Files Opus 4.1 (42.90%), and baseline OpenAI o3 (23.30%).
- The Power of Scaffolding & Sandboxed Execution: Frontier models alone struggle with exact coordinate mapping and formula syntax. Catalyst’s schema-first deterministic JavaScript runtime and AG Grid orchestration give it a significant edge over purely generative approaches.
- Closing the Gap with Dedicated Suites: While enterprise suites with decade-old spreadsheet calculation engines (like Kingsoft’s WPS AI and Google Sheets) lead the pack, Catalyst demonstrates that an independent, modern web-native agent can compete at a global standard.
3. The Benchmark Verification Journey
Submitting to an official benchmark like SpreadsheetBench requires strict rigor and full reproducibility across thousands of test cases. After running the evaluation across all 912 benchmark instructions and iterating through the official evaluation pipeline, I submitted my complete execution traces, output workbooks, and evaluation logs to the SpreadsheetBench team. Following review and verification by the benchmark maintainers, Catalyst was officially posted on the public SpreadsheetBench Leaderboard in the #10 spot.
4. Why Catalyst Performs: The Engineering Behind the Score
In my previous post on Building Catalyst: Architecture of a Schema-First Conversational Spreadsheet Intelligence Agent, I outlined the core architectural decisions that make Catalyst unique. Those same principles directly contributed to this benchmark performance:
1. Schema-First Isolation (No Context Choking)
Rather than dumping thousands of raw spreadsheet rows into the prompt context (which leads to attention degradation and hallucinated coordinates), Catalyst extracts a compact schema:
- Detected column types & structural boundaries
- Table header hierarchy & non-standard offsets
- Anonymous 3-row data sample
This keeps the LLM’s attention focused on semantic intent and coordinate logic rather than noisy cell values.
2. Code Synthesis over Token Prediction
Catalyst does not guess numbers. When an instruction demands Calculate compound growth for Q3 excluding null rows, Catalyst generates verifiable JavaScript / spreadsheet transformation scripts that run in a controlled client-side sandbox. The code executes against the actual in-memory grid dataset, returning mathematically exact results.
3. Coordinate-Aware Cell & Sheet Manipulations
Real-world sheets have irregular shapes. Catalyst’s prompt scaffolding enforces strict bounds checking, dynamic range targeting, and formula syntax validation, preventing off-by-one errors when writing back to target cells.
5. Key Engineering Lessons from Benchmarking at Scale
Running Catalyst against the entire 912-instruction SpreadsheetBench evaluation set reinforced several crucial principles for building reliable spreadsheet AI:
- Deterministic Execution Beats Probabilistic Guessing: Relying on an LLM to directly output calculated cell values or numerical answers leads to compounding error rates. Giving the agent a sandboxed JavaScript runtime to manipulate tabular structures deterministically is what bridges the gap from demo to production.
- Context Efficiency Is Key to Accuracy: Minimizing token overhead through schema extraction and selective data sampling doesn’t just reduce latency—it actively prevents the model from losing track of column relationships in noisy datasets.
- Robust Range & Coordinate Bounds Checking: In real-world workbooks with non-standard offsets, empty rows, and multi-level headers, prompt scaffolding must strictly enforce coordinate verification before executing writes to prevent unintended cell overwrites.
Conclusion
Rigorous, reproducible evals are essential for transitioning AI agents from exciting demos into reliable daily tools. The SpreadsheetBench evaluation validated my architectural thesis: combine frontier language models with deterministic code execution and schema-first data pipelines.
Try out Catalyst today, inspect the GitHub Repository, and explore the live SpreadsheetBench Leaderboard to see how modern agents are reshaping the future of data manipulation.
