We Benchmarked EBITDAI Against Claude for Excel. Here Are the Real Numbers.
Ten real financial modeling tasks. Thirty runs per tool. Identical seeded workbooks, identical pass/fail assertions. Claude for Excel passed 100% of tasks. So did EBITDAI running Kimi K3, at roughly eleven cents of model cost per task. This post is the full results, the methodology, and the one failure the benchmark found in our own product. Updated September 2, 2026 with a fourth column: Gemini 3.8 Flash, added to the Pro plan the day it launched, passed 93% at about ten cents a task.
Why we're publishing this
Every AI Excel tool claims to be good. Almost none of them show you a number, and the ones that do rarely show you how it was measured. So we built a benchmark and ran it against the strongest competitor in the space, Anthropic's Claude for Excel, under conditions designed to be fair to them, not to us.
Up front, the framing: this is not a "we beat Claude" post. Claude for Excel is excellent. It passed every task we threw at it. The story in this data is not quality superiority. It's that a specialized doctrine layer lets much cheaper models match that quality, which changes how much modeling you can afford to do.
The headline results
10 scenarios, 3 runs each per tool, scored by automated assertions against the final workbook (never by eyeballing). "Cell checks" are the individual assertions behind those pass/fail grades: does the Year-3 revenue tie out, is the balance sheet balanced, is the interest expense right.
| Tool | Tasks passed | Cell checks passed | Model cost / task |
|---|---|---|---|
| EBITDAI + Kimi K3 | 30/30 (100%) | 261/261 (100%) | $0.11 metered; under $0.02 on the $19 sub* |
| Claude for Excel | 30/30 (100%) | 246/246 (100%) | est. $0.12 to $0.20 on the $20 Pro plan* |
| EBITDAI + DeepSeek V4-flash | 26/30 (87%) | 254/261 (97.3%) | $0.03 metered (no sub needed) |
| EBITDAI + Gemini 3.8 Flash added Sept 2 | 28/30 (93%) | 257/261 (98.5%) | $0.10 metered (launch pricing) |
| ChatGPT for Excel | not yet run, see caveats below | ||
Claude's cell-check denominator is smaller because checks that only EBITDAI can satisfy (loading our modeling guidance, following our house sign and formatting conventions) are skipped for competitors rather than counted as failures. That asymmetry deliberately favors the competitor. Metered costs are the AI provider's list price per task, measured from actual token counts across all 30 runs. *Subscription figures are estimates; the assumptions are spelled out in the cost section below.
The ten tasks
These are the jobs finance people actually hire an Excel agent for, not toy prompts:
- Build from scratch: a SaaS operating model, a DCF on Tesla, a simple LBO, a debt schedule, a working capital schedule, a cash runway model.
- Work with what's there: budget vs. actuals analysis, editing an existing model, fixing a deliberately broken model, cleaning up messy data.
Every run starts from an identical seeded workbook. Every run is graded by the same automated assertions on the final spreadsheet: specific values, live formulas, structural requirements. No human judgment in the scoring loop.
How we kept it fair
Benchmarks published by vendors deserve skepticism, including this one. Here is exactly what we did to keep our thumb off the scale:
- Claude ran in real Excel. The competitor column drives Anthropic's own add-in inside real Excel Online, with automation clicking through its actual UI. It used its own tools, its own prompts, its own workflow.
- Scoring is layout-agnostic. The scorer finds labeled rows wherever a tool chose to put them. This matters: early versions of our scorer penalized correct competitor models for using a different layout, at one point grading a perfect Claude LBO as four failures because the model started in column E. We fixed every one of those bugs in the competitor's favor and re-scored from stored workbook dumps.
- EBITDAI-only checks don't count against competitors. Requiring Claude to follow EBITDAI house doctrine would be an automatic fail, so those checks are skipped for it. The published numbers are the apples-to-apples comparison.
- Harness crashes don't count as model failures. A browser dying mid-run is our problem, not the model's. Crashed runs were retried, not scored.
- No speed comparison, on purpose. The two setups measure time differently (our rig measures agent time; the competitor rig measures browser wall clock including UI round-trips and permission dialogs), so any speed claim would be misleading. We're not publishing one.
- Every run stored its evidence. Each of the 123 runs saved a full workbook dump plus the complete tool log, so any scoring change can be re-verified offline against the original artifacts. If you want to dig into the raw dumps, email us.
The real story: usage per dollar
With Kimi K3 and Claude both at 100%, quality doesn't separate them on this suite. Cost does. Priced from the actual token usage recorded across all 30 runs, at each provider's metered list rates:
| Model | Cost per task | Tasks per $5 of AI spend |
|---|---|---|
| DeepSeek V4-flash (off-peak) | $0.033 | ~154 |
| Kimi K3 (metered) | $0.107 | ~47 |
| Gemini 3.8 Flash (launch pricing) | $0.098 | ~51 |
For perspective: running the entire 30-task benchmark cost $0.97 on DeepSeek, $2.95 on Gemini 3.8 Flash and $3.20 on Kimi. A month of heavy daily modeling on either is pocket change next to a per-seat AI subscription, and that's before prompt caching, which our benchmark conservatively priced at zero benefit. Real steady-state DeepSeek cost is likely materially lower than the figure above.
What a $20 subscription actually buys
Metered rates are what our benchmark rig paid, but most real users are on flat plans: Claude for Excel draws from a $20/month Claude Pro plan, and most EBITDAI Kimi users run the $19/month Kimi Code Moderato tier. Neither vendor publishes token quotas, so what follows is an estimate, built from third-party reports of each plan's allowance plus the per-task usage we measured in this benchmark. Treat the ranges, not the midpoints, as the claim.
| Claude Pro ($20/mo) | Kimi Code ($19/mo) | |
|---|---|---|
| Reported allowance | ~45 messages per 5-hour window | ~300 to 1,200 calls per 5-hour window |
| Model calls per build | est. 6 to 10 | 6 (measured median) |
| Builds per window | ~5 to 7 | ~50 to 200 |
| Monthly ceiling (one window per working day) | ~100 to 165 builds | ~1,100+ builds |
| Effective cost per build at ceiling | ~$0.12 to $0.20 | under $0.02 |
| Implied cost per million tokens | ~$9 to $15 (est.) | ~$1.30 or less |
How to read this. The Claude estimate assumes a modeling build consumes 6 to 10 message-equivalents of the window allowance, in line with our measured agents (6 and 8 model calls per task, median) and with Claude for Excel's approval step before each workbook write; Anthropic's actual accounting is compute-based and not observable from outside, and reviewer reports of Pro exhausting in a heavy session or two are consistent with this range. The Kimi allowance figure is reported for the tier below Moderato, so it understates what the $19 tier allows. Both columns assume the entire plan goes to modeling; in reality a Claude Pro plan is shared with everything else you use Claude for, which lowers its modeling ceiling further. The implied per-token figures divide the plan price by the ceiling times our measured ~13,500 tokens per build; for Claude, whose token appetite per build we cannot see, treat that row as an order-of-magnitude estimate only.
One inversion worth knowing: flat plans only win at volume. Below roughly 180 builds a month, metered Kimi ($0.107 per build) is cheaper than its own subscription, and DeepSeek is cheaper than both at any volume. The subscription math above is the heavy-user case, which is exactly the user a $20 Claude plan serves worst.
That's the actual pitch. Not "our AI is smarter than Claude" (it isn't, and on this suite it doesn't need to be). It's that EBITDAI's guidance layer gets frontier-level results on financial modeling tasks out of models that cost cents, so you never have to meter your usage against a plan limit.
Where DeepSeek failed, and the bug it found in our product
DeepSeek's four failures deserve more attention than its 26 passes, because all four were silent wrong answers: every formula live, the sheet looking right, and the numbers wrong. That's the most dangerous failure class in financial modeling, and it's the one a pass rate hides.
Three of the four failures were the same mistake on the same task. On the debt schedule, DeepSeek computed first-year interest on a $2M loan drawn at the start of the year as roughly $64k, where the correct answer is $144k to $160k. When we dug in, the root cause wasn't really the model. Our own written guidance defined interest as the average of beginning and ending balances, and separately defined the beginning balance as the prior period's ending balance. For a loan drawn at the start of period one, following both rules literally gives you interest on an average that includes a zero. Internally consistent, materially wrong. Kimi noticed the ambiguity and resolved it correctly on its own; DeepSeek did exactly what the instructions said.
Full disclosure on how we handled it: the defect was found mid-benchmark, and we deliberately did not fix it until measurement was complete, so the results couldn't be tuned to the test. After the benchmark closed, we fixed the guidance to state that interest must be charged on the post-draw opening balance. In targeted re-testing, DeepSeek went from 0/3 to 3/3 on the debt schedule, now building an explicit post-draw opening balance row and computing $160k, which is correct. We expect a full re-run to land DeepSeek around 97%, but until we've actually run it, the published number stays 87%.
The remaining failure: one cash runway run where DeepSeek ignored the seeded assumptions and invented its own. Wrong opening cash, wrong growth rate, wrong opex. This is why every EBITDAI output still deserves the same skim you'd give a first-year analyst's work, and why our doctrine makes assumptions explicit and auditable instead of buried in formulas.
Caveats, because every benchmark has them
- Three runs per task is a small sample. A 3/3 versus 0/3 swing on one scenario is suggestive, not proof. We publish the per-run counts rather than just percentages so you can weigh that yourself.
- Two tools hit the ceiling. Kimi and Claude both scoring 100% means this suite no longer discriminates at the top. Separating them would take harder tasks, not more runs.
- ChatGPT for Excel isn't in the table yet. Its column requires a store install and account login we haven't completed. The rig is built and the column is planned; we'll update this post when it runs.
- DeepSeek costs assume off-peak rates and zero cache hits. All runs happened off-peak; peak rates are double. Caching cuts the other way and could reduce real cost several-fold. Both effects are disclosed rather than netted.
- Gemini's cost is the first column that includes cache hits. 27% of its input tokens hit Google's cache and are priced at the cached rate, so its figure is slightly more favorable than the full-rate DeepSeek and Kimi figures. Google's launch pricing also doubles on January 1, 2027, which would put it near $0.19 a task.
Update, September 2, 2026: Gemini 3.8 Flash
Google shipped Gemini 3.8 Flash today and it joined the Pro plan the same day, so we ran it through the identical suite: 10 tasks, 3 runs each, same seeded workbooks, same assertions, on our in-process rig. It passed 28 of 30 tasks (93%) and 257 of 261 cell checks, at $0.098 of metered model cost per task, and it was by far the fastest model on our rig at a median of 18 seconds per task against 50 for DeepSeek and 134 for Kimi. (That is a same-rig comparison; we still publish no speed claim against Claude, whose rig measures something different.)
Both misses were the SaaS model's Year-1 revenue, and both were silent wrong numbers: one run billed the average customer count for the year ($720k instead of $600k), one billed the ending count ($840k, which also flipped the Year-1 EBITDA from a loss to a profit). The third run got it exactly right, and every other scenario was 3 for 3, including the debt schedule that tripped DeepSeek.
One scoring correction, disclosed in the same spirit as the earlier ones: all three of Gemini's Tesla DCF runs first scored as missing a WACC input. Each workbook had one, labeled "Weighted Average Cost of Capital (WACC)". Our check only accepted a label that began with "WACC" or "Discount Rate", skipped that row, and read the sensitivity table's header instead. We widened the pattern and re-scored the stored workbook dumps rather than re-running the model; the re-score reproduced every previously published result exactly, so no other column moved.
What this means if you're choosing a tool
If you already pay for Claude and model occasionally, Claude for Excel is good, and this benchmark confirms it. If modeling is a daily part of your job, the math changes: EBITDAI with Kimi K3 matched Claude's pass rate on every task in this suite at about eleven cents of metered model cost per task, and DeepSeek covers most of that ground at three cents. On flat plans the gap is starker: the same $20 that buys an estimated 100-something builds a month on Claude Pro buys a ceiling of a thousand-plus on a Kimi Code subscription. That's the difference between metering an AI seat and just using it every time a model needs building.
We'll keep re-running this benchmark as models and our guidance evolve, and we'll keep publishing the misses along with the hits, because the misses are where the product gets better. The debt schedule fix exists because a benchmark was unflinching enough to find it.
Related Articles
Best AI for Excel in 2026
Claude, Copilot, ChatGPT, and specialized tools compared for Excel work.
Read moreAI Model Pricing for Financial Modeling
Kimi vs DeepSeek vs OpenAI vs Anthropic: benchmark and cost comparison.
Read moreEBITDAI Is Now Powered by Kimi and DeepSeek
AI usage included on every paid plan, or bring your own AI account: same CFO-grade models.
Read more