Frontier finance benchmarks

The arc of LLM progress now runs through an enormous, unprecedented physical buildout. AI capex this year is on the order of $1T across hyperscalers, labs, and the broader supply chain. Though there is debate on timing, it is consensus that this investment must eventually earn a return, and that it must come from deploying models across the economy, with knowledge-work verticals first.

Finance in particular presents a core focus, given its outsized influence on economic function, business creation, and capital allocation. It is also an unbounded knowledge-work field. Beyond the work financial firms staff today, there is nearly limitless demand in doing more, from deeper research coverage to broader economic modeling to government policy budgeting and planning.

As such, we sought to evaluate the state of model progress within finance, to see both how far we have come and where we are heading. One of AGI's end goals is to automate critical portions of finance, not only for frontier-lab ROI and enterprise efficiency, but also to further one of America's greatest strengths, its robust and liquid capital markets. Greater visibility into that development is then critical.

We propose, to our knowledge, the first meta-analysis of this kind.

Model card chronology

In particular, we first look toward the set of benchmarks that frontier labs themselves cite on their model cards. Figure 1 presents this timeline on a per-lab basis. Though not every model release includes a finance score, it is clear that finance is an increasingly important priority for the frontier labs. Of the 9 flagship launches from July through December 2025, only Anthropic's cite a third-party finance benchmark. From February 2026 onward, finance evals appear in launches from most major labs.

Figure 1.

Finance benchmarks in frontier launches

22 flagship launches from 5 frontier labs, July 2025 through July 2026.

Saturation since release

In addition to the finance benchmarks directly cited by frontier lab model cards, we believe there to be a broader set of evals equally representative of model performance and interest, albeit not as publicly (Figure 2). We also measured their best published model score on a per-benchmark basis from initial release to today.

Figure 2.

Best published score at release and now

Light bars show the best at-release score for each eval. Solid bars show the best current score for each eval, observed July 24, 2026. Note that metrics differ by benchmark and are named on each bar. Benchmarks with one bar have no separate at-release score to show.

First, every benchmark's published score has increased since release, in some cases by double digits within months. Second, and we believe more importantly, each benchmark remains relatively unsaturated. Even with the best frontier models, the best published scores still hover around 50 to 60%, with no leaderboard having crossed 70% on their headline metric. Models have improved quickly on these evals without coming close to completing them.

Benchmark landscape

We found that, despite these finance evals being nominally finance-specific, they all still highly vary in their construction.

The underlying reason is that finance itself is a broad, load-bearing term, with many individual sub-industries comprising many different jobs, each involving many various tasks. One way to begin defining finance is top-down, delineating between “buy-side” and “sell-side,” then across investment banking, equity research, private equity, and hedge funds, and then across asset classes within each. Another way is bottom-up, as tasks themselves can be stratified across the nature of their expected input-output relationship. Certain tasks involve just retrieval-based Q&A, for instance, whereas other longer-horizon tasks involve completed deliverables in applications like Excel or PowerPoint. Thus, finance benchmarks by construction have to be opinionated on what exact types of workflows they seek to approximate and measure model progress by.

We believe that there is little open discussion comparing such nuances of finance benchmarks, despite their active use in hillclimbing frontier models and thus deployment across financial firms. As such, we hope to provide this write-up as a valuable source of said context.

Primary differences

Of the many criteria that separate workflows within finance, two broad axes dominate. The first axis separates deterministic answers from open-ended answers. The second axis separates Q&A and retrieval-based questions from Excel, PowerPoint, and Word-based deliverables.

With this framing in mind, we plotted the set of finance benchmarks below, illustrating the primary design differences (Figure 3). It's clear that these evals, despite being finance-focused, are all quite different. The implication is that eval score comparability is only useful conditioned on further grouping or clustering based on sub-characteristics.

Figure 3.

Benchmark entries by output type and answer determinism

We assign each position from public descriptions of tasks and outputs.

Qualitative map of 7 finance benchmarks that have public materials The horizontal axis runs from question answering and retrieval, through text deliverables, to working files. The vertical axis runs from deterministic answers to open-ended ones. Open-ended Deterministic Q&A and retrieval Text deliverables Excel, PowerPoint, and Word files Vals v1 / v1.1 May 2025 / Feb 2026 Vals v2.0 May 2026 Rogo BFB May 2026 Hebbia FSB September 2025 DiligenceBench July 2026 Mercor APEX, IB January 2026 Handshake BTB April 2026 live public sources prepared documents or workspace runner selects research method

What's more is that one might expect Q&A and retrieval-based evals to eventually give way toward longer-horizon, deliverable-based evals, with the latter being more complex. However, this has yet to be the case. Vals' v2.0 Finance Agent, Rogo's BigFinanceBench, and Thoughtful Lab's DiligenceBench (Q&A and retrieval-based benchmarks) all released after Mercor's APEX-Agents and Handshake's BankerToolBench (deliverable-based benchmarks), to similarly low levels of model saturation. At the moment, the output axis of whether a deliverable is involved or not does not seem to proportionally affect task difficulty.

Detailed overview

Importantly, the characteristics distinguishing finance benchmarks are many, far beyond just the two axes presented above. Below, in Table 1, we detail each benchmark in further context. Variables include: the information provided to the model at test-time, the tools and harness available, the required output, the grading method and verifier, and the released task-level artifacts.

Note that we encourage readers to read the original source documents provided by each benchmark constructor to get the clearest, task-level context of what models are being tested on.

Comparison of seven finance benchmark entries by release, corpus, work, information, tools, output, grading, and released artifacts.
Attribute Vals Finance Agent v1.0 / v1.1 Hebbia Financial Services Benchmark Mercor APEX-Agents, IB Handshake BankerToolBench Vals Finance Agent v2.0 Rogo BigFinanceBench DiligenceBench
Release May 2025v1.1 refresh Feb 2026 September 2025 January 2026 April 2026 May 2026 May 2026 July 2026
Task corpus 537 total337 test, 50 public 600+ questions 160 investment-banking tasks 100 tasks 927 total450 test, 27 public 928 total50 public 150 tasks
Work Research questions on public filings and web sources Extraction, summarization, and reasoning on finance documents Long investment-banking tasks in prepared workspaces Complete investment-banking tasks in a fixed data room Research and multi-step questions on public evidence Open-book research questions with auditable calculations Open-ended equity research and investment diligence
Information Live web and SEC EDGAR Supplied document sets Ten prepared workspaces with staged files Prepared files and date-locked data services Live web, SEC filings, and market data Live web and SEC EDGAR Public evidence, method selected by the runner
Tools Web search, EDGAR search, page parsing, and retrieval Not disclosed Nine applications and 63 tools Three data services, LibreOffice, and Python Search, filings, documents, calculator, and market data Web search, EDGAR, URL retrieval, Python, and submission No fixed harness or tool set
Required output Researched text answer, number, or verdict Text answers, summaries, and analyses Console messages and saved files133 + 27 across 160 tasks Excel, PowerPoint, Word, PDF, and CSV files Researched text answer Number or conclusion with a calculation Equity-research memo
Grading method LLM judge, per-question criteriaMetric: final-answer accuracy LLM evaluator consensusWeighted per-criterion scores LLM judge, task criteria and workspace stateMetric: Pass@1 Agentic verifier opens files and checks formulas Question criteria give partial-credit and all-pass results Two judges, weighted rubric and separate final answers LLM judge, weighted criteria for facts, reasoning, and risks
Released task-level artifacts Public question subset and code repositories Methodology paper and sample questions Tasks and harness repository Tasks and repository Public question subset and code repositories Tasks, repeated trajectories, and two-judge grades Tasks and repository

Table 1. Benchmark entries detailed via public materials. Vals v1.1, which was published as an update in February 2026, uses the same 537-question benchmark with updated data, harness, and evaluation.

Task distribution

Each benchmark report also includes a published task distribution by their categorization of choice (Figure 4). For example, Rogo's largest workflow is KPIs and unit economics, Mercor's is sensitivity analysis, and Handshake's is financial modeling and scenario work. It is important to note that the classification schemes themselves differ, with Vals classifying by question type, Mercor and Rogo by workflow, Handshake by product, and DiligenceBench by sector. Since we do not have access to the private datasets for each benchmark at this time, we rely on each constructor's own definitions rather than an independently derived, shared taxonomy.

Figure 4.

Task distribution, as published by each benchmark

Self-published task distributions. Vals v2.0 and Hebbia do not publish per-category task counts.

Case studies

This collection of finance benchmarks was released with varying degrees of public transparency, with Rogo's BigFinanceBench and Vals' Finance Agent being the most auditable. As these two evals do not require any private workspaces, documents, or harnesses to conduct tests on, we provide preliminary case studies below.

More specifically, since Rogo's benchmark releases per-model trajectories and LLM-judge grades, we audit those directly. And for Vals' benchmark, we additionally generate the trajectories ourselves by evaluating and grading the 27 public v2.0 tasks.

Rogo: Public dataset (50)

Rogo's public task set includes 50 tasks (of a total of 928), alongside 3 trajectories for each of 10 models on every task. Grades are also included from 2 LLM judges, Gemini 3.1 Pro and Claude Opus 4.7.

Since Rogo released per-model trajectories and judge traces, our focus is secondary analysis: the summary statistics below, a source audit of every task against the primary filings, and descriptive labels on the traces of the 42 post-QC tasks.

Note that the 10 models evaluated cover a wide ability range. On Rogo's full 928-task leaderboard, they span 22.4-58.8% on rubric-score and 6.6-44.3% on final-answer accuracy.

For each task, we counted how many of the 10 models earn majority-vote final-answer credit under Rogo's Gemini judge, with credit meaning a correct final answer in at least 2 of 3 trials. Below, Figure 5 illustrates this dispersion across the 50 tasks, including the 8 tasks that failed our QC tests.

Figure 5.

Task distribution by number of models receiving majority-vote final-answer credit

Tasks are distributed by the number of models that earned final-answer credit in at least 2 of 3 trials under Rogo’s Gemini judge. In specific, the lefthand cluster includes 29 tasks where only 0-2 models received majority credit.

The distribution splits toward the tails, weighted toward tasks that either most models pass or most models fail. Importantly, for 24 of the 50 tasks, no model earns any credit. And even after controlling for 8 source-audit failures, there are still 16 zero-credit tasks among the retained set of 42.

Rogo: QC findings (8)

Additionally, our quality-control analysis surfaced 8 tasks that were misspecified. In other words, there was either prompt-rubric misalignment, or the reference answer was incorrect. Figure 6 details all 8 examples.

Figure 6.

Excluded tasks post-quality control (n=8)

Rogo: Failure modes

For the remaining 42 retained tasks, we analyzed the released trajectories and assigned each task one primary failure mode. Figure 7 below illustrates the failure modes.

Figure 7.

Failure modes across the 42 retained tasks

For model trajectories on each task, we consolidated shared error paths into primary failure modes.

From our analysis, we found that Misunderstood definition or basis was the primary failure mode on 20 of the 42 tasks, more than double any other error path. And when including for model trajectories where Misunderstood definition or basis is a secondary or tertiary failure mode, there are a total of 31 tasks. To elaborate, this failure mode includes definitional points like gross versus net revenue, reported versus adjusted, and period-end versus period-average. Broken calculation chains, missed line-item builds, and missed off-filing sources all are other examples of common failure modes.

For more detail on how a model exhibits one of these failure modes specifically, we walk through one retained task end to end.

Rogo: Task deep dive

The task below asks for the change in Remitly's fraud loss as a share of net revenue from FY 2024 to FY 2025. Rogo's reference answer is 0.8%, which its rubric treats as a change of +0.8 percentage points. Rogo's bundle releases 30 trajectories on this task, 3 from each of the 10 models. 14 of the 30, spread across 6 models, return +0.6 points instead, including all 3 trials from each of Claude Opus 4.7, Claude Sonnet 4.6, and Gemini 3.1 Pro. Most models make the same mistake: dividing fraud losses by reported revenue instead of net revenue. Figure 8 illustrates the full prompt and both calculation paths.

Figure 8.

Example of misunderstood definition or basis failure mode

Dollar values are in thousands. Each ratio is calculated directly from values in Rogo's reference material.

Rogo task bf-77a2cbc3ee“What is the change in fraud loss as a share of net revenue from FY 2024 to FY 2025 for Remitly? Answer in % with only the final answer rounded to one decimal place.”

Repeated model path

Use reported revenue

  1. FY 202584,166 ÷ 1,635,147 = 5.147%
  2. FY 202457,929 ÷ 1,263,963 = 4.583%
  3. Change0.564 points

Rounded answer: +0.6pp

Rogo reference path

Calculate revenue minus transaction expenses

  1. FY 202584,166 ÷ (1,635,147 − 549,480) = 7.752%
  2. FY 202457,929 ÷ (1,263,963 − 431,604) = 6.960%
  3. Change0.793 points

Rounded answer: +0.8pp

Remitly's FY 2025 10-K labels the deducted line transaction expenses and never defines net revenue.

The important nuance is that (1) the task never defines net revenue, and (2) neither do Remitly's cited filings (which never include a line item as "net revenue"). A model dividing by reported revenue therefore reaches +0.6 on internally consistent arithmetic. Rogo's rubric, however, derives the denominator the way a real-world analyst would be expected to, implicitly understanding to net out transaction costs, yielding +0.8.

We call this a definition fork, the point where two internally consistent calculations separate on an unstated definition. Neither path contains an arithmetic mistake. The fork sits upstream of the math, at the moment the model decides what the words in the prompt mean. Importantly, we still consider this a valid, logically consistent task. Part of a real analyst's job is resolving exactly this kind of context without it being spelled out, and 14 of the 30 trajectories failing to do so is a real capability signal to hillclimb against.

Vals v2.0: Public dataset (27)

Vals' released dataset only includes the tasks themselves, so we produced our own trajectories and LLM-grades. Specifically, we evaluated all 27 public v2.0 tasks with MiniMax M3 (n=3 trials per task), alongside GLM 5.2 as the LLM-judge. In sum, we produced 81 solver traces and 81 judge grades.

MiniMax M3 passed 40 of the 81 graded trials with a mean rubric score of 0.77. GLM 5.2 attributed the failures to wrong interpretation on 12 trials, calculation error on 11, incomplete answers on 9, no answer at all on 5, and wrong sources on 3. Task-level results vary significantly: 7 of the 27 tasks pass all 3 trials, 8 pass none, and no comparables or financial-modeling task passes all 3.

We audited all 27 task specifications against primary sources, the same QC audit we applied to Rogo's public 50. 13 tasks are marked valid, 4 have minor defects, 9 have material defects, and 1 is fatally misspecified. Interestingly, the defects belong to different parts of the correctness distribution here. On Rogo's benchmark, every broken task belonged to the zero-credit bar. On Vals' benchmark, 2 of the 7 all-pass tasks have defects, 1 material and 1 fatal, implying that for QC purposes, a task a model always solves can still be misspecified.

Vals v2.0: Failure modes

We read all 81 traces and coded the retained 20 failing tasks with the same failure modes used for Rogo (Figure 9).

Figure 9.

Primary failure modes across the failing Vals tasks

One primary failure mode per failing task, across 20 retained tasks in total.

We found that Misunderstood definition or basis is the most common failure mode across both benchmarks, appearing as the primary source of failure on 20 of Rogo's 42 retained tasks and on 7 of the 20 failing Vals tasks. On Vals' benchmark, the most frequent definitional failure modes included model errors on interpreting lease rent versus financing interest and gross versus net interest.

Dataset access

It's clear that finance is an important domain for model capabilities to improve on. Our findings suggest, however, that doing so requires substantial nuance and rigor in understanding both (1) what subcharacteristics or latent workflows in finance are actually being represented in training data, and (2) what primary failure modes are inducing low levels of benchmark saturation in the first place.

Our mission is to understand these limits and to ensure that frontier models are hillclimbing on the proper signals. Ones that portray the best of human experts. To access our off-the-shelf datasets built in the shape of common failure modes on known finance evals, reach out to us directly at team@watermark.ac.

Sources

Hebbia Financial Services Benchmark release postpaper
Handshake BankerToolBench paperdatasetrepository
Rogo BigFinanceBench release postpaperleaderboarddataset
DiligenceBench release postdatasetrepository

← All research