Finance Eval Meta-Analysis

Measuring the state of frontier model progress in finance.

Frontier finance benchmarks

The arc of LLM progress now runs through an enormous, unprecedented physical buildout. AI capex this year is on the order of $1T across hyperscalers, labs, and the broader supply chain. Though there is debate on the timing, it is consensus that this investment must eventually earn a return, and that it would come from deploying models across the economy, with white-collar verticals first.

Finance in particular presents a core focus, given its outsized influence on economic function, business creation, and capital allocation. It is also an unbounded white-collar field. Beyond the work financial firms staff today, there is nearly limitless demand in doing more, from deeper research coverage to broader economic modeling to government policy budgeting and planning.

As such, we sought to evaluate the state of model progress within finance, to see both how far we have come and where we are heading. One of AGI's end goals is to automate critical portions of finance, not only for frontier-lab ROI and enterprise efficiency, but also to further one of America's greatest strengths, its robust and liquid capital markets. Greater visibility into that development is then critical.

We propose, to our knowledge, the first meta-analysis of this kind.

Model card chronology

In particular, we first look toward the set of benchmarks that frontier labs themselves cite on their model cards. Figure 1 presents this timeline on a per-lab basis. Though not every model release includes a finance score, it is clear that finance is an increasingly important priority for the frontier labs. Of the 9 flagship launches from July through December 2025, only Anthropic's cite a third-party finance benchmark. From February 2026 onward, finance evaluations appear in launches from most major labs.

Figure 1

Finance benchmarks in frontier launches

22 flagship launches from 5 frontier labs, July 2025 through July 2026. Click any marker to open the source launch post, model card, or evaluation report.

Scores are the labs' self-reported results under their own harnesses.

Saturation since release

In addition to the finance benchmarks directly cited by frontier lab model cards, we believe there to be a broader set of evals equally representative of model performance and interest, albeit not publicly (Figure 2). We also measured their best published model score on a per-benchmark basis from initial release to today.

Figure 2

Best published score at release and now

Light bars show the best result each publisher reported at release. Solid bars show the best current result, observed July 24, 2026. Note that metrics differ by benchmark and are named on each bar. Benchmarks with one bar have no separate at-release value to show.

Two observations follow from Figure 2. First, every benchmark with a published starting point has climbed since release, in some cases by double digits within months. Second, and we believe more importantly, each benchmark remains relatively unsaturated. Even with the best frontier models, the best published scores still hover around 50 to 60%, and no leaderboard has crossed 70% on their headline metric. Models have improved quickly on these evals without coming close to completing them.

Benchmark landscape

We found that, despite these finance evals being nominally finance-specific, they all still highly vary in their construction.

The underlying reason is that finance itself is a broad, load-bearing term, with many individual sub-industries comprising many different jobs, each involving many various tasks. One way to begin defining finance is top-down, delineating between “buy-side” and “sell-side,” then across investment banking, equity research, private equity, and hedge funds, and then across asset classes within each. Another way is bottom-up, as tasks themselves can be stratified across the nature of their expected input-output relationship. Certain tasks involve just retrieval-based Q&A, for instance, whereas other longer-horizon tasks involve completed deliverables in applications like Excel or PowerPoint. Thus, finance benchmarks by construction have to be opinionated on what exact types of workflows they seek to approximate and measure model progress by.

We believe that there is little open discussion comparing such nuances of finance benchmarks, despite their active use in hillclimbing frontier models and thus deployment across financial firms. As such, we hope to provide this write-up as a valuable source of said context.

Primary differences

Of the many criteria that separate workflows within finance, two broad axes dominate. The first axis separates deterministic answers from open-ended answers. The second axis separates Q&A and retrieval-based questions from Excel, PowerPoint, and Word-based deliverables.

With this framing in mind, we plotted the set of finance benchmarks below, illustrating the primary design differences (Figure 3). It's clear that these evals, despite being finance-focused, are all quite different. The implication is that eval score comparability is only useful conditioned on further grouping or clustering based on sub-characteristics.

Figure 3

7 benchmark entries by output type and answer determinism

We assign each position from public descriptions of tasks and outputs. Each marker shape shows how information reaches a model.

Qualitative map of 7 finance benchmarks that have public materials The horizontal axis runs from question answering and retrieval, through text deliverables, to working files. The vertical axis runs from deterministic answers to open-ended ones. Open-ended Several acceptable answers Deterministic One expected answer Q&A and retrieval Text deliverables Excel, PowerPoint, and Word files Vals v1 / v1.1 May 2025 / Feb 2026 Vals v2.0 May 2026 Rogo BFB May 2026 Hebbia FSB September 2025 DiligenceBench July 2026 Mercor APEX, IB January 2026 Handshake BTB April 2026 live public sources prepared documents or workspace runner selects research method

What's more is that one might expect Q&A and retrieval-based evals to eventually give way toward longer-horizon, deliverable-based evals, with the latter being more complex. However, this has yet to be the case. Vals' v2.0 Finance Agent, Rogo's BigFinanceBench, and Thoughtful Lab's DiligenceBench (Q&A and retrieval-based benchmarks) all released after Mercor's APEX-Agents and Handshake's BankerToolBench (deliverable-based benchmarks), to similarly low levels of model saturation. At the moment, the output axis of whether a deliverable is involved or not does not seem to proportionally affect task difficulty.

Detailed overview

Importantly, the characteristics distinguishing finance benchmarks are many, far beyond just the two axes presented above. Below, in Table 1, we detail each benchmark in further context. Variables include: the information provided to the model at test-time, the tools and harness available, the required output, the grading method and verifier, and the released task-level artifacts.

Note that we encourage readers to read the original source documents provided by each benchmark constructor to get the clearest, task-level context of what models are being tested on.

Comparison of seven finance benchmark entries by release, corpus, work, information, tools, output, grading, and released artifacts.
Attribute Vals Finance Agent v1.0 / v1.1 Hebbia Financial Services Benchmark Mercor APEX-Agents, IB Handshake BankerToolBench Vals Finance Agent v2.0 Rogo BigFinanceBench DiligenceBench
Release May 2025v1.1 refresh Feb 2026 September 2025 January 2026 April 2026 May 2026 May 2026 July 2026
Task corpus 537 total337 test, 50 public 600+ questions 160 investment-banking tasks 100 tasks 927 total450 test, 27 public 928 total50 public 150 tasks
Work Research questions on public filings and web sources Extraction, summarization, and reasoning on finance documents Long investment-banking tasks in prepared workspaces Complete investment-banking tasks in a fixed data room Research and multi-step questions on public evidence Open-book research questions with auditable calculations Open-ended equity research and investment diligence
Information Live web and SEC EDGAR Supplied document sets Ten prepared workspaces with staged files Prepared files and date-locked data services Live web, SEC filings, and market data Live web and SEC EDGAR Public evidence, method selected by the runner
Tools Web search, EDGAR search, page parsing, and retrieval Not disclosed Nine applications and 63 tools Three data services, LibreOffice, and Python Search, filings, documents, calculator, and market data Web search, EDGAR, URL retrieval, Python, and submission No fixed harness or tool set
Required output Researched text answer, number, or verdict Text answers, summaries, and analyses Console messages and saved files133 + 27 across 160 tasks Excel, PowerPoint, Word, PDF, and CSV files Researched text answer Number or conclusion with a calculation Equity-research memo
Grading method LLM judge, per-question criteriaMetric: final-answer accuracy LLM evaluator consensusWeighted per-criterion scores LLM judge, task criteria and workspace stateMetric: Pass@1 Agentic verifier opens files and checks formulas Question criteria give partial-credit and all-pass results Two judges, weighted rubric and separate final answers LLM judge, weighted criteria for facts, reasoning, and risks
Released task-level artifacts Public question subset and code repositories Methodology paper and sample questions Tasks and harness repository Tasks and repository Public question subset and code repositories Tasks, repeated trajectories, and two-judge grades Tasks and reference runner

Table 1. 7 benchmark entries with public materials. Vals v1.1, a February 2026 refresh, uses the same 537-question benchmark with updated data, harness, and evaluation. The released-artifacts row records what each publisher had posted on July 24, 2026.

Task distribution

Each benchmark report also includes a published task distribution by their categorization of choice (Figure 4). For example, Rogo's largest workflow is KPIs and unit economics, Mercor's is sensitivity analysis, and Handshake's is financial modeling and scenario work. It is important to note that the classification schemes themselves differ, with Vals classifying by question type, Mercor and Rogo by workflow, Handshake by product, and DiligenceBench by sector. Since we do not have access to the private datasets for each benchmark at this time, we rely on each constructor's own definitions rather than an independently derived, shared taxonomy.

Figure 4

Task distribution, as published by each benchmark

Self-published task distributions, with the 6 largest categories shown and the rest grouped. Vals v2.0 and Hebbia do not publish per-category task counts.

Bar width is the literal share of that benchmark's tasks in the category.

Case studies

This collection of finance benchmarks was released with varying degrees of public transparency, with Rogo's BigFinanceBench and Vals' Finance Agent being the most auditable. Since these two evals do not require any private workspaces, documents, or harnesses to conduct tests on, below, we provide preliminary case studies.

More specifically, since Rogo's benchmark releases per-model trajectories and LLM-judge grades, we audit those directly. And for Vals' benchmark, we generate the trajectories ourselves, running its 27 public v2.0 tasks.

Rogo: Public dataset (50)

Rogo's public task set includes 50 tasks (of a total of 928), alongside 3 trajectories for each of 10 models on every task. It also includes grades from 2 LLM judges, Gemini 3.1 Pro and Claude Opus 4.7.

Since Rogo ran the models and graded every answer, everything we add is secondary analysis of the released files. Our additions are the summary statistics below, a source audit of every task against the primary filings, and descriptive labels on the traces of the 42 post-QC tasks.

The 10 models cover a wide ability range. On Rogo's full 928-task leaderboard they span 22.4% to 58.8% on the rubric metric and 6.6% to 44.3% on final-answer accuracy. Both ranges reflect Rogo's questions, tools, rubrics, and harness rather than just finance ability in the abstract.

For each task we counted how many of the 10 models earn majority-vote final-answer credit under Rogo's Gemini judge, with credit meaning a correct final answer in at least 2 of 3 trials. Figure 5 plots this count across the 50 tasks and displays the 8 tasks that failed our QC tests.

Figure 5

Models earning final-answer credit per task

A model earns credit on a task when Rogo's Gemini judge marks its final answer correct in at least 2 of 3 trials. The marked area holds the 8 source-audited QC flags.

Among 42 retained tasks, 21 earn credit from at most 2 of the 10 models and 14 from 8 or more. Only 7 land in between.

The distribution splits toward the tails, weighted toward tasks that either most models pass or most models fail. Importantly, no model earns credit on 24 of the 50 tasks, and even though we found 8 source-audit failures sit in that group, removing them still leaves 16 zero-credit tasks among the 42 we retain.

Rogo: QC findings (8)

After our quality control audit on Rogo's 50 public tasks, we discovered 8 tasks where we believe there was prompt-rubric misspecification and/or errors within the reference answers themselves. For example: rubric lines that contradict each other, a quarter pulled from the wrong fiscal calendar, or a cost ratio graded as a gross margin. Figure 6 steps through all 8, with one condensed proof per card.

Figure 6

8 excluded tasks, one proof each

Use the arrows to step through the 8 source-audited QC flags.

Rogo: Failure modes

For the remaining 42 retained tasks, we analyzed the released trajectories and assigned each task one primary failure mode. Figure 7 below illustrates the failure modes.

Figure 7

Failure modes across the 42 retained tasks

Each task carries one primary failure mode. The light part counts tasks where the failure appears anywhere.

Failure modes are our descriptive coding of the released traces.

Our analysis marks Misunderstood definition or basis as the primary failure mode on 20 of the 42 tasks, more than double any other condition. Each task carries exactly one primary failure mode, and counting the tasks where it appears at all, primary or not, raises that to 31 of the 42. This failure mode covers choices the prompt leaves open, including definitional points like gross versus net, reported versus adjusted, period-end versus period-average, or which entity the question intends. Broken calculation chains, missed line-item builds, and missed off-filing sources all are other examples of common failure modes.

For more detail on how a model exhibits one of these failure modes specifically, we follow one retained task end to end.

Rogo: Task deep dive

Below, we take a closer look into one Rogo task in detail. The task below asks for the change in Remitly's fraud loss as a share of net revenue from FY 2024 to FY 2025. Rogo's reference answer is 0.8%, which its rubric treats as a change of +0.8 percentage points. Rogo's bundle releases 30 trajectories on this task, 3 from each of the 10 models. 14 of the 30, spread across 6 models, return +0.6 points instead, including all 3 trials from each of Claude Opus 4.7, Claude Sonnet 4.6, and Gemini 3.1 Pro. Most models make the same mistake, dividing fraud losses by reported revenue. Figure 8 details the full prompt and both calculation paths.

Figure 8

One task, two defensible calculations

Dollar values are in thousands. We calculate each ratio directly from values in Rogo's reference material.

Rogo task bf-77a2cbc3ee“What is the change in fraud loss as a share of net revenue from FY 2024 to FY 2025 for Remitly? Answer in % with only the final answer rounded to one decimal place.”

Repeated model path

Use reported revenue

  1. FY 202584,166 ÷ 1,635,147 = 5.147%
  2. FY 202457,929 ÷ 1,263,963 = 4.583%
  3. Change0.564 points

Rounded answer: +0.6pp

Rogo reference path

Calculate revenue minus transaction expenses

  1. FY 202584,166 ÷ (1,635,147 − 549,480) = 7.752%
  2. FY 202457,929 ÷ (1,263,963 − 431,604) = 6.960%
  3. Change0.793 points

Rounded answer: +0.8pp

Remitly's FY 2025 10-K labels the deducted line transaction expenses and never defines net revenue.

The miss is a model error, but a specific kind of one. The task never defines net revenue, and neither do Remitly's cited filings, which report no line under that name. A model dividing by reported revenue therefore reaches +0.6 on internally consistent arithmetic. Rogo's rubric, however, derives the denominator the way a real-world analyst would be expected to, implicitly understanding to net out transaction costs, and that answer yields +0.8.

We call this a definition fork, the point where two internally consistent calculations separate on an unstated definition. Neither path contains an arithmetic mistake. The fork sits upstream of the math, at the moment the model decides what the words in the prompt mean. Importantly, we still consider this a valid, logically consistent task. Part of a real analyst's job is resolving exactly this kind of context without it being spelled out, and 14 of the 30 trajectories failing to do so is a real capability signal.

Rogo's printed intermediate ratios carry small arithmetic slips of their own, 7.756% and 6.956% where direct calculation gives 7.752% and 6.960%. Both versions round to the same +0.8 change and decide nothing. The narrower point is about reading a leaderboard. A miss here is a definitional miss rather than a calculation miss, and separating the two takes the sources and trajectories behind the grade rather than the grade alone.

Vals v2.0: Public dataset (27)

Vals published no run evidence, so we produced our own. In July 2026 we ran all 27 public v2.0 tasks with MiniMax M3 as the solver, GLM 5.2 as the judge, and 3 trials per task. That produced 81 solver traces and 80 judge grades, 1 grade missing.

MiniMax M3 passed 40 of the 80 graded trials with a mean rubric score of 0.77. GLM 5.2 attributed the failures to wrong interpretation on 12 trials, calculation error on 11, incomplete answers on 9, no answer at all on 5, and wrong sources on 3. Task-level results vary significantly: 7 of the 27 tasks pass all 3 trials, 8 pass none, and no comparables or financial-modeling task passes all 3.

We audited all 27 task specifications against primary sources, the same audit we applied to Rogo's public 50. 13 tasks come through valid, 4 carry minor defects, 9 carry material defects, and 1 is fatally misspecified. Interestingly, the defects belonged to different parts of the correctness distribution here. On Rogo, every broken task sat in the zero-credit bar. On Vals, 2 of the 7 all-pass tasks carry defects, 1 material and 1 fatal, so a task the model always solves can still fail the audit.

Vals v2.0: Failure modes

We read all 81 traces and coded the 20 failing tasks with the same failure modes used for Rogo. Figure 9 tallies the results.

Figure 9

Primary failure modes across the failing Vals tasks

One primary failure mode per failing task, 20 tasks in total. Rows and order match Figure 7 so the two benchmarks read side by side.

Both sets are our coding, read from Rogo's released trajectories on one side and from our own runs on the other.

Misunderstood definition or basis leads both tallies, primary on 20 of Rogo's 42 retained tasks and on 7 of the 20 failing Vals tasks. On Vals, the most frequent failure modes included model errors on interpreting lease rent versus financing interest, gross versus net interest, and more generally, which entity perimeter a question intends. These are the same species of fork the Remitly task isolates.

What's important to note is that across two benchmarks sharing no tasks, models, judges, or harness, the same leading failure modes are present. A leaderboard gap on frontier model cards can mean any of a capability gap, a definition mismatch, or simply a defective task, and distinguishing the three requires task-level artifacts that most publishers do not release.

Limitations

This article is v1.0 of a versioned reference, published in July 2026 from publisher materials we accessed on July 24, 2026. Later readings, Rogo's July 28 leaderboard figure, Hebbia's materials accessed July 29, and the Figure 1 launch-artifact checks completed July 29, are dated where they appear. Our landscape map and comparison table will refresh in later versions under the same definitions. Both case studies stay frozen, one on Rogo's May 2026 public-50 release and one on our July 2026 Vals runs, and later versions will note changes rather than edit them silently. Cite this article as Finance Eval Meta-Analysis, v1.0, July 2026.

Rogo produced and owns the released model runs and grades. We did not run or grade any model on Rogo's tasks, and no part of our Rogo analysis independently replicates Rogo's results. Figure 1 quotes each lab's launch materials as written, with sources recorded per entry in the model-card record, so its scores are the labs' self-reported results rather than the benchmark publishers' gradings. Figure 2's at-release values likewise come from each publisher's launch materials. We placed each benchmark in Figure 3 by editorial judgment from public task and output descriptions, without measuring anything. The task-distribution panels in Figure 4 restate each publisher's own published category counts. Hebbia's landscape entry rests on its September 2025 release post and methodology paper, accessed July 29, 2026, the only materials it publishes. We assigned the failure modes by reading trajectories, and we established no cause experimentally.

Our Vals case study rests on runs we executed ourselves in July 2026, with MiniMax M3 as the only solver, GLM 5.2 as the only judge, and 3 trials per task. We kept the run manifests in the repository. Those runs give us model-failure evidence, and only that. They do not support comparison with Vals' published leaderboard, which uses Vals' agent, judges, and configuration, none of which we reproduce.

Retained means we can use a task for analysis, not that the task is free of every issue. Our retained Rogo set contains 26 valid tasks, 9 with convention or clarity notes, and 7 with rubric errors we judge nondecisive. Some export rows share a model, task, and trial key, and where a matching grade existed we kept the completed retry over the infrastructure-error row. That choice changes no majority-vote count in Figure 5.

Dataset access

Our datasets target the layer this article ends on, the definition forks and adjacent failure conditions that headline scores cannot show. The first datasets hold research and derivation tasks in the style of Rogo and Vals, and a later offering can extend to workspace and artifact deliverables. Every task and environment is authored independently. We offer the datasets privately, with no public release. Get in touch [here] to access our off-the-shelf finance post-training datasets, or reach out to us directly at [email].

Sources

Hebbia Financial Services Benchmark release postmethodology paper
Handshake BankerToolBench paperrepositorydataset
Rogo BigFinanceBench current leaderboardpaperrelease postdataset