Abstract
Financial AI benchmarks have multiplied faster than our agreement on what they should measure. Most inherit the shape of a human job, scoring a system on how faithfully it reproduces a legacy work product such as a research report, a valuation model, or a compliance memo. We argue this is the wrong target. The analyst's workflow is a contingent solution to human bounded cognition, limited bandwidth, and institutional control. It is an organizational technology, not the essence of finance, so a benchmark faithful to it certifies competence at the artifact rather than at the economic function the artifact was built to serve. Our position is that evaluation should be organized around that function, which we formalize as capital–asset matching: the alignment of capital-side demand with asset-side supply along six shared dimensions, namely return, risk, horizon, liquidity, payoff form, and constraints, under institutional restrictions. Matching is the keystone of a small set of invariant economic functions that endure as workflows change, and we develop it in full while mapping the rest under the same lens. From it we derive a finance-native framework whose three capability layers close one decision loop: understand what the capital needs, read assets by economic substance, and communicate an auditable verdict in which a justified no-match is a first-class outcome.
Framework Overview
Figure 1. The Capital–Asset Matching framework. Left: legacy workflow-replication benchmarks. Center: the matching predicate aligning capital-side demand with asset-side supply along six dimensions. Right: the finance-native benchmark design with layered scoring architecture.
Key Contributions
Capital–Asset Matching
A formalization of the core economic function: aligning capital-side demand with asset-side supply along six dimensions (Return, Risk, Horizon, Liquidity, Payoff Form, Constraints).
Three-Layer Framework
A finance-native benchmarking framework: (A) demand understanding, (B) broad asset analysis by economic substance, (C) auditable decision communication.
Asset Coverage Map
A six-class asset taxonomy read as a coverage map rather than a label set, providing principled design for benchmark item construction.
Layered Scoring Architecture
Four separate scoring axes — correctness, suitability, robustness, and compliance — that treat automated judges as measured instruments rather than oracles.
Empirical Probe
Evidence via 2025–2026 results dissociation and pre-registered paired-framing experiments that workflow competence does not imply matching competence.
Research Agenda
A prioritized research agenda covering demand elicitation, cross-asset comparability, fiduciary reasoning, and real-money evaluation safety.
Paper Structure
- 1 Introduction: The Benchmarking Problem in Financial AI
- 2 Financial Workflows as Historical Artifacts of Human Constraints
- 3 The Functional Map of Finance: What Endures When Workflows Do Not
- 4 The Economic Core of Finance: Capital–Asset Matching
- 5 A Finance-Native Benchmarking Framework
- 6 Asset Taxonomy and Benchmark Coverage
- 7 Benchmark Task Design
- 8 Evaluation Metrics and Scoring Architecture
- 9 Illustrative Benchmark Cases
- 10 Empirical Probe: Does Workflow Competence Imply Matching Competence?
- 11 Anticipated Objections
- 12 Limitations, Risks, and Open Research Questions
- 13 Research Agenda
- 14 From Blueprint to Platform: A Living Benchmark Observatory
- 15 Conclusion
- A The Matching Schema: Demand, Supply, and Verdict
- B The Four-Axis Scoring Rubric
- C The Asset-Dossier Tool Interface
Citation
@article{lu2026rethinking,
title={Rethinking Financial AI Benchmarks: From Workflow Replication to Capital--Asset Matching},
author={Lu, Jiacheng and Wang, Sinuo and Zhao, Wentao and Sun, Rui and Luan, Beidi and Song, Tao and Li, Jing and Jiang, Daxin and Hua, Cheng and He, Yijia and Wang, Weijian and Yang, Qixuan and Li, Nan and Tu, Jun and Jin, Yu and Wu, Zhengze and Guan, Haibing and Bai, Zuo},
year={2026}
}