Rethinking Financial AI Benchmarks

From Workflow Replication to Capital–Asset Matching

Position Paper · 2026

Jiacheng Lu1,2, Sinuo Wang4,2, Wentao Zhao5,2, Rui Sun2,†, Beidi Luan2, Tao Song1, Jing Li2, Daxin Jiang2, Cheng Hua1, Yijia He6, Weijian Wang1, Qixuan Yang1, Nan Li1, Jun Tu1, Yu Jin7, Zhengze Wu1,7, Haibing Guan1,†, Zuo Bai2,3,†

1Shanghai Jiao Tong University   2StepFun   3Finstep   4University of Adelaide   5Tsinghua University   6Peking University   7Foresight Fund

Corresponding Authors  ·  Contact: hbguan@sjtu.edu.cn, baizuo@stepfun.com

View PDF

Abstract

Financial AI benchmarks have multiplied faster than our agreement on what they should measure. Most inherit the shape of a human job, scoring a system on how faithfully it reproduces a legacy work product such as a research report, a valuation model, or a compliance memo. We argue this is the wrong target. The analyst's workflow is a contingent solution to human bounded cognition, limited bandwidth, and institutional control. It is an organizational technology, not the essence of finance, so a benchmark faithful to it certifies competence at the artifact rather than at the economic function the artifact was built to serve. Our position is that evaluation should be organized around that function, which we formalize as capital–asset matching: the alignment of capital-side demand with asset-side supply along six shared dimensions, namely return, risk, horizon, liquidity, payoff form, and constraints, under institutional restrictions. Matching is the keystone of a small set of invariant economic functions that endure as workflows change, and we develop it in full while mapping the rest under the same lens. From it we derive a finance-native framework whose three capability layers close one decision loop: understand what the capital needs, read assets by economic substance, and communicate an auditable verdict in which a justified no-match is a first-class outcome.

Framework Overview

Capital–Asset Matching Framework: From legacy workflow replication to a finance-native benchmarking framework organized around capital-side demand understanding, broad asset analysis by economic substance, and auditable decision communication.

Figure 1. The Capital–Asset Matching framework. Left: legacy workflow-replication benchmarks. Center: the matching predicate aligning capital-side demand with asset-side supply along six dimensions. Right: the finance-native benchmark design with layered scoring architecture.

Key Contributions

Capital–Asset Matching

A formalization of the core economic function: aligning capital-side demand with asset-side supply along six dimensions (Return, Risk, Horizon, Liquidity, Payoff Form, Constraints).

Three-Layer Framework

A finance-native benchmarking framework: (A) demand understanding, (B) broad asset analysis by economic substance, (C) auditable decision communication.

Asset Coverage Map

A six-class asset taxonomy read as a coverage map rather than a label set, providing principled design for benchmark item construction.

Layered Scoring Architecture

Four separate scoring axes — correctness, suitability, robustness, and compliance — that treat automated judges as measured instruments rather than oracles.

Empirical Probe

Evidence via 2025–2026 results dissociation and pre-registered paired-framing experiments that workflow competence does not imply matching competence.

Research Agenda

A prioritized research agenda covering demand elicitation, cross-asset comparability, fiduciary reasoning, and real-money evaluation safety.

Paper Structure

Citation

@article{lu2026rethinking,
  title={Rethinking Financial AI Benchmarks: From Workflow Replication to Capital--Asset Matching},
  author={Lu, Jiacheng and Wang, Sinuo and Zhao, Wentao and Sun, Rui and Luan, Beidi and Song, Tao and Li, Jing and Jiang, Daxin and Hua, Cheng and He, Yijia and Wang, Weijian and Yang, Qixuan and Li, Nan and Tu, Jun and Jin, Yu and Wu, Zhengze and Guan, Haibing and Bai, Zuo},
  year={2026}
}