Quantitative ResearcharXiv q-finSIGNAL 92E056

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

This study introduces PortBench, a benchmarking framework designed to evaluate large language models (LLMs) on portfolio management tasks. Unlike existing benchmarks, PortBench covers six heterogeneous asset classes and includes a static QA dataset and a dynamic five-stage allocation pipeline, along with two novel metrics to assess correlation and error compounding, providing a more comprehensive evaluation for real-world financial decision-making.

01 ABSTRACT

PortBench is a new benchmark for evaluating LLMs in portfolio management. The authors constructed a static QA dataset of 6,269 questions across seven templates and a dynamic five-stage allocation pipeline. They proposed two metrics: a dual-layer correlation score for inter-asset hedging and intra-class concentration, and CEPS to quantify error compounding across stages. Experiments with ten frontier LLMs across four market periods showed that only 32.5% of 120 evaluations beat equal weighting on Sharpe, indicating strong QA performance does not guarantee superior portfolio performance. The benchmark supports real-time evaluation to mitigate pretraining contamination, but results depend on chosen models and market windows.

02 KEY FINDINGS

  1. Proposes PortBench covering six asset classes (2015-2025) with static QA (6,269 questions) and dynamic five-stage pipeline.
  2. Introduces dual-layer correlation score (cross-class hedging and intra-class concentration) and CEPS metric for error compounding.
  3. Evaluates under three stress windows and three risk profiles, supports real-time evaluation to reduce contamination.
  4. Evaluation of ten LLMs shows only 32.5% of runs beat equal weighting on Sharpe.
Return to the primary source

AI GENERATED SUMMARY / DISCOVERED BY ARXIV Q-FIN

Read original