PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
This study introduces PortBench, a benchmarking framework designed to evaluate large language models (LLMs) on portfolio management tasks. Unlike existing benchmarks, PortBench covers six heterogeneous asset classes and includes a static QA dataset and a dynamic five-stage allocation pipeline, along with two novel metrics to assess correlation and error compounding, providing a more comprehensive evaluation for real-world financial decision-making.
01 ABSTRACT
PortBench is a new benchmark for evaluating LLMs in portfolio management. The authors constructed a static QA dataset of 6,269 questions across seven templates and a dynamic five-stage allocation pipeline. They proposed two metrics: a dual-layer correlation score for inter-asset hedging and intra-class concentration, and CEPS to quantify error compounding across stages. Experiments with ten frontier LLMs across four market periods showed that only 32.5% of 120 evaluations beat equal weighting on Sharpe, indicating strong QA performance does not guarantee superior portfolio performance. The benchmark supports real-time evaluation to mitigate pretraining contamination, but results depend on chosen models and market windows.
02 KEY FINDINGS
- Proposes PortBench covering six asset classes (2015-2025) with static QA (6,269 questions) and dynamic five-stage pipeline.
- Introduces dual-layer correlation score (cross-class hedging and intra-class concentration) and CEPS metric for error compounding.
- Evaluates under three stress windows and three risk profiles, supports real-time evaluation to reduce contamination.
- Evaluation of ten LLMs shows only 32.5% of runs beat equal weighting on Sharpe.
AI GENERATED SUMMARY / DISCOVERED BY ARXIV Q-FIN