What Survives Honest Evaluation? Leakage-Safe, Search-Aware Assessment of LLM-Driven Trading Strategy Discovery
ORIGINAL / What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery
A new framework structurally corrects common methodological flaws in LLM-driven trading strategy research, namely look-ahead bias and unaccounted search intensity. By excluding look-ahead via registry-validated tools and recording all evaluations to deflate performance, it provides a more realistic assessment of LLM-discovered strategies and quantifies evidence thresholds for passive benchmarks and human rules.
01 ABSTRACT
The paper presents a strategy-discovery system that structurally avoids look-ahead bias and logs all search attempts. Empirical results show that no LLM-discovered strategy survives honest evaluation, while passive benchmarks pass. The authors argue that pre-registered hypotheses require lower evidence bars than brute-force search and emphasize the large sample sizes needed for credible certification.
02 KEY FINDINGS
- Registry-validated tools structurally eliminate look-ahead, and statistical corrections are insufficient without them.
- All strategy evaluations are logged and used to deflate reported performance, highlighting the impact of search intensity.
- On two datasets (453 US stocks and 39 ETFs), honest evaluation rejected all LLM-discovered strategies and certified passive benchmarks.
- A leaky oracle with Sharpe 35 survives Deflated Sharpe and PBO tests, showing the need for structural guards.
- Quantifies the large sample sizes required for credible certification of moderate effects and compares human trading rules.
AI GENERATED SUMMARY / DISCOVERED BY ARXIV Q-FIN