“runnable artifacts are not the same as faithful operationalization”
LLMs can write plausible prediction-market trading strategies, but a strategy that runs is not one that means what its author intended. This benchmark pairs 581 source-grounded strategy descriptions with 1.28 million Polymarket markets and 183.6 million executions, then asks whether a model can turn a coarse economic idea into faithful, executable code. Strict end-to-end validity is rare, and fees and liquidity reshape trading paths rather than merely taxing returns.
Extensive technical background assumed
Platforms mentioned: Polymarket