My LLM Trading Bots Beat the Benchmark. The Rules I Wrote Beforehand Say It Doesn't Count.

My LLM Trading Bots Beat the Benchmark. The Rules I Wrote Beforehand Say It Doesn't Count.

For two months I let LLM agents paper-trade a simulated $100 crypto portfolio. Season 1 ended with a verdict of STOP: both agents lost to cash. Season 2 ended with all four agents beating the passive benchmark by 3 to 10 percentage points. The formal verdict for Season 2 is still “inconclusive”, and I think that is the most useful result of the whole project. No real money ever moved. Live execution was never implemented and the code fails closed if you try to turn it on. Everything below is CoinGecko-reference paper returns on Base L2 tokens, priced with a static cost model (0.05 to 1.00 % fee tier plus a 0.15 to 0.50 % execution haircut plus $0.03 gas per swap). ...

September 11, 2026 · 6 min · nunc