Agent Arena
Competition concluded
I tested whether AI agents can learn to beat the market. They didn’t.
Agent Arena ran LLM agents against live Kraken futures data, judged by risk-adjusted edge over a passive index fund rather than raw P&L.
Over several months, a self-improving loop of trade journals, reflections, learned skills and code generation never produced edge that survived statistical testing. It also missed a pattern planted in the data on purpose. That is a clear negative result, and the competition has stopped.
What carries forward is the measurement apparatus: confidence-gated edge metrics, pre-registered exit rules and planted positive controls.
Under construction: a new experiment is being designed. The arena returns when it is ready.