Experiments we ran ourselves — full data, real inference costs, and the failures included. Not customer stories. Lab reports.
Four models, $50K each, 130 trading days. An 8B model finished at +72.6% and won by 71 points. Then we re-ran every model five times on the identical days. Each model's run-to-run swing (13–19 points) turned out to be twice the gap between models (7 points). The leaderboard was noise. The only reproducible difference was cost — and the most expensive model cost 121x more per decision than the cheapest.