Can ChatGPT trade profitably? Since January 2026, TradeRank Arena has tested that question continuously: give the newest models simulated capital, live prices and full autonomy, then let them trade against one another for a month. Across seven seasons, 49 models and 2,686 trades, it offers the closest controlled evidence we have on whether a GPT trading bot — or any frontier-model AI trading agent — can make money.
The models improved dramatically over those months; the trading record did not. Only 41.5% of model-seasons finished profitable — fewer than half. No model stayed good: Season 1's winner finished last in Season 2. In Season 5, closed June 2026, the lighter, cheaper tiers swept the podium while the flagship reasoning models — GPT-5.5, Claude Opus and Grok — finished 8th, 6th and 7th.
TradeRank runs the arena with paper trading — fills at a fetched price, no market impact — so if anything it flatters the models. Whatever separates winning from losing here, it isn't intelligence.
What separated the best AI trading agents from the rest
TradeRank's post-mortems across 656+ trades point to three factors — not model intelligence — that divided winners from losers:
Patience. MiniMax M2.5 won a season trading 16 times. The model that traded 100 times finished last.
Sizing. In one season a 17% win rate beat a 90% win rate, because the size was in the right trades. In another, GPT-5 Mini posted the highest win rate in the field — 55% — and still lost money.
Costs. Gemini burned ~$50 of a $10,000 book on fees in two weeks; the busiest models gave up to 1.1% of their portfolio to transaction costs. In a simulation with zero slippage. Live is worse.
Patience, sizing, exits and cost control all sit in the execution layer. That is precisely the layer a chat model does not possess on its own. The useful question is therefore not which model reasons best, but what trading system surrounds it.
The same pattern appears beyond one arena. When Agents Trade (ACM Web Conference 2026) tested four agent architectures across five model backbones. Architecture drove materially different trading behaviour; the backbone explained far less of the outcome variation. Swap the model and little changes. Swap the scaffolding — sizing rules, action frequency and exit logic — and everything changes.
Why LLM trading does not turn reasoning into returns
The thesis is the easy part. A ChatGPT trading bot can produce a sharper trade rationale than most desk research, and an AI crypto prediction can sound entirely defensible. Accounts rarely fail because the thesis was badly written; they fail because losers are held out of hope and winners are cut out of fear. That is the execution gap the arenas keep exposing.
Everyone gets the same answer. When much of the market prompts the same few models on the same public data, conclusions converge and the edge closes. An inference anyone can reproduce for a few cents is consensus, not private information — and this worsens as models improve.
You can't validate before risking money. Systematic trading rests on out-of-sample testing: hold back data the strategy never saw. Impossible with a frontier model — it trained on data containing the outcomes. Ask about March 2020 and it knows how March 2020 ended. The lookahead bias is in the weights.
How AlphaNet engineers the execution layer
If execution determines returns, engineering goes there.
AlphaNet treats this as a quantitative trading problem. Signals come from a hybrid engine: researcher-designed market-structure factors competing alongside deep learning models trained on structured market data, with sizing and exits governed at the strategy layer. Order execution runs on adaptive deep-learning and reinforcement-learning algorithms that work orders into the book rather than paying the spread.
Three things follow.
Validation before deployment. Our models train on data we control, so data can be held back. Walk-forward testing on genuinely unseen periods, net of fees and modelled slippage. Not proof, but a real out-of-sample process, which no frontier-model agent can have.
A private edge. Proprietary models on a proprietary feature set. What they find isn't for sale by the token.
Verifiable results. Strategies trade on our DEX, so every fill is on-chain. Users can check the live trades and performance without trusting us.
It's also why the strategies run as Autopilot — fully systematic, end to end. The arenas show exactly where discretionary judgement fails: entries, sizing, exits. Those are precisely the layers our strategies never hand back to a human mid-trade.
Are AI trading bots reliable? Three questions to ask
What handles sizing and exits?
If it's the same model writing the thesis, the layer that decides outcomes wasn't engineered.
How was it validated before going live?
A backtest of a model trained on the outcomes is not validation.
Can I verify the live results myself?
A wallet address is an answer. A dashboard screenshot isn't.
Sources: TradeRank Arena season records and post-mortems (traderank.ai — note TradeRank operates the arena it reports on); Nof1 Alpha Arena Season 1 (nof1.ai, on-chain on Hyperliquid); "When Agents Trade: Live Multi-Market Trading Arena for LLM Agents," ACM Web Conference 2026. Model architecture and validation methodology: Section C of the AlphaNet whitepaper.
Get AlphaNet in your inbox.
Get the latest AlphaNet research, strategy reports, and market briefs delivered straight to your inbox.
