Published on: 2026-09-02
Updated on: 2026-09-02
AI trading bots can now read filings, scan thousands of instruments and write trading code, yet the evidence that these capabilities produce better returns remains mixed. Machine intelligence has improved trading inputs, including research, speed, and automation, without reliably improving the output that matters: risk-adjusted returns after costs. That gap shows where human judgment may still add value.

AI has improved sharply at market analysis, coding and automation, but consistent, benchmark-beating trading performance remains unproven.
AIEQ, one of the longest-running AI-driven equity ETFs, returned 4.5% annualised over the five years to June 2026 versus 13.4% for the S&P 500 Total Return Index, according to Morningstar data.
Backtests can flatter machines: overfitting, data leakage and ignored trading costs can make a losing system look like a winner on paper.
Human judgment can still add value in trade selection, risk control and deciding which questions are worth answering, with AI handling more of the repetitive analytical work.
Three capabilities have moved fastest.
Processing: Current models can read financial statements, news flow and price data at a scale no individual analyst can match.
Research: They can compare scenarios, write trading code and help test ideas rapidly.
Autonomy: Agent-style systems can plan tasks, call tools and take actions instead of only answering prompts, allowing them to run more of a workflow from end to end.
The improvement is measurable. In the July 2026 Backtrader-Bench preprint, tool-augmented frontier models reached 90% accuracy on a 30-question algorithmic-trading benchmark, compared with 73% for the strongest no-tool baseline. That demonstrates better trading-related reasoning and tool use, not better investment returns.
Solving trading problems better is not the same as making profitable trades. That claim needs different evidence.
Long-term track records for fully autonomous AI trading agents remain scarce, so the best public evidence comes from several places: AI-driven funds, controlled LLM-agent experiments and institutional backtests.
AIEQ offers one of the longest public records for AI-driven stock selection. As of June 30, 2026, its NAV had returned 4.5% annualised over five years, compared with 13.4% for the S&P 500 Total Return Index. Since its 2017 inception, the gap was smaller at 10.0% versus 11.3% annualised.
JPMorgan researchers used LLM agents to classify macroeconomic regimes and allocate between assets accordingly. In historical tests published in 2026, all eight agents reportedly beat a conventional 60/40 portfolio on a risk-adjusted basis, with the strongest exceeding it by about 0.7 percentage point annually while taking less volatility.
The important detail is that the agents operated inside a structured macro-regime allocation framework rather than receiving an unrestricted mandate to trade anything they wanted.
So the answer is more nuanced than either side of the AI debate suggests: AI can improve a trading system. Current evidence does not show that a smarter model automatically becomes a better trader.
Three reasons, and none of them is fixed simply by making the model larger.
Markets adapt. Any exploitable pattern one AI can discover may also be discovered by well-funded quantitative firms and competing systems. Once enough capital trades the same opportunity, prices adjust, and the edge can weaken.
The backtest is also easier than the market. AI is extremely good at finding patterns, including relationships that only worked by chance. Overfitting, historical-data leakage and unrealistic execution assumptions can all make performance look stronger than it really is.
Small trading costs can become significant in a high-turnover strategy. A system generating a thin theoretical edge across hundreds of trades may see much of that edge disappear once spreads, slippage and fees are applied repeatedly.
Direction is only part of the job. A model can be right more often than wrong and still lose money because position size, exit timing, average loss versus average win and execution costs determine the final result.
Software can scan thousands of instruments, monitor markets continuously, automate calculations and repeat defined workflows without fatigue. A human will not out-read a system capable of processing hundreds of reports while simultaneously monitoring price data and news.
That changes the useful question from:
How can a trader process information faster than AI?
to:
Where can human judgment still improve the trading process?
CFA Institute has identified hybrid human-machine decision-making as an emerging model for investment processes. The human role increasingly lies in deciding which information deserves attention, whether an opportunity is attractive enough to justify the risk, and when changing conditions make a previously valid assumption unreliable.
You will not beat AI at the things AI already does well. The opportunity sits where speed matters less and judgment, selectivity and risk control matter more. Four practices follow.
Being the five-thousandth participant to react to the same inflation print is not an edge. Better questions are: what is already priced in, which assumption could be wrong, what would kill the thesis, and whether the reaction is temporary or structural. Retrieving information is becoming cheaper. Choosing the right question to answer is still valuable.
More signals do not require more trades. A well-designed bot can stay flat too, but the trader still has to decide what evidence deserves capital and when changing conditions make an otherwise valid signal unreliable. Define the thesis, entry condition, invalidation and exit before committing money, and skip anything that cannot satisfy all four.
Before entry, know how much you are prepared to lose, what proves the idea wrong and whether several positions share one underlying exposure. Decide in advance when the strategy itself no longer deserves your trust. AI can generate a view. It cannot remove the consequences of allocating too much capital to a wrong one.
Instead of asking AI only to generate a trade, use it to challenge one. Ask for the bearish case against your bullish view, the assumptions hidden inside your thesis, alternative explanations for the same data and a testable version of your trading rule.
Robeco’s August 2026 work on agentic AI makes a similar argument: access to a model alone is unlikely to create alpha. Differences may instead emerge from data quality, validation history, portfolio constraints, mandate boundaries and human sign-off.
Any bot pitch or backtest screenshot should survive five questions before it influences a trading decision.
Question |
Why it matters |
What did it beat? |
Profit means little without an appropriate benchmark. |
Are costs included? |
Spread, slippage and fees can erase small edges. |
Was it tested on unseen data? |
Testing on familiar data increases the risk of overfitting or memorisation. |
What was the maximum drawdown? |
Returns can hide how much risk produced them. |
Has it traded live? |
Paper and backtested results are not the same as real execution. |
A 40% backtested return tells you surprisingly little until you know all five answers.
Not on the evidence available so far. Individual systems can outperform over particular periods, but persistent gains after costs against a fair benchmark remain difficult to demonstrate. For example, AIEQ returned 4.5% annualised over the five years to June 30, 2026, versus 13.4% for the S&P 500 Total Return Index, although its since-inception gap was much smaller.
At rapidly processing large amounts of information, monitoring markets and repeating defined analysis, often yes. Human judgment can still contribute by interpreting unusual conditions, questioning assumptions, deciding how much confidence a signal deserves and determining when evidence is too weak to justify risk.
AI can be useful as a research assistant for summarising information, coding, testing rules and stress-testing a thesis. Allowing it to execute financial decisions autonomously is a different threshold. FINRA’s 2026 oversight report highlights AI agents because they can plan, decide and act with greater autonomy, while stressing the importance of oversight, guardrails and monitoring.
AI trading bots are getting smarter, but greater intelligence does not automatically produce better trading returns. Markets adapt, backtests can mislead, and even an accurate forecast can become an unprofitable trade when execution and risk are poorly managed.
Trying to outrun AI at information processing is unlikely to be a useful contest. Use the machine to research faster, test assumptions and argue against your own thesis, while keeping responsibility for trade selection and risk sizing with you. The advantage is less about being smarter than AI and more about using it without confusing better analysis with a guaranteed trading edge.