We ran about 90 strategies on Kalshi for ten days. No edge held up.

Last updated: October 6, 2026

From August 27 to September 6, 2026, we ran about 90 trading strategies on Kalshi, each one an AI agent with a single written strategy. They made 4,200 bets that settled. Five strategies lost by enough to say so with statistics, and no early lead lasted as its sample grew. All of it was on paper, so no money was at risk.

How the test worked

A paper test is only worth something if the fills are honest. Each agent read the live Kalshi order book and placed paper orders that filled against that real book, with Kalshi's real fee schedule, and its positions settled when Kalshi settled the markets. When a fill was disputed, we checked it against Kalshi's own record of trades, and every one matched. The paper trading page explains how the fills work.

We also ran a control: an agent that picked sides with a coin flip. Over more than 250 bets it lost about 0.4¢ a contract, which is about what the fees cost. That told us the measuring stick was straight.

What lost

  • Buying 89–95¢ “sure things” in thin entertainment markets: −36¢ a contract (q = 0.009 after correcting for the number of strategies tested). The worst idea in the test: one miss erases many wins.
  • Favorites in the last hour before a game: priced at 87%, they won 80% of 139 bets, for −6.5¢ a contract (p = 0.0497).
  • Momentum, mostly on the 15-minute crypto markets: −3.35¢ a contract over 781 bets (p = 0.035), counted on day six. The momentum write-up has the details.
  • Trading order-book imbalance: lost.
  • Resting wide bids to catch dips: lost.

All five lost at p < 0.05 on their own. Only the first survives a correction for testing many ideas at once, so call the other four very probably losers.

What looked like a winner and wasn't

  • Bitcoin momentum was up $16 on its first day. The momentum family it belonged to is in the list above.
  • Late favorites on ether showed ten points of edge after 16 bets. It did not last.
  • Betting the leader while a game was in play won 85% of 145 bets. Its lead shrank toward zero as the sample grew.
  • The last one standing was betting the leader in cricket matches in play: it won 94% of 32 bets against 86% implied by the prices (p = 0.075), in a market too small to matter. That is plausibly luck.

So some individual agents did finish with positive paper P&L, and the leaderboard on our home page shows agents that are ahead. Each lead rests on a small sample, like the cricket agent's 32 bets. With about 90 strategies running, a few will be ahead by chance alone. What the test found is that no early lead held up as its sample grew, and that the five results strong enough to pass p < 0.05 were all losses.

What we took from it

Count the bets before you trust a win rate. The early leaders looked good at 16 bets. As their samples grew, every one of them drifted back toward zero. If a strategy is up after a week, the first question is how many bets that is.

High hit rates hide the tail. A contract bought at 90¢ wins 10¢ and loses 90¢, so one loss gives back nine wins, and the fee comes on top.

The liquid markets are fair to within fees. Against price patterns, order-book structure, resting orders, public settlement sources, and an AI model's judgment layered on any of them, we found nothing. If there is an edge left, it is thinner than about 2¢ a contract, which is below what this test could measure at a reasonable cost.

Backtests break at the fill. On a 15-minute market the price can move between the moment a signal fires and the moment an order arrives. A backtest that fills at the price it saw assumes a trade that may not have been there.

Caveats

  • It was exploratory. The strategies went in over three waves, not as one fixed plan, and on day six we paused 20 strategies that were losing. Their bets stay in the totals above.
  • Only one result survives a multiple-testing correction: the “sure things”. With about 90 strategies, a few results under p = 0.05 are expected by chance alone.
  • The momentum p-value is weaker than it looks. Bets in 15-minute windows close together in time are not independent, and we did not cluster the test for that.

What the academic data says

Constantin Bürgi, Wanying Deng and Karl Whelan of University College Dublin studied Kalshi's prices from its start in 2021 through April 2025, in markets open at least 24 hours: 313,972 contract prices from 46,282 contracts (“Makers and Takers: The Economics of the Kalshi Prediction Market”, CESifo Working Paper No. 12122, January 2026). They found a favorite–longshot bias. After fees, contracts bought for 10¢ or less lost over 60% of the money put into them on average, and contracts above 70¢ earned small positive returns. Before fees, the average contract returned −20%. Makers did better than takers, with average returns of −9.64% against −31.46%.

Test your own idea

We ran all of this on Windmill. You describe a strategy in plain English and an AI agent trades it on paper against the live order book, with Kalshi's fees, and every run shows what it read, what it reasoned and what it ordered. If you have a strategy you think is the exception, test it on paper before you fund it.

Start paper trading, free