Experiment 1 Concluded · Mar 17 – Jun 8, 2026 · 83 Days

Model A wins. Here's what we learned.

3 models. 83 days. 600 trades. The Crowd beat The Contrarian by $65. The Arb never got off the ground. Full results and post-mortem below. Test 2 launches Jun 15.

WINNER
Model A · Final
The Crowd
Follow consensus: YES ≥65%, NO ≤35%
$1,995.94
Return
+99.6%
Record
383–65
Trades
505
Model B · Final
The Contrarian
Fade longshots: always NO when crowd says ≤30%
$1,931.37
Return
+93.1%
Record
69–2
Trades
80
Model C · Inconclusive
The Arb
Trade Poly vs Kalshi gaps when they diverge ≥10pts
$1,000.00
Return
0.0%
Record
0–0
Trades
6
Post-Mortem
Model A · The Crowd
WINNER
Final Balance
$1,995.94
Return
+99.6%
Win Rate
85.5%
Trades
513

The Crowd won because it traded volume. 513 trades across 83 days. 69% of them were sports markets, where conviction is cleanest: favorites at 83% average entry probability win 86% of the time. The math checks out.

The edge here is real but narrow. Average win was $19.57 on a $100 bet. The conviction gate (65/35) filters out the noise. What's left is a steady drip of small, high-probability wins, punctuated by occasional $100 losses when the favorite collapses.

What to watch in Test 2: We tighten the gate to 75/25 as the new challenger. Fewer trades, higher confidence. Does stricter selection improve the return-per-trade, or does it just cut volume without improving quality?

Model B · The Contrarian
Final Balance
$1,931.37
Return
+93.1%
Win Rate
97.2%
Trades
81

97% win rate sounds better than it is. At an average entry of 13.9% probability, Model B was risking $100 to win $16. Every trade. The two losses wiped out the gains from 12 wins each. That's the structural trap of fading longshots with flat sizing.

The thesis isn't wrong. Crowds do overprice unlikely events. But you need position sizing that reflects the actual edge, not a flat $100 regardless of odds. Betting $100 to win $16 at 13.9% probability has almost zero expected edge above breakeven.

What to watch in Test 2: Model B gets retired as a standalone model. The insight lives on in a new variant that pairs the contrarian signal with Kelly-adjusted sizing instead of flat bets.

Model C · The Arb · Inconclusive
Final Balance
$1,000.00
Return
0.0%
Win Rate
N/A
Trades
6

The Arb never had enough to work with. The strategy requires Polymarket and Kalshi to disagree on the same market by 10+ points. That happens, but rarely. The Spread surfaces 1 to 5 qualifying pairs per day on a good day, zero on most days.

6 trades placed, all still open, all in long-duration political markets that never approached resolution. The concept is valid. The data coverage wasn't there. Model C is inconclusive, not disproven. We'll revisit when Kalshi coverage improves.

Experiment 1 · The 3 Models
Model A · The Crowd
Follow the conviction
  • ▸ Markets resolving within 14 days
  • ▸ YES when prob ≥65%
  • ▸ NO when prob ≤35%
  • ▸ $100 flat sizing
The thesis: smart money has already priced it in.
Model B · The Contrarian
Sell the longshots
  • ▸ Markets resolving within 14 days
  • ▸ Always NO when prob ≤30%
  • ▸ Pick the most liquid qualifying market
  • ▸ $100 flat sizing
The thesis: crowds overprice unlikely events. Sell the expensive tickets.
Model C · The Arb
Trade the disagreement
  • ▸ Only when Poly and Kalshi gap ≥10pts
  • ▸ YES on the lagging (cheaper) platform
  • ▸ Up to 30-day resolution window
  • ▸ $100 flat sizing
The thesis: when two liquid markets disagree by 10+ points, one is wrong.
Experiment Log
Mar 11, 2026
Test 1 (single model) begins.
$1,000 starting balance. 65/35 conviction gate. Short-duration only (≤14 days). Flat $100/trade.
Mar 17, 2026
3-model experiment launches.
Test 1 archived (1W/0L, +$6.50). Three models start simultaneously at $1,000 each: Model A (Crowd), Model B (Contrarian), Model C (Arb). Each follows a distinct strategy. Results reviewed Apr 17, 2026.
Jun 8, 2026
Experiment 1 concluded. Model A wins.
83 days. 600 trades. Final: Model A +99.6% ($1,996), Model B +93.1% ($1,931), Model C 0% ($1,000 inconclusive). Full post-mortem published. Test 2 launches Jun 15.
Jun 15, 2026
Test 2 launches Jun 15. The Crowd vs. two new challengers
Model A (Control) defends. New challengers: Model D (tighter 75/25 gate) and Model E (crowd conviction + momentum confirmation). $1,000 each. Same rules. 30 days.
Test 2 · Launching Monday Jun 15

Model A is the new control. Two challengers are trying to beat it. Same $1,000 start. Same 30-day window. Different hypotheses.

Model A · Control
The Crowd
  • ▸ YES when prob ≥65%
  • ▸ NO when prob ≤35%
  • ▸ Resolves within 14 days
  • ▸ $100 flat sizing
Unchanged from Experiment 1. The benchmark to beat.
Model D · Challenger
The Sharp Crowd
  • ▸ YES when prob ≥75%
  • ▸ NO when prob ≤25%
  • ▸ Resolves within 14 days
  • ▸ $100 flat sizing
The thesis: tighter conviction filters out noise. Fewer trades, better quality.
Model E · Challenger
The Momentum Crowd
  • ▸ Same 65/35 gate as Model A
  • ▸ Also requires 3pt+ move today in trade direction
  • ▸ Resolves within 14 days
  • ▸ $100 flat sizing
The thesis: crowd conviction plus fresh price movement beats stale conviction alone.
Full Trade Ledger
📋

No trades logged yet. First picks will appear after the next pipeline run.

How This Works
01
30-day test periods. Rules can change.
Each test runs for 30 days with a fixed ruleset. After 30 days we review the results, publish the findings here, and update the rules for the next test. The system gets smarter over time. You can see exactly how.
02
No retroactive edits. Ever.
Every trade is logged at the moment it's picked, before the outcome is known. The system runs on autopilot via GitHub Actions. We don't touch the ledger manually. Win or lose, it stays in the record.
03
P&L closes when the market resolves.
When a market resolves YES or NO, the trade closes automatically. Wins scale with the odds. A bold call at 70% pays less than a contrarian call at 30%. Losses are capped at $100. The goal: find systematic edge, not luck.

All trades are paper (hypothetical) and for informational purposes only. This is not financial advice and is not an invitation to copy trades. Prediction markets carry risk of total loss. We're publishing this experiment to be transparent about how the system works and to hold ourselves accountable to a verifiable track record. The rules are documented above. The ledger is unedited.

Test 2 Launches Friday
Follow the next experiment live.
Get the full Test 2 launch breakdown in the newsletter, then watch the ledger update in real time.
No thanks, just show me the experiment