1,440 backtests, three noise worlds, one honest verdict

MRdata · methodology note · 2026-08-16

Every backtest is a choice among configurations: which desks must be buying, how big, entered when, held how long, gated how. Pick the configurations after seeing the data and you can prove nearly anything. This week I did the opposite, in public discipline if not in public view: declared the whole grid first, ran all of it once, and reported the distribution — with the top of the leaderboard deliberately unnamed.

The setup

1,440 configurations of the accumulation pattern — every cross of desk test, flow floor, entry session, five holding periods, two quiet windows, three RSI bands, three volume gates, and a gap filter — over the lake's full window (2026-05-14 → 08-15), equal-weighted, no stop, with realistic costs ($5 a side plus 0.3–1.5% spread depending on price band, at a $10,000 account scale). Then the control that matters: the entire grid re-run three more times on placebo worlds — the same tickers, with every flow date shifted by +7, +13 and −9 sessions. Whatever the placebo worlds "earn" is what pure noise manufactures here.

The verdict

  • The median configuration nets −0.16% before costs — and −6.46% after. At a $10,000 scale, spreading across dozens of micro-cap names is mostly a fee machine: 49% of configurations were positive before costs; 17% after.
  • 5.7% of real configurations beat the noise envelope. By construction, chance alone delivers 5%. That gap is the edge's entire showing across the whole grid, and it rounds to nothing.
  • One placebo world produced a +53.65% configuration from scrambled dates. Anything a real backtest shows on this window has to be read against that: spectacular cells are free here.
  • 0.4% of noise configurations turned $10,000 into more than $30,000 by sequentially compounding one position at a time. Tripling an account in three months sits comfortably inside this window's noise distribution. Achieving it proves nothing; targeting it teaches nothing.
  • The configuration my own recent exploration had converged on lands at the 56th–63rd percentile of the real curve. Middle of the pack. Reported as a position, not a result.

What survives this

Not nothing — but nothing a backtest can bless. The dials that looked strongest in-sample almost all sit inside their own noise envelopes. The one level that cleared its envelope did so at the minimum reportable sample size, one among roughly fifty tested — about what chance yields — and goes into the study queue with a declared mechanism before anyone looks at it again.

What actually survives is the method this site already runs on: rules registered in advance, scored forward, wins and losses in public. The live receipt on the track record is slow, small-n, and often unflattering — and after watching scrambled dates produce a +53% configuration, slow and unflattering is exactly what credible looks like. The sweep didn't find the edge. It found, precisely and at scale, why this site refuses to go looking for edges that way.

One email per market day. Zero when nothing happened.

Just the letter. No promotion, unsubscribe anytime.