Daily Bilbo Boxes

Compression box → buy-stop at box high → 50 EMA trail
judged vs risk-matched random entries, on return AND max drawdown
daily bars 2000–202617 + 30 namesYahoo, split-adj2026-07-26
Paper portfolio & live watchlist — positions, setups arming now, every 2026 trade

Plain-English summary

Sometimes a stock goes quiet: its daily swings shrink and price coils into a tight range for days or weeks. This study asks a simple question — when price finally bursts out of the top of that quiet range, is it worth buying?

Yes, measurably. Buying the upside breakout of these “compression boxes” made more money than buying the same stocks on randomly chosen days with identical risk controls — our test for “is this skill or luck?” It kept working on 30 stocks that were never used to design the rules. The reverse trades don’t work: when price breaks out the bottom, neither betting on more downside nor bargain-hunting beats a coin flip.

What it feels like to trade: about two out of three trades lose small — typically 1–3%, because the exit plan cuts losers fast — and a handful of trades each year run for months and pay for everything (examples below include +156% in TSLA and +110% in INTC). If small frequent losses bother you, this system will too.

The rules, in one breath: wait for the quiet spell; place a buy order just above its range top, only when the stock is already above its 21-day average; if the trade fires, the worst-case exit is the range bottom; as the trade works, follow it up with the 50-day average and sell when price falls through it. A second, faster version watches during the day to catch ranges that break out the moment they end. And the account is always flat when a company reports earnings — one overnight earnings surprise taught us that lesson below.

What this is not: a prediction. It’s a measurement of what these rules did in the past, before fees, in a simulation. The forward test on a paper account is running now.

Validated — one-look pass on fresh names

The upside box break carried +0.5–0.9% per trade over risk-matched random entries — on the 17 discovery names and again on 30 names it had never seen.

Every long variant beat all 1,000 random-baseline draws on return, with max drawdown below 87–100% of them. On the 30 fresh megacaps the random baseline earns roughly zero, so nearly all of the strategy’s +0.76%/trade there is signal, not drift.

The same box is dead in every other direction: shorting the down-break, buying the down-break, and buying the 21 EMA reclaim after a down-break all measured at or below random.

+0.90%
excess per trade vs risk-matched random, 17 names
1,269 trades · PF 2.08
PASS
one-look validation, 30 fresh names
ret pctile 100 · DD pctile 0
35%
win rate — a trend system: many small stops, few huge runners
median trade −1.2%
163pp
max drawdown, 17-name book — below all 1,000 random draws (median 351pp)
sum-of-returns curve
The exact rules

Box: when the daily Phase Oscillator (Saty Mahajan’s PO) enters compression, the box is the high/low of the first 5 compression days (fewer if the squeeze ends early).

Entry: resting buy-stop at box high, working for 20 trading days after the box locks — but only on days where the prior close is above the daily 21 EMA. If price touches the box low first, no trade. Long only.

Initial stop: box low.

Trail: each day the stop ratchets up to yesterday’s daily 50 EMA (never down). Gap-honest fills: stop fills at min(open, stop).

Scratch variant (scr3): additionally exit if the close is back below the box high within 3 days of entry. Cuts book max-drawdown ~35% for ~27% less total P&L — the flavor for drawdown-punished accounts.

The measuring stick: every result is compared to 400–1,000 draws of random entries on the same names, gated the same way (close > 21 EMA), with stop distances drawn from the strategy’s own risk distribution and the same exit engine. A draw records its mean return and its equity-curve max drawdown — the strategy has to beat random on both.

The idea, from the top

1 · Compression: the market coiling

Saty Mahajan’s Phase Oscillator measures where price sits relative to its 21 EMA, scaled by ATR. Its compression flag fires when Bollinger bands squeeze inside the ATR envelope — daily ranges shrinking, buyers and sellers agreeing on price. Squeezes end. The question a trader cares about is which way, and whether the resolution is tradeable after honest costs and honest baselines.

2 · The box: making the squeeze tradeable

The Bilbo box is the high and low of the first 5 compression days. That range converts a fuzzy “squeeze” into two hard prices: break the top → long trigger; touch the bottom first → stand aside (we measured the down-break every way we could think of — it’s dead, see the NO-GO list below). The entry is a resting buy-stop at the box high, gated by the prior close being above the 21 EMA so we only take breaks in names already trending up. Nothing about the entry needs hindsight: the order can physically rest at the broker the moment the box locks.

3 · The exit: lose small, let the 50 EMA decide

Initial stop at the box low — if the breakout was false, the box itself says so. From there the stop only moves up, ratcheting to yesterday’s 50 EMA each day. There is no profit target: 65% of trades lose about a percent, and the whole system is paid for by the few that trend for months. The optional 3-day scratch (out if the close falls back inside the box within 3 days) is for drawdown-punished accounts: same engine, softer equity curve.
Three trades, start to finish

Discovery — 17 liquid names, 2000–2026

Strategy vs 200 random-entry books

Cumulative sum of per-trade returns (1 unit per trade, entry-ordered). Gray band = 10th–90th pctile of random draws with identical gates, risk and exits. +2.25%/trade vs +1.36% random.
up-break long (ema50 exit)random medianrandom p10–p90
Everything else the box says — nothing
NO-GOShort the down-break: −0.57%/trade across 1,119 trades — below even random shorts (7th pctile), with the worst drawdown of anything tested.
NO-GOBuy the down-break: +20-day forward return +1.82% vs +1.95% on random days (34th pctile). With a stop it drops to the 2nd pctile — 53% of fades stop out before the recovery arrives.
NO-GOBuy the 21 EMA reclaim after a down-break: +0.91%/trade sounds fine but sits at the 0.5th pctile vs random; adding it to the book grew drawdown faster than P&L and blocked 202 higher-value up-break entries.
Exit sweep — 16 exits, same entries

Excess return per trade vs each exit’s own random baseline

The entry’s edge survives every one of 16 exits (return pctile ≥ 98 in all). The 50 EMA trail extracts the most entry-specific alpha; the 3-day scratch trades some of it for a smaller drawdown. Fast ATR trails and fixed targets amputate the right tail that pays for this system.
The box-length question (added 2026-07-26)

Does the box need exactly 5 bars? Yes, keep it — and a lookahead confession

The 5-bar box was inherited from the original lower-TF Bilbo Box study, so we tested it. The first sensitivity table made the full-episode box (box = the whole squeeze range, locked when compression ends) look clearly better — +3.35%/trade vs +2.25%. That number was wrong. “Compression ended today” is only knowable at that day’s close, but the flawed scan let the entry fill intraday on that same day, gated by the day’s own closing state. A resting order cannot do that, and most squeezes break out on exactly that day. The armable version — order can only rest from the following session, episodes that break on the resolution day are missed entirely — keeps just 384 of 1,295 trades:
The first-5 box wins on every metric once fills are honest, because its trigger is finalized mid-squeeze and a stop order can sit at the broker waiting for the resolution day — the day the episode box structurally cannot trade. The live spec is, and stays, the validated first-5 box. The invalid row is kept above deliberately: it was briefly promoted to the live config before the error was caught the same day, and this catalog doesn’t hide its mistakes.

2022–2026: the two specs, week by week — honest fills only

Cumulative realized P&L (scr3 exit, 24-name live universe, booked at trade exit, from 2022-01), both variants armable-only. The first-5 box ends the window at +954pp vs +176pp for the armable episode box, which trades a quarter as often and lost money in 2022–23:
first-5 box (live)episode box, armable
Why the flawed version looked so good: the trades it invented — intraday fills on the very day each squeeze resolved — are the best trades in the family. Remove them and the episode box keeps only late breakouts, a quarter of the flow. The general lesson, learned here for the price of a same-day config revert: any signal whose trigger level is finalized by the entry bar’s own close cannot be traded with a resting order, and backtests that pretend otherwise flatter exactly the trades that matter most.
Validation — 30 names the spec never saw

Frozen spec, one look

Spec and read-bars fixed before the run (ret pctile ≥ 95, DD pctile ≤ 20). Universe rule fixed too: top 30 by 2-yr dollar volume, ex the 17. Shown: the deployment flavor (scr3). +0.47%/trade vs −0.04% random; both bars cleared by both exits.
up-break long (scr3 exit)random medianrandom p10–p90
Per-trade average by era — improving, not decaying:

Where it worked, name by name (30 fresh names, scr3)

High-beta names carry it; slow defensives are dead weight. Vol gradient is directional (high-vol half +1.46%/tr vs +0.72%) but not statistically settled at n=30.
The combined book — resting orders + intraday nowcast

Two entry paths, one sleeve

The INTC-class trade — a squeeze that breaks out on the very day it resolves — can’t be captured by a resting daily order (the box isn’t final until that day’s close). The fix is the same one the hourly system uses: a fast execution layer. At each 5-minute close, today’s provisional daily bar is rebuilt and the daily compression flag is nowcast; a market buy fires the moment the squeeze is releasing and price crosses the running episode high. The resting first-5 order (K5) keeps priority — its trigger always sits at or below the nowcast’s — and the nowcast (NC) only ever trades the episodes K5 already died on. One trade per name per squeeze.
combined K5+NC, cumulative per-trade returns
24-name live universe on intraday-derived daily data, 2003–2026: 1,914 trades, +1.51%/trade, PF 1.89, +2,888pp. On the 17-name discovery set the union added +42% more total P&L than the resting-order spec alone at equal risk-efficiency (sum/DD 24.9 vs 24.3) — the two paths fire on different squeezes by construction. The NC leg’s own showcase: INTC’s April-2026 squeeze resolution, +110% in 62 trading days, entered intraday at the 5-minute close that confirmed the release — with no knowledge of the daily close required. Both paths run in the live paper ledger from 2026-07-27; NC entries log a counterfactual scratch exit and a flicker tag so forward data settles the open design questions.
Never hold through earnings

One overnight gap taught the lesson; 1,914 trades confirmed it costs nothing

In the 2026 paper simulation, a UNH trade entered on a clean breakout (Jan 22) and was up comfortably — then UNH reported earnings before the open on Jan 27 and the stock opened 16% lower. The protective stop can’t help against an overnight gap: the exit fills where the market opens, not where the stop sits. A −7% worst case became −16.6%.
policyavg/tradePFtotalmaxDD
hold through earnings+1.51%1.89+2,888pp209pp
flatten, stay out+1.39%1.95+2,657pp158pp
flatten, re-enter next day (adopted)+1.55%2.00+2,974pp189pp
21% of all trades in the 24-name history held through at least one earnings report. Exiting at the close before every report and buying back the next day (only if the trend is still intact) beat holding on every measure — because mid-trend earnings gaps cut both ways about evenly, while the catastrophic ones only ever hurt the holder. The long runners survive the one-day pause: INTC’s +110% keeps +70 under the rule. Sleeve policy: flat by the close before every report, re-enter the next close if price still holds above the trailing stop.
Honest limits
Gross returns — no spread/swap/commission modeled; at 150–225bps/trade edge and 11–18 day holds the friction is real but secondary.

Survivor universe — both name lists are drawn in 2026; the defensible claim is beating random within the same names, which the baselines control for, not the absolute P&L.

No position cap — the book averages ~4 concurrent positions (max 13) at one unit each; a capped live version will differ.

Ranking noise — 16 exits on one dataset: the scr3-vs-ema50 sum/DD ordering flipped on the fresh names, exactly as the noise caveat predicted. Treat those two exits as equivalent; scr3 = lower absolute DD, ema50 = more P&L.

History spent — the 17 names’ full history was used in discovery; the 30-name run was the one-look validation. What remains is forward.