▌ STRATEGIES
dashboardfindings

The findings — 60 strategies against buy-and-hold

A concluded programme, not a live page

Core programme concluded 11 August 2026; entries added since are dated in the text below. The ledger as of this build: 60 ideas — 1 pass, 16 mixed, 41 fail, 1 invalid, 1 rejected.

Everything below describes that fixed ledger and does not update with the dashboard around it. The one surviving rule (btc-trend) runs on as a live strategy at a fifth of the account — a PAPER account, no real money anywhere in this project — and its ongoing record lives on its own page, not here. A personal learning project, not financial advice.

The short version

The short version of FINDINGS.md, written for the site. Every number comes from a permanent ledger (registry.json) where failures are never deleted.

Over six batches I wrote down 60 trading ideas — each one specified completely before testing, each traced to a published paper where one existed — and scored every one against the laziest possible alternative: buying and holding. The bar was fixed in advance: on data the rule never saw during design, beat every benchmark on return per unit of pain (MAR ≥ 1.30×), depth of worst loss (≤ 0.75×), and raw growth (≥ 0.80×), after costs.

The ledger as of this build: 1 pass, 16 mixed, 41 fail, 1 invalid, 1 rejected. (These counts are generated from the ledger when the page is built, never typed.)

The one survivor

Hold Bitcoin while it closes >2% above its own 100-day average; sell below; nothing in between. Test window: +36.3%/yr with a −27.4% worst fall, versus +22.2%/yr and −53.0% for just holding Bitcoin. It also worked on five coins it was never designed on.

The asterisk: a different classic trend rule on the same coin (buy 55-day highs) failed the same test. So this is not "trend-following works" — it is one specific rule on one asset, which is also why it runs at only a fifth of the (paper-trading) account.

Three attempts to overturn the zero

  • More rules — 34 published anomalies (calendar effects, dip-buying, momentum, volatility filters). Zero survivors. The consistent killer: every rule that cut drawdown did it by being out of the market, and that costs more return than it saves.
  • Better data — I extended the price-only dataset with volume, the VIX family, interest rates, and company fundamentals from SEC filings, each series integrity-tested against events that unmistakably happened (the VIX had to spike 6× in March 2020; it did). Ten more published anomalies became testable. Zero survivors.
  • Combination — maybe many mediocre rules diversify into one good portfolio? Components were chosen by structural rules written in advance, never by their scores. Zero survivors.

Two later entries (2026-08-25) probed the remaining directions and confirmed the zero: Tom Hougaard's add-to-winners pyramiding lost to the plain momentum rule it decorated (same drawdown, less return), and short volatility via a contango-gated SVXY — the third stream the full write-up had named as the missing bet — underperformed simply holding the ETF, because the VIX curve is in contango 93% of days and the signal adds only whipsaw. Full detail in the dated addenda of the complete findings.

The most useful number: 1.8

After the portfolio tests, I measured the correlation between all 30 US components. Average: +0.55, which works out to 1.8 effective independent bets across 30 strategies. Four "different" registry entries correlate at +0.94 to +0.98 — four names, one strategy. Nearly everything in the ledger is the same bet — long the market, throttled — wearing different costumes. That is why combination couldn't work, and why earlier "near-misses" were fewer than they appeared.

I then built a candidate specifically to fix this number: the classic multi-asset trend rule (Faber 2007, followed to the letter) across ten ETFs — stocks, Treasuries, gold, commodities, REITs, currencies — with the success bar written down first: the three-strategy set had to measure at least 2.5 independent bets. It measured 1.97. Its returns were fine (a quarter of the market's drawdown at near-market Sharpe); it failed at its purpose, because four of its ten sleeves are still gated long equity and move with everything I already run. The first entry in the ledger to fail for a portfolio reason rather than a returns reason — and direct evidence that a wider basket of the same mechanism is not diversification.

Why I believe the zero

A null result is only worth publishing if the instrument could have detected a positive:

  • Adversaries attacked the scoreboard first and got worthless strategies to pass 14 documented ways (a rule that never trades; a rule reading tomorrow's price; "hold Solana every Wednesday"). Every route is closed and pinned by a regression test.
  • A junk portfolio of pure calendar noise was scored before the real ones. Averaging garbage produced a better Sharpe than the market (dilution flatters every risk-adjusted ratio) — and still failed, on the growth gate. The method can't manufacture an edge from nothing.
  • A one-day data error would have been the best result in the ledger. A deliberately worthless rule ("buy the day after the VIX falls") scores +2.6%/yr on correct data — and +66.4%/yr if the VIX feed is shifted by a single row. It failed on the real data, which is how I know the feed is aligned.
  • One soft spot is on the record: a benchmark that earns almost nothing is a free line. Against a currency ETF that made +0.3%/yr, even calendar noise "wins" every ratio (11× the growth of nearly zero is still nearly zero). It cannot fake a full pass — that requires beating every benchmark, and the real ones held — but it can pad a mixed verdict, so mixed verdicts get read line by line.

The cleanest lesson in the whole ledger arrived late, from a take-profit overlay on the leveraged-Nasdaq gate: on the held-out test years it earned +50.8%/yr with half the drawdown — the best numbers anything ever posted — and on the training years the identical rule lost to the gate it decorates (+21.8%/yr vs +36.1%/yr). When the two halves of history disagree, you are looking at a coin flip that landed well in the half you happened to hold out. Anyone judging on the test window alone would have shipped it. That is what most backtest-driven trading looks like from the inside.

The closest miss

The 30-rule combination: best risk-adjusted numbers of anything tested (−11.9% worst fall vs the market's −24.5%, Sharpe 1.14 vs 0.97) — and still a fail, because it earned 0.64× the market's growth and underperformed on sectors its components never saw. The drawdown diversifies away; a third of the return goes with it.

What I couldn't test — and the nastiest trap I found

Individual US stocks: no free price source serves the companies that died (0 of 11 famous bankruptcies), so any strategy tested there flatters itself automatically. Worse: two dead companies' tickers (BBBY, SBNY) now serve other companies' prices under the same symbol — data that passes every quality check because nothing is wrong with it except whose prices it is. Meanwhile company fundamentals are genuinely clean (SEC filings are point-in-time — IBM's Q3-2021 revenue: $17.6bn as filed, $13.3bn as restated a year later, and you can know which number was knowable when) — but clean fundamentals on a survivor-only universe is the same lie with better provenance, so I registered no fundamental strategies at all.

What this doesn't prove

Not that markets are efficient, and not that nothing works. Sixty ideas, daily bars, long-only, free data, one specific bar. Individual stocks, intraday, options, bonds and short-selling were never tested. It proves exactly this: these 60 ideas, on this data, under this bar, did not beat buying and holding — with one narrow exception.

What would change my mind on the portfolio question: a component with a genuinely different return stream — short volatility, long-duration bonds, cross-asset carry — rather than another long-only gate on the same index. The test is whether it moves that 1.8 number, and the multi-asset attempt above is the first entry to have been graded on exactly that test. It failed it.

The long version, and the ledger itself

The full account — every term explained, every number traced to the ledger, the 14 documented ways a worthless strategy passed before the holes were closed — is served beside this page: FINDINGS.md (plain text, exactly as written). Next to it sits registry.json, the permanent ledger itself: all 60 ideas with their rules as registered before testing, every verdict, every recorded read of every test window, failures never deleted. Both are copies made when this page was built — they are the evidence this page summarises, not links that can rot.