# Sixty strategies against buy-and-hold: the findings

*Written 2026-08-11 at the end of the core research programme in this directory, and updated 2026-08-18 for the entries registered since (the dated addenda at the bottom carry their full detail). The headline counts on this page are generated from the ledger when the site is built, never typed. Every number
in this document is in `registry.json` or in a run recorded there — if a figure is not in
the ledger, it is not in this write-up.*

*Audience: me in six months, and anyone who might otherwise spend a month repeating this.
Terms are explained the first time they appear, because I was a beginner when this
started and the write-up should not require more than I had.*

---

## 1. The answer

**Sixty trading ideas were written down in full before being tested, tested once
each against a fixed bar, and one survived.**

The count: **1 pass, 16 mixed, 41 fail, 1 invalid, 1 rejected.**

The first fifty concluded the core programme on 2026-08-11. Five later entries
(2026-08-18, all MIXED) did not change the answer, but two of them sharpened it:
the take-profit overlay is now the ledger's cleanest demonstration of the
coin-flip signature (§4), and the multi-asset trend rule is the first candidate to
fail for a portfolio reason rather than a returns reason (§8).

A few terms, once:

- A **backtest** replays a trading rule against historical prices to see what it would
  have earned. Backtests lie easily, which is what most of this document is about.
- **Buy-and-hold** means buying the asset once and doing nothing. It is the laziest
  possible alternative, and it is the opponent every idea here had to beat — because if a
  rule cannot beat doing nothing, the rule is the wrong place to put effort.
- **CAGR** (compound annual growth rate) is growth per year. All CAGRs here are **net of
  trading costs**, because costs flipped real conclusions in this project — one Vietnam
  strategy earned +270% before costs and lost everything after them.
- A **drawdown** is a fall from a high point to a low point. **Max drawdown** is the worst
  one — the deepest hole you would have had to sit through without panicking.
- **MAR** is CAGR divided by max drawdown: return earned per unit of worst-case pain.
- **Sharpe** is return divided by day-to-day wobble (volatility) — a second opinion on
  the same question, reported so it can disagree with MAR.

**The bar.** Every idea was scored on data it had never seen (see §4) and had to beat
**every** benchmark — the traded asset's own buy-and-hold, the market index, and an
equal-weight basket both rebalanced and left alone — on all three of: MAR at least
**1.30×** the benchmark's, max drawdown no worse than **0.75×**, CAGR at least **0.80×**.
The margins are the point: "slightly better risk-adjusted return" is what noise looks
like. A 1980s calendar effect got MAR 0.50 against the market's 0.46 and was correctly
killed by exactly that margin.

### The survivor: `btc-trend-100d`

Hold Bitcoin while its price closes more than 2% above its own 100-day average; sell when
it closes more than 2% below; do nothing in between. Decided at each daily close, traded
at the next open, 10 basis points cost per side, on Binance data (the venue the order
would actually rest on).

On the test window (2023-11 to 2026-08, never touched during design):

| | CAGR | max drawdown | MAR |
|---|---|---|---|
| the rule | **+36.3%/yr** | **−27.4%** | **1.32** |
| just holding Bitcoin | +22.2%/yr | −53.0% | 0.42 |

It also worked on five coins it was never designed on (it lost −10.2%/yr there while
holding those coins lost −26.8%/yr in a falling market — losing much less than doing
nothing is the edge working).

**The honest asterisk, and it matters:** `donchian55-btc` — a different classic
trend-following construction (buy 55-day highs, sell 20-day lows) on the *same coin, same
window, same costs* — **failed** (+14.3%/yr, MAR 0.39). If "trend-following works on
Bitcoin" were true as a general claim, both should have passed. One construction passing
while its sibling fails means the edge is **specific to this rule on this asset**, not a
law of nature, and the two must never be cited as confirming each other.

The other early result, `trend3x-150d` (hold a 3× leveraged Nasdaq fund while the Nasdaq
is above its 150-day average), is recorded as **MIXED**: it crushes holding the leveraged
fund itself (+31.7%/yr at −40.4% vs +17.3%/yr at −81.7%) but its drawdown is *deeper*
than simply holding the unleveraged index (−35.1%) or the S&P (−24.5%), and on four
sectors it was never designed on it lost −3.8%/yr while holding them made +6.9%/yr. It
reduces one specific fund's drawdown; it is not a general edge.

---

## 2. Three attempts to overturn the answer

The first 34 ideas (batches 1–2) produced zero survivors. Rather than stopping there,
the programme made three deliberate attempts to prove that result wrong.

### Attempt 1: more rules (34 ideas, zero survivors)

Every major family of published price-based ideas: calendar effects, mean reversion
(buying dips), momentum, volatility filters, trend rules, cross-sectional ranking. Each
traced to a published academic source, each written down in full before testing. The
consistent killer was the **CAGR gate, not the drawdown gate**: every rule that reduced
drawdown did it by being out of the market, and being out of the market costs more return
than it saves. That sentence is the single most repeated finding of the whole programme.

### Attempt 2: better data (10 ideas, zero survivors)

The sweep's own conclusion was that the binding constraint was the *data*, not the rules:
the price cache stored only open/high/low/close, so the strongest published anomalies —
volume effects, volatility-index signals, interest-rate regimes, fundamentals — could not
even be registered. So the data layer was extended: trading volume for 43 instruments,
the VIX family (an index of how expensive stock-market insurance is — high VIX means
fear), Treasury and credit interest rates, and company fundamentals from SEC filings.

Every new series was integrity-tested before use (§4 and §6 cover what those tests
caught). Ten ideas were registered against the new data — the high-volume return premium,
volume-conditioned reversal, illiquidity ranking, VIX term-structure timing, the variance
risk premium, VIX-spike buying, yield-curve and credit-spread regime gates — each citing
its published source. **All ten failed or scored mixed.** The closest single-rule miss of
the first fifty was here: `credit-spread-gate-spy` (hold the S&P while corporate borrowing
spreads are below their recent median) posted MAR 0.47 vs the market's 0.46 with a little
over half the drawdown (−19.0% vs −33.7%) — and still failed, because +8.9%/yr is 0.58×
the market's return and the bar is 0.80×. Cutting the hole in half while giving up 42% of
the growth is not beating buy-and-hold.

### Attempt 3: combination (3 portfolios, zero survivors)

Several failed rules cut drawdown while giving up return — raw material, in theory, for a
**portfolio**: combine many mediocre-but-different rules and diversification does the
rest. Three equal-weight combinations were tested, with the components chosen by
*structural* rules written in advance (every fully-scored US candidate; every batch-3
rule trading exactly the S&P; every fully-scored crypto rule) — never by their scores,
because picking components for having scored well is selection wearing a diversification
costume. Zero survivors. The best of them is the closest miss of the programme (§5), and
the reason combination cannot work with this pool is the independence finding (§3).

---

## 3. The intellectual core: 30 components, 1.8 independent bets

**Correlation** measures whether two things move together: +1 means always, 0 means no
relationship. A portfolio's whole promise rests on its parts *not* moving together —
when one is down, another is up, and the ride smooths out.

After the combination verdicts were recorded, every component was simulated standalone
and the correlation of daily returns measured between every pair. Across the 30 US
components the average pairwise correlation was **+0.55**. There is a standard way to
turn that into "how many genuinely different bets is this": N/(1+(N−1)·ρ̄). The answer:

> **30 components. 1.8 effective independent bets.**
> (The 6-component and 9-component combinations: 1.8 and 1.9.)

In plain terms: this ledger contains, in sixty costumes, roughly **two** bets — "be long
the stock market, throttled" and "be long crypto, throttled". Adding 24 more components
to the US portfolio added *zero* diversification, because a rule that is long the S&P
on Mondays and a rule that is long the S&P when the VIX falls are the same bet with
different trigger words.

This retroactively explains most of the ledger:

- **Batch 2's seven MIXED verdicts** were reported as seven near-misses. Three of them
  "passed" only against XLE, an energy fund that fell 60.6% — and the correlation matrix
  now shows the four sector-rotation entries (`low-vol-sectors`, `max-effect-sectors`,
  `idio-vol-sectors`, `52w-high-sectors`) correlate at **+0.94 to +0.98** with each
  other. Four registry rows, one strategy. The honest count of distinct near-misses in
  batch 2 was never seven.
- `halloween-spy` and `halloween-sectors` correlate at **+0.98** — the same calendar bet
  aimed at one index versus twelve funds.
- `xs-momentum-crypto` and `high-volume-premium-crypto` correlate at **+0.89** — coin
  momentum and coin volume-spikes pick the same coins at the same times.
- And it is why the survivor's sibling failing (§1) mattered so much: **counting costumes
  as confirmations is the easiest way to talk yourself into an edge.**

One measurement caveat, recorded with the result: sleeves that are in the market less
than a quarter of the time show *low* correlation mechanically — they are mostly cash —
so the true redundancy is, if anything, understated.

---

## 4. Why the null result is credible

"Sixty ideas, one survivor" only means something if the measuring instrument could
have detected a survivor, and would not invent one. Most write-ups skip this part. It is
the part that makes the rest worth reading.

**The instrument was attacked before it was trusted.** Adversarial reviewers were pointed
at the scoreboard with one instruction: get a worthless strategy to PASS. They succeeded
repeatedly, and every route is now closed and pinned by a regression test — **14 documented
false-PASS routes**, each of which produced a clean pass on real data before it was
fixed. Highlights, because the shapes are instructive:

- A strategy that **never traded** passed: a flat account has no drawdown, so its MAR was
  infinity, which beats any benchmark.
- A rule that read **tomorrow's price** passed every gate — there was originally no
  lookahead defence at all.
- **"Hold Solana every Wednesday"** — pure calendar noise — passed, because against a
  benchmark that *lost* money every ratio-based gate degenerates and the bar collapsed to
  "make any money at all". The fix: the training window gets a veto, and a losing
  benchmark triggers an absolute floor instead of a ratio.
- Scoring with `record=False` left **no trace**, so forty variants could be tried and one
  reported as a first-attempt success. Every read of the test window is now logged
  permanently — a pass on attempt thirty reads as a pass on attempt thirty.

**Data lies were tested for with known answers.** Every new series had to contain events
that unmistakably happened: the VIX had to spike 6× peak-to-trough in March 2020 (it
did: 13.68 → 82.69); the 2-year Treasury yield had to rise ~4 points across 2022 (it
did: 0.77 → 4.72); S&P volume had to explode in the March 2020 panic (it did: 6× the
2019 median); Apple's earnings dates had to land on the exact known days (3 of 3 did).
Events are measured **peak-to-trough, not endpoint-to-endpoint** — an earlier check
measured an oil spike by its endpoints, saw the leftovers instead of the event, and the
lesson was pinned.

**Two deliberate canaries, both with proven teeth:**

- *The junk portfolio.* Averaging many noisy things mechanically smooths the result, so a
  ratio-based bar could in principle be gamed by combining enough garbage.
  `combo-junk-calendar` — nine sleeves of pure calendar noise and already-refuted
  seasonality — was registered with "THIS MUST FAIL" as its success criterion and scored
  *before* any real combination. It produced a **better Sharpe than the market** (0.97 vs
  0.93) and 0.44× the drawdown — dilution really does flatter every risk-adjusted number —
  and it failed anyway, on the CAGR gate and the 1.30× margin. The method does not
  manufacture edges from nothing.
- *The shifted-VIX canary.* The one failure the lookahead check structurally cannot see
  is a feed that is off by one day — the check perturbs data *by* date, and the error *is*
  the date. So `vix-down-day-spy` (hold the S&P the day after the VIX falls — worthless
  by construction, since VIX changes track the *same day's* market move) was registered
  as a must-fail. On correct data: +2.6%/yr, MAR 0.10, dead on arrival. On a VIX series
  deliberately shifted one row early: **+66.4%/yr, MAR 3.24 — an unqualified pass.** A
  single-row data error would have manufactured the best result in the ledger. The canary
  failed on the real data, which is evidence the feed is aligned — and the shift test is
  now a permanent part of the data-layer verification.

**A soft spot found later, recorded rather than smoothed over: the near-zero
benchmark.** Sibling of the losing-benchmark degeneracy above. The floor that stops a
rule "beating" a benchmark that lost money only triggers when the benchmark actually
LOST — and against a line that earned almost nothing, every ratio gate degenerates
legitimately (11× the CAGR of nearly zero is still nearly zero). The calibration
canary for the multi-asset batch exposed it: calendar-parity noise "passed" against
holding the euro ETF, which earned +0.3%/yr over the test window, while failing
against every real benchmark. This **cannot mint a false PASS** — a PASS requires
beating every line, and the hard lines held — but it **can inflate a MIXED**, so any
MIXED on a universe containing a flat asset has to be read by asking WHICH lines
passed. Lines against near-flat benchmarks are decoration.

**The tells behaved.** Momentum and reversal — opposite selections over identical names
and dates — did not both pass (both passing is the signature of a harness that deletes
losers). All four entries registered as expected-failures failed exactly as registered.

**The coin-flip signature, demonstrated by a live candidate.** The cleanest teaching
example in the ledger arrived after the core programme: `trend3x-tp50`, a take-profit
overlay on the trend3x gate (sell at +50%, re-enter after a full gate reset),
registered because a real paper-account trade had given back a $1,000 open profit. On
the test half it looks like the best discovery here: **+50.8%/yr at −22.4% drawdown,
MAR 2.27**, against plain trend3x's +31.7%/yr at −40.4%. And on the training half the
same overlay **loses to the rule it decorates**: +21.8%/yr against +36.1%/yr, MAR 0.58
against 0.64. The two halves of history disagree about whether the overlay helps at
all, which is what a coin flip looks like when one side of it happens to land in your
test window. The mechanism is mundane: a +50% cap sells a decade-long trend early
(the 2010s training years) and shines in a choppy, crash-containing window (the test
years). Any process that judged candidates on the test window alone would have
shipped it. This is why the training window has a vote, and why a spectacular test
number should raise suspicion before excitement — the full detail is in the
2026-08-18 addendum.

None of this proves the instrument is perfect. It proves the instrument can detect a
positive (the junk and canary results are, mechanically, detected positives on corrupted
inputs) and that the known ways to fake one are closed.

---

## 5. The closest miss, honestly

`combo-us-all`: the equal-weight combination of all 30 fully-scored US rules, failures
included. On its test window (2020-07 to 2026-07):

| | CAGR | max drawdown | MAR | Sharpe |
|---|---|---|---|---|
| the combination | +10.8% | **−11.9%** | **0.91** | **1.14** |
| hold the S&P | +16.7% | −24.5% | 0.68 | 0.97 |

The best MAR and the best Sharpe anything posted in sixty tests. Less than half the
market's drawdown. It passed outright against three *real* benchmarks (developed
international, emerging markets, healthcare — not the fell-60% energy fund that flattered
batch 2's MIXEDs). And its MAR ratio against the S&P was 1.33× — comfortably above the
junk portfolio's 0.89×, so this is **not** the dilution artifact the calibration was
built to catch.

It is still not a pass, for two reasons that survive any amount of goodwill:

1. **It earned 0.64× the market's return** against a bar of 0.80×. Same epitaph as the
   other forty-nine: the smoothness is bought by being partially out of the market, and
   that costs more growth than it saves in pain.
2. **It did not transfer.** On the four sectors none of its components were designed on,
   it earned +7.0%/yr while just holding them earned +10.3%/yr.

Verdict as recorded: MIXED. The honest sentence: *the drawdown diversifies away, and a
third of the return goes with it.*

---

## 6. What was not tested, and why

**US individual stocks — blocked by survivorship, and worth understanding.**
**Survivorship bias**: if your price history only contains companies that still exist,
every strategy tested on it looks better than reality, because the bankruptcies are
missing from the record. The check was direct: ask the free price sources for eleven
famous corpses (Enron, Lehman, WorldCom, SVB, First Republic, Bed Bath & Beyond, …).
**Zero of eleven were served.** Nine are simply absent. The other two are the most
deceptive thing this project found: **ticker reuse**. Yahoo serves "BBBY" and "SBNY" as
healthy series running to today — because after Bed Bath & Beyond and Signature Bank
died in 2023, *other companies took over their ticker symbols*. Nothing is wrong with the
data as data; it passes every missing-value, gap and staleness check ever written. It is
simply a different company's prices under the name you asked for. The only check that
catches it is "does the series *end* when the company died" — a question about history,
not about data quality.

The bitter part: company fundamentals are *fine*. SEC EDGAR serves every number ever
filed, with the date it was filed — **point-in-time**, meaning you can know exactly what
was public on any given day. The proof is in the verification suite: IBM's Q3-2021
revenue was **$17.618bn as first filed** (2021-11-05) and **$13.251bn as restated** a
year later — a −24.8% rewrite (a spin-off reclassification; GE's 2023 revenue moved
−48% the same way). A backtest reading today's restated number on the original date is
reading the future, and most free fundamental sources serve exactly that. EDGAR doesn't.
But clean fundamentals applied to a survivor-only price universe is the same lie with
better provenance, so **no fundamental strategy was registered at all** — that was a
pre-declared gate, and it held.

**Three families remain out of reach:** value/accruals/profitability ranking and
post-earnings drift (both need survivorship-free individual-stock prices; not obtainable
free), and pre-Fed-announcement drift (tradable ~3% of days — below the 5% exposure floor
under which this harness refuses to judge anything, the same floor that made the Santa
rally INVALID rather than a verdict).

**Vietnam.** Tested honestly and closed early: nothing cleared the bar against buying
the HOSE names once and never touching them, after Vietnam's ~1% round-trip costs. The
best VN idea (`xs-momentum-vn`, +34.4%/yr at −23.0%, MAR 1.50 standalone) beat several
individual names and still failed against the buy-once basket — more return (1.20×) but
a deeper drawdown (1.06×) and nowhere near the MAR margin (1.13× against 1.30×) — and
its edge did not transfer to the eight HOSE names it was never designed on (+8.8%/yr
there vs +21.4%/yr for holding them). An earlier VN trend rule made the same point more
bluntly: it beat the monthly-rebalanced benchmark and lost to buy-once-and-drift, +508%
vs +930% — the choice of benchmark alone flipped the verdict, which is why buy-once is
in every benchmark set here. Two hard local lessons are recorded in the ledger: the only VN index
fund's Yahoo feed served **one frozen price through the entire 2022 crash** (a stale
print passes every missing-data check — flatness needs its own check), and Vietnam's
T+2.5 settlement means shares cannot be sold for ~3 sessions after buying, which
mechanically **rejected** a reversion rule that was legal on paper (114 violations). No
VN combination exists because only one VN idea was fully scored, and one thing cannot
be combined.

---

## 7. What this does not prove

Stating the limits as plainly as the findings:

- **It does not prove markets are efficient.** Sixty ideas is sixty ideas. The strongest
  published anomalies in individual stocks were never tested here (§6), intraday
  behaviour was never tested (daily bars only), and options, futures, bonds-as-an-asset
  and short-selling were never tested (long-only, cash instruments, no leverage unless
  declared).
- **It does not prove "nothing works".** It proves these 60 rules, on this data, at
  these costs, under this bar, did not beat buying and holding — with one specific,
  narrow exception.
- **The survivor itself is one draw.** `btc-trend-100d` passed a hard test once, its
  edge transferred to unsearched coins, and its sibling construction failed — which is
  exactly what a real-but-specific edge looks like, and also not far from what a lucky
  draw looks like. It runs live at a fifth of the paper account precisely because the
  honest confidence level is "worth running small", not "proven".
- **The bar is a choice.** 1.30×/0.75×/0.80× are judgement calls fixed in advance. A
  different bar would pass different things — which is why the junk portfolio and the
  calendar-noise attacks matter more than the thresholds themselves: they show what this
  particular bar lets through (nothing, so far) and what it was designed to stop.
- Costs are modelled (2bp US, 50bp Vietnam, 10bp crypto per side), not lived. Slippage,
  taxes, and the experience of actually sitting through a −27% drawdown are not in any
  number here.

---

## 8. What would make the portfolio question worth asking again

The combination attempt failed for a measured reason: the pool contains ~2 independent
bets. The fix is not more rules of the same kind — a 51st long-only gate on the S&P
would join the +0.55 correlation blob and change nothing. The portfolio question becomes
interesting again only when a component would arrive with a **genuinely different return
stream**:

- **Short volatility** — earning the insurance premium (selling options or their proxies)
  pays in calm markets and loses in crashes: near-mirror timing to a long-equity gate.
- **Duration** — long-term government bonds respond to growth and inflation, not
  earnings; the 10-year yield data to build the signals is already cached.
- **Cross-asset carry** — hold what pays you to hold it, across currencies, commodities
  and bonds; the classic uncorrelated-return family in the literature.

Each of these needs new instruments or new data (options-based funds, bond futures or
their ETF proxies, currency pairs), and each would face the same registry, the same bar,
the same calibrations. The correlation matrix — not any single verdict — is the test of
whether it was worth adding: a real diversifier would drag the effective-bets number off
1.8. That is the number to watch.

**The first attempt at exactly this has now been made, and it failed — which is itself
a finding.** `gtaa10-faber` (2026-08-18): the canonical published multi-asset trend
rule (Faber 2007, verbatim) on a ten-ETF universe declared in advance across
equities, Treasuries, gold, commodities, REITs and currencies — built for one
purpose, to be a third independent bet next to trend3x and btc-trend, with the bar
written into its registry entry before any number existed: effective bets for the
three-strategy set had to reach 2.5 of 3. It is the first candidate in this ledger to
fail for a **portfolio** reason rather than a returns reason. Its returns did what
the literature promises (a quarter of the market's drawdown at a near-market Sharpe;
the gates it failed were the familiar CAGR ones). But its correlation with trend3x is
+0.49 daily, +0.43 monthly, and the three-strategy set measures **1.97 effective
bets of 3** — four of its ten sleeves are gated long equity, and a gated long-equity
book moves with a gated leveraged-Nasdaq gate whatever the bonds and currencies do.
The rule built specifically to drag the ~2-bets number upward did not move it. The
lesson sharpens the paragraph above: breadth alone is not diversification when the
breadth is still mostly equity — the missing bet has to come from a different
MECHANISM (short volatility, carry, duration), not a wider basket of the same one.
Full numbers in the second 2026-08-18 addendum.

---

*The equipment survives the programme: `fetch.py --verify` (data integrity, 21 checks),
`score.py --selftest` (6 measurement traps + 14 pinned false-PASS routes),
`registry.py` (the ledger — 60 entries, failures never deleted, every read of every test
window logged). Anyone extending this work should run all three before believing
anything, including this document.*

---

## Addendum, 2026-08-18: take-profit and stop-loss overlays on trend3x

*Prompted by a real trade: the paper account's TQQQ position was up about $1,000 open
profit and gave nearly all of it back while the gate stayed on. The question was whether
an exit overlay would have kept that money. Three overlays were registered
(`trend3x-tp50`, `trend3x-trail15`, `trend3x-trail15-reentry`), parameters fixed at
registration, and scored on the standard rig. All three came back MIXED, and none is
shippable. Verdict: the paper account keeps plain trend3x.*

The head-to-head on the test window (2021-08-23 to 2026-08-07) looks, at first glance,
like a case for the take-profit: sell at +50% and the test half returns +50.8%/yr at
-22.4% drawdown, against plain trend3x's +31.7%/yr at -40.4%. But the train half says
the opposite: there the same overlay earned +21.8%/yr against plain's +36.1%/yr, with a
worse MAR (0.58 vs 0.64). The two halves disagree about whether the overlay helps at
all, which is the same coin-flip signature this ledger has already used to fail other
ideas. What actually happened is mechanical: a +50% cap costs you most of a decade-long
trend (the 2010s train window) and happens to sell near local tops in a choppy,
crash-containing window (the test half). Whether the next five years look like the
first kind of decade or the second is exactly the thing nobody knows.

The trailing 15% stop shows the same disagreement in milder form (train MAR 0.53 vs
plain's 0.64, test 0.99 vs 0.78) and gives up a third of the return in both halves.
The stop-with-re-entry variant is worse than plain everywhere and worse than its own
no-re-entry sibling, because re-entering on a 20-day high buys back into chop
repeatedly during 2022.

Two further nails, both from the standard gates rather than from judgement. First, none
of the three cleared the fixed bar: every one failed the drawdown gate against SPY, and
the re-entry variant failed against QQQ too. Second, the holdout: pointed at the four
never-searched sectors, all three overlays LOST money (about -3.9%/yr) while holding
those sectors made +6.9%/yr, the identical trap-4 failure already recorded against
plain trend3x. An overlay cannot fix a rule's transfer problem; it inherits it.

The practical answer to the original question: the give-back was not a bug, it is the
cost the trend gate pays on purpose, and the overlay that would have kept that one
trade's profit would have cost more than it saved over the train decade. The one
overlay finding worth remembering if drawdown comfort ever matters more than return:
the plain trailing stop did cut the daily test drawdown from -40% to -22% at a price of
about 10 points of CAGR. That is a trade you could choose to make with open eyes; it is
not an improvement, and it did not pass.

---

## Addendum, 2026-08-18 (2): multi-asset trend following (batch 6)

*The first candidate that is not a long equity rule in disguise: Faber's timing model
(2007, SSRN 962461, rules quoted verbatim in the registry entry) on a ten-ETF universe
declared in advance across equities, government bonds, gold, commodities, REITs and
currencies. Registered as `gtaa10-faber` with a calibration canary `gtaa10-junk-months`
scored first. Rejected instruments and reasons are written in the entry (USO/UNG for
structural roll decay, UUP for double-counting the euro, SLV, AGG/BND, GSG, BDRY).
Moskowitz, Ooi and Pedersen (2012) is the futures-native canonical form; its 40%/sigma
vol sizing needs leverage and shorting this simulator refuses to price, so the
long-or-cash published form was used unmodified rather than inventing a hybrid.*

**The calibration came back MIXED rather than FAIL, and the reason is a finding.**
Calendar-parity noise on the same universe and sizing beat exactly one benchmark line:
holding FXE, which earned +0.3%/yr over the test window. Against a near-zero POSITIVE
benchmark every ratio gate collapses (11x CAGR of almost nothing is still almost
nothing), and the 0.50 MAR floor that guards against losing benchmarks does not apply
because +0.3% is not a loss; the same floor correctly killed the FXY line next to it.
Structurally this cannot mint a false PASS, since PASS requires beating every line and
the hard benchmarks (SPY, DBC, both basket flavours) all held. But it inflates MIXED:
a flat benchmark is a free line, and any MIXED on a multi-asset universe has to be
read by asking WHICH lines passed. The noise did not survive the gate, so the
machinery is fit to score the real candidate.

**The candidate: MIXED, and the shape of the miss matters.** Test window (2020-09 to
2026-08): +5.4%/yr at -6.9% daily drawdown, MAR 0.79, Sharpe 0.91. The published
promise (equity-like Sharpe, bond-like drawdown) shows up exactly: drawdown one
quarter of SPY's, Sharpe within a tenth of SPY's. What fails is the same gate that
killed most of the fifty: CAGR. It earned 0.32x SPY's return against a 0.80x bar, and
lost the return race to every risk asset it times (EFA, EEM, GLD, DBC, VNQ). It also
failed MAR vs SPY (1.13x against the 1.30x margin). It cleared the bar only against
the monthly-rebalanced basket, the two bond ETFs, and the two flat currency trusts.
The test window is a near-uninterrupted bull market; Faber's own case is built on
1973-2008, which contained multi-year bears this window did not offer. The holdout
does not transfer either: the same rule on the four never-searched sectors made
+4.7%/yr while holding them made +9.7%/yr, trap four's exact signature, same as
trend3x before it.

**The purpose test failed, and that is the headline.** This was built to be a third
independent bet next to trend3x and btc-trend. Measured on overlapping days 2017-2026:
correlation +0.49 daily (+0.43 monthly) with trend3x, +0.15 (+0.24) with btc-trend.
Effective independent bets for the three-strategy set: **1.97 of 3** against the
registered bar of 2.5. Four of its ten sleeves are gated long equity, and a gated
long-equity book moves with a gated TQQQ whatever the other six sleeves do. It is
roughly half a new bet, not a third bet. The registry entry said in advance that high
correlation with the existing pair is failure of purpose even on a gate pass, so:
failure of purpose. Caveat noted in advance and repeated here: btc-trend is in the
market only 53% of days, so its low correlations are partly mechanical.

**What this run could not verify.** The ETF window opens in 2007-02 (FXY inception),
so the strategy was never scored through a 1970s-style multi-year bear, which is the
regime the literature says pays for everything. A verdict on that costs futures data
this project does not have. And the correlation was measured on simulated curves over
nine overlapping years, not on live records.

**Where this leaves the search.** The gap Harry named (an uncorrelated return stream)
is still open. Multi-asset trend as implementable in this account is not it: the
long-only ETF form keeps enough equity beta to be half the same bet. Streams that
could actually be a third bet remain the ones named at the end of the programme:
short volatility, carry, duration as its own signal. All are mechanisms other than
trend, which the mechanism-balance warning has been asking for anyway.

---

## Addendum, 2026-08-25: add-to-winners pyramiding on SPY momentum (`momentum-pyramid`)

*Registered from Tom Hougaard's principle (tradertom.com; "Best Loser Wins", 2022): add
to winning positions, never to losing ones. Hougaard publishes the principle, not
parameters, so the 1/3 increments, the new-month-end-high add trigger and the 1.0 cap
were fixed at registration, before any measurement, and the monthly cadence is a
translation of his intraday discretion onto this ledger's TSMOM framework. Verdict:
FAIL, and the pre-declared head-to-head against plain TSMOM was not cleared either.*

The declared null hypothesis was not buy-and-hold but the plain full-size 12-month rule
— the claim under test was the management principle, so the bar was relative: MAR at
least 1.10x plain, drawdown no worse, CAGR at least 0.80x. On the test window
(2016-07-13 to 2026-08-07) the pyramid earned +11.0%/yr against plain TSMOM's +12.1%/yr
with an IDENTICAL -33.7% daily drawdown: MAR 0.91x, DD 1.00x, CAGR 0.91x. The failure
mode was named in the registration notes before the run: scaling in gives up
early-trend gains, and the 12-month exit still leaves the position at full size at
tops, which is exactly where the drawdown lives — so the pyramid pays the entry cost of
the sizing and collects none of its protection. The principle's whole case is that
failed signals are cheap at 1/3 size, and SPY's monthly signal simply does not whipsaw
often enough for that discount to matter.

The standard gates add the usual two nails. Against hold-SPY it failed everything (CAGR
0.72x, MAR 0.72x, DD 1.00x), and on the four never-searched sectors it made +3.9%/yr
while holding them made +10.0%/yr — the same trap-4 non-transfer already recorded
against the trend3x family. A management overlay inherits its base rule's transfer
problem; this is now the second demonstration.

One caution on what this does and does not refute. It refutes THIS translation —
monthly cadence, 1/3 steps, SPY — not Hougaard's intraday practice, which is
discretionary and therefore not testable on this rig at all. But the direction of the
failure is instructive: the principle is sold on cutting losers cheaply, and at monthly
resolution on an index that grinds upward, the cost of arriving late to every trend is
larger than everything the small entries save.

---

## Addendum, 2026-08-25 (2): short volatility, the pre-named third stream (`svxy-contango`)

*Section 8 named short volatility as a stream that could be a genuinely different bet,
and this entry tested it: the Simon & Campasano (2014) contango signal — already FAILED
here when pointed at SPY — pointed instead at the instrument the paper actually means, a
short-VIX-futures ETF (SVXY, fetched for this test). Signal wording copied verbatim from
the failed sibling so the entries differ only in what is traded. Verdict: FAIL, the
57th idea and the 56th non-pass.*

The number that kills it is not the drawdown, it is the gate's uselessness: the VIX
curve sits in contango 93% of days, so the rule is nearly always long and the signal
adds nothing but whipsaw. On the test window (2022-02 to 2026-07) it earned +8.8%/yr at
-43.7% against +17.4%/yr at -46.4% for simply holding SVXY — the filter cost half the
return and kept essentially all of the pain — and against hold-SPY's 0.62 MAR the
strategy's 0.20 is not close. The train half (+38.7%/yr) is carried by the pre-2018
-1x era, an instrument that no longer exists; the entry declared that leverage change
in advance and the test half answers for the instrument that does exist.

The honest reading for §8: 'short volatility' as implementable in this account — a
long-only position in a half-leverage ETF, timed by a signal that is almost always on —
is not the diversifying stream the section hoped for. The premium the literature
documents lives in futures spreads and options, instruments this programme's account
does not trade. That closes the second of the three named directions on the same
grounds as multi-asset trend: not falsified as an idea, unreachable in this
implementation universe.
