AXOFUND

Correction of record — 2026-08

The four-gate chain below is the published record of a closed era. On 2026-08-08 we measured the survivability score against the forward window it claimed to predict and found rank correlation 0.98: the score restated its outcome. The chain was retired. The current protocol is pre-registered per-family adjudication on realised forward evidence: every book over 27 cells (three windows, three direction arms, three objectives) at leverage one, verdicts frozen before data contact, demotion on forward evidence only. The exposition below stays published because the reasoning that built it, and the measurement that killed it, are the same discipline.

Method

Four gates, in order. No idea skips one.

A candidate earns capital by clearing tests that answer different questions. P: the edge is specifically this book's, not a factor in costume. S: the result would survive being traded rather than being the best draw of a search. L: the size it can actually carry. A: the money came from the mechanism the book named, and not from something else the book happened to be holding. Ranking is lexicographic, P first, and that ordering is a measured result rather than a preference: on the only forward record we hold for it (n = 44 windows, one asset complex), passing P is the strongest predictor of forward survival we have measured; candidates that cleared S and failed P rank materially below it; and the top of the in-sample leaderboard ranks below both. A high composite score with a P-failure is an anti-signal, not a recommendation. The ordering and the sample are published; the effect sizes are withheld as calibration. We treat the result as promising rather than proven: it comes from one complex, and its out-of-corpus replication is still owed.

Nothing a model emits adjudicates here. Models author strategy code and propose hypotheses at volume; every verdict on this page is arithmetic, and every model opinion in the stack is advisory by construction.

The chain

Four legs, and the one we added late

The fourth leg is younger than the other three. Attribution became a gate on 2026-07-30, after an audit and a placebo failure described below. Every verdict this house issued before that date cleared three legs, not four, and we date the change rather than absorb it.

Gate architecture as of 2026-08-02. Attribution was added 2026-07-30, after an audit and a placebo failure described below.
Leg The question What arbitrates A failure means
P — specificity Is the edge this book's, or a factor wearing its name? A matched control ensemble, plus a calibrated edge-free generator Something real earned, and it was not the claimed mechanism
S — survivability Would this result survive being traded? Seven pass/fail floors, then a score The winner is a curve fit until proven otherwise, and it was not
L — leverage What size can it carry at the worst window? Bootstrap sizing against engine margin, netted across shared legs Not that the book is wrong, but that it is smaller than it looked
A — attribution Did the P&L come from the mechanism the book named? A placebo built to null the mechanism, not merely to match composition Pattern-only. It reclassifies as a control

Gate S

Survivability — guilty until proven robust

The gate runs after grading and treats the winner as a curve fit until it proves otherwise. That inversion is the whole design. An optimiser's job is to find the parameters that scored best on the data it was given; the harder it searches, the better that score and the more of it is luck. The grade-A backtest is therefore not the safest candidate. It is the most suspect one, because it is the one the search worked hardest to inflate.

Every candidate runs unlevered at fixed size: sizing is an L-question, and compounding artefacts have manufactured large mirages out of losing books. The gate is floors-first: seven pass/fail checks, and only a candidate that clears every one earns a continuous score. A large edge cannot buy its way past a fatal flaw. The floors, without their thresholds: the edge is distinguishable from chance; it is unlikely to be a net loser under block bootstrap; it survives execution-cost jitter; it clears a multiple-testing noise floor that rises with the number of trials tried; the tuned winner does not collapse below its own authored defaults out of sample; the winner sits on a plateau rather than an isolated parameter spike; and no walk-forward fold liquidates.

P( mean return ≤ 0 ) under stationary bootstrap resampling of the daily path

Ranked by how often they cut a graded winner, the top two killers are knife-edge isolation and tuning overfit, which are exactly the two failure modes a shop that deploys its best backtest ships without noticing. The ordering is published; every threshold value and the per-floor kill counts are withheld as calibration.

Which statistics actually predict out-of-sample survival is a separate question, and we measured it on our own corpus rather than assuming it. Taking an archived sweep cohort whose selection timestamps give a clean forward split, 43 of 690 archived candidates reproduced exactly under the current registry; each in-sample gate leg was scored by its discrimination against survival over the first 30 out-of-sample days. Every magnitude statistic we tested (headline grade, raw Sharpe, deflated Sharpe, the permutation p, the cost-jitter check) ranked inversely. Everything that certifies "this edge is large and real" landed below a coin flip. Only two legs ranked correctly, and neither of them measures the size of the edge. The inversion is published, because it is a negative and it generalises. Which gate legs the two are, and the discrimination values of our own instruments, are withheld as calibration.

Then the same measurement humbled itself. Re-run at a corrected, dimensionality-scaled trial budget across roughly 216 graded cells, the leg that had ranked best under the earlier flat budget inverted. Same metric, same computation, opposite sign. A metric's calibration is a property of the pipeline that produced it, not of the metric alone, so any leg validated on a mis-specified budget is suspect until re-validated on the real one. We now re-check a leg every time the budget or the search space moves.

So the gate ranks on consistency and emergence rather than on magnitude. A front-loaded book that made most of its P&L early in the window and then reversed is floored regardless of its headline number: that shape is a rejection signature in our own forward record, and no headline figure overrides it. The floor's parameters, and the size of the forward penalty, are withheld.

Gate P

Specificity — the load-bearing filter

Each candidate faces a control ensemble built to null exactly the mechanism it claims. A basket strategy is re-run on random same-size baskets; if random selections do as well, the chosen basket carried no information. Every control is a full re-run through the same engine, and the ensemble size is pre-registered before any of it is read.

p = ( 1 + #{ controls ≥ treatment } ) / ( N + 1 )   N full re-runs, N pre-registered

Timing strategies face a different null (a regression of forward returns on the market factor, testing the residual), because a shuffle null destroys beta and mints false positives. Per-window p-values combine by Fisher's method (−2Σln p ~ χ²) and the family-wise verdict applies Benjamini–Hochberg at a pre-registered false-discovery rate, with consistency required across independent windows.

The gate proves itself on planted controls. Every market we enter gets a deliberate factor strategy seeded among the candidates, expected to clear S and die at P. The G10 carry basket did exactly that; the full exhibit is in the 2026-06 working paper Enforced structure ports. Statistical structure does not. If a market's planted control ever passes, every verdict from that market is void until the gate is fixed. That rule has been invoked.

How often P fires is the number the page owed and never printed. Of 265 candidates that accumulated enough independent windows to face it, 19 survived. Roughly nine in ten of the things that reach the specificity gate die there, and each of those candidates was already the surviving parameterisation of a several-thousand-trial search.

A matched control answers the wrong question

A matched control answers whether a book beats other books like it; it cannot answer whether the book beats nothing at all. A calibrated edge-free generator therefore runs beside the control ensemble, and a book must clear both. The incidents that forced that correction, and the doctrine reversal we published with it, are set out in Two nulls, because they answer different questions (2026-08).

Gate L

Leverage — what size can it actually carry?

L* = min( bootstrap-sized, engine-margin ) taken at the worst window

Most survivors cannot carry even 1x. That is a distribution, not a failure of the books: it is what honest sizing returns when the worst window binds. The distribution itself is withheld.

Books whose relationship is enforced (a rebalance mandate, a defended band, share-class identity, contract convergence) take scenario-sized leverage instead. A bootstrap cannot contain a break that never occurred in its sample, so a discontinuity the size of the largest peg abandonment on record is imposed by hand on any band book rather than inferred from history. Every enforced book also registers a same-day kill trigger on its enforcement mechanism, because the risk that ends an enforced relationship is the enforcer stopping, and that arrives as news rather than as a drawdown.

Capacity is measured with shared legs netted. Two books that trade the same instrument do not add capacity; they divide it. Netting runs across the whole book set before any sleeve is sized, and it always nets down.

The fourth leg

Attribution — did the mechanism earn it?

Beating your controls is not the same as earning from what you claimed. A book must attribute its P&L to the mechanism it named. This became a gate leg on 2026-07-30 for two reasons, both of them ours.

The first was an audit. Of a twenty-book set under review, fifteen were pattern-only: no named enforcer at all, just a relationship that had held. "These series co-move" is a screening observation, not a mechanism, and a candidate that cannot name an obligated party and the trigger that binds them now pre-classifies as a control at authoring time rather than being discovered to be one later.

The second was the placebo that passed. A placebo which shuffles the data but leaves the claimed leader intact does not null the mechanism; it nulls nothing, and it will pass alongside its treatment. So the placebo is now specified against the mechanism, named at authoring, and a mechanism field is machine-readable on every seed. Enforcement alone is also not sufficient: a real, dated, contractual obligation can be fully priced, discharged exactly as written, and pay nobody. We measure the dislocation against round-trip cost before authoring, not after.

FIG. 1 — ATTRITION AT THE TWO ARITHMETIC GATES
S — entered 620 S — cleared 35 P — entered 265 P — cleared 19
Two censuses, drawn to their own denominators. Top pair: of 620 graded optimiser winners in the survivability census of 2026-07-11, 35 cleared all seven floors, roughly one in twenty. The largest single cause of rejection is a winner sitting on an isolated parameter spike rather than a plateau; the second is a tuned winner losing to its own untuned defaults out of sample. Bottom pair: of 265 candidates that accumulated enough independent windows to face the specificity gate, 19 survived: per-window permutation p-values combined by Fisher's method under Benjamini–Hochberg false-discovery control at a pre-registered rate, with consistency required across windows. The two pairs are different populations at different vintages. They are not a chain and they do not multiply. Every threshold value, and the per-floor kill counts, are withheld as calibration. The authored → tested → survived funnel is published on the index and is not restated here.

Evidence protocol

Every test runs in a point-in-time cell: it freezes its universe and its parameters at a dated T0 and is read on at least three independent forward windows. One window is not evidence: a single-window cell is recorded as untestable, not as a result. Stands on: What a campaign declares before it runs, the charter

A result counts only on event-driven fills. Every stage (the first optimiser trial, the Monte-Carlo sidecars, the forward fleet, live) runs the same strategy object over the same event stream. There is no port, so there is no port drift. The exhibit is a counterexample: two engines given identical formulas disagreed not on noise but on which trades exist, because arrival-order cross-sections, event-count clocks and decoupled decision-and-fill get silently rewritten in a wide-frame port. The price is far slower sweeps per trial, and we paid it in engineering rather than in correctness. No published item yet; the counterexample sits in the internal record.

A one-bar entry lag is enforced by the harness. A signal measured on bar t fills at t+1, universally, and it is a property of the harness rather than a parameter, because anything left to the author will eventually be got wrong by one author. Both magnitudes attach to results we corrected: a same-bar act-at-close measurement inflated a steelman edge roughly fivefold, five-sixths of it stale-print and bid-ask noise reversion; and a share-class pair that was positive in every calendar year at same-bar entry was negative in every calendar year lagged one bar, with a block-bootstrap probability of a non-positive mean above 0.999. Review rule: "enters one bar after it decides" must be proved from the code, never from the prose. Stands on: Enforced structure ports, §3

Search honesty means random search at production-parity budgets; Bayesian optimisation was retired after we measured it indistinguishable from random out of sample. Trial budgets scale with a strategy's dimensionality and the optimiser is allowed to overfit as hard as it likes, because the search is priced downstream: every trial raises the bar the winner must clear. Capping the search caps the discoverable edge; pricing it does not. The honest caveat from our own corpus is that this makes aggressive search safe, not productive; those are separate findings. Stands on: We measured our own instrument, §2

Every tuned winner runs against its own untuned defaults, out of sample: identical code, identical data, identical window; the only variable is the parameter dictionary. We owe this lens to a bug: an audit found that none of our forward books were trading optimiser winners (every one ran authored class defaults), which handed us the control arm we had never built. In the first probe, five of six strategies did worse tuned than untuned, two of them already holding deployment slots. Below a minimum out-of-sample trade count the verdict is recorded as direction, not evidence. No published item yet; the probe sits in the internal record.

Before any real verdict is read, the whole pipeline runs on calibrated, edge-free synthetic data per market to measure its own false-discovery rate. At full search width it produced a best-of-search survivor on every edge-free panel across six venue calibrations: the manufacture happens upstream of the gate, the downstream nulls are what remove them, and the cost being reported is optimisation, gate and review capacity consumed, not verdict validity. The same instrument measures the other side: a kill verdict without a stated minimum detectable edge is not a finding, and our search's detection floor currently sits above the dislocation most genuine contractual obligations actually pay. The generator's own out-of-sample moment-match validation is outstanding, and until it lands the harness is not validated and neither is any rate quoted from it. Stands on: We measured our own instrument

Charters are pre-registered: the kill bar, the success metric and the stop rule are written before the run, not after the result. Campaigns carry a calendar stop (the current one is 21 days) and no extension without a re-charter. Pre-registered designs pin the exact artefacts they will read; "the newest run" is not a specification. Stands on: What a campaign declares before it runs, the charter

Negatives are published, and no model verdict is ever a floor. Refuted ideas are archived with their evidence and appear in the regime library. Model opinions are advisory throughout; consolidation lives in code, and the gate is arithmetic. Cross-references on this site are hand-maintained; a link check runs before publish.

Standing item

The leakage standard is house rules, not reviewer taste. Strict trailing windows, so warm-up bars return nothing rather than a number. Causal labels only: any label whose definition consumes data after its timestamp is leakage, including cross-time terciles and full-window percentile ranks. The universe is frozen point-in-time and includes the names that died; a universe of "the instruments that still exist" has already excluded every failure. And no selection on a forward outcome: not the fold, not the instrument, not the window, not the parameter set.

Everything learned from data fits after the split, not only models: imputation medians, encoders, quantile cut-offs and normalisers all fit inside the fold. Choosing a threshold over all history and then testing on part of that history means the test already knows the answer. Fold boundaries carry an embargo and a purge: the embargo equals the forward horizon, so a label that looks h ahead drops the h straddling the boundary, and any event crossing the seam is purged, so a trade opened in training and closed in test is not counted twice. Resample-bin semantics are commented at every call site. A left-labelled, left-closed resample is safe; a floor-to-frequency can read up to a full bin of future data into the current bar. One wrong bin is enough to leak, so every resample carries a comment saying which way it points.

The data itself leaks. Unadjusted corporate actions manufacture edges that pass naive screening: on unadjusted bars, three ex-dividend sawtooth pairs out-scored the genuine share-class mechanism they were meant to sit below. One daily-bar market in our store records zero delistings in four years across more than four thousand tickers: that is survivorship absence, not survivorship bias, and it independently blocks every cross-sectional book on that venue. Data repair outranks new research when the store is manufacturing candidates faster than the gate can refute them.

The honest limit: an event-driven engine makes execution lookahead structurally hard; it does nothing about a leaky feature. This discipline sits beside the engine, not inside it. Stands on: the leakage standard, a standing item