We do not claim capability from machine learning. Models author strategy code and generate hypotheses at volume in this shop, and nothing a model emits is permitted to adjudicate. The gate is arithmetic.
What follows is an architecture description and three measurements taken against ourselves: that our own language-model judges herd, that a learned scorer out-ranks the hand-built gate we still refuse to let it replace, and that the dominant failure mode of a machine-learning estate is not a crash but a component returning success while doing nothing. Every claim here is negative or self-limiting. That is deliberate. A capability claim invites a proof we cannot give without publishing the thing that pays us.
§1
Ideas enter as queued experiments: from scouts reading outside literature, from a regulatory and rulebook feed, and from the system's own follow-up proposals off contradictions in its record. A novelty gate decides whether a proposal is a genuinely new mechanic or a re-parameterisation of one we already hold; a re-parameterisation is what the optimiser does anyway, so it never triggers authoring. What survives that is written as a strategy module by a tiered cascade of models. Every authored module, from any tier, faces the same admission test: it must import cleanly and fire trades on real bars. Code that does not fire is discarded before it costs a single optimisation hour.
From that point on, no model opinion enters any verdict. The claim is verifiable by structure rather than by benchmark: no model output is an input to any leg of any gate, and that is checkable by reading the gate. It is a governance property, not a performance one, and it is the opposite of a capability boast: a firm that has wired a model into its decisions cannot make it.
| Machines propose | Statistics dispose |
|---|---|
| Hypothesis and mechanic generation | Survivability floors on the bootstrapped return path |
| Strategy code authoring | Specificity p-values against matched control ensembles |
| Novelty screening against the existing book | Attribution of profit to the named mechanism |
| Naming, summarising, triage of the record | Leverage taken as the worse of bootstrap and engine margin |
| Quality and duplication flags on written output | Promotion, retirement and every kill verdict |
The scarce resource is not compute but the supply of genuinely new mechanics, and the tokens spent writing code for them. A sweep searches the parameters of one mechanic; it cannot invent a second one. A busy cluster is not evidence of progress.
§2
Language models are widely used as reviewers, and the assumption underneath that practice, that a second model is a second opinion, is testable. We tested it on our own tooling twice.
The first test asked two judges of different lineages to call survive or fail on high-scoring candidates from in-sample information alone, with the forward outcome held back. Both rejected everything. They agreed with each other completely and trivially, and scored exactly the base rate. There is no diversity to be had between two reviewers when neither has any discrimination to disagree about.
The second was an instrument built for the purpose, scoring paired judge verdicts against ground-truth forward outcomes. It returned three findings in the same direction: the judges' errors are strongly correlated, so they fail together; their discrimination against survival is indistinguishable from chance; and the second judge adds essentially nothing over the first. The direction and the sample are published here; the correlation and discrimination values for our own review instrument are withheld, as calibration (see Redaction, below).
§2.1
The result is only worth publishing because it had a consequence. Judges no longer touch survival verdicts; they do novelty, quality and duplication work, where their output is checkable by a human in seconds. Where a panel is still used, it is built against its own tendency to converge.
The panel uses role-diverse lenses rather than putting the same question to N models, because asking one question of five models samples one opinion five times. A reviewer is required to produce the strongest case against the position it just took. Agreement on a question known to be difficult is treated as a defect signal, not a confirmation. The strongest lever is an external, non-model anchor: every model verdict is cross-checked against the statistical gates, because models drift together and statistics do not. A house ontology injected into every prompt raises consistency and raises anchoring with it, which makes it good for naming, bad for independent judgement, and it is kept out of judgement.
The public literature reached the same place before we did, which is why our reading list carries a shelf on crowds, herding and correlated error. We cite it rather than claim the finding as ours; what is ours is the measurement on our own instrument, and the demotion that followed.
§3
Our survivability gate is a conjunction of thresholds over a fixed set of in-sample legs. We asked whether a model over exactly the same legs (no new information, no new features) predicts forward outcomes better than the thresholds do. It does, and the ordering is stable across two independent bake-offs.
The ordering we publish: a tabular foundation model ranks first; pruned tree ensembles come next; our threshold conjunction and a logistic regression over the same legs tie at the bottom. The mechanism is in that last fact. A linear model over the legs only matches the gate, so the lift is not information the gate lacked but nonlinear interaction between legs: a weak reading on one leg is tolerable conditional on a strong reading of another, which an AND-of-thresholds cannot represent by construction. Which legs interact, and in which direction, is withheld as calibration. Our conclusion at the time was that the gate was not wrong, only the wrong shape. The retirement of 2026-08-08 went further: mean-based scoring cannot see returns carried by the tail, whatever its shape.
The measurements: grouped cross-validation by strategy on n = 685 forward-evaluated cells, replicated in a full bake-off at n = 708–764 with every candidate scored against a 50-times shuffled-label null, and checked walk-forward by training on one month and testing on the next. Every candidate cleared the null; the discrimination values, and the margin between them, are withheld as calibration of our own instrument.
The actual content of this section is the policy. The scorer is advisory, never a floor. It may re-rank candidates; it may not admit or reject one, and no promotion decision has ever been taken on it. The reason is that it is trained on the corpus it would gate: gating on a model fitted to the same population is in-sample selection wearing a machine-learning hat, and ranking is the only honest use of it. Promotion to a hard floor requires forward evidence from outside its training set, a held-out strategy set or a later campaign, and until that exists the answer is no. The artefact manifest records the training corpus, so that requirement stays auditable by someone who does not trust us. The scorer is retrained per campaign and never trained on the cells it scores.
A self-denying ordinance published in full is worth more than an accuracy number, and it is the only part of this section a reader can hold us to. Related reading: our shelves on tabular deep learning and evaluation metrics.
§4
That programme had already been declared dead. On an early pilot we concluded the legs were exhausted and carried no learnable signal, and we stated it twice. The conclusion was wrong, for two nameable reasons: the pilot was small, and a defect in the feature pipeline was nulling one of the inputs, so a leg that existed was being measured as if it did not.
The standing lesson is not "we were wrong once". It is that dramatic results on this material, dramatically good or dramatically dead, are almost always small-sample or leakage, and must be verified before they are believed. That applies to our own scepticism as forcefully as to a strategy that looks like a discovery. The re-measurement was therefore run against a shuffled-label null before it was accepted, which is the same treatment any candidate edge gets here.
§5
On 2026-07-11 we retired a pre-trained time-series forecasting programme, having falsified all three of the uses we had built for it (as an entry gate, as a per-trade direction classifier, and as an input to regime detection) across five days of testing. One mechanism explains every failure: the model sets its forecast band from the realised volatility of the context it is fed, so it reprices the recent past rather than anticipating the future. Its lead-lag profile against realised volatility peaks at zero to one day forward, measured over 607 four-hour bars. A quantity that tracks the past coincidentally cannot lead it.
The server was stopped and the scheduled prefetch removed the same day, the hardware was reclaimed, and the one branch we did not falsify (a purpose-built fine-tune judged on incremental information over the backward volatility feature) was left open and unbuilt, because nobody has built it. A published retirement with a date on it is rarer than a published launch, and it is better evidence of how a research process actually runs. Full record: the decommission notice A forecasting programme, retired (2026-07).
§6
The characteristic failure of a machine-learning estate is not a crash. A crash is loud, dated and fixed the same day. The characteristic failure is a component that returns success while doing nothing: a health endpoint answering "ok" for a service that has lost its accelerators, a scheduler reporting a finished job that never ran, a feed returning a valid empty response. Every log line looks normal. The output is plausible. Nothing alarms.
We now treat this as a class rather than a run of bad luck, because in the week to 2026-08-02 a routine operations review turned up six instances at once, in six different subsystems.
None of these corrupted a verdict on its own. Together they establish the risk class, and they are the reason the engineering is built to fail loudly rather than plausibly.
§6.1
Given a choice between a component that breaks visibly and an approximation that is wrong invisibly, we pay for the first. The clean case is our own execution layer: when it broke, the books submitted 700 orders over a twelve-hour parallel window and filled none of them. Every one was denied or rejected at the risk and matching layers. The defect was undeniable, diagnosed and fixed the same day, and the honest conclusion recorded at the time was that the comparison the run was meant to settle could not be made at all. A vectorised approximation of the same books would have returned a confident, well-formed, entirely fictitious result.
The operational form of the doctrine travels further than the anecdote. Every data source carries a declared minimum expected item count, so a successful empty response is an alert rather than a no-op. A source that is legitimately empty and one that is silently empty are different states and are recorded as different states. Success with zero rows is an error. A plausible empty result is never allowed to pass as a finding.
§6.2
One more failure worth publishing, because it is the most transferable thing on this page and it costs us nothing to give away. In an earlier era of our search, one gate leg was the single best out-of-sample predictor we had. We then corrected the trial budget underneath the search (the same leg, computed the same way, on roughly 216 graded cells at the corrected budget) and it inverted, from strongly predictive to worse than a coin flip. The leg is not named and the two discrimination values are withheld; the direction of the flip is the publishable part.
In the earlier era the metric had been measuring an artefact of an under-powered search. When the search got the budget it needed, the artefact went away and took the apparent signal with it. The rule that follows is standing: a metric's calibration is a property of the pipeline that produced it, not of the metric alone. Re-check every leg's discrimination whenever the trial budget or the search space moves, and treat any leg validated under a mis-specified budget as suspect until re-validated.
§7
Inference runs local-first on hardware we own. The panel is deliberately multi-lineage, so cross-review is not one model agreeing with itself; §6 records what happens when a lineage silently drops out. Neither fact is a claim about quality. They are structural properties, chosen so that the failure modes we care about are the ones we can see.
The authoring inversion is the part of this we would defend hardest, and it runs against the trend. We do not ask the author model to be clever. We ask it to remember. Three memories, keyed by strategy family, go into every authoring call: the code of strategies that passed the gate, as exemplars; a registry of mechanics that were tried and died, with the reason, as anti-examples; and semantic retrieval over our own research corpus, so the author sees what we already know about the kind of edge it is being asked to build.
The evidence is an A/B against an ungrounded control on the same model, so the only variable was memory: grounding cut the rate of authored strategies that never trade by roughly twentyfold. The two rates themselves are withheld. The model did not change. The memory did.
The honest frontier is not the one a capability story would choose. A bigger author is mostly the wrong move. Of the ideas we discard, essentially none fail on code; they fail on logic, as untradeable or seldom-firing mechanics, or they are correctly rejected as duplicates. The number of survivors is bounded by how many real edges the market offers, not by the author's fluency. What compounds is the corpus of our own survivors and corpses, which exists only because we kept the failures.
The boundary, stated plainly: this is a description of what we built. It is not a claim about what it can find. We publish no benchmark of our own stack.
§7.1
The computing layer under all of this is one private cluster, owned outright: commodity CPUs and consumer GPUs, a shared file store, and a batch scheduler with three priority lanes, in which live-book health checks preempt calibration and calibration preempts research. The node, core and GPU inventory is withheld. The bridge that matters is not in the hardware: the estimator, the strategy that consumes it, and the scheduler entry that runs it are maintained together, so when a result is wrong the fix lands in whichever layer was actually at fault, the same day.
Scale, with its apparatus. On 2026-08-07 this cluster ran a pre-registered case-control study (54 analyses: 17 equity exchanges plus one crypto venue, across three event thresholds) and a 177-cell paired backtest tournament, both submitted, verified and adjudicated inside the same day. Optimisation campaigns run at production trial depth on the same scheduler; the store they read is described on data, and the models the GPUs serve are described in §1 and §7. This is high-performance computing at commodity prices, and the point of owning it is not speed. It is that the whole research loop, from bar ingest to a published figure, runs on hardware in the building, with no dependency that can change its terms.
The engineering lesson we would repeat to anyone: the limiter has never been the hardware. Each time throughput stalled, the cause was our own code or configuration. A parameter sweep gained roughly 20x from removing an interpreter lock and adding a bar cache (2026-07-11): the wall was code. A scheduler affinity default on hybrid-core silicon pinned every job to the same eight hardware threads; removing it doubled measured node throughput (2.02x), and our first estimate of that gain, 2.7x, was itself wrong (measured on an overloaded node) and is corrected here the way any other number would be. The operating doctrine that followed: an idle queue is investigated as a fault.
§8
We will not claim that models find edges. Models propose; every edge on our books must survive a control ensemble on realised forward evidence; the survivors of the gate retired on 2026-08-08 have not yet done so and are requeued as unproven. We will not claim predictive capability of any kind from a model: the one forecasting programme we deployed was falsified three ways and retired, and we published that instead. We will not publish benchmark numbers for our own models, judges, or scorers, because their discrimination values are calibration, and calibration is the asset. We will not offer vendor names as credentials: which model authored a strategy is an implementation detail with a shelf life of months, and it is not evidence about anything.
The reason is one sentence. Any capability claim invites a proof we cannot give without leaking, and a claim that cannot be falsified is marketing.
The decommission notice A forecasting programme, retired (2026-07) cites §5 of this paper as the architecture's one published retirement. Cross-references are hand-maintained; a link check runs before publish.
2026-08 — first publication, status current. Title and date are fixed at publication and are how this item is cited. It will not be silently edited; a correction is issued as a new item that cites this one.
2026-08-10 — an exception to the line above: passages describing the gate chain as current were corrected in place, each carrying its date, after the chain was retired on 2026-08-08. The correction of record is on method.
Withheld from this paper, deliberately and by standing policy: the correlation and discrimination values for our own judges and scorers; the identity of the gate leg that inverted in §6.2 and its two values; the trial budgets, thresholds and floors of the gate; model identities, sizes and serving locations; the size of our survivor and failure corpora; and the contents of the mined family recipes, which are distilled tradeable parameters. Where a value is withheld, the sentence says so.