Confidential
01
↓ / → to advance
Isard Labs Isard Labs

Technical documentation · Level 3 — Deep

The internals, named.

The methods, ladders, risk machinery and governance behind the pipeline — described precisely enough to interrogate, for reviewers who push on every claim.

Most sensitive tier. Strictly confidential.

How to read this tier

We name the methods. We don’t hand over the recipe.

What’s here

The machinery

Named techniques, ladder structures, signal families, isolation and risk design, and the governance that enforces them. Enough to evaluate the engineering.

≠
What stays private

The edge

Exact gate thresholds, the proprietary research that sets them, specific signal formulas, and live strategy specifications. The discipline is shown; the alpha is not.

The claim being defended here is the rigor, not any single strategy.

01 / 06 · platform

εὕρεσις — “discovery”

HEURESIS

Search wide · prune hard

HEURESIS · deep

Search wide, prune hard, never validate the same thing twice.

  • Grid enumerates the space cheapest-first from a frozen YAML template; a pruner turns red results into a skip-set that shrinks the next batch’s search space — a combinatorial feedback loop.
  • Pre-registered batch screens are the working mode today. An experiment identifier is issued at run start, the pre-registration is frozen by hash, and the tool path plus the code commit — with a dirty-tree flag — are stamped into the output artifact. Amendments are possible only through an explicit, noted amendment; a pre-registration edited without one turns the bar red.
  • Four-layer de-dup — in-session fingerprints, a cross-session knowledge base, the validator’s own run history, and the local pre-flight gate — means compute is never spent on a known answer.
  • An LLM proposal pipeline exists and is not in service. Five stages, ending in a deterministic resolver that makes zero model calls. Its semantic critic checks that a chosen primitive’s actual semantics match the described intent — catching, for instance, a rate-of-change primitive wired to a concept it does not measure. The mode is currently cold, by decision.

HEURESIS · containment & discipline

A creative engine that cannot cheat — and is measured on discipline.

  • Hallucination containment. A proposal for a signal that doesn’t exist becomes a structured request into AGORA — it can’t reach into the validator to make itself pass.
  • Staged inference, and the rule that makes it honest. Multiple-testing control runs within a stage and never across it: a screen row may disclose but can never bank a claim, and only a confirmation row on data disjoint from the screen carries one. When a batch measured zero of fifty-five confirmations surviving on untouched hold-out, the earlier stage-level admissions were closed on that evidence — and the banked total is published as zero rather than omitted.
  • One predicate, two callers. A screen may not report until both its controls have fired, and the acceptance decision must be a single predicate invoked by the live path and by the negative control alike — because a control evaluating a different predicate than the run cannot fail on the dimension that matters.
  • Research spirit, enforced statically. An auditor checks label provenance, undeclared acausal constructs, full-sample statistics inside screens, and private metric implementations — the last as a shrink-only ratchet over known historical sites. Rules it cannot yet check are marked skipped with a stated reason, not silently passed.
  • Green bar as code, with named falsifiers. A property must name its falsifier: a check that ran and produced a value proves nothing unless something establishes the value could have come out differently. Its dashboard check alone ships more than two dozen, each shown to fire and to stay silent.
  • Type debt as a bounded ratchet, not a claim. Its type checker stands red and has never been green; a separate gate fails on any growth and on unrecorded paydown. A red bar honestly red beats a green bar made green by deleting the check.

02 / 06 · platform

αὐτόματα — “self-acting”

HEURESIS AUTOMATA

The experiment layer · provenance by construction

AUTOMATA · the three axes

Where a number came from, what it is, and what it may decide.

A value can be impeccably sourced, correctly labelled, fully documented — and still be doing a job nobody authorised. So every number carries three independent axes, and the third is enforced by the type system rather than by review.

AxisAsksValues
ORIGINHow was it produced?observed · computed · reproducibly stochastic · externally reported · interpretive
KINDWhat authority does it carry?canonical · stated default · derived
ROLEWhat is it allowed to decide?gate · disclose · resolve · diagnostic
  • Role is a type, not a field. A gating value and a disclosure-only value are different types. A promotion predicate accepts only the first, so passing the second is an error from the type checker before anything runs — the failure mode where a disclosure quietly hardens into a verdict becomes unrepresentable rather than merely forbidden.
  • An LLM cannot author an empirical number. Not enforced by a flag — the type a model is permitted to return has no numeric field at all.
  • Reproducibility is checked against its own requirements. Only a computed value reproduces from code alone. An observed value needs its snapshot, a stochastic one its seed manifest, an externally reported one its source version — and the ledger refuses a value recorded without its companion reference.
  • Absence is visible. A missing measurement rendered as 0.0, False, [] or pass reads exactly like a reassuring result. Unspecified and default are different values, and the validator fails closed.

AUTOMATA · the constitutional suite

Seventeen synthetic worlds where the right answer is known.

The release gate is not a test of the strategies. It is a test of the research process itself: build a world whose ground truth you control, then check that the apparatus reaches the correct conclusion about it.

no alphaembedded alphalook-ahead trapfactor disguiseparameter lotteryconditional regimecost cliffprompt injectionstrategy decayexecution impossibilityaccounting corruptionmultiplicityunderpoweredcherry-pickinglane violationbenign absencestatic role enforcement

Every world is paired with its foil. A system that refuses everything passes “no alpha” trivially — so the embedded-alpha world exists to prove the apparatus can still find a real edge. A cost model that reports every strategy as dead passes the cost-cliff world exactly as vacuously. Each world therefore carries a second arm whose only job is to fail if the first is being satisfied by pessimism.

That is the whole doctrine in one artifact: a control that cannot fail is not a control, applied to the validator rather than to the strategy.

AUTOMATA · the experiment record

An experiment you can replay without trusting the person who ran it.

  • The Experiment IR. A typed envelope around the strategy spec: hypothesis, null, as-of date, snapshot identity, splits, minimum detectable effect, baselines, controls, trial family, seeds, fidelity class, independence unit. Frozen and hashed; immutable once it has seen evaluation data. A research screen is never a script — if something cannot be expressed as a record, that is a capability request, not permission to write another one-off.
  • Leakage refused at compile time. Static temporal analysis of the expression tree rejects a look-ahead construct before data is read, rather than after a validation run has been paid for.
  • Power before compute. The minimum detectable effect is computed — and requested from KAIROS, who own that instrument — before compute is allocated, so an underpowered experiment is refused rather than run and then interpreted.
  • A declared differential oracle. A second, independent execution and accounting engine exists solely to disagree with the imported one; its only permitted output is a divergence report. The most recent run compared twenty measurements across four cases with zero unexplained differences and eight declared ones.
  • Sequential error control. The search is adaptive — results are inspected, branches chosen, families abandoned — so a fixed-horizon significance rule is invalid. Trial families carry versioned policy epochs with online error control, and thresholds cannot be retroactively re-based.
  • Impact invalidation, never deletion. A data or code incident marks downstream evidence as pending recomputation and keeps it. Negative results stay searchable; the record of a wrong number includes the fact that it was wrong.
  • It survives being killed. The durability gate sends a real SIGKILL mid-generation and requires no duplicate, no loss and no mutation — with the event chain intact across the restart. Its foil, run alongside, loses a trial.

03 / 06 · platform

ἐπιστήμη — “demonstrable knowledge”

EPISTEME

The eleven-rung ladder · the only certifier

EPISTEME · the ladder

Eleven rungs. Climb in order. First red halts the run.

123 456 789 1011 +5 sub-rungs default on

1 data quality · 2 causality · 3 costs · 4 signal · 5 walk-forward · 6 statistics · 7 path risk · 8 generalization · 9 adversarial · 10 Monte Carlo · 11 deployment. The canonical eleven always run and halt on first red. Five default-on sub-rungs — attribution, two cross-engine accounting parities, regime provenance and per-regime metrics — bring a default green run to sixteen scored records. Microstructure, live-forward certification and live-execution fidelity remain opt-in and are off by default today.

EPISTEME · signals

A hand-built signal library — no off-the-shelf indicators.

200+ primitives across more than twenty families, each implemented causality-first rather than pulled from a generic technical-analysis package.

trendmomentummean-reversionvolatilityvolumemicrostructureregimespectralcointegrationfundingopen-interestimplied-volcomposite

Selection is three-stage: domain pre-filtering, then false-discovery-rate control across many tests, then a fresh-data holdout — so a feature can’t survive on multiple-testing luck.

EPISTEME · statistics & path risk

Significance and survival, separated.

  • Multiple-testing control — false-discovery-rate (FDR) control plus permutation p-values, with a ledger so the testing burden is accounted for, not hidden.
  • Path risk — bootstrap and permutation Monte Carlo separate edge quality from path luck; trade-overlap scoring measures independence.
  • Ruin probability — the real stop threshold is derived from the probability of ruin at the current allocation, not chosen by feel.

EPISTEME · adversarial & Monte Carlo

The strategy is attacked before the market gets to.

Adversarial battery

Four categories + live

Signal-validity (reverse direction, shuffle labels, ablate features), edge-fragility (slippage sweeps, freshness decay), structural stress, and robustness (block-bootstrap ruin) — plus a live-execution category.

+
Monte Carlo

Synthetic worlds

Geometric Brownian motion, Merton jump-diffusion for fat-tailed crises and a GARCH overlay — a 3×3 volatility×regime grid, hardened with a GARCH-EVT tail (expected-shortfall) gate.

Each subtest returns machine-readable diagnostics, so discovery can learn from how a candidate failed.

EPISTEME · generalization & lifecycle

One asset is an anecdote. Several is a claim.

  • Generalization gate. A strategy must pass on a floor of ≥3 instruments and clear a pass-rate of at least half the tested universe. (The ratio was deliberately eased from a stricter setting to remove a cliff edge — the 3-instrument floor remains.)
  • Lifecycle doesn’t end at deploy. A CUSUM monitor watches live returns for decay; a breach steps a strategy down through reduced allocation toward retirement.
  • Re-deployment is gated by five separate gates — an enhanced strategy must clear all of them before it earns capital back.

EPISTEME · the agent & governance

The LLM is on a very short leash.

  • Two nodes, both schema-bound. One turns a natural-language idea into a spec; one explains a red gate. Neither can alter the validator or a verdict.
  • Local inference only. Runs against a local model — no external telemetry on proprietary research.
  • Code generation cannot reach the core. The generator runs in three modes behind an allow-list that can never write to the causality, statistics, path-risk, data, spec, core, run-store, ladder or knowledge-base packages. The validator core is human-authored by construction, not by policy.
  • The dependency graph is CI-enforced. An import-linter contract keeps 40+ packages strictly layered; 270+ codified rules each bind a gate to a library, a function and named tests, machine-checked in CI; reproducibility is checked twice on every run.
  • A governed exception, stated rather than hidden. A registered halting gate can be softened to annotation-only — but only with a decision record, pre-committed soft and hard deadlines, and automatic reversion to strict when the hard deadline passes. A temporary exception that cannot expire is a permanent one.
  • 240+ ADRs record why every non-obvious decision was made, and 25 CI jobs plus 66 pre-commit hooks enforce them — including checks for silent exception handlers, orphaned tools, and lint suppressions without an expiry date.

EPISTEME · honest certification

Certifying a searched strategy without fooling yourself.

  • Search and certification windows must be disjoint. For any searched spec the gate asserts the two windows do not intersect. Where a search touched all available history it returns no disjoint window possible rather than inventing a clean slice, and a spec with undeclared provenance fails closed instead of being given the benefit of the doubt.
  • A cohort-level false-discovery ledger. Beyond per-strategy significance, the corpus is asked a harder question: of the strategies we currently call product, how many should we expect to be false? It is deliberately a ledger and not a gate — a number to be looked at, not one to be passed.
  • A contamination register. Windows in which a live book was not actually following its strategy are declared and excluded from evidence — with no enable flag, no override and no force path, because a safeguard that an environment variable can disarm is not a safeguard.
  • A null-hypothesis benchmark asset. A synthetic zero-structure random walk with a pinned seed, ten years and nine timeframes, volatility-matched to the real universe. If a strategy wins on it, it is fitting noise — full stop.
  • Cross-engine accounting parity, named precisely. Positions are re-run through two third-party engines and reconciled to a correlation of 1.0000 and 0.9968 with exact trade-count agreement. That is fill-and-accounting fidelity, not independent re-derivation of the signal — the compiler that would make it independent is designed and unbuilt, and every productization receipt carries a machine-readable field saying engine_independence: false.
  • Claims get retracted. When the cross-engine work was briefly overstated as a “moat,” the claim was withdrawn on the record rather than defended. The receipt field above exists so the overstatement cannot recur silently.
Stated plainly: live-forward certification on post-registration data is built and currently off, awaiting the execution platform’s burn-in export. Today’s greens are certified on search-disjoint historical windows, not on forward data.

04 / 06 · platform

καιρός — “the opportune moment”

KAIROS

Regime manufacture & validation · the ecosystem’s instruments

optional · elective per strategy

KAIROS · the optional fork

It only enters the picture if a strategy asks for it.

a strategy spec no regime gate validated as-is — the regime rung is skipped entirely opt-in regime gate when_regime: … KAIROS-blessed stream validated regime labels

KAIROS manufactures the classifiers and validates them on its own ladder, publishing blessed label streams that any strategy may consume — a research factory with an independent line. A sibling may also author a classifier and submit it behind a documented protocol; that path is supported and, to date, essentially unused in production. The skip is literal: a spec declaring no regime inputs never runs the regime rung.

KAIROS · the gating ladder

Thirty rungs, six families — and only two of them can decide anything.

FamilyChecksGates the verdict?
R · identificationIs the state real and causal?yes
D · durationIs the timing of changes right?yes
E · evaluationAgreement vs oracles (look-ahead)no — diagnostic
F · forecastingPredictive skill of the labelsno — diagnostic
A · aux-contractReliability class of the outputno — diagnostic
G · strategistDownstream analyticsno — diagnostic

Duration only runs if identification is green, so a classifier can pass identification and still fail on timing — and you will know exactly which.

A second ladder was then built to audit this one. Seven rungs, pre-registered at a named commit before the module existed, after the team proved the original could be climbed by a model that ignores its inputs. Its organising distinction is the sharpest idea in the platform: a hygiene rung is necessary but carries no evidence of quality — a pocket calculator that always returns 7 passes every hygiene rung — while an informative rung can only be passed by a model actually reading the data. Conflating them is what lets a near-perfect score built mostly from structural passes read as quality.

Every prior blessing was re-scored against it. Under 2% survived — against the team’s own pre-registered prediction of under 15%, published before the run.

KAIROS · doctrine

Two gating verdicts, eight principles, and a wall around diagnostics.

  • Identification before duration. Top-level green requires both; a classifier can pass identification yet fail on timing — and you’ll know exactly which.
  • Diagnostic rungs cannot gate. A test reads the orchestrator’s own source and fails if the evaluation family so much as appears in the verdict path — and the same wall extends to the forecasting, aux and benchmark surfaces. Enforced in code, not convention.
  • Eight codified principles — every label leaks; filtered, not smoothed; stationarity is the null; verdicts are per instrument-and-period; and, added later from measured failures, directional asymmetry and selection-window circularity — each checked automatically.
  • Ladder and benchmark are separate, and must be. Real markets carry no observable regime label, so recovery, lead-time and transition-detectability cannot be gating rungs. They are scored on a synthetic known-regime generator that never gates and never blesses — a measurement surface deliberately walled off from the verdict.
  • Registries are derived, not declared. A hand-written map is a claim about what emitted a stream; the manifest field is a record written at emit time by the emitter. When they disagree the manifest wins — and any taxonomy that cannot be resolved is named as unresolved rather than omitted from the list.
  • Adversarial classifiers ship as permanent controls. Several kinds exist specifically to be rejected, alongside a corpus of known-bad fixtures. A validator that never rejects anything is not a validator.
  • Forecast reliability is a published label. Each taxonomy carries a graded honesty class — direction-usable, persistence-reliable, or descriptive-only — shipped alongside the regime stream, so a consumer knows what the labels can and cannot support.

KAIROS · models & extensibility

A broad model zoo, one validation contract.

  • Classical to deep. Hidden Markov and semi-Markov models, change-point and Bayesian online detectors, K-means and Wasserstein clustering, random forests, wavelets, hierarchical Dirichlet-process HMMs, Markov-switching GARCH, particle filters, deep state-space and temporal-fusion models, and a topological (persistent-homology) detector — dozens of classifier kinds behind one protocol.
  • Sibling-authored adapters, a supported side door. Another team may ship a classifier as a module behind a runtime-checked protocol and have KAIROS validate it with no per-candidate engineering. It is a real, documented path — and essentially all blessed streams in service today came from KAIROS’s own factory, not through it.
  • The instruments hat. Minimum detectable effect, null design, empirical null calibration, effective sample size and capture denominators are owned here, published under a service contract, and served to siblings on request — so the question “could this sample even detect that effect?” is answered by someone with no stake in the answer.
  • HMAC-chained blessed ledger. Every passing classifier’s blessing is recorded in a tamper-evident, append-only ledger — with a real key-rotation on record.
  • Attribution surface slices live strategy performance by regime stream and label — diagnostic only, never a gate.

05 / 06 · platform

πρᾶξις — “putting into practice”

PRAXIS

Risk machinery · execution

PRAXIS · the circuit breaker

A two-layer circuit breaker built for fat tails.

position intent Layer 1 · drawdown ramp warn kill Layer 2 · EVT-GPD stop stop fat tail sized exposure + parked stop

Layer 1 scales exposure down a ramp as drawdown grows (warn → kill, then cooldown). Layer 2 fits a generalized Pareto tail to recent losses to size the stop for the 6-plus-sigma moves crypto actually makes. Both checkpoint to disk and survive a restart.

PRAXIS · reconciliation & isolation

Reconciled against reality on every cycle.

PRAXIS local state exchange reconcile · 7 checks every cycle drift < tolerance → auto-correct local drift ≥ tolerance → halt + alert

Six-layer per-strategy isolation contains any single failure. Determinism is re-checked boot over boot — a bundle’s trace-hash must match its previous boot under an identical program key, and the guarantee only strengthens with each restart.

PRAXIS · the capital gate

The safety that stands between a green bundle and real money.

  • Fail-closed start-up. With live mode set but safety config unset, the process refuses to start — no accidental live runs.
  • Pre-trade order-rate limiter — the explicit guard against a Knight-Capital-style runaway — plus a stop attached to every entry.
  • Dead-man’s switch + flatten-on-kill, and a halt that outlives the process: it takes an operator ack, not a restart, to clear.
  • Staged capital ramp behind a burn-in floor, advanced only on both a minimum-days floor and a gate-pass streak, and stepped back on a hair trigger.
  • A second gate: the production-green receipt. Live start-up reads the validator’s productization receipts and refuses any spec that fails engine-independence, named-mechanism, out-of-regime-survival or corpus-orthogonality — re-checked every liveness cycle, with kill reasons for a receipt going stale or flipping mid-flight.
  • The gate proves it can fire, at every start-up. PRAXIS builds a throwaway risk manager from the real config, submits orders engineered to breach each ceiling, and refuses to start unless every one is rejected.
  • Declared absence is a pass; undeclared absence is a failure. Turning off the reconciler, a safety rule or a deploy check requires a written reason recorded in config — and an empty reason is rejected, because a declaration that declares nothing is absence wearing a permission’s clothes.
  • Today the venue-connected book runs on an exchange sandbox with virtual funds, and the core runtime has never routed an order to a production endpoint. Crossing to real capital is an explicit operator step, never automatic.

06 / 06 · platform

ἀγορά — “the public marketplace”

AGORA

The contract, mechanised

AGORA · deep

The contract, mechanised.

  • Tracked vs ephemeral. Bundles, schemas, ledgers and the journal are permanent; candidates and results are working files. A two-tier garbage collector reclaims the latter and refuses, by configuration, to touch protected paths.
  • 50+ typed journal events. Every actor emits them under file locks; the log is append-only and committed daily — the cross-platform source of truth.
  • A packet broker runs cross-team requests through an explicit state machine (open → acknowledged → replied → closed) with schema-linting in CI.

AGORA · governance & safety

The commons audits itself.

  • Derived analytics, regenerable. A queryable database is built from the journal — the markdown, YAML and JSONL remain the source of truth.
  • A written constitution, and a metric constitution beside it. A charter fixes each platform’s lane, what counts as success — “a period with zero survivors and high disciplined throughput is a good one” — and ten standing gates, with an amendment clause that invites contradiction. A companion document holds one canonical meaning per metric: find a second implementation, and the rule is to converge it and verify bit-identical agreement.
  • Sibling-admission protocol. Joining the bus requires a written proposal, a non-duplication table against every incumbent, a declared read/write surface and a self-audit — exercised in the open this month for the ecosystem’s sixth platform.
  • A standing audit panel runs numeric-claim verification and unjustified-bypass detection, under a charter rule of “verify before crediting: read the code or run it; never credit a report.” Its numeric-claim verifier re-derives asserted figures from source, and exists because a cycle-close report once quoted two different totals for the same quantity.
  • Bundles carry a SHA-256 seal over their sealed file list, which the executor verifies byte-for-byte using the validator’s own hasher before loading.
Open, and stated as open: AGORA’s own bundle-integrity check currently verifies an older, empty hash map and passes vacuously — the real seal is enforced on the executor side, not here. It is one of four ecosystem controls this month’s audit found unable to fire. See the next slide.

The machine earning its keep

Concrete catches — the apparatus preventing an expensive mistake.

  • A dozen edges that were one bet. A set of strategies that looked independent was shown by the effective-bets / correlation apparatus to be barely more than a single bet — over-diversification caught before any capital was concentrated into correlated risk.
  • Single-asset flukes, rejected. The ≥3-instrument generalization floor kills a strategy that shines on one instrument and collapses on three — the difference between an anecdote and a claim.
  • Fitted noise, unmasked. The adversarial battery reverses a strategy’s own signal and shuffles its labels; “edge” that survives the sabotage is exposed as curve-fitting, not signal.
  • 1,600+ proposed signals, roughly one in twenty adopted. Disciplined triage kept the rest — duplicates and noise — out of the validator entirely.
  • A validator that failed its own test. KAIROS proved its prior ladder could be climbed by a model ignoring its inputs, built a stricter one, and re-scored every earlier result against it — with under 2% surviving, against a pre-registered prediction of under 15%.
  • A metric with six private implementations. Before writing a seventh, someone checked — and found the canonical function already published, plus six private copies inside the validator’s own libraries, splitting into two families whose answers diverge by more than fivefold at realistic sample sizes. Filed as a finding rather than worked around locally.
  • A capacity formula at war with its own documentation. The docstring described one calculation and the body computed another; the two disagree by roughly three orders of magnitude on the most liquid pair in the inventory, and the output feeds a flag that blocks promotion. Both readings were stated as defensible and the question referred, rather than a local coefficient being quietly chosen.
  • Two counts rendered as one fraction. Three separate surfaces displayed “rungs passed / total” as 16/11 — a main-rung total against a count that includes sub-rungs. Individually correct, not commensurable, and nonsense as a ratio.
  • A quadratic hiding in a linear ledger. A sibling’s algebraic correction — that a net Sharpe is exposure-invariant under proportional costs — led to the discovery that a gross-PnL field carried a spurious exposure factor, making dollar PnL quadratic while every charge against it stayed linear. Invisible to two thousand passing tests, because the screen that exercised it held a position size where the bad factor is exactly one.

None of these reached capital. That is the value: the mistakes are caught here, on purpose, before they cost anything.

Open findings

Four controls that currently cannot fire — named, owned, unfixed.

A document that only lists caught defects is describing a marketing position, not a discipline. These are open as of this month’s audit. They are here because publishing them is cheaper than being found with them — and because a reader who cannot see the open list has no way to judge the closed one.

  • The commons’ bundle-integrity check passes vacuously. It verifies a hash map that the current sealing convention no longer populates, so it checks zero files and exits clean. The real byte-level seal verification happens on the executor side; the commons-side gate is presently theatre, and is being reconciled to the convention that actually shipped.
  • The gate-effectiveness audit reads an event nobody emits. It tabulates gate failures from the journal to find gates that never catch anything unique. The event name it consumes has never been written, so it always reports “no underperforming gates” — the exact failure mode it was built to detect, in itself.
  • The sibling-admission check is one-directional. It walks the roster and confirms each entry appears in the schema enums. It has no reverse pass, so a platform admitted to the enums but missing from the roster is invisible: it reports “PASS: 5 siblings” while a sixth has been trading packets for six days. Found and reported by the panel whose own ratification duty it guards.
  • The journal schema-lint is failing against its own enum. The broker emits events under an actor name that was added to the packet vocabulary and never to the journal vocabulary — so a slice of the log does not validate against the schema that governs it. Two copies of one vocabulary, drifting, which is the defect class the ecosystem has now paid for repeatedly.

The pattern connecting all four: a control over a list must derive that list from outside itself. A check whose subject comes from the same place as its answer will always agree with itself.

The doctrine that runs through all six

A control that cannot fail is a defect — equal to a wrong answer.

“Everything measured was green; everything unmeasured was red. Principles do not hold a line; runnable commands do.”— the ecosystem charter, amended after an audit proved the point against itself

  • A test suite that passed hundreds of times — without running. Found, and fixed. Vacuous passes are hunted deliberately, and every tool must ship a foil it is proven to catch.
  • “The blessing carries no evidence.” KAIROS declared most of its own prior output unproven, built a stricter ladder to demonstrate it, and re-scored everything rather than defend it.
  • A moat claim, retracted. EPISTEME withdrew an overstated cross-engine claim on the record — and then shipped a machine-readable field on every certified spec so the overstatement cannot recur silently.
  • An experiment closed by its own negative result. HEURESIS ran an autonomous agent-driven discovery mode for eight weeks, measured no graduations across roughly 1,600 candidates, and switched it off — keeping the infrastructure and publishing the result.
  • Unrun is not green. A bar records the time each check last executed, so a check that has never run renders as unknown — deliberately treated as worse than red, because red is at least a measurement.
  • The open list is published too. Four controls that currently cannot fire are named on the previous slide. A discipline that only reports its catches is a marketing position.

The ecosystem, in motion

Not a snapshot — a coordinated, self-correcting system.

400K+
lines of production code, six platforms, one contract
5,000+
permanent cross-team hand-offs on record
350+
architecture decision records across the stack

Coordinated through dated hand-offs, a packet broker and a standing audit panel — with data adapters and cost models spanning nine asset classes, though the certified corpus itself remains crypto. Built since April 2026 on three years of prior research; the sixth platform was proposed, admitted and productive inside a single week. The architecture is exercised continuously, not just at design time.

Plain-language glossary

The acronyms, defined.

CI — continuous integration: automated checks on every change.
ADR — architecture decision record: a numbered design rationale.
WFV — walk-forward validation: test on the unseen next window.
CPCV — combinatorial purged cross-validation: leakage-safe testing.
FDR — false-discovery-rate control: guards multiple-testing luck.
GBM / MJD — geometric Brownian motion / Merton jump-diffusion: price models for Monte Carlo.
GARCH / EVT — volatility-clustering model / extreme value theory: heavy-tail stress.
GPD — generalized Pareto distribution: fits the loss tail to size stops.
HMM / HSMM — hidden (semi-)Markov model: regime classifiers.
CUSUM — cumulative-sum monitor: detects live edge decay.
HMAC — hash-based signing: tamper-evident records and requests.
JSONL — JSON Lines: the append-only journal format.

The standard

Search like a lab.
Reject like a fund.
Doubt your own results first.

A validation apparatus rigorous enough to invalidate its own work — and honest enough to say so. That is the moat. Questions are welcome; that is rather the point.

Confidential — most sensitive tier. Do not distribute. · Figures describe engineering footprint, not investment performance.