Confidential
01
↓ / → to advance
Isard Labs Isard Labs

Technical documentation · Level 2 — Architecture

The machine, opened up.

Per platform: stack, module shape, data flow, what makes it distinctive, and how mature it is. Then how six independent codebases coordinate without ever importing one another’s source.

Figures describe engineering footprint, not investment performance.

Architecture at a glance

Five engines around one commons.

AGORA · the commons files, schemas and hash-pinned wheels — no platform imports another’s source two manufacturers, one gate HEURESIS discover · at rate AUTOMATA prove · with provenance EPISTEME validate · the only certifier PRAXIS execute candidates experiments green bundles KAIROS optional regime layer blessed regimes a strategy may gate on — opt-in

Four trust boundaries hold it together: no RPC and no source imports (coupling is files, git and published wheels), validated bundles are frozen (machine-built, never hand-edited), signals can’t skip the gate (a new primitive must be requested and implemented, not smuggled in), and no manufacturer certifies itself (a local ladder run is preflight, and preflight is not green).

Lifecycle data flow

One idea, traced through the filesystem.

StepActorWrites to AGORA
1 · proposeHEURESIScandidates/<batch>/*.yaml
1b · pre-registerAUTOMATAjournal/ — a frozen, hashed experiment record before any data is read
2 · validateEPISTEMEresults/<batch>/*.scorecard.json
3 · promoteHEURESISbundles/<strategy>/ (git-committed)
opt · timeKAIROSregimes/<taxonomy>/*.parquet — only if a strategy gates on it
4 · executePRAXISjournal/YYYY-MM.jsonl (fills)

Every row is append-only or content-addressed. The journal threads them together, so any fill can be walked back to the search that produced it.

01 / 06 · platform

εὕρεσις — “discovery”

HEURESIS

The idea engine · discovery at rate

HEURESIS · role & stack

Exhaustive enumeration, then pre-registered screens.

Grid mode enumerates a parameter space deterministically, cheapest-first, from a frozen YAML campaign config whose hash is recorded — same config, same candidate sequence. Pre-registered batch screens are the current working mode: a named hypothesis, its controls and its acceptance predicate are registered and hashed before the run, with the tool path and code commit stamped into the output artifact.

An LLM-driven proposal mode exists in the tree and is not currently in service — see maturity.

Python 3.12Pydantic v2 (frozen)Typer CLIstructlogSQLite (WAL) KBClaude / Ollama / Mock

HEURESIS · architecture & differentiators

It reaches the validator through a subprocess, a folder and a published wheel.

  • The engine never imports EPISTEME. It shells out to the epi CLI and reads and writes AGORA files; the source tree contains no sibling imports and a gate keeps count.
  • Canonical metrics come from EPISTEME’s published wheels, not from reimplementation. A metric defined twice is a metric with two answers — so the definition is installed, never copied.
  • Deterministic template path. Campaign configs are frozen, thresholds load from the shared schema, and a config hash is recorded.
  • Four-layer de-duplication. In-session fingerprints, a cross-session knowledge base, the validator’s own run history and the pre-flight gate — so no candidate is ever paid to validate twice.
  • Staged inference, and a ledger that can say zero. Multiplicity control runs within a stage and never across: a screen result may disclose but can never bank a claim, and only confirmation on data the screen never touched counts. When a stage’s admissions did not survive that rule, they were closed on the evidence and the banked count published as zero.
  • Hallucination containment. A proposal for a signal that doesn’t exist becomes a structured request into AGORA — it cannot edit the validator to make itself pass. Of 1,600+ such requests filed to date, roughly one in twenty was adopted.

HEURESIS · maturity

Small, sharp, and governed by its own green bar.

~20K
lines of source — focused, not bloated
1,300+
tests · ≥90% branch-coverage gate
green-bar
discipline enforced as code, not convention

It runs pre-registered discovery “marathons” with cost gates, positive and negative controls and significance testing. It is judged on rate and discipline, not hit-rate.

An autonomous, agent-driven campaign mode was built, run for eight weeks, and switched off. Across roughly 1,600 machine-proposed candidates it produced no graduations, and the failure concentrated in a single rung — cost sensitivity — consistently across four different local models and every context size tried. The infrastructure was kept; the operational mode was closed, and the negative result written up rather than quietly left running.

02 / 06 · platform

αὐτόματα — “self-acting”

HEURESIS AUTOMATA

The experiment layer · provenance

AUTOMATA · role & stack

A second manufacturer, judged on a different thing.

HEURESIS is judged on rate of disciplined candidates. AUTOMATA is judged on whether a conclusion can be walked back to the run, the spec, the snapshot and the code that produced it — and on whether the search burden behind it is stated honestly. Both submit to the same gate. Neither can certify itself. Running them side by side is the point: it makes the comparison measurable instead of arguable.

It imports canonical metrics from EPISTEME’s published wheels rather than reimplementing them, asks KAIROS for regime distinctions rather than building regime models, and hands execution to PRAXIS. What it builds is the layer none of them owned.

Python 3.1218 enforced lanesmypy --strictSQLite event storecontent-addressed artifactslocal LLM, proposal-only

AUTOMATA · architecture & differentiators

The rules live in the type system, not in a review checklist.

  • An Experiment IR, frozen and hashed. A research screen is never a script. It is a typed record — hypothesis, null, as-of date, snapshot, splits, minimum detectable effect, baselines, controls, seeds — frozen and immutable once it has seen evaluation data. It wraps EPISTEME’s strategy spec rather than replacing it.
  • Look-ahead refused at compile time. Static temporal analysis walks the spec’s expression tree and rejects a leaking feature before any data is read — rather than catching it after a validation run has been paid for.
  • Every number carries an origin and a role. Origin says how a value was produced; role says what it is allowed to decide. Because a gating value and a disclosure-only value are different types, using a disclosure to make a promotion decision fails type-checking before anything runs.
  • An LLM cannot author a number. The rule is not enforced with a flag. The type a model is permitted to return has no numeric field at all — a flag can be set wrongly; a type that cannot hold a float cannot carry a fabricated one.
  • A declared differential oracle. A second, independent execution engine exists solely to disagree with the imported one. Its only permitted output is a divergence report.
  • An adaptive trial ledger. Search is adaptive — results are inspected, branches chosen, families abandoned — so error control is sequential and versioned, and thresholds cannot be retroactively re-based.
  • Nothing is deleted. A data or code incident marks downstream evidence as pending recomputation; the evidence graph is append-only, and negative results stay searchable.

AUTOMATA · maturity

Seven days old, and the strictest bar in the stack.

2,200+
tests, and ~26K lines of test code against ~31K of source
0
type errors under --strict across ~300 files, with no exemption list
29
named checks in one recorded bar — every one green, and dated

Its bar separates gates from disclosures: a gate blocks, a disclosure reports and is never allowed to harden into a gate by rhetoric. It publishes its own incompleteness — of 61 formal done-criteria, ten are declared not built rather than quietly omitted. And a run of the bar is itself recorded, so a check that has never executed renders as unknown, not as green.

03 / 06 · platform

ἐπιστήμη — “demonstrable knowledge”

EPISTEME

The gate · the only certifier

EPISTEME · role & stack

An eleven-rung ladder that halts the moment a candidate fails.

The largest and most important platform, and the only thing in the ecosystem permitted to certify. A spec climbs eleven canonical rungs — data quality, causality, costs, signal selection, walk-forward validation, significance, path risk, generalization, adversarial stress, Monte Carlo and deployment readiness — plus default-on sub-rungs covering attribution, cross-engine accounting parity, regime provenance and per-regime metrics, which bring a default green run to sixteen scored records. Further sub-rungs for microstructure and live-forward certification remain opt-in. On green, it emits an executable bundle.

Python 3.12 · pinned lockfile44 atomic packagesPydantic v2pandas · numpy · scipyscikit-learnlocal LLM (Ollama)

EPISTEME · architecture

Forty-four single-purpose packages in an enforced layering.

  • Atomic libraries. One concern each — features, labels, cross-validation, portfolio, microstructure, regime, Monte Carlo, lifecycle — composed into the ladder.
  • The dependency graph is law. An import-linter contract forbids upward imports; a violation fails CI. The architecture can’t silently rot.
  • Fail loud. Typed exceptions only; a failed gate halts the run rather than quietly compensating later.
  • The validator core is human-authored. Only two narrow, schema-bound LLM nodes (intake and red-gate explanation) ever touch a live run.

EPISTEME · differentiators

The verdict is trustworthy because it can’t be argued with.

  • Pre-committed, immutable thresholds. Every pass/fail line is fixed in advance and audited. You cannot tune the gate to admit a marginal strategy.
  • Byte-identical reproducibility. The same spec and fixtures produce identical artifacts down to the byte — the two-process check runs on every CI run.
  • Cross-engine accounting parity. A strategy’s positions are re-run through two third-party backtest engines and the accounting reconciled — correlation of 1.0000 against the first and 0.9968 against the second, with trade counts matching exactly across the canonical cohort. This is fill-and-accounting fidelity, not independent re-derivation of the signal: the compiler that would make it independent is designed and not yet built, and the receipt shipped with every spec says so in a machine-readable field.
  • Institutional memory. Every run is content-addressed and kept; the archive of what didn’t work is itself an asset.

EPISTEME · maturity

The most heavily engineered platform in the stack.

40+
atomic packages
500K+
lines of code & tooling
13K+
tests · ≥90% branch gate
240+
architecture decision records

Driven by dozens of consumer-feedback cycles. Multi-asset support is shipped, not aspirational — cost presets and calendars for spot and perpetual crypto, equities, ETFs, indices, FX, commodities, futures and prediction markets, with point-in-time universe reconstruction, corporate actions and macro release calendars. The certified corpus itself is still entirely crypto; the machinery to move beyond it is already in place.

04 / 06 · platform

καιρός — “the opportune moment”

KAIROS

The regime factory · and the ecosystem’s instruments

optional · a strategy uses it, or doesn’t

KAIROS · an elective capability

Off to the side of the core pipeline — by design.

a strategy spec no regime gate validated as-is — the regime rung is skipped entirely opt-in regime gate when_regime: … KAIROS-blessed stream validated regime labels

KAIROS builds the classifiers and validates them on its own ladder, then lends the blessed labels to whoever wants them — a research factory with an independent line, not a service desk. A sibling may also submit its own classifier behind a documented protocol; that path is supported and, so far, rarely taken. The skip is literal: a spec that declares no regime inputs never runs the regime rung at all.

KAIROS · architecture & differentiators

EPISTEME’s discipline, pointed at regime models.

  • Two gating verdicts, and only two. Of six rung families, exactly two can decide anything: a classifier must first prove it identifies the right state, then that it gets the timing of state changes right — identification before duration, and duration only runs if identification is green.
  • Diagnostic rungs can never gate — enforced, not agreed. Whole rung families that use look-ahead are allowed for insight but mechanically barred from the verdict. A test reads the orchestrator’s own source and fails if the evaluation family so much as appears in the verdict path.
  • Thresholds are provenance-tagged — literature-cited, stated-default, or derived — and immutable within a pinned version; drift fails CI.
  • The blessed ledger is hash-chained. Every blessing is an HMAC-chained entry linked to its predecessor, so the record of what was approved cannot be quietly rewritten.
  • Sibling-authored adapters. Another team may ship a classifier as a module behind a documented protocol and have KAIROS validate it with no per-candidate engineering. A supported path, distinct from the factory that produces almost everything in service.
  • The second hat: the ecosystem’s instruments. Minimum detectable effect, null models, effective sample size and capture denominators are owned here and served on request — so “could this sample even detect that effect?” is answered by someone other than the claimant.

KAIROS · maturity

The most self-critical platform in the stack.

40+
packages · a broad model zoo, classical to deep
100+
architecture decision records
200+
fixtures, including classifiers built to be rejected

It built a second, stricter ladder specifically to audit its own gate — pre-registered at a named commit before the code existed — after proving the original could be climbed by a model that ignores its inputs. It then re-scored every prior result against the stricter bar rather than keeping the flattering ones, and under 2% survived, against its own pre-registered prediction of under 15%.

It also ships classifiers designed to fail, permanently, as controls — a validator that never rejects anything is not a validator.

05 / 06 · platform

πρᾶξις — “putting into practice”

PRAXIS

The executor · disciplined runtime

PRAXIS · role & stack

Exchange-agnostic, and paranoid by design.

PRAXIS loads validated bundles, re-verifies them, and runs them. Strategies are loaded as self-contained modules — PRAXIS holds no signal logic of its own; it is the disciplined runtime around someone else’s proven edge. Reaching real capital is an explicit operator action; today it runs on paper and exchange sandbox.

Python 3.12Pydantic v2pandas · pyarrowhttpx (HMAC-signed)FastAPI + React/StreamlitDocker

PRAXIS · architecture & differentiators

One stack per strategy, total isolation.

  • A fifteen-gate loader. Nine numbered gates plus six sub-gates: conformance, fixture replay, cross-boot determinism, seal coverage, a byte check against the seal using the validator’s own hasher, cost-model drift and annualisation cross-validation. A bundle that fails any of them is quarantined, never run.
  • A green must be a grade, not a word. A bundle that declares itself green without a scorecard behind it is refused — and running one anyway requires a written election with a stated reason, recorded in the config.
  • Per-strategy crash isolation, six layers deep — module namespace, exception, state, capital, risk and audit. Each strategy now also runs as its own container stack on its own network, so one blowing up halts only itself.
  • Two-layer circuit breaker. A portfolio drawdown ramp plus a fat-tail-aware stop, sized to respect the moves crypto actually makes.
  • Reconciled against reality. Eight checks against the exchange — at boot and on a scheduled sweep in the deployed runner, and every liveness cycle under the supervisor. Small drift auto-corrects from the venue; drift beyond tolerance halts trading.
  • Fidelity reconciliation at the decision level. The exact bars the live runner saw are replayed back through the strategy offline and diffed bar by bar, with typed divergences that can halt a book — the honest answer to “does the live system actually do what the backtest said?”

PRAXIS · the capital gate

The safety that stands between a green bundle and real money.

  • Fail-closed start-up. With live mode set but safety config unset, the process refuses to start — no accidental live runs.
  • Pre-trade order-rate limiter — the explicit guard against a Knight-Capital-style runaway.
  • Dead-man’s switch + flatten-on-kill, a stop attached to every entry, and a halt that survives a restart (it takes an ack, not a reboot, to clear).
  • Staged capital ramp behind a burn-in floor — exposure scales up only after a strategy has proven itself in-market, and steps back on a hair trigger.
  • The gate proves itself before it guards anything. At start-up PRAXIS builds a throwaway risk manager from the real config, constructs orders guaranteed to breach each ceiling, and asserts every one is rejected — refusing to start if any gets through. A safety check that has never been shown to fire is not a safety check.

PRAXIS · maturity

Production-grade engineering, at deliberately trivial size.

~43K
lines of engine source, plus ~17K of operator surface
~1 : 1
slightly more test code than engine code
2,800+
tests — unit, conformance, property

Containerised, with runbooks for deploy, promote, halt, reconcile and key rotation, and a real operator product on top — a FastAPI service and a React dashboard per strategy, covering fidelity, capital ramp, kill switch, execution quality and stress tests.

The core runtime has never routed an order to a production exchange endpoint. Nine books run today: eight simulated, one on a venue sandbox with virtual funds. Of nineteen enforced gates, several point at the running fleet rather than the source tree — every deployed book must be running, or declare why not.

Scope: separate PRAXIS-branded deployments trade small prop-firm accounts and are outside this document. No client or fund capital is deployed in any of them.

06 / 06 · platform

ἀγορά — “the public marketplace”

AGORA

The commons · coordination

AGORA · role & surfaces

A message bus with no server — just files and git.

HEURESIS AUTOMATA EPISTEME KAIROS PRAXIS AGORA file surfaces schema · candidates · results · bundles · regimes · signal-requests append-only journal · 50+ event types · the cross-platform source of truth

AGORA · architecture & differentiators

The boring choice that makes everything auditable.

  • Tracked vs ephemeral. Bundles, schemas and the journal are permanent; candidates and results are working files, swept by a garbage collector that refuses to touch protected paths.
  • The journal is the spine. An append-only event log — one source of truth for “what ran on a given day” that no single platform’s local state can answer.
  • A packet broker gives cross-team requests a real state machine (open → acknowledged → replied → closed) instead of ad-hoc messages — and its state is derived from the journal rather than stored, so it cannot disagree with the record.
  • The message format enforces honesty. A cross-team packet is schema-linted at commit time and must carry a pre-mortem, a count check, a cost class, and two honest numbers — at least two named numeric claims the packet is prepared to be held to, or an explicit opt-out. The format itself makes you state your numbers.
  • Bundles carry a SHA-256 seal over their sealed file list, and the executor verifies bytes against it with the validator’s own hasher before loading. AGORA’s side of that check is currently the weaker one — see the open findings.
  • A written constitution. A charter defines each platform’s lane, what counts as success, and ten standing gates — with an amendment clause that invites contradiction: a charter nobody argues with is one nobody reads. It binds by convention; the subset that is wired is each repo’s named green bar, which CI runs.

AGORA · maturity

Small surface, heavy guarantees.

50+
event types in the journal vocabulary
5,000+
permanent cross-team hand-off records
every commit
safety cycle in CI, with deeper presets on longer cadences

A standing audit panel runs numeric-claim verification and unjustified-bypass detection, and is the busiest correspondent on the bus. Its numeric-claim verifier exists because a cycle-close report once quoted two different totals for the same thing — so the commons now re-derives asserted numbers from source rather than trusting the typing.

Two static gates hunt vacuous controls by name: every tool must have a test that defines a real assertion, verified by parsing the test rather than its filename; and every lint or type suppression must carry an expiry date, enforced strictly in CI.

How they coordinate

The whole contract, on one page.

SurfaceProducerConsumer
schema/EPISTEMEHEURESIS · KAIROS · PRAXIS · AUTOMATA
artifacts/EPISTEMEall siblings — canonical metrics as hash-pinned wheels
candidates/ · results/HEURESIS · AUTOMATA · EPISTEMEeach other (ephemeral)
bundles/EPISTEMEPRAXIS
regime_requests/HEURESIS · AUTOMATAKAIROS (optional)
regimes/KAIROSEPISTEME · PRAXIS (opt-in)
signal_requests/HEURESISEPISTEME
docs/handoffs/all sixall six — schema-linted packets
journal/all sixoperators · audit

No row in this table is a function call. Every one is a file in a git history.

Plain-language glossary

The acronyms, defined.

CI — continuous integration: automated lint, type and test checks run on every change.
ADR — architecture decision record: a dated, numbered rationale for a design choice.
LLM — large language model (here: local or Claude, on a tight leash).
KB — knowledge base: a platform’s record of prior runs and candidates.
WFV — walk-forward validation: train on the past, test on the unseen next window.
CPCV — combinatorial purged cross-validation: leakage-safe model testing.
FDR — false-discovery-rate control: guards against multiple-testing luck.
EVT / GPD — extreme value theory / generalized Pareto: fat-tail stop sizing.
HMM — hidden Markov model: a common regime classifier.
CUSUM — cumulative-sum monitor: detects live edge decay.
HMAC — hash-based signing: makes records and requests tamper-evident.
JSONL — JSON Lines: the append-only, one-event-per-line journal format.

Evidence, not assertion

The discipline, caught in the act.

  • Effective-bets audit. A portfolio of apparently independent strategies collapsed, under the correlation apparatus, to barely more than a single effective bet — the over-diversification illusion, caught before capital was concentrated.
  • Generalization & adversarial gates. The ≥3-instrument floor rejects single-asset flukes; the adversarial battery reverses and shuffles a strategy’s own signal, so anything that only worked as fitted noise is exposed.
  • False-green audits. Internal checks found controls that could not fail — a test that “passed” hundreds of times without running; a gate that certified hundreds of scorecards while being structurally unable to fail — and fixed them.
  • Signal triage. Of 1,600+ machine-proposed new signals, disciplined triage adopted roughly one in twenty — the rest kept out of the validator rather than diluting it.
  • A control whose subject was never reached. A fidelity reconciler covered one of nine live books for weeks while its scheduler printed “nothing to do” and exited clean every hour. The fix was to make the gate enumerate from the running fleet rather than from config — a subject list drawn from config agrees with config and disagrees with reality, which is precisely why the hole was invisible.
  • Open, and stated as open. This month’s audit found four ecosystem controls that currently cannot fire — including the sibling-admission gate, which reports “PASS: 5 siblings” and is structurally unable to notice a sixth. They are named, owned and being fixed. Publishing them is cheaper than being found with them.

Every example above is the safety net doing its job — and the last one is the safety net catching itself.

Why the architecture is the moat

Anyone can write a backtest. Almost nobody audits the auditor.

  • Determinism end-to-end — discovery is config-driven, the validator is byte-reproducible, the executor re-verifies its own determinism.
  • Hard gates over judgement — thresholds are pre-committed and immutable in two independent validators (strategies and regimes).
  • Self-adversarial by habit — the teams actively hunt controls that can’t fail, treat a vacuously-passing gate as a defect equal to a wrong answer, and retract their own claims when the evidence doesn’t hold.
  • Two manufacturers, one gate — a second research factory was admitted deliberately, optimised for provenance rather than rate, precisely so the claim “this way of working is better” becomes a measurement instead of an opinion.
  • No source coupling — six projects that fail, evolve and deploy independently, sharing definitions as published artifacts and joined only by an auditable commons whose whole history is evidence.

Go deeper

You’ve seen the machine.
Now the internals.

L3 · Deep → names the methods — the rungs, the circuit breaker, the signal families, the governance — for reviewers who want to push on every claim.

Confidential — not for distribution. · Figures describe engineering footprint, not investment performance.