Mechanism Interferometry Modularity certificate
Read the paper

A causal-modularity calculus for factorial regimes

Outcome effects need not add.
Mechanism changes do.

Two autonomous causal changes can interact violently in outcome space. In log-density-ratio space they add exactly. Wherever that additivity fails, the failure itself is a gauge-invariant curvature field, and its magnitude separates nonlinear response from real coupling between mechanisms and from curvature manufactured by incomplete observation.

3observable certificate conditions
3exact conservation laws
10conditions that fail the run closed
The intervention square Four regime distributions at the corners of a square. Two primitive mechanism changes act along the edges. The cycle closes exactly when the mechanism curvature of the face is zero.
κAB(x) = log pAB(x) p0(x)pA(x) pB(x)

The idea without the notation

What a causal claim is actually claiming

Most causal claims get read as "A raises B." The claim underneath is stranger and much stronger: that the world is assembled. Behind the tangle of correlations sits a workshop of separate small rules, each responsible for producing one thing out of a few others, and each independently replaceable. You could reach in, swap one out, and the rest would keep running unchanged, not out of loyalty but because they never depended on it to begin with. That is the entire content of the word mechanism. Everything else on this page answers a single question: how could a creature standing outside the machine ever check that picture, and find the seams?

Never study a world. Study the difference between two.

A single world is bottomlessly complicated. Two worlds that share almost everything are simple relative to each other.

Change one rule and the consequences cascade. Every downstream quantity shifts, and the perturbed world can look globally, dramatically different from the original. Now ask a narrower question, observation by observation: how much more at home is this particular event in the changed world than in the old one? All the shared machinery cancels out of that comparison. What survives is a fingerprint living only at the seam, on the changed rule and its immediate inputs.

Effects are global. The difference in explanation is local.

Which is why interferometry is the right word here rather than decoration. An interferometer never observes an arm of the apparatus directly. It observes the fringe pattern where two nearly identical beams fail to match, and everything the beams share cancels. In this setting the two beams are worlds, and the cancellation is the whole method.

The contrapositive has teeth. If the fingerprint refuses to localize, if explaining the difference between the worlds keeps requiring appeals to everything at once, then either the seams are not where you drew them or you are not seeing the whole machine. Locality of evidence works as a test here rather than as an assumption, and it is the first condition of the certificate.

The arrow points where the accounting lives.

A rule is a complete account of the thing it produces. Whatever its inputs hand it, it has to say what happens next, and its possibilities have to exhaust; the bookkeeping is anchored on its output. Its inputs it merely consumes. It owes them nothing, and replacing the rule leaves their statistics untouched, because they were never that rule's responsibility.

Inside the small fingerprint region, then, cause and effect come apart by an almost childish procedure. Remove one item at a time and look at what is left. The changed variable is the one whose removal leaves the remainder indistinguishable from the unperturbed world, and it can vanish without a trace precisely because everything else in the region was upstream of the edit.

The arrow points at the coordinate that deletes cleanly.

The procedure also fails out loud. When nothing deletes cleanly, the region is contaminated with downstream debris. When several things do, the perturbation was too symmetric to break the tie and a different tilt is needed. Plenty of methods will silently point the arrow backwards; this one mostly declines to answer, which in causal inference is a virtue bordering on luxury. You can watch it decline in the deletion audit.

Interaction is a fact about your ledger, not about the world.

Turn two knobs wired to genuinely separate parts of a machine. The outcomes can show violent synergy, two dials that each do nothing alone and everything together, because the machine's response is curved. In nature, curved responses are the norm.

Plausibility multiplies across independent parts, so its bookkeeping adds. In the one ledger kept over the complete state of the machine, two edits to separate rules compose the way two patches to disjoint files compose: the combined diff is just both diffs. The program's behaviour may change explosively. The patch algebra does not care.

So the question that biology, medicine and economics all treat as one question is really three. Is the response curved? That is ubiquitous, and fine. Are the two knobs secretly attached to the same part? Or are you simply not seeing enough of the machine? Ordinary interaction analysis mashes the three into a single number.

Mere nonlinearity produces no signal at all in the mechanism ledger, however spectacular it looks in the outcomes.

That last claim is not rhetorical. The interferometer below is an experiment with maximal outcome synergy whose complete-state curvature is exactly zero at every setting you can dial.

Your own coin flips are the probe.

Flip two independent coins to decide two independent tweaks. However chaotic the machine, a genuinely complete view of it can never make your own coin flips look correlated, because the fully seen state screens each tweak off from the other. So if, after reading every instrument you own, learning that the first coin came up heads tells you something about the second, your instruments are not the whole story.

Independence in, independence out, or something is hidden.

The failure comes with an exact anatomy: the leftover correlation between the two fingerprints that the instruments failed to pull apart. That narrows the culprits to a short list, and every item on it is worth finding. A shared part. An unmeasured piece of state. A filtered dataset. An intervention that changes character once it is combined with the other one.

Better still, the books balance. Averaging over a hidden variable cannot create coupling out of nothing. Whatever coupling the hiding conceals pointwise comes back, in perfect balance on average, as an opposite correlation between the visible fingerprints. Coarse-graining relocates dependence; it can neither mint it nor destroy it. That turns incompleteness from a lament into a compass. Propose a new measurement, add it to the view, and ask whether the interference dies. The variables that kill it are the missing state, and the conservation laws are the exact statement of the balance.

Two quieter points give the thing its shape.

There is no privileged normal world. Like altitude, an individual fingerprint depends on where you put sea level, but the steepness between two towns does not, and everything real here is steepness. The verdict survives your choice of which lab setting to call baseline.

And the evidential style is double-entry bookkeeping rather than hypothesis testing. Modularity implies an overdetermined web of identities across every combination of perturbations, so cooking one entry breaks several others. Confidence comes from the coherence of the whole ledger rather than from one significant coefficient, which is a different and considerably older epistemology than most of statistics runs on.

The reframe

Causal variables are whatever makes your actions compose

The certificate never pronounces on the world in itself. It certifies modularity relative to an interface: a repertoire of actions you know how to take, and a resolution at which you can see. Pass the audit and a workshop-of-separate-rules picture of your data provably exists, with replacement rules you can construct explicitly. Fail it and you learn where, and by how much.

That inverts the usual ontology. Instead of assuming nature arrives pre-diced into variables and asking which of them cause which, causal variables get defined by what they do to the algebra of your actions. They are the coordinates in which doing two things at once is exactly doing them one at a time; in which effects are global but evidence is local; in which each rule accounts for its output and owes nothing to its inputs; in which your independent choices stay independent under full observation.

Finding causes is finding the interface that makes agency compositional. The arrow of causality is where the accounting lives, the substance of causality is the additivity of independent edits, and the boundary of what you causally understand is drawn, sharply and measurably, exactly where that additivity breaks.

Four answers, one number

An interaction coefficient is four answers wearing one coat.

A single interaction statistic cannot tell you which of four different things you are looking at. Mechanism interferometry moves the object of study away from a chosen response summary and onto the full regime law, where the four separate.

Response nonlinearity

The map from state to outcome manufactures synergy on its own. The generating mechanisms stayed autonomous the whole time.

Mechanism coupling

The combined intervention really did change how a primitive mechanism operates. This is genuine density curvature, and the only case that deserves the word interaction.

Missing state

Marginalizing a variable away makes exactly flat latent factors conditionally dependent. The curvature belongs to your observation set rather than to the world.

Implementation drift

The joint arm was never the composition of the two single arms. When a negative control with no interference returns anything but zero, the apparatus is the finding.

The local likelihood-ratio hologram

re(x) = pe(x)p0(x) = qe(xtxpa(t))pt0(xtxpa(t))

A soft intervention that replaces only the conditional mechanism of X_t leaves every other factor to cancel out of the regime likelihood ratio. The perturbation can move every descendant marginal it likes; the full joint ratio still lives on the target and its direct parents alone. Localization does not let you escape conditional-independence testing. It relocates the search from relations between states to relations between the design and the state.

The full taxonomy runs to eight causes. Beyond the four above, nonzero curvature can come from two perturbations that share a target, from regime-dependent measurement, from poor overlap or ordinary estimation error, and from a mismatch between the intervention you named and the one the apparatus delivered.

The publishable core

A finite biconditional certificate

Relative to a proposed graph and target assignment, three observable properties are necessary and sufficient for a modular soft-intervention representation of the whole regime family. There is nothing else left to check.

Locality

Each primitive ratio depends on one target family only: the node and its parents.

rj(x) = rj(xtj, xpa(tj))

Conditional normalization

The ratio integrates to one over the target, conditional on its proposed parents.

E0[rjXpa(tj)] = 1

Square flatness

Every two-dimensional face of the intervention cube closes. Order two already forces flatness at all higher orders.

κjkb(x) = 0

locality targetwise normalization square flatness

if and only if

a context-invariant modular soft-intervention representation exists

Order two suffices

On a complete cube, vanishing square curvature makes every higher-order Möbius coefficient vanish by itself. Zero face curvature holds each edge increment fixed under flips of the other coordinates; the remaining cube is connected, so each increment equals its reference-corner value and the log-densities become exactly additive. Higher-order tests survive as diagnostics under estimation error, never as extra requirements.

Gauge invariance

Add any field of the form a(x) + sum_j s_j b_j(x) to the log-densities and every mixed difference of order two or higher survives untouched. The primitive potentials depend on which corner you called the reference; the curvature does not. Reversing a coordinate orientation flips the sign of a mixed difference without disturbing the claim that it is zero. A reference corner is a coordinate convenience, not a scientific baseline.

Existence, not identification

The theorem says a proposed graph and target assignment admit a modular causal realization of the observed regime family. It does not say the proposal is the only one that would. Inferring the proposal uniquely from data still requires faithfulness and adequate intervention coverage, and the certificate never pretends otherwise.

When corners are missing, the testable content is a null space

Drop corners from the design and the question stops being which squares you can see. The complete testable content is the lack-of-fit space of a main-effects-only design matrix with a function-valued response. Toggle corners below; the ranks are computed exactly as mic-design computes them.

Observed corners

full cube

Observed corners8
Main-effects rank4
Lack-of-fit dimension4
Complete square faces6
Squares span lack of fityes

Design geometry

fully testable

The six-corner case is the one to remember. Remove 000 and 111 from a three-factor cube and not a single complete square face survives. Six observations against rank four still leave two genuine flatness restrictions, one of them h110 - h100 - h011 + h001 = 0. An implementation that only enumerates squares would report that nothing is testable, and it would be wrong. Where pairwise curvature is the target, plan a resolution-V fraction.

Interactive running example

Same experiment, different observation set, different verdict

X₁ and X₂ are independent Rademacher sources, Y = X₁X₂ + ε, and the two perturbations tilt the two source means. Drag the tilts and watch the panels disagree. On the left is the complete state, where the mechanism ratios multiply exactly and the curvature is identically zero at every setting you can reach. On the right is the same experiment seen through Y alone.

Design parameters

paper fixture
0.60
0.50
0.80

Noise sets how sharply the curve turns, not how far it travels. The range is fixed by ab alone.

Outcome synergy 0.300
Full-state κ 0.000
Marginal κ range [-0.3567, 0.2624]
Moment battery passes

The defaults reproduce the paper fixture exactly: a = 0.6, b = 0.5, sigma = 0.8, synergy 0.3, and a marginal curvature spanning log(1-ab) to log(1+ab).

Complete state (X₁, X₂, Y)

modular

rAB = rArB κ = 0

Observed Y only

curvature

κAB(y) = log[1 + ab tanh(y/σ2)]

A nonlinear outcome can show maximal synergy while the complete generating mechanisms stay exactly autonomous. Marginalization built the interaction field you are looking at. No mechanism coupling was involved.

The same example is the canonical blind spot of the cheap statistic. In outcome-only space both primitive ratios are identically one, so the scalar moment battery E_0[r_A r_B] = 1 passes trivially while the curvature is nowhere zero. Its weighted mean vanishes because the baseline is symmetric and the curvature is odd. A battery that passes has told you nothing.

Four exact fixtures

What curvature means depends on which way it broke

Each scenario below is an exact-population fixture from artifacts/simulations/exact_results.json, pinned to fourteen decimal places by the Rust conformance tests and reproduced in the paper figures. Three of the four are failures, which is why they are here.

Flat composition

kappa = 0

Two independent tilts on two independent sources. The primitive ratios multiply exactly at every point of the complete state, every square face closes, and the certificate holds. Held-out pair distributions follow from the singles alone.

This is the only one of the four in which a certificate may be issued. The other three are measured against it.

Curvature0.000000
Cov(r_A, r_B)0.000000
Certificatemay pass

Composition check

exact algebra

Curvature from hidden state

Three conservation laws

Whatever conditional coupling marginalization creates between flat latent factors has to come back, with opposite sign on average, as correlation between the observable primitive ratio fields. Marginalization relocates the coupling; it never removes it.

Law 1 · curvature balance

Cov0(rA,rB) =E0[rArB(eκ − 1)]

The cheap moment statistic turns out to be a signed, ratio-tilted average of curvature. Its blind spot is quantitative: a sign-varying field can integrate to zero.

Law 2 · latent conservation

Cov0(rA,rB) =E0[Cov0(LA,LBX)]

Coupling hidden by observation returns with opposite sign as correlation between observable ratio fields. Neither side is directly the other; only the average is exact.

Law 3 · nested state expansion

CX = E[CXWX] + Cov(rAXW,rBXWX)

Add a candidate measurement. If the residual curvature disappears, the old curvature was variation of now-separated mechanisms across the new coordinate. Curvature becomes a search criterion.

Missing Fisher information

For infinitesimal perturbations the observed curvature equals the conditional covariance of the complete-state intervention scores. That quantity is the off-diagonal missing-information term in Louis's decomposition, and therefore the negative of the matching observed-information cross entry whenever the complete cross-Hessian vanishes. State expansion recovers the missing mechanism information in stages by the tower property.

The tidy two-term split is special to covariance. At orders three and above, Brillinger's conditioning formula brings in set-partition sums.

Flatness is never a reward

Curvature survives any invertible transformation of the state, and any learned representation that satisfies design sufficiency, which is what makes representation search legitimate. Rewarding a representation for flatness inverts that: an encoder scored on low measured curvature can score well by discarding the coordinates where the curvature was visible.

The prescribed order avoids that trap. Learn the representation without rewarding flatness, freeze it, check that raw observations carry no residual regime information, test flatness on held-out data, and prefer sparse local ratio supports.

Orientation without ratio estimation

Delete a coordinate. See what survives.

Within a localized family, delete each coordinate in turn. The target is the coordinate whose deletion leaves the remaining marginal law invariant. Estimation reduces to a handful of low-dimensional two-sample equivalence tests instead of conditional means of a calibrated density ratio.

The master marginal identity

E0[re(X) ∣ XA] = pe(XA)p0(XA)

This identity is what makes deletion testing legitimate. The conditional mean of a high-dimensional ratio equals a low-dimensional marginal ratio, so the ratio itself never has to be estimated.

Deletion evidence

equivalence bounds

Scenario
0.20

A deletion is certified invariant only when the whole interval sits below the tolerance, and certified changed only when it sits entirely above.

Failing to reject a discrepancy is not evidence of invariance. The decision uses two-sided equivalence bounds on the normalized discrepancy R_v = D_v / (D_full + eta), and an interval that straddles the tolerance yields the honest answer rather than the convenient one. A target is returned only when exactly one coordinate is certified invariant and every competitor is certified changed.

Deletion pass count1
State machineUNIQUE_TARGET

Normalized discrepancy by deleted coordinate

oriented

Exactly one coordinate is certified invariant and every competitor is certified changed. The family is oriented.

Undercoverage degrades gracefully. When the proposed parent set is incomplete, an affected descendant generically drives the pass count to zero instead of quietly reversing the arrow. Given the family sets, the graph can be reassembled by peeling in topological order even without target labels.

Multiple passes end the run. An ambiguous orientation is reported as ambiguous. A proposal adapter may rank which follow-up experiment to run next, and the admissible class contains only replacement conditionals for the proposed target, but nothing in that ranking is permitted to resolve the ambiguity by itself.

Inference without category mistakes

Two tracks, matched to the sampling design

Density curvature is a functional of four normalized regime laws. A residual-product test characterizes zero curvature only when the pooled design odds are one. Confusing the two does not weaken the test; it silently changes the null.

Four-law functionals

For arbitrary known, state-independent regime quotas.

Mw = E0[w(rABrArB)]

Estimable as E_AB[w] - E_A[w r_B] with ratio nuisances normalized on held-out baseline data. It needs overlap control but no assignment model. Under positivity, vanishing for every bounded witness is equivalent to zero curvature almost surely.

The sampling-odds anchor

κAB(x) = log q11(x)q00(x)q10(x)q01(x) log ρ11ρ00ρ10ρ01

Curvature is excess conditional design log-odds once the known pooled-design odds have been subtracted. Equal-corner balancing is safe when it follows from the sampling contract, and equal empirical counts on their own establish nothing: counts that happen to look balanced are not an assignment mechanism. Arbitrary per-corner quotas are not safe at all. Under pooled odds rho, zero conditional covariance characterizes kappa = -log OR(rho) rather than kappa = 0, and a free-intercept multinomial fit can bury the offending term completely if the analyst reports a centered interaction instead of reconstructing the full posterior odds and subtracting.

Grade 1 · verified

product_odds_verified

A completed sampling-odds audit, carrying its audit identifier and a sha256: fingerprint of the declared allocation source. The constructor recomputes the pooled log-odds and refuses the object outright if it is not product.

Grade 2 · reweighted

reweighted_to_product

A completed reweighting audit with its own artifact fingerprint. The distinction that matters is tense: this records a reweighting that happened, never one that is planned.

Grade 3 · neither

diagnostic_only

Everything else, carrying the reason it could not be graded higher. A projection built on this grade can inform, and it can never certify.

The evidence travels with the estimate. Every GCM projection takes a design-evidence object as an argument and serializes it alongside the result, so a residual-product number can never circulate without the grounds that made it admissible. Deserialization re-runs the constructors and bit-compares the recomputed probability sum and log-odds ratio, which means a hand-edited product_odds_verified object cannot acquire authority by skipping the constructor.

Preflight the design before you touch the data

Track eligibility is decided from the manifest alone. Change the corner quotas and the selection contract below and watch the verdict, which follows the same branches mic preflight takes. The reason codes are the ones in mic-audit.

Experiment manifest

corner quotas

Requested inference track
Within-regime selection contract
What was randomized?

In exploratory mode nothing is ever serialized as a passing certificate. Every affected result is watermarked diagnostic instead.

Preflight report

READY
Pooled design log-odds log OR(rho)0.000000
Product samplingyes
Four-law eligibleyes
Product-factorial eligibleyes

The report opens with assumptions and abstentions, never with a table of p-values.

From a public square to a fail-closed report

Can I put my experiment in here?

Four-law mode does not need product assignment. The gate is a complete square, a declared randomization unit above the measurement, and known state-independent quotas. The three public designs below already have that shape. None of their microdata are bundled; a missing file is an abstention, not a fake certificate.

GEO GSE133344

Norman 2019 Perturb-seq

Control, singles, and doubles are a square. Same-complex pairs should look epistatic on phenotype and flatten once the complex is observed. Linear pathways should uniquely orient. Cells are measurements; replicate_id is the unit.

Track: four_law until construct assignment is product.

NCI-ALMANAC / DrugComb

Drug combination screens

Vehicle, A, B, and combo, when all four arms exist. Shared-target pairs should curve until target engagement is added. Incomplete squares are dropped, never imputed. Cluster at the plate or replicate, not the well.

Track: four_law. Combo screens are almost never product-factorial.

Graduation / cash × training

Factorial anti-poverty RCTs

Does cash work through training, or are they autonomous mechanisms? That is locality plus deletion, not an interaction coefficient on consumption. Cluster at household or village. Person-period rows are measurements.

Track: four_law unless assignment odds at that unit are product.

The histogram four-law path is a diagnostic. It fingerprints the table, refuses clusters that span regimes, reports raw normalizer residuals, and draws a coarse κ field. It does not localize a family, does not orient a target, and never serializes certificate_status: passed. Templates and the mapping are in docs/DATASET_ELIGIBILITY.md.

Proposal adapters

What is allowed to count as evidence

Passive DAG learners, parsimony searches, residual heuristics and previous audit runs are all welcome here, connected as proposal adapters. The boundary they sit behind is narrow, and the code enforces it.

It may decide what to test next. It may not decide what is true.

The epistemic boundary, as specified in docs/PROPOSAL_ADAPTERS.md
Proposal outputPermitted useProhibited interpretation
Candidate support or DAG edgeOrder localization and deletion testsConfirmed parent, target, or causal edge
Parsimony-frontier inclusion frequencyStability diagnostic and search priorityCalibrated edge probability
Residual or model-importance scoreExploratory witness or candidate generatorEvidence of causal direction
Candidate measurement blockQueue a held-out state-expansion comparisonExplanation of curvature
Candidate follow-up tilt or design cornerPlan a new randomized experimentResolution of an existing ambiguity

Naming rule

Uncalibrated selection frequencies, model importances, Bayes-free score ratios and normalized vote counts are recorded as score, frequency or priority. The words probability and confidence are reserved for quantities that were actually calibrated, against a stated population.

Discovery and confirmation

Strict mode accepts two routes: freeze the proposal before any confirmatory data is observed, or allocate outer folds at the true randomization unit before fitting the proposal. Inner cross-validation does not repair the reuse. With no independent confirmation sample left, the run is marked diagnostic_only.

Parsimony frontier

Candidate supports are ordered by cardinality first and only then by a learner-specific complexity measure, because model byte size and parameter count are not comparable description lengths across learner families. Inclusion frequencies are descriptive rates over a designed ensemble.

Three authority tiers, and one wall that does not move

Point the system at a pile of tabular data with no manifest, no declared design and no prior graph, and it will recover what it honestly can. Autonomous mode is a crawler over the certificate machinery rather than a new inference principle, and it weakens nothing. Three of the four missing contracts have partial data-driven resolutions: regimes can be discovered from context columns, sampling proportions can be estimated from corner counts, and randomization units can be guessed conservatively and audited at the coarsest plausible clustering. The fourth does not move.

State-independent selection cannot be established from the data alone.

No pattern in observed rows proves that inclusion did not depend on state within regime. That is the permanent wall between surveying and certification, and every artifact a survey emits carries the sentence saying so.

Authority tiers, from docs/AUTONOMOUS_MODE.md
TierRequiresSerialized as
A · certified Declared sampling, selection and clustering contracts, supplied by a human or an upstream system certificate_status from the strict ledger
B · corroborated diagnostic Every internally checkable audit passes: design estimability, four-law form, overlap, lens battery, conservation cross-checks, and negative controls where they exist diagnostic_only plus a typed checklist
C · proposal Passive-learner output, quarantined behind the boundary above proposal_only

Why tier B is the interesting one

Far stronger than a discovery algorithm's edge weight, and deliberately weaker than a certificate. It carries no scalar score and no total ordering, so a reader always sees which gates passed and which were inapplicable. The crawler can be unleashed without supervision because overclaiming is a type error rather than a matter of restraint.

The testability atlas

Before any estimation, design discovery reports which context pairs form usable interferometers on this data, which are aliased, and which need more corners. Graph assembly then proceeds under abstention, where the holes are the product rather than a defect, and curvature drives the search for what to measure next.

The data audits itself

The algebra is overdetermined, so conservation acts as an internal consistency check on estimates the survey computes anyway. It only counts as estimator evidence when the two sides come through separate held-out routes; derived from the same fitted ratios it can pass by construction, and is recorded as algebra-consistency only.

The estimator lens battery

Run the same projection through several estimator families and compare them. The gate is deliberately one-sided. Disagreement blocks a strict certification, since a projection that depends on the learner is not a property of the data. Agreement earns nothing at all: the families share folds and data, so the comparison is a preregistered sensitivity heuristic and never a calibrated joint test.

Estimator families

estimate and standard error

3.0

Each pairwise gap is scaled by the root sum of squared standard errors. A non-positive or non-finite standard error is rejected outright, and a battery of fewer than two families is rejected as well.

Lens battery audit

AGREES
Max scaled gap0.000000
Worst pair

Ten ways a strict run refuses to answer

Strict mode returns no certificate when any of these holds. Exploratory mode continues, and every affected result carries a diagnostic watermark that stops it from ever being serialized as certificate_status: passed.

  1. Within-regime selection is unknown, or depends on the state.
  2. Common support is inadequate.
  3. GCM was requested under non-product sampling with no reweighting.
  4. A design contrast is aliased or unidentifiable.
  5. Deletion gives no unique orientation.
  6. The cluster unit is missing or inconsistent.
  7. Model calibration falls outside tolerance.
  8. The self-normalization residual exceeds policy.
  9. Negative controls show implementation curvature above threshold.
  10. A data-adaptive proposal reuses its discovery observations as confirmatory evidence.

A unified audit pipeline

Two statistical primitives, applied repeatedly

The whole protocol runs on weighted residual products and low-dimensional two-sample equivalence tests, over shared cross-fitting and cluster-resampling infrastructure.

  1. Validate design

    Sampling odds, selection contract, overlap, randomization unit.

  2. Localize

    Design-IAMB with cross-fitted residual-product independence tests.

  3. Orient

    Simultaneous MMD and energy-distance equivalence tests over deletions.

  4. Interfere

    Weighted GCM projections and joint multinomial interaction fields.

  5. Predict

    Self-normalized primitive-ratio products on held-out combinations.

  6. Expand state

    Find measurements that flatten reproducible curvature blocks.

Cross-fitting graph

Outer folds carry the scientific evaluation and keep adaptive witnesses away from their own tests. Inner folds tune nuisances. The cluster bootstrap supplies uncertainty at the assignment unit and never at the row.

Compositional prediction

Held-out pair distributions are predicted by the product of primitive ratios. The raw normalizer residual is reported alongside the self-normalized estimate, since self-normalization can conceal the very failure under test.

Longitudinal caution

Ratios are estimated at the transition level. Whole-trajectory ratio products are barred as a primary estimator because their variance grows exponentially in the horizon, and the temporal graph is unrolled rather than pretending a static DAG has no cycles.

Safe-Rust reference architecture

An evidence-producing audit system, not a modeling framework

Unsafe code is forbidden across the workspace. Every stochastic operation takes a deterministic seed recorded in the evidence ledger, and every optimization has to preserve both the exact-algebra conformance fixtures and the strict-mode failure behaviour. Neural components are optional; every core theorem and exact simulation runs without them.

FrankenPandas

Typed ingestion, Parquet and Arrow interchange, fold tables, grouped summaries, reproducible report frames.

FrankenNumPy

Dense arrays, deterministic random streams, broadcasting, Gram matrices, vectorized ratio and kernel evaluation.

FrankenSciPy

MMD and energy statistics, optimization, bootstrap and permutation engines, stable special functions.

FrankenTorch

Joint multinomial models, differentiable density-ratio fields, representation sufficiency audits, CPU and Metal execution.

mic-coreexact algebra, curvature, conservation, compensated summation
mic-designfactorial geometry, face enumeration, rank and null space
mic-datamanifests, std CSV ingest, cluster fingerprints
mic-statsGCM, MMD, energy distance, parsimony frontier
mic-modelposterior-odds reconstruction, hierarchical logit
mic-auditevidence ledger, reason codes, fail-closed policy
mic-enginepreflight, histogram four-law, mic-tabular
mic-simexact scenarios and closed-form fixtures
mic-clisimulate, design, validate-manifest, preflight
mic-conformancecross-crate golden journeys, no runtime API

Reproduce the fixtures

# exact-population scenarios, written to JSON
mic simulate all --output artifacts/simulations/rust_exact_results.json

# the design gate that this page simulates
mic preflight examples/configs/nonproduct_sampling_demo.json

# pooled corner odds on their own
mic design odds 0.10 0.20 0.30 0.40

# cluster-weighted histogram four-law (std CSV; not Packet 1)
mic-tabular report examples/configs/four_law_discrete.json --base-dir .

The mic surface is simulate, design, validate-manifest, preflight, orient, propose-tilt, version. The std-CSV path lives in a second binary, mic-tabular, with ingest, four-law, report and survey, until that wiring lands on the reserved CLI. Unknown mic commands exit 2, while an abstention does not: refusing to answer is the honest outcome, not a process failure.

What a run leaves behind

run/
  manifest.resolved.json   provenance.json
  design_audit.json        proposals.json
  overlap.json             localization.json
  orientation.json         moments.json
  curvature.json           composition.json
  state_expansion.json     negative_controls.json
  report.html  report.md   evidence.jsonl

Every node in the execution graph records input hashes, code revision, feature flags, seed, wall time, numerical policy and output hashes.

Machine reason codes emitted by mic-audit
CodeRaised when
non_product_sampling_for_gcmA residual-product track was requested where a face lacks product pooled odds.
state_dependent_selectionWithin-regime inclusion is unknown, or depends on state and is unmodeled.
selection_model_unvalidatedA selection model was declared with no validated evidence attached.
overlap_failureCommon support is too thin to compare the regimes.
orientation_unresolvedDeletion testing produced no unique target.
design_contrast_aliasedThe requested contrast cannot be separated in the observed design.
no_testable_flatnessThe observed corners leave no lack-of-fit degree of freedom beyond main effects.
non_square_contrasts_requiredObserved square contrasts do not span the testable lack-of-fit space.
negative_control_curvatureA negative control shows implementation curvature above the floor.
estimator_family_disagreementEstimator families differ beyond the preregistered gap tolerance.
cluster_spans_regimesA randomization unit appears under more than one regime, so the assignment unit is not a regime unit.
missing_regime_dataA declared regime has no included clusters in the ingested table.
empty_common_supportNo cell of the projection carries positive mass in all four corners.

Empirical path

Validate cleanly, then scale scientifically

First

Factorial feature-flag pilot

Independent assignment, rich telemetry, negative-control pairs drawn from physically separate modules, and known deployment units make engineered systems the cleanest first validation. Randomization at the deployment unit means inference at the deployment unit. Record what was actually delivered as well as what was assigned.

Success looks like square closure on the negative controls, one strong outcome interaction whose curvature is equivalent to zero, and one deliberately shared-resource pair whose curvature shrinks once the resource state is added.

Second

Combinatorial Perturb-seq

Compare conventional epistasis against full-distribution curvature, treat guide competition as an implementation effect rather than a biological one, and predict held-out pair distributions from the singles. Alternate guides, guide-order swaps, non-targeting controls and dose matching all carry weight here.

The unit of inference is the replicate, construct or batch. Cells are measurements, not independent interventions.

localitywhich causal family changed

deletion invariancewhich family member is the target

square curvaturewhere autonomous composition fails

conditional score covariancehow much joint mechanism information was lost