Risks and open decisions
26. Risks and Mitigations¶
| Risk | Consequence | Mitigation |
|---|---|---|
| Tolaria policy creep | Training infrastructure begins ranking or rejecting candidates | No utility imports; neutral protocols; namespec and import tests |
| Mainline–branch divergence | Counterfactual results do not describe live training | One step engine; parity tests; explicit approximation error |
| Momir mode collapse | Best-of-\(K\) candidates are functionally identical | Explicit latent, winner-take-all objective, functional diversity metrics |
| Elesh overreach | Structural gate pre-judges utility and biases the pool | Deny reward/future-utility inputs; rule-driven hard checks |
| Tezzeret semantic drift | Compiled artefact differs from canonical design | Canonical hash, manifests and Urabrask runtime conformance |
| Urabrask judicial creep | QA begins issuing admission recommendations or tokens | QualityReport schema excludes verdicts; forbidden imports; tests |
| Augustin evidentiary creep | Judge alters tests or gathers favourable evidence | Augustin consumes immutable reports only; no Tolaria dependency |
| QA overfitting | Test plans are tuned to known candidate families | Pre-versioned plans, source blindness, sealed regression fixtures |
| Adjudication overfitting | Thresholds are tuned after seeing confirmatory results | Validation-only calibration and frozen policy versions |
| Provider leakage | Candidate origin influences tests or judgement | Blinded Sarpadia views and source-absence tests |
| Telemetry underspecification | Different deficits appear identical | Orientation-bearing gradients, temporal context and information ablations |
| Moving target | Candidate becomes stale before integration | Latency budgets, staleness curves and re-qualification |
| Counterfactual nondeterminism | Branch differences reflect runtime noise | Academy-exact causal reference, divergence localisation, measured Field uncertainty and escalation |
| QA cost dominance | Evidence costs more than the adaptation it protects | Predeclared budget and QA coverage–cost Pareto curve |
| Survivorship bias | System cannot learn refusal or failure modes | Retain structural rejects, QA failures, no-op and long-term regressors |
| Host co-adaptation | Same-host ablation exaggerates value | Separate no-op and re-adaptation branches |
| Install–lyse oscillation | Boundary growths churn as execution noise moves the estimate | Threshold hysteresis (INV-33, ADR-0005): admit strictly above retain by a versioned band sized against measured σ_exec; churn metrics; cooldowns as frequency limiter only; pre-registration |
| Tamiyo micromanagement | Strategic controller becomes local policy | Slow cadence, aggregate inputs and interface prohibition |
| Narset budget escape | Tactical controller creates ungoverned capacity | Envelope validation in Leyline, Augustin and Kasmina |
| Narset co-design / editorial angle | Tactical policy encodes diagnosis, topology or ancestry into the assignment | Narrow GrowthIntent; schema-forbidden fields; direct Nissa-to-Momir route |
| Telemetry mediation | Momir sees Narset's interpretation rather than the host observation | One canonical Nissa publication with shared observation identity |
| Covert request channel | Narset and Momir encode designs through continuous budgets, aliases or candidate count | Coarse enums, deterministic canonical resolution, invariance and anti-collusion tests |
| Bootstrap ceiling | Momir becomes a blueprint selector or mutation table | Parent-relative and reference-frontier objectives; ancestry dropout; de novo gate |
| Permanent scaffold dependence | Production generation fails without stock reference seeds | Explicit null ancestry, withdrawal schedule, held-out scaffold-free evaluation |
| Permanent bitwise burden | Exactness requirements prevent realistic kernels, scale or hardware evolution | Treat Academy exactness as a retained metrology profile; calibrate Field execution rather than requiring universal bitwise identity |
| Premature execution withdrawal | Field noise changes rankings or no-op decisions before it is understood | Decision-aware Field gate, uncertainty margins and Academy escalation |
| Lockstep scaffold withdrawal | One subsystem loses support because another subsystem is ready | Independent three-axis ScaffoldState and separate gate ownership |
| Multi-scaffold confounding | A failure after simultaneous withdrawal cannot be attributed | One-axis confirmatory transitions and declared interaction experiments |
| Hidden scaffold correlation | Fixed seeds, exact execution and stock ancestry make one another look stronger than they are | Selected scaffold interaction matrix and final fully withdrawn corner |
| Kasmina legacy blueprint creep | Host physiology quietly regains a preferred design catalogue | Reference population lives in Sarpadia/controls; Kasmina imports no blueprint library |
| Sarpadia leakage | Related branches cross train/test boundaries | Group split by base host trajectory |
| Oona control coupling | UI or logging changes training behaviour | Read-only events and isolation tests |
| Codename opacity | New contributors cannot find responsibilities | Plain-English README header, glossary, typed record names and diagrams |
| Namespec drift | One codename accumulates multiple meanings | ADR, package ownership manifest and compatibility sunset dates |
| Toy-task non-separability | Methods appear equal because the space is too small | Sweep width and grammar complexity before broad conclusions |
| Graph grammar explosion | Design and verification become intractable | Staged grammar levels and explicit ceilings |
| Asynchronous compilation staleness | Candidate is obsolete before Tezzeret finishes | Compilation budget, caching and re-qualification |
27. Open Design Decisions¶
The subsystem names and their principal authorities are not open decisions. Namespec 1.0 is locked. The following implementation choices remain open.
27.1 Project-level name — decided¶
Decided 2026-08-08 (ADR-0003, PDR-0006): the name is locked as Simic —
repository, package (src/simic/), and presumptive publication name. The
predecessors (ESPER, ESPER LITE) present as lineage history behind a clean
seam. This decision never affected the subsystem names, which are locked
with Namespec 1.0 and reaffirmed unamended in ADR-0003.
27.2 Default maturation mode¶
Should the ecological default be one-shot generation, isolated nursery training, or a mixed Narset policy after both modes are characterised?
27.3 Winning-branch deployment¶
Should live operation adopt the winning branch state directly or restore and replay it? The answer may depend on hardware placement, branch latency and checkpoint cost.
Conditionality note (§6/§8 of docs/concept/reviews/2026-08-08-esper-pivot-peer-review.md, ruled at the decision gate): fast landing (flash-clone) is not required for the pivot — ordinary blending is sufficient — but if it is ever pursued, two collisions bite. A candidate matured in a branch is co-adapted to that branch, so copying it into a live host that followed a different trajectory is the transplant §14.6 forbids; fast landing therefore requires resolving this decision toward branch adoption (restore-and-replay instead pays the replay cost and reopens the staleness window). And §12.4's minimum blend and holding windows assume gradual alpha; a one-or-two-step landing trips them, so that mechanism would need re-deriving for a regime where alpha is not the thing taking time.
27.4 Growth grammar¶
What is the smallest safe grammar materially more expressive than a residual microcell without becoming unrestricted architecture search?
27.5 QA horizon and evidence floor¶
What horizon captures trajectory value before branch-divergence noise overwhelms the intervention signal? What evidence-completeness threshold should force retest?
27.6 Urabrask–Augustin contract¶
Which measurements are raw, which are certified derived facts, and which hard QA statuses make a candidate ineligible by policy? The separation is locked; the exact report surface is not.
27.7 Commitment semantics¶
Does commitment retain a named removable growth indefinitely, or may a later authorised consolidation merge it into the host while preserving lineage?
27.8 Retrieval similarity¶
Should Sarpadia retrieve by telemetry distance, learned state embeddings, gradient alignment, functional effect, task context, lineage history or a calibrated mixture?
27.9 Augustin–Emrakul maintenance boundary¶
Should Augustin issue only a tenancy verdict, or also a bounded class of permitted maintenance actions? The preferred direction is verdict plus constraints, with Emrakul selecting the safe physical schedule.
27.10 Oona transport boundary¶
Should Oona own the event bus implementation or only durable projections and operator adapters? In either case, Leyline owns schemas and training remains independent of Oona availability.
27.11 Request-channel granularity¶
Fix the smallest set of coarse GrowthIntent classes that gives Narset useful tactical authority without creating a high-bandwidth covert design channel to Momir.
27.12 Bootstrap reference population¶
Fix the initial reference families, canonical graph forms, mutation radii, ancestry-dropout schedule and scaffold-withdrawal gate. The corpus must be broad enough to teach structural literacy without becoming the permanent ceiling.
27.13 Tolaria integration boundary¶
Should Tolaria call a generic Kasmina host protocol, or should a thin integration adapter live outside both domains? The result must preserve infrastructure neutrality and one execution path.
27.14 Tolaria Field-withdrawal gate¶
Fix the acceptable Field-to-Academy selection regret, accept/no-op disagreement, uncertainty coverage, tail-failure rate, calibration expiry conditions and mandatory escalation policy for each assurance class.
27.15 Scaffold interaction budget¶
Fix which two-way and three-way scaffold interaction cells are required at toy, image and scaled stages. The programme must preserve interpretability without committing to an unnecessarily exhaustive Cartesian product at every scale.