Two documents written for different worlds — organizing LLM-enabled robotic forces and governing LLM worker minds' memory — turn out to be solving the same problem: how to grant a fluent machine substantial initiative inside defined boundaries without surrendering accountability, auditability, or control. This audit measures the as-built system (Tasks 13–15, staging-verified) against every principle in the military document, honestly.
independent derivations of the same maxim
Centralize accountability, standards and authorization; decentralize bounded machine employment.
Floors in code, taxonomy in data, prose at the edge — stability from immutable history plus validated change; flexibility from class-gated autonomy at the level where the knowledge lives.
The convergence is structural, not cosmetic. Both designs refuse the two tempting simplifications — machines as personnel, and machines as mere tools — and land on the same third thing: delegated-action systems with centralized certification and decentralized, bounded, evidence-expandable employment. Where the military document is ahead of us (authorization expiry, reversible certification, an uncertainty grammar, adversarial testing), it reads as a roadmap, not a rebuke — see §05.
their institution ↔ our mechanism, with the implementing artifact
| Military concept | Governed Mind Spec mechanism | Implementing artifact |
|---|---|---|
| Human–Machine Team (one human leader, machines with a control suite) | Worker + human owner + consent gates (accept ≠ provision) | chief-worker card escalation.humanOwner; routes/chief.ts consent flow |
| Approved autonomy profile per machine | MindSpec + rendered Operation Manual (constitution/working) | memory-plan.ts · manual-render.ts |
| Task–environment–authority matrix (§2.3) | ViewSpec: view × write × provenance × retention × index × metabolic class | memory-plan.ts ViewSpec |
| Three chains: operational / technical / force-authorization (§5) | OrgBlueprint (acts) / substrate invariants (117) / floors + write postures | blueprint-rules · mind-tools CONVENTIONS · memory-plan-rules |
| Independent certification authority (§4.6, §10) | Deterministic compiler/validator — floors code-injected, never model output | blueprintToMemoryPlan (D-7/D-8); amendments gateClassFor |
| Responsibility stack — no accountability gap (§2.2) | Operation journal + derivedFrom provenance + G-2 authority accounting | memory-journal.ts · insight front matter · auditor.ts |
| Autonomy expands only after demonstrated reliability (§10.4) | G-8: feasibleNow holds targets; promotion earned by frozen evidence | promotionCriteria · draftPromotionAmendments |
| No unrestricted self-modification; changes logged, bounded, reversible (§7.5) | G-9 metabolic classes + the amendment fast loop (event-sourced revisions) | amendments.ts · amender.ts · chief_spec_amendments |
| Configuration identity; software as ammunition (§12) | card.lock integrity, mind.json seed ledger, immutable pool history, drift detection | drwn lock/ledger · 117 contract · spec-store fail-closed loader |
| LLM as advisory interface; deterministic policy engine checks (§2.4, §7.1) | Strict draft schemas — the model proposes bounded drafts; the app stamps identity, paths, time | memory-schema.ts · distill.ts snapshot-bound provenance |
| Machine-readable orders with boundaries and abort rules (§7.2) | The constitution: locked views, never-do list, gates, evolution meta-rules | manual M0–M6 sections |
| Stop-the-system culture; reward justified intervention (§9) | G-3 overrides with mandatory reasons; locked walls escalate rather than fail silently | OverrideObservationDraft · AMENDMENT_LOCKED 409 + escalation payload |
| Dissent channels outside the command chain (§9) | G-6 dissent memory: perspectives preserved, never flattened into consensus | DissentPerspectivesSchema · dissentRequired |
| No-blame incident learning (§9, §13) | G-7 experiment observations — informative failure distilled win-or-lose | ExperimentObservationDraftSchema |
| Multiple independent brakes (§15) | Consent gates, fail-closed parsers, last-placement guard, human-only retirement, operator-gated resets | provision 409 · parseMemoryPlan · unplace guard · reset --remote refusal |
| Governance metrics: envelope compliance, audit completeness (§17) | Auditor three-edge: plan↔reality, replay↔spec, blueprint↔spec + cache coherence | auditor.ts verifyReplay/blueprintConsistency/verifyCacheCoherence |
| Modular acquisition; replaceable model (§14) | Substrate composed never forked; model is an env slug; safety lives outside the model | 115 doctrine · DEFAULT_MODEL_SLUG · Zod/rules/journal |
| Machines are equipment with team interfaces, not comrades (§9) | Minds are card-defined, versioned, auditable configurations — persona is voice, not personhood | card/persona/beliefs machinery · checkpoint lineage |
click a row for the evidence; filter by verdict
where the texts could be swapped and still be true
"Deployed systems should not rewrite core mission or safety logic through open-ended learning. Adaptation can occur in controlled layers… Changes affecting behavior should be logged, bounded and reversible. Model updates should require configuration control."
G-9 splits every field into locked / experimental / open. Open adapts autonomously (taxonomy — their "route optimization"); experimental needs an operator gate; locked never changes in the fast loop. Every amendment is an event-sourced revision with frozen evidence, a decisions-view log entry, and a replay check that detects any bypass.
"1. human communicates intent; 2. language model generates a proposed structured plan; 3. deterministic policy engine checks the proposal; 4. operator reviews…; 5. authorized task controller executes; 6. independent monitors detect boundary violations."
The model may produce only bounded semantic drafts — strict Zod schemas reject any identity, path, or timestamp it invents; provenance must resolve inside the supplied snapshot; the journaled executor writes; the auditor monitors drift. The live distill incident proved the boundary: when the model followed a stale skill contract, the parser failed closed — nothing wrong reached storage.
"'Certified autonomous robot' is too broad a category. A more defensible statement is: This system configuration is approved for these tasks under these conditions with this level of human supervision."
No mind is "autonomous" as an identity. Each carries a rendered constitution naming its locked views, write gates, promotion criteria, and evolution rules — derived from one plan revision, replaced only through gated change. The manual literally is the approval statement, per configuration, per mind.
red = new findings this lens surfaced; amber = already on the register
The §2.3 matrix asks "for how long does authorization remain valid?" Our postures, grants, and promotions never lapse — once granted, autonomy persists until actively amended.
§10.4: "Certification should be reversible. A serious incident… may reduce or suspend authorization." G-8 promotes on evidence; nothing demotes on evidence. The auditor reports incidents but cannot suspend a posture.
§7.3 demands separating verified observation / inference / report / hypothesis. Our observations carry provenance (source, refs) but no epistemic status; the refinery's evidence_type (explicit/implied/structural) never crossed into the Worker Mind schema.
§10.2: test with deceptive inputs, contradictory orders, prompt injection. Our gates cover faults, tampering, and schema violations — but no adversarial corpus attacks the ingest/distill boundary (poisoned pasted text steering distillation).
§11.2 warns against one foundation model. We run one model slug and one storage service — and lived a common-mode event when provider drift broke every distill identically (root-caused in Task 13). The fail-closed boundary contained it; diversity would have avoided it.
§4.6/§14: the fielding office must not self-certify. Our validator is code (good: not the model) but lives in the same codebase and actor as the operator suite. The three chains are logically separated, not institutionally.
§11.1: predetermined behavior on comms loss, mission-specific. Our de facto behavior is fail-closed (MEMORY_UNAVAILABLE, journaled resume) — a sane default ("stop, preserve data"), but chosen implicitly, not declared per mind in its manual.
§7.4: machines must not argue "trust me." Our skill texts instruct "mark uncertainty rather than overstating," and G-6 preserves dissent — but nothing structurally strips persuasive self-advocacy from model output.
§11.3: isolate compromised machines, revoke credentials, preserve evidence. γ's grants projection makes revocation expressible (per-worker identities, folder grants), and immutability preserves evidence — but no quarantine runbook exists.
ordered by leverage; the first four are the new-finding closures