Developer documentation
Changelog
Last reviewed 31 August 2026
All docs
CAIN Trust Fabric Changelog#
All notable changes, architectural milestones, cryptographic primitives, and enterprise releases for the CAIN Trust Fabric are documented here.
⚡ View Live Interactive Telemetry & Changelog Feed
Operational proof -- cain-mr-02 (written every 30 minutes by the cluster check itself)#
- proof 62 at 2026-09-28T07:08:07Z: OPERATIONAL · 4/4 replicas, agreement yes · write committed at sequence 610 with a quorum certificate signed by 3 members (atl, lax, sjc), verified: True
- signed proof · verify the whole chain
Operational proof (written every 30 minutes by the cluster check itself)#
- proof 69 at 2026-09-28T07:23:09Z: OPERATIONAL · 4/4 replicas, agreement yes · write committed at sequence 17773 with a quorum certificate signed by 3 members (atl, lax), verified: True
- signed proof · verify the whole chain
Live soak (running, updated hourly by the soak itself)#
- checkpoint 33 at 2026-09-28T06:42:10Z · 33.045 of 72.0 h · height 17427 · 16334 committed, 102 refused · 93 replica kills / 93 restarts · divergences 0 · anomalies 0 · MCPGate authorization enforced (ALLOW + refused replay)
- signed checkpoint · verify all checkpoints
[42.38.1] — 2026-09-28 (All evidence published on all three sites; Engine B authority verifier + conformance/bypass reports)#
- Every claim's evidence is now served identically by all three sites. The two evidence roots (
platform-gateway/frontend/proof/bundleandclawx-site/evidence) were mirrored so the union of every published file is served by cainstudio.online, mcpgate.online and clawx.click (cross-site one-root-only 443 → 0, NOT_SERVED 0). The signed claims registry, the full public evidence inventory, the signed machine index and the one-command verifier are now linked from the/proofand/evidencepages, andverify_all.pyre-derives every claim's artifact SHA-256 from every site: 186/186, INTACT. - Engine B independent verifier (Prompt 2, Part 41).
scripts/cain42_l5/verify_authority_bundle_engine_b.py— a second clean-room verifier that imports nothing from CAIN, alongsideverify_authority_bundle.py. It independently re-derives canonical JSON, domain-separated Ed25519 signatures, expiration, revocation, delegation scope, authority intersection, policy binding and action binding forAgentIdentity/AuthorityGrant/DelegationGrant/AuthorizationDecision.tests/test_cain42_l5_authority_verifier.py— 25/25 attack cases refused (forged / altered / replayed / expired / revoked / key-substituted identity; scope expansion; constraint widening; delegation depth; delegation outliving its parent; expired and revoked grants; trust-floor bypass; policy downgrade; self-authorization; policy-scope escape; action substitution; cross-tenant access; single-use replay; over-authorization). Published at/proof/bundle/l5-authority-2026-09-28/on all three sites (verifier.py.txt, worked example, conformance, bypass report, forensic inventory). - Independently verified from more than one source (2026-09-28). The published evidence is re-verified by five independent verifier implementations that share no code —
verify_all.py(every claim and artifact, from all three sites), two separately-written clean-room L5 authority verifiers (verify_authority_bundle.pyandverify_authority_bundle_engine_b.py, both VALID on the same bundle),verify_l5_bundle.py, and the multi-source driver's own registry check — and they all agree:python3 scripts/cain42_l5/verify_multisource.py→VERIFIED_MULTI_SOURCE. Report at/proof/bundle/l5-authority-2026-09-28/CAIN42_MULTISOURCE_VERIFICATION.json. Scope: multi-implementation verification by the same project; no third party has reviewed CAIN-42; not a certification. - Honest status: INCOMPLETE, not A+. The verifier is self-attested (same operator; no third party). The L5 runtime layer is delivered separately (see [42.38.0]). Performance is NOT_VERIFIED. The bypass report declares messaging and email NOT_IMPLEMENTED and the secrets broker NOT_IMPLEMENTED, and states the unconfined-process (RAW) gap: a process that does not route through CAIN or MCPGate is not seen.
[42.39.0] — 2026-09-28 (Prompt 3: continuous trajectory governance + independent verification published to both evidence roots)#
- Trajectory governance (
cain45/l5/trajectory_governance.py).AgentTrajectorywith hash-chained signed events and an RFC-6962-style Merkle root;ObjectiveEnvelopedrift detection; deterministic explainableTrajectoryRiskStatebands;AuthorizationLeases that invalidate on expiry/revocation/trust-floor/policy/context/risk breach;detect_material_change; a formal state machine; governed containment;ContinuousAuthorizer(restriction-only). 15 executable invariants and an 18-test adversarial suite. - Independent verification published. Three clean-room verifiers (no shared code, no CAIN imports) all return VALID: unified 22/22, authority 22/22, L5 gate 10/10. Transcripts + per-file SHA-256 manifest on both evidence roots:
platform-gateway/frontend/proof/bundle/l5-governance-2026-09-28/,clawx-site/evidence/l5-governance-2026-09-28/. - Claim accuracy. Independent *implementation* verification, not third-party review. CAIN-42 is not claimed to be "fully independently verified from more than one source". No L5 claim.
[42.38.0] — 2026-09-28 (Prompt 2: Agent Identity & Dynamic Authority Fabric — hosted wiring, two clean-room verifiers VALID)#
- Identity fabric (
cain45/l5/agent_identity.py). A canonical, signed, versionedAgentIdentity(model + runtime identity, owner, capabilities, public key, policy binding, trust reference, delegation root) with lifecycle CREATE → ROTATE → REVOKE → REISSUE and signed identity statements. IDENTITY ≠ AUTHORITY. - Dynamic authority fabric (
cain45/l5/authority_fabric.py). Capability/resource/operation-scoped, temporal, trust-bounded, budget-bounded grants; non-escalating delegations;AuthorizationDecision(ALLOW/DENY/ESCALATE/CONTAIN) with machine-readable reason codes; deterministic intersection of the narrowest valid constraint; revocation cascade; single-use anti-replay; explainable decisions. 10 executable invariants. - Enforcement & hosted wiring. Restriction-only wiring into
cain45/hypervisor.pyandcain/mcp_proxy.py; caller↔subject authentication; gateway serviceplatform-gateway/cain_l5_gateway.py+/l5/*API, opt-in viaCAIN_L5_GATEWAY=1, fail-closed. - Evidence. 90 new tests pass (206 more in the hypervisor/MCPGate regression group); 30 executable invariants all hold; two clean-room verifiers (no CAIN imports) return VALID —
verify_l5_bundle.py(10/10),verify_authority_bundle.py(22/22); 3,000-step soak with 0 governance violations; L5 gate 53.6 µs/op ALLOW, 45.3 µs/op DENY. - Defects fixed. Single-use replay on nonce reuse; delegation narrowing not applied to evaluation; revocation invalidating the signed grant signature.
- Limits. TESTED library, not the hosted default path; no real LLM agent; no hardware root; one provider; no third-party review; bypass audit found 57 ungoverned paths. No L5 claim is made.
[42.35.0] — 2026-09-28 (Evolution #9: six authority fabrics — intelligence may propose, only a quorum-certified decision authorizes)#
- New governed fabrics, all library-level, none able to authorize.
cain45/trajectory.py(Trajectory Firewall),cain45/stopping.py(Stopping Intelligence — a stop can only make an action *less* permissive),cain45/swarm.py(Swarm Authority Fabric — CHILD ≤ PARENT, SWARM ≤ POLICY, votes are informational),cain45/tooling.py(Governed Tool Creation + Dynamic Discovery — a new tool starts with NO AUTHORITY),cain45/memory.py(Governed Memory Fabric + Calibration — memory cannot become policy),cain45/economics.py(Resource Authority — time/scope/resource/policy/trajectory-bound, revocable, non-escalating). - One rule. LEARNING, PREDICTION, MEMORY, SIMULATION, DISCOVERY and AGENT CONSENSUS may change capability but never create execution authority. 419 new unit tests plus
tests/test_cain42_e9_security_invariants.py, one executable matrix binding every Part XXVIII invariant to a check. - Limits. The six fabrics are TESTED unit libraries, not VERIFIED; no published independent evidence bundle exercises them and they are not wired into the hosted enforcement path. World model is tabular (not neural); memory is not a vector store. The hosted runtime remains PRE_PRODUCTION.
[42.34.0] — 2026-09-28 (Governed policy evolution: learning can propose, only the cluster and a human can authorize)#
- Policy changes are governed. A tenant's policy (which capabilities a ZoD may hold, the largest budgets it may get) now becomes active only when the live cluster cain-mr-01 commits its activation. Widening it needs a registered human who is not the proposer; narrowing it needs no one, because authority may always go down. On the live cluster: an agent's first policy was approved by a human and activated (QC 16148); the agent then proposed an expansion and approved it itself — refused, and the cluster was never asked; a failure lab's restriction was activated (QC 16151) and a running ZoD lost its authority on its next call; a ZoD above the ceiling was refused.
Verify: e8-governance-2026-09-28 — python3 verify_e8_governance.py . -> 19/19 checks, VERIFIED, Python + cryptography, no CAIN code.
- Predictions, simulations, agent votes and memories are not authority. Four things an intelligent system produces were presented to the hypervisor as the basis for a ZoD: a world-model prediction with 0.99 confidence, a simulated ALLOW citing a real certified sequence, ten agents' unanimous signed vote, and a memory replaying a real decision with the verdict changed. Each was refused because the live cluster had not certified it; across the run exactly one ZoD was ever authorized.
- Limits. Scripted identities, not a real LLM agent. CAIN does not contain a world model, digital twin or learning memory: the run shows their outputs cannot become authority. The governor runs as a library on the gateway host. The registry now also names seven older files on the sites whose self-asserted statuses (CERTIFIED, PRODUCTION_HARDENED, OPERATIONAL_PROVEN...) no evidence supports; they stay for history, marked superseded.
[42.33.0] — 2026-09-28 (Fail-open defects found and fixed; the authority check is no longer quadratic)#
- Four ways authority could survive what should end it, found by an independent probe, all fixed. Moving the clock back revived an expired grant (now: a stored time high-water mark refuses a clock that goes backwards). A tool could run while the evidence log was unwritable (now: the intent is recorded before the effect, so no log means no execution). A failing trust service, and a cluster that errored during authorization, raised instead of refusing (now: recorded refusals). Tests: 4 former gaps now pass, and fail on the previous code.
- Faster. Every action re-verified every signature since the ZoD began. Now each entry is re-hashed but its signature verified once: gate p50 with 100 ZoDs in the store went from 113 ms to 18.7 ms (1,000 ZoDs: 36 ms; p99 about 0.3 s). Measured on the gateway host.
- Limits. Library-level fixes in the hypervisor on the gateway host; the p99 tail is not yet explained.
[42.32.0] — 2026-09-28 (Authority lapses on policy change, cluster membership/epoch change and spent risk/blast-radius budgets)#
- Evolution #7: the five conditions Evolution #6 left open, on live authority. Every ZoD was authorized by the live cluster cain-mr-01 (decision QC 15876, ZoD QCs 15880-15892), made one successful tool call, and was refused on the next call with the tool never running once: the policy root changed; the policy source became unreadable (unknown is refused, never read as unchanged); the risk budget was spent; the blast-radius budget was spent (actions now carry consequence classes C0 read to C4 security/infrastructure); and a delegate spent its parent's budget — delegates are charged up the whole chain, so splitting work across children cannot multiply authority. Each ZoD is bound to the cluster's real membership configuration (hash recomputed, a quorum of replicas agreeing), re-read before every action (36 live reads in this run); a changed epoch, a changed membership and an unreachable cluster were each refused. Also fixed: a retry after a replica committed but timed out used to be refused as 'not committed'; the client now takes the committed sequence from the replica's cached reply and accepts it only if that certificate verifies over the exact request.
Verify: e7-lease-2026-09-28 — python3 verify_e7_lease.py . -> 60/60 checks, VERIFIED, Python + cryptography, no CAIN code.
- Limits. Stated, not hidden: the three membership/epoch changes were INJECTED into the hypervisor's view (the live cluster was not re-keyed); the policy and budget trips are real. Enforcement is the hypervisor library on the gateway host. Only tool-call budgets were exercised live.
[42.31.0] — 2026-09-28 (Agent Hypervisor (ZoD) authorized by the live cluster; authority leases that expire, revoke and cascade)#
- Agent Hypervisor / ZoD runtime, authority from the live cluster. An agent never holds execution authority; it acts only inside a ZoD (Zone of Decision), and only after the live 4-server cluster cain-mr-01 has committed that ZoD's authorization with a quorum certificate (3 of 4 Ed25519 signatures) that the hypervisor checks itself. Code runs under real confinement (bubblewrap namespaces + cgroup v2, no network, host tree invisible). Two ZoDs were authorized at cluster sequences 13379 and 13380; 10 attacks were refused, each as a signed DENIED entry; 45 evidence entries, hash-chained.
Verify: cain45-zod-live-2026-09-27 — python3 verify_cain45_zod.py . -> 10 PASS, VERIFIED (about 0.3 s), Python + cryptography, no CAIN code.
- Authority leases, tripped live (Evolution #6). A ZoD is a temporal authority lease. For each of 9 conditions a fresh ZoD was authorized by cain-mr-01 (decision QC 15746, lease QCs 15748-15759), one tool call succeeded under that live authority, the condition was tripped, and the next call was refused with the tool never running: TTL expiry, trust below floor, agent identity swapped, tool schema changed (rug-pull), security context changed, trajectory fork, explicit revocation, parent quarantined (child loses authority with it), required evidence deleted. Four of these were holes found and closed this release: before the fix a child ZoD kept acting after its parent was quarantined, a tool whose schema changed after authorization was still called, a re-registered (swapped) agent identity kept acting, and deleting the evidence log did not stop execution.
Verify: e6-live-lease-2026-09-28 — python3 verify_e6_lease.py . -> 48/48 checks, VERIFIED, Python + cryptography, no CAIN code.
- Limits. Stated, not hidden: the invalidation logic runs in the hypervisor library on the gateway host, not on the cluster nodes; what comes from the cluster is the authority being invalidated. The evidence-deletion row is SELF-REPORTED: its result is signed by the run's hypervisor, but the log that would prove it is the one deleted. Not implemented yet: invalidation on policy, epoch or membership change, risk and blast-radius budgets. A separate 4-node 'authoritative state' layer in the code is SIMULATED (one process holds all 4 keys) and is not used for any of this evidence. seccomp, egress allowlists and hardware attestation are not established.
[42.30.0] — 2026-09-27 (Customers can verify their own decision records offline; disk and memory monitoring)#
- Offline verification of your own decisions:
GET /fabric/decisions/{id}/signed-record(your API key) returns the decision exactly as stored and signed. Check it with the published verifier, which contains no CAIN code, using a key fetched from a different site:
python3 verify_decision_record.py record.json --key https://mcpgate.online/fabric/decision-signing-key. Verifier and a worked example: decision-signing-2026-09-27. The tests use that published verifier: an exported record verifies, and an edit with a recomputed digest fails on the signature.
- Disk and memory monitoring: on 2026-09-27
/tmpon the Atlanta server (a 1.7 GB in-memory filesystem) filled to 100%, silently breaking commands and test runs, and nothing monitored disk space. A watcher now reads disk and memory on all four servers and alerts on any change. An unreachable server counts as failing, never as passing. Its only automatic action is removing stale test folders from a nearly full/tmp. Latest reading: every server is below 80% disk use. /fabric/statusnow describes the live mode, enforce by default, instead of always describing shadow mode.
[42.29.0] — 2026-09-27 (Gate X: every hosted decision record is signed)#
- Gate X PASS (5 of 9 extra gates). Each hosted decision record was already protected by a digest over all its fields and every stage, but the digest was unkeyed: anyone able to write the database could edit a record, including the stages after consensus or the final verdict, and recompute it. Every record is now also Ed25519-signed by a key kept in a separate file outside the database, and the integrity check requires both a matching digest and a valid signature.
- Verified live: a decision recorded on cainstudio.online verifies as signed.
- Public key: /fabric/decision-signing-key. Anyone can check a record's signature offline.
- Tests: an edit with a recomputed digest is caught; so are a signature moved between records and a key readable by other users. Older records are reported as unsigned, never as signed.
- Not covered: someone with root on the gateway server, which holds both the key and the database.
- Consensus client hardened: checking the cluster's commit certificate no longer throws on a cluster record without member keys or on a malformed certificate from a replica. It refuses cleanly, as documented. A test that expected an uncertified "ALLOW" to count as committed now asserts the opposite.
[42.28.0] — 2026-09-27 (13 test failures traced and fixed; a tampered-evidence flag traced and corrected)#
- The previous full run's 13 failures are fixed. Latest full run: 5453 passed. Its 8 failures are in a separate change still in progress: the claims registry grew from 16 to 24 claims, and module mirrors are not yet synced. They are not claimed as fixed. Every one of the 13 was traced, and none was fixed by weakening a check:
- The consensus tests built proposals whose digest did not match the operation they carried. The engine has rejected those since 2026-09-23, so the tests now build them correctly (173 consensus tests pass).
- The OPA policy was changed by a reviewed fix on 2026-09-23 but never re-signed, so its signature check reported a mismatch. It is now re-signed and verifies.
- The enclave test still expected the label "INTEL_SGX". There is no TEE on this host, and the test now requires the honest "software-simulated, not hardware-attested" label.
- The Verification Center flagged 1 of 16 signed claims as TAMPERED, correctly. On 2026-09-26 we fixed a bug in the 72-hour soak verifier and overwrote the published copy that the claim binds by hash. The signed bytes are restored, and the fix is published beside them as
verify_soak.v2.py.txt. All claims verify again. - That soak's real outcome: it stopped at 24 of 72 hours, and the fixed verifier returns FAIL: several checkpoints lack MCPGate enforcement evidence. Consensus stayed consistent throughout: 0 divergences, and every certificate verified. See
soak-72h-2026-09-25/VERIFIER_UPDATE.txt. The multi-region soak is a separate run. - The daily signed sync now measures the live 4-region cluster. It previously checked the retired single-host containers. Every replica's health, state hash and running image are read on its own server.
llms.txtcorrections: the retired single-host cluster is marked historical, the outdated statement that the hosted pipeline does not use the cluster is removed, the withdrawn installer is no longer advertised, and directory links now point at real files.
[42.27.0] — 2026-09-27 (Storage-loss recovery, live MCPGate enforcement, hosted certificate check, formal models)#
- Two-replica storage loss, rebuilt from off-host backups. On the live 4-server cluster, the Miami and Silicon Valley replicas lost their storage at the same time and were rebuilt only from backups held in other regions. While both were down the cluster committed 0 of 4 writes; afterwards 0 decisions were lost and all four were identical 10.3 s after restart. verify 416 certificates
- MCPGate enforcing live consensus. Authorizations committed by the live 4-server cluster; every tool call went over HTTP through the MCPGate proxy to a separate MCP server process. 5 authorized calls ran (per the server's own log); 12 attacks were blocked, and each caller received the gate's signed denial. verify with one command
- Hosted decisions: the gateway now checks the quorum certificate itself. Security fix: the hosted consensus stage used to accept a replica's word that a decision was committed, so one lying replica could have authorized a decision with no quorum. The gateway now verifies 3 pinned Ed25519 signatures over the digest it computes for that decision, and shows the check on every decision. verify a decision yourself
- Formal models checked. TLA+ models of the PBFT commit/view-change rules and of the MCPGate gate, checked exhaustively by TLC: 0 violations in 7.3 million distinct states, and all 8 deliberately broken variants caught. An earlier run had been recorded as incomplete because the checker stopped at the model's normal end state; fixed. models and results
- Claims registry rebuilt from current evidence. 24 signed claims. It had still said one host and no formal verification, and still called the failed single-host 72-hour soak 'in progress'; it now records that soak as FAILED and the multi-region soak as running. Gates: 14 of 15 (A-O) and 3 of 9 (P-X); hardware attestation is blocked (no TPM, SEV or TDX on any server). verify the registry
- One-way network partitions. On the live 4-server cluster: a replica that can talk but not listen, a one-way link between two backups, and a replica that can listen but not talk. The cluster kept committing in each case, went through view changes, and all four replicas held identical decision chains after each heal. verify 968 certificates
- Daily restore validation. Every day each replica's newest off-host backup (8 replicas, 2 clusters) is fetched from the server in another region that holds it and proven to be a quorum-signed prefix of the live history; tampered backups fail even with a re-hashed manifest. verify
- Signed evidence index for crawlers and AI agents. /cain42-evidence-index.json on all three sites: every claim with its status, limits, artifact hashes and verification command, every gate, and the live endpoints, generated from the signed registry and signed with the evidence-root key; llms.txt carries the same, generated. claims registry
- Degraded network: safe, but slow. 10% packet loss, 120±40 ms jitter, 5% duplication and reordering on all four replicas of the live 4-server cluster for 4 minutes: no fork, but throughput fell from 1.76 to 0.16 commits/s and 22 of 61 writes timed out; it recovered fully afterwards. The evidence publisher now runs the privacy firewall before anything reaches the sites (a bundle was briefly public with an internal subnet in its description). verify 2,728 certificates
- Engine 948b189 on the 4-server cluster: fewer view changes, same throughput under loss. View-change backoff now resets only when a view commits, and a fresh view is not accused. Released reproducibly (two independent builds, identical image ID; signed release manifest 17/17), upgraded replica by replica under live traffic (each caught up in 11-31 s, no quarantines). Re-running the same degraded-network test: view changes fell from 14 to at most 6, but throughput under 10% loss stayed about the same (0.16 -> 0.18 commits/s). The storm was not the bottleneck; message delivery under loss is. Safety held (4,448 certificates, no fork). before/after, verify 4,448 certificates · release manifest
[42.26.0] — 2026-09-27 (SDK guard is now default-deny)#
- Default-deny in the self-hosted SDK and MCP proxy (
cain/guard.py,cain/mcp_proxy.py): an action that is not declared is now held for approval instead of allowed. Declare authorized tools withguard(allow=[...]),MCPTransparentInterceptor(allowed_tools=[...])orCAIN_ALLOWED_ACTIONS(glob patterns). The held decision names the action and says how to declare it. Destructive and sensitive calls are still held or denied even when allowlisted. - Opting out is explicit and recorded:
CAIN_UNKNOWN_ACTION_POLICY=allowrestores the old behaviour, and every decision made that way says the caller explicitly chose it. CLAWX passes it on purpose, because its own capability gates are its authorization layer. - Tests: the 70 test files that depend on the guard pass (1516 passed, 0 failed). The test suite declares its tools by name and never uses
*. The tests for default-deny itself clear the allowlist and prove the hold. - Homepage correction: the "known gap" card still said the 5 audit bypasses were unfixed. They were fixed on 2026-09-26 (282c7a1), and the card now says so.
[42.25.0] — 2026-09-27 (Per-customer enforcement; reproducible release for the 4-server cluster; audit fixes)#
- Shadow to start, enforce to pay: every customer starts in shadow mode, free. CAIN records every verdict on their real agent traffic and blocks nothing except identity, entitlement, the kill switch and approval holds. Customers on paid plans can switch themselves to enforce, where denials actually block, with
POST /fabric/settings {"mode": "enforce"}. Changes are logged, one customer's mode never affects another's, and every decision records which mode applied. - Reproducible release for cain-mr-02: two independent, from-scratch builds produced exactly the image running on all four servers. The signed release manifest verifies 17/17, including the reproducibility checks. Release manifest Production gates: 13 of 15.
- SDK guard: allowlists (
guard(allow=[...])/CAIN_ALLOWED_ACTIONS) and an opt-in strict default-deny mode (CAIN_UNKNOWN_ACTION_POLICY=require_approval). Under the default policy, a decision now says when an action was allowed only because nothing objected rather than because it is allowlisted. - Audit fixes:
- The broken install link is replaced.
- Unsupported claims are removed from the insurance, EU AI Act and drift pages.
- The source code is backed up to three other servers daily, and a restore has been tested.
- Prometheus and the legacy cluster ports are closed to the internet.
[42.24.0] — 2026-09-27 (Four servers, four regions: any single server can fail)#
- Silicon Valley joined: the fourth server (8 CPUs, 32 GB) is on the encrypted mesh. It uses a new address range, because its k3s already uses the old one; the change was made live without restarting anything.
- New cluster cain-mr-02: 4 PBFT replicas on 4 servers in 4 regions (Atlanta, Los Angeles, Miami, Silicon Valley), one each, with keys generated on each server. It was built reproducibly (the same image ID on all four) and deployed beside cain-mr-01, whose soak continues.
- Whole-server loss, each in turn:
- Atlanta, Los Angeles, Miami and Silicon Valley each went fully offline: 6/6 writes committed every time. The leader was on the stopped server every time, so there were four successful leader changes.
- Losing Los Angeles used to stop the old layout. Now it doesn't.
- Two servers down: 0/3 committed, and those requests never entered history.
- 264 certificates; the standalone verifier passes 40/40 from the public URL. Verify →
- Computed resilience (/api/v1/live-cluster/resilience?cluster=cain-mr-02): the cluster survives losing any 1 replica, any 1 server or any 1 region. The only single point of failure left is the provider (all four servers are Vultr).
- Live API for either cluster: add
?cluster=cain-mr-02to /api/v1/live-cluster/status, /health, /membership, /qc/{seq} and /resilience. - Production gates: 12 of 15 (gate A passes; multi-provider is not claimed).
[42.23.0] — 2026-09-26 (Bit-for-bit reproducible builds; off-host backups; test suite green)#
- Reproducible builds: two independent, from-scratch builds of the same commit produced the same image ID (
sha256:bc8f22c0…), with all 11 filesystem layers bit-for-bit identical. The base image is pinned by digest, every Python package is pinned including dependencies, and timestamps come from the commit. Evidence The live cluster moves to the reproducibly built image at the next rolling upgrade, after the soak. The release manifest will claim "reproducible" only then, and its verifier checks the claim. - Off-host backups: every replica's hourly backup is also stored on another region's host, with checksums verified on arrival.
- Test suite: 870 passed, 0 failed.
[42.22.0] — 2026-09-26 (Byzantine replicas handled across three regions; proof every 30 minutes)#
- Byzantine-replica tests across Atlanta, Los Angeles and Miami on the production image:
- Forging replica: it sent votes claiming to come from another member. Every honest replica rejected them, 6 each, and 6/6 writes committed.
- Equivocating primary: it sent two conflicting, correctly signed proposals for one sequence. All three honest replicas proved it from the primary's own two signatures, quarantined it, changed view (0 to 1) and committed 6/6 under the new primary.
- 102 certificates plus the proof; the standalone verifier passes 29/29 from the public URL. Verify →
- Run on a disposable cluster with the same placement, because the production cluster refuses fault injection by design. Scope: f=1 and two behaviours.
- Operational proof every 30 minutes: a real write committed through PBFT, with its quorum certificate re-verified and every replica's state, signed and hash-chained. The homepages and the block above show the latest proof, and the proof expires after 1 hour. Verify the chain
- Test suite fully green: 870 passed, 0 failed. The earlier failures were fixed:
- A real approval-workflow bug: the trust stage held actions from new principals even after a human approved that exact action.
- Two stale copies in the pip package: it had drifted from the live code, including this week's containment fix.
Test report Production gates: 11 of 15.
[42.21.0] — 2026-09-26 (Signed release manifest; published test report)#
- Signed release manifest for what the live cluster runs: the same image ID on every host and every replica, all 1,461 image files byte-identical to the git commit (software measurement), a 47-package software bill of materials, and an Ed25519 release-key signature. A standalone verifier passes 16/16 from the public URL. Release manifest Still missing: reproducible builds and a pinned base image.
- Published test report: 865 passed and 4 failed, generated from the run's JUnit output with the counts unedited. Every consensus-engine suite passes. The 4 failures predate this work and are listed by name: 2 hosted-pipeline tests and 2 source-mirror copy mismatches. Test report
- Production fault injection stays off: the live cluster's fault-injection endpoint answers 403, by design. Byzantine-replica tests run on a separate, disposable cluster across the same three regions.
[42.20.0] — 2026-09-26 (Network partitions on the live cluster; anti-entropy fix; operational drill)#
- Network-partition tests on the live cluster: one host's encrypted link was taken down while its replicas kept running. The link was then restored by a timer on that host.
- P1, Miami cut off: the other three replicas committed 6/6. The isolated replica refused both writes sent to it and committed nothing.
- P2, 2|2 split: Los Angeles against Atlanta + Miami. Neither side has a quorum, and neither committed anything, so there was no split brain.
- All four agreed within about 3 s of each heal. 370 decisions on each of 4 replicas, identical chains, 2,960 certificates, standalone verifier 32/32. Verify →
- Defect found by the first partition run, fixed in
81bdf84and rolled out live: after a heal, an idle replica waited for the next client request before catching up. Replicas now run anti-entropy every 15 s. They catch up only through certified state transfer and quorum-signed NEW_VIEWs, and never on fewer than f+1 peers' claims. - Operational drill on the live cluster:
- Load: 1 client, 2.0 commits/s (p50 0.51 s); 4 clients, 4.9/s (p95 1.2 s); saturating near 3.3/s.
- Rolling restart of all four replicas under continuous writes: 83/83 committed, 0 client-visible failures.
- One replica's total disk loss: rebuilt from its peers in 13.6 s with an identical quorum-signed history.
- Verify 2,832 certificates →
- Two live rolling upgrades (
9fedb49,81bdf84): one replica at a time, with the primary last. Each replica caught up in 10–14 s. - Resilience profile computed from the real placement: /api/v1/live-cluster/resilience. The cluster survives the loss of any one replica, or of the Atlanta or Miami host. Losing the Los Angeles host or the single provider stops progress safely. This is shown, not hidden.
- Restore from backup on the live cluster: the Los Angeles replica was restored from an online backup taken 12 decisions earlier, with checksums matching the backup manifest. It reached identical height and state in 8.3 s while writes continued, and 0 decisions were lost. 3,096 certificates, 32/32. Verify → Still missing: off-host backups and a drill where several replicas are lost at once.
- Rollback drill on the live cluster: rolled back to the previous engine and forward again, one replica at a time with the primary last. Every replica caught up in 10–14 s, the cluster was HEALTHY 4/4 after each direction, and no replica was quarantined. Rollback record
- 72-hour soak on the live multi-region cluster started at 21:39 UTC, with signed hourly checkpoints and the live block above. Hourly backups are scheduled. Production gates: 9 of 15.
- Production gates: 7 of 15 pass (partition and recovery added).
[42.19.0] — 2026-09-26 (Operational drill on the live cluster: a real defect found and fixed; production gates published)#
- Production gates on every homepage: the fifteen gates a multi-region production claim needs (failure domains, cross-region consensus, regional failure, partition, Byzantine node, convergence, evidence continuity, recovery, clean-room verification, benchmark, invariants, supply chain, rollback, disaster recovery, independent reproduction). Each gate shows its status and links to its evidence. 5 of 15 pass today. A live line under the gates shows the cluster's state as your browser measures it now.
- Defect found on the live cluster (fixed in
9fedb49): a load test ran while the Atlanta host was swapping. The Atlanta replica lost primacy while waiting, signed a proposal anyway, then accused the honest replicas of equivocation using a proof whose two halves had different signers. It quarantined them and stopped catching up. The honest replicas rejected the bogus proofs, so safety held. Four fixes: primacy is re-checked under the signing lock; equivocation needs both statements from the accused; a deviation needs a real leader proposal; stored proofs are re-verified at startup. Tests reproduce the incident. Rolled out the same day as a live rolling upgrade: one replica at a time with the primary last, each caught up in 10–14 s. The Atlanta replica re-verified its stored proofs at startup, released its peers and caught up. All four replicas agree at the new height. - Operations tooling:
/api/v1/live-cluster/health(HEALTHY or DEGRADED returns 200, HALTED returns 503, for uptime monitors); a rolling-upgrade tool (one replica at a time, primary last, rollback containers kept); online backups (integrity-checked snapshots of every replica); and a drill harness for load, rolling restart and disk loss. - Capacity: 21 stale containers stopped on the Atlanta host with the owner's approval, which ended its swap thrash.
[42.18.0] — 2026-09-26 (CAIN-42 PBFT cluster live in three regions)#
- Live multi-region cluster
cain-mr-01: 4 PBFT replicas on 3 hosts in 3 Vultr regions (Atlanta, Los Angeles ×2, Miami), n=4, f=1, quorum 3. The replicas connect over a WireGuard mesh and are never exposed publicly. Each replica's Ed25519 key was generated on its own host; members are configured, with no trust-on-first-use. All replicas run the same image (sha256:95d0ddfb…), built from commite1cf69c. Measured RTT: Atlanta–Miami 15 ms, Atlanta–Los Angeles 51–61 ms, Los Angeles–Miami 65 ms. - Live cluster page, on all three sites: live status of all four replicas. Your browser fetches any decision from every replica and verifies each Ed25519 vote itself. It can also audit the history continuously, walk the whole decision chain from genesis, and run a tamper lab (edit a real certificate and watch it get rejected). The page is read-only.
- Fault injection on the live cluster, with real replicas stopped on their own hosts:
- Miami down: 6/6 committed.
- One Los Angeles replica down: 6/6 committed.
- The whole Los Angeles host down (2 of 4): 0/3 committed; it refused, as required.
- Atlanta down, including the primary: view change, then 6/6 committed with no Atlanta signature.
- Full recovery: all four replicas at sequence 42, same state.
- Evidence: 336 certificates. The standalone verifiers pass 49/49 checks. The signatures confirm the schedule: no decision taken during an outage carries a signature from a stopped replica, and no refused request ever entered any decision chain. Verify in your browser → · REPRODUCE.txt
- Measured commit latency (client to committed ALLOW, sequential): baseline p50 502 ms, p95 797 ms, max 1480 ms (n=20). The Atlanta host is a 2 vCPU / 3.4 GB VM under memory pressure, which dominates the tail.
- Engine fix
e1cf69c: the 72-hour soak found 2 replicas in view 39 and 2 in view 40, with zero commits for 25 hours. Two causes: there was no progress timer, and a NEW_VIEW sent as a reply was dropped. Both are fixed. The regression test fails 4/5 on the old engine and passes 5/5 now; the full BFT suite passes 455/455. - Deployment finding: in the first layout, two replicas on one host could not reach each other through the host's published ports (Docker DNAT plus ufw). The replicas now use host networking bound to the overlay only.
- Scope, unchanged: one operator and one provider run all three hosts. Placement is stated by the operator. The hosted decision pipeline does not route decisions through this cluster yet. There is no hardware attestation and no third-party review. The source code is not published.
[42.17.0] — 2026-09-26 (Public truth layer: three homepages rebuilt around the current CAIN-42 state)#
- New homepages on cainstudio.online, mcpgate.online and clawx.click, built from one template (
scripts/cain42_site/build_pages.py), sharing one stylesheet, one script and the same favicon. Every status badge is fetched by the browser from its source (/fabric/status, the signed node state proof, the soak's latest checkpoint, the signed claims registry). A source that does not answer shows as NOT LIVE. - /now.html and /now.json: current commit, gateway start time, the enforcement flags of the running process, live pipeline mode, cluster size and quorum, conformance summary, claims by status and known limitations.
scripts/cain42_site/build_now.pygenerates it; nothing is typed by hand. - Deployment mode stated plainly: the hosted pipeline runs in shadow mode (
FABRIC_ENFORCEoff). Identity, entitlement, the kill switch and execution holds enforce; policy, risk and verification are recorded. The PBFT authorization layer is verified on disposable clusters and is not in the live path. - Disclosed on mcpgate.online: the self-hosted proxy's default policy allows by default, and a failed quarantine lookup is skipped. This is an open defect.
- Removed: the ticker items "architectural monopoly", "$1B valuation roadmap", "32 production features" and "court-admissible WORM" from the homepages, proof, sandbox and mcpgate-proof pages, and an "ISO/IEC 42001 certification" line. No such certification exists.
- Previous homepages kept as history under
/proof/bundle/history/homepages-2026-09-25/and/evidence/history/homepages-2026-09-25/, marked superseded and noindex. - Operational fix: the gateway watchdog killed every fresh gateway process before it finished starting under heavy load, which kept all three sites at 502. It now grants a 300 s startup grace. clawx.click now serves its PNG favicons.
[42.16.0] — 2026-09-25 (72-hour adversarial soak started, with live signed hourly checkpoints)#
A dedicated 4-node CAIN-42 cluster (fast path and DAG on) is crash-restarted every 10 minutes for 72 hours under continuous PBFT and DAG traffic. Every hour it publishes a checkpoint, hash-chained and signed with the evidence-root key after the privacy firewall clears it. Each checkpoint holds:
- quorum certificates taken from the running nodes;
- a cross-node agreement check;
- a real MCPGate-enforced authorization and its refused replay;
- cumulative counters.
Checkpoints and verifier. verify_soak.py shows VERDICT: RUNNING until the run completes, then PASS or FAIL. The verdict is pending, and the claims registry says so. It is a dedicated cluster on one host, not the production cluster.
[42.15.0] — 2026-09-24 (Final evolution: signed public claims registry and the Verification Center)#
Status: self-attested.
- Verification Center: one signed registry of all 16 public claims. Each has a status (VERIFIED, TESTED, BENCHMARKED, SIMULATED, UNVERIFIED or NOT_IMPLEMENTED), an evidence level from 0 to 6, SHA-256 digests of its artifacts, and its limits. Your browser checks the evidence-root Ed25519 signature and re-hashes every artifact (VALID or TAMPERED); you can also drop any file to check it.
verify_claims.pydoes the same from the command line and compares the key across all three sites. - Executable invariants: 28 of 28 pass on the real code, 22 with full coverage and 6 partial, each with its gap named. Examples: no quorum means no consensus, no authorization and no execution; agents, memory, the DAG, delegation, trust scores and AI predictions cannot create authority.
- Long-horizon composition and budget firewall: individually authorized steps cannot exceed an agent's aggregate policy (salami attacks), and action budgets are enforced outside the model.
- Public-evidence privacy firewall: keys, tokens, credentials, private IPs, internal URLs, server paths and source code block publication. All published bundles pass.
- Still NOT_IMPLEMENTED (and listed as such in the registry): independent failure domains, hardware attestation, formal verification, a 72-hour soak, third-party review or certification.
[42.14.0] — 2026-09-24 (Evolution 6: agents propose, consensus decides)#
Status: PARTIALLY VERIFIED, SIMULATED. Agents are scripted, not LLMs; attestation is simulated; the run is in-process.
- Agentic gate at MCPGate. An agent's tool call runs only with a PBFT-committed authorization *and* all of these:
- a valid identity certificate from the pinned authority;
- an intact agent-signed trajectory that declares the call;
- the same plan and model version as at consensus;
- the same pinned tool definition;
- a valid, non-escalating delegation when another agent acts.
Agreement among agents or reviewers counts for nothing, and memory can never authorize.
- Found and fixed:
- Memory: re-submitting a trusted memory's id with a bad signature overwrote it.
- A rug-pulled tool still presented its original identity.
- Trajectory forgery was classified as context drift, so the agent was never quarantined.
- 1,000 trajectories: 1,000 of 1,000 ended as expected, 380 attacks were blocked, and 620 of 620 allowed actions have complete verified proof chains. In a 1,000-step trajectory, 0 stale authorizations were accepted after plan changes. MCPGate agentic check p50 9.7 ms.
- CAIN42-AGENT-PROOF-PACKAGE: verified by one command.
- Not implemented: LLM agents, budget certificates, trust decay, a 10,000-action stress test, a 72-hour soak, the gate in the live MCPGate.
[42.13.0] — 2026-09-24 (Evolution 5: MCPGate enforces consensus; proof-carrying authorization)#
Status: Disposable cluster on one host. The live MCPGate does not run the consensus gate yet.
- MCPGate enforces consensus. Every earlier report listed this as not implemented. A tool call now passes MCPGate only with a consensus authorization that satisfies all of these:
- its quorum certificate verifies;
- it is time-bound;
- the call is exactly the authorized action;
- the call is inside the authorized scope;
- the caller is the authorized identity;
- the security context at execution time equals the one bound into consensus;
- it has not been used before.
Every decision is a signed, hash-chained EnforcementProof or DenialProof ("why not?"), and every execution gets an ExecutionProof.
- CAIN42-PROOF-PACKAGE: 4 authorized calls executed and verified; 9 attacks blocked, each with the right reason code. One command,
cain_proof_verify.py, gives VERIFIED on all 8 sections, and INVALID or INCOMPLETE when any byte is changed. - Defect found and fixed: ordering bias. The Evolution 4 DAG order put node 1 first in 71% of rounds and node 4 last in every round (chi-square 542). The new tie-break is derived from the committed PBFT history: 23–27% first-position rates, chi-square 1.75.
- Verifier gap closed: the standalone verifier now re-derives scope, expiry and the security-context hash from the committed request.
- Refinement: in 160 view-change decisions from randomized real-engine runs, the engine's choice equals the reference model's every time. In 9 of them, only the vote-history rule kept the committed value.
- Crash consistency: each of the 4 nodes crashed at 7 points, fast path off and on: 56 of 56 safe, and the restarted node caught up.
- Not implemented: key rotation, protocol-version certificates, epoch transitions, a 72-hour soak, and the gate in the live MCPGate.
[42.12.0] — 2026-09-24 (PBFT Evolution 3 fast path; Evolution 4 DAG layer)#
Status: Disposable clusters on one host. Neither feature is enabled on the live cluster.
- Evolution 3 fast path (off by default). A replica commits without the COMMIT round only with all four members' votes, never fewer. It is safe because every view change reports each member's vote history. An exhaustive model check (N=4, f=1) found the naive fast path and a "latest vote" variant unsafe, with counterexamples; the implemented vote-history rule holds. Fast-path certificates: 152 certificates from a real run, 41 fast, verified 30/30.
- Defects found and fixed while building this:
- One member could grow a replica's memory without bound by naming future view numbers.
/pbft/messageparsed request bodies of any size before any check.- Fast commits received by state transfer were served as invalid certificates.
- A restarted honest node was recorded as a DAG equivocator.
- DAG anchors were ordered out of sequence after a restart.
- DAG history gaps from downtime were never filled.
- Measured honestly: in paired A/B runs on this host, neither the fast path nor the removal of a 50 ms poll produced a measurable latency difference (95% intervals include zero). Latency here is CPU-bound.
- Evolution 4 DAG layer: signed vertices, availability certificates (3 of 4 attestations), causal proofs, and a deterministic order. PBFT stays the only authority: replicas refuse to vote for an anchor without a valid availability certificate. DAG evidence: 40 requests ordered identically on all four nodes, including a crash-restarted one; 18/18 checks, and the verifier recomputes the order itself.
- Not implemented: request batching, DAG checkpoints and pruning, ordering-fairness measurement, a 24-hour soak, MCPGate consuming the AuthorizationCertificate.
[42.11.0] — 2026-09-24 (PBFT Evolution 2: quorum certificates you can verify in your browser)#
Status: Self-attested by one operator; the evidence run used one host.
Verify the certificates in your browser → · REPRODUCE.txt · AI_VERIFY.json
- Engine (commit
364f1bf). Up to 4 proposals can be in flight, and sequence numbers are never skipped. A pacemaker adapts timeouts but can never authorize anything. AnAuthorizationCertificateis valid only with a valid COMMIT quorum certificate and a request whose intent, proposal and action hashes match. Leader-health, censorship and fairness telemetry are diagnostics only.pbft_*Prometheus gauges were added. - Fixed. The post-commit authorization proof hardcoded
trust_state=HIGH,risk_state=LOW,risk_score=0.1and a fixed trust vector, none of which consensus ever evaluates. It now reportsNOT_EVALUATED. - Tests. 135/135 PBFT Evolution 1+2, authorization, progress, tool and omega; 167/167 wider Byzantine and cluster suites.
- Public evidence. A disposable 4-node cluster was built from
c188442. It committed 12 decisions, its primary was killed, the survivors changed view (0 to 1) and committed more, and the old primary restarted and caught up: 20 decisions in total. All 160 COMMIT/PREPARE quorum certificates are published with the original Ed25519-signed votes, plus the view-change certificate and 4 AuthorizationCertificates. The standalone verifier (no CAIN code) and the in-browser page each return 29/29 checks. Those include 7 tamper controls that must be rejected: a forged signature, below quorum, COMMIT votes presented as PREPARE, a different digest, cross-cluster replay, cross-view replay, and a vote altered after signing. The bundle is byte-identical on all three sites (sha256f96088a8edf710b3…). - Not claimed. Independent failure domains. Public verifiability of the live cluster. Later on 2026-09-24 the live
cain-vccluster was upgraded to this build, node by node (rehearsed first on a disposable copy, including rollback). State was preserved on all 4 nodes, a smoke write committed, and rollback copies were kept. Its first new decision's certificates and AuthorizationCertificate verify on all 4 nodes. But its first two decisions predate certificates, and its API is not publicly reachable.cain-clusterandcain-sbxstill run older images. That MCPGate consumes the AuthorizationCertificate (it does not yet). Third-party review.
[42.10.4] — 2026-09-22 (Last-mile enforcement wired into a live route; infrastructure hardening)#
Status: Not production in the business sense — no customer traffic, same-operator infrastructure, no third-party review.
cain_agi_control_boundary.ControlBoundary.submit_proposal() — the one pipeline in this repo that calls MCPGateLastMileEnforcer (in-flight parameter-mutation defense: the dispatched action is re-checked against its own signed Proof-Carrying Decision immediately before execution) — was fully built and covered by tests/test_agi_*.py, but reachable only from pytest: no live route ever constructed a ControlBoundary or called submit_proposal() (grepped main.py/routers/*.py/cain_private_api.py; none existed). Now live at POST /fabric/agi/propose (auth-gated), plus GET /fabric/agi/sandbox-tools and GET /fabric/agi/evidence-chain. Execution is scoped to a small registered sandbox tool set (sandbox.read_file/write_file/send_notification under a per-tenant directory) — this does not open a general tool-dispatch surface. 5 new endpoint-level tests; all 52 pre-existing test_agi_*.py tests still pass.
Also shipped: cain_agent_trust_passport.py (composes existing identity attestation, evidence-backed trust state, and capability-delegation-chain verification into one signed artifact; fails closed on a quarantined issuer or unverifiable chain rather than fabricating a trust score or scope).
Also fixed: the BFT certification harness (scripts/cain42_bft/maximum_assurance.py) hardcoded "state_transfer_implemented": false in every rejoin-stage result regardless of actual outcome. The protocol is real (cain_pbft_engine_33.py, cluster_api.py's /pbft/state-transfer + /pbft/catch-up); the field now reflects a live, side-effect-free probe instead of a constant, and the stage now explicitly retries the same production catch-up endpoint before giving up. Verified live on a fresh disposable cluster: rejoin converged via the existing automatic mechanism alone. The full 12-stage certification pipeline's rejoin stage still reports FAIL; that discrepancy is unresolved and is most likely certificate/sequence state accumulated by earlier stages (view-change, partition, rollback) interacting with the state-transfer protocol's strict no-gap validation, not a simple startup timing race.
Also removed (twice — written by two different concurrent sessions, on two different days): a draft evidence page under clawx-site/ claiming "108/108 tests, 4/4 nodes OPERATIONAL" for modules (cain_mcpgate.py, cain_actionproof.py, cain_trust_state_engine.py) that do not exist anywhere in this repository. Never linked from anywhere, never deployed — caught before publication both times.
Infrastructure status (verified, distinct from the claims above): DNS for all three domains resolves to this host; valid auto-renewing TLS; the gateway runs under systemd with Restart=always and has been stable for the observed uptime. This is a narrower claim than "production" in the business sense — see AI_VERIFY.json's infrastructure_status field for exactly what is and is not claimed. Still open: multi-node BFT rejoin certification has not passed the full pipeline, and most of the larger "autonomous agency" roadmap this codebase is periodically asked to build (economic authority engine, agency graph, agent population governor, and similar) remains intentionally unbuilt — a single session's scope was last-mile enforcement wiring and honest status reporting, not the full spec.
[42.10.3] — 2026-09-21 (Verification Lab: in-browser verification and a live sandbox)#
Status: Same-author evidence, one host, no third-party review.
New page /lab.html on cainstudio.online, mcpgate.online and clawx.click (cain42-lab). It (1) fetches the 4-node PBFT fault-test bundle, checks its SHA-256 against AI_VERIFY.json, and verifies every Ed25519 state proof and quorum-certificate signature in the visitor's browser with WebCrypto (a port of verify_cluster_evidence.py; tested to give 33/33 on the published bundle, the same as the Python verifier, and to reject a one-bit signature change); (2) fetches a live node's freshly sealed signed state proof and verifies it in the browser; (3) drives the keyless /fabric/try decision sandbox (fixed scenarios, throwaway tenant, 20 per hour per address).
Known defect shown on the page, not hidden: the policy (OPA) and risk (fuzzer) stages report unavailable in this deployment because their backing services are not running here, so scenarios that advertise a denial by those stages (for example prompt-injection) return REQUIRE_APPROVAL instead. The lab prints the advertised expectation next to the observed verdict. Not fixed yet.
Not proven by anything here: independent failure domains, partition behaviour, security, third-party review. The public /api/v1/cluster/status field byzantine_f1_readiness reads "PROVEN" but is a membership-count topology check only (its own basis field says so); it is not a fault-tolerance result.
[42.10.3] — 2026-09-21 (Code secrecy enforced on all three sites; public lab built, awaiting gateway restart)#
Status:
Live now: implementation source withdrawn from public serving#
Earlier on 2026-09-21 several evidence bundles published implementation source as downloadable files. That contradicted the owner's requirement that the code stay secret. All of it was removed from every served directory and returns 404 on cainstudio.online, mcpgate.online and clawx.click. Bundles now publish only SHA-256 *commitments* to the code that produced them, recorded results, and small standalone verifiers that import nothing from CAIN. A guard (scripts/check_public_ip_exposure.py and a test) fails if implementation source reappears in a served directory, and the builders that a daily job runs were changed so they cannot republish it. Consequences you should know: the bundles can no longer be *re-run* from public material (they can still be verified: signatures, hashes, recorded results); and copies fetched while the files were public cannot be recalled. A scan of the served directories found no private keys, environment files, databases or live credentials.
Built and tested, NOT live until the gateway is restarted: the public lab#
/lab on all three sites lets anyone exercise a dedicated 4-node PBFT sandbox twin (own keys and state, separate from the test cluster) and verify every signature in their own browser (WebCrypto Ed25519). They can submit authorization requests, crash up to two nodes, and watch a quorum certificate verify or fail; they can also verify the live test cluster's four signed state proofs, and tamper with the recorded fault test to see verification fail. The lab API accepts only validated fixed-shape input, runs docker stop|start with fixed arguments on the four sandbox containers only, caps simultaneous crashes at two, auto-restarts nodes left down, is rate-limited per client and globally, has a kill switch, and returns only whitelisted fields. 19 API tests and 3 browser-logic tests (the page's own JavaScript is run under node against the real signed data and against tampering, and its canonical JSON is compared byte for byte with the Python implementation; that comparison found and fixed a mismatch on the DEL character). Until the running gateway is restarted, /lab, /api/v1/lab/* and the homepage verify strip do not exist on cainstudio.online and mcpgate.online. clawx.click already serves the static lab page (/lab/index.html); its live sections report errors until the gateway restarts.
Not proven#
Everything here is same-operator evidence on one host; the sandbox demonstrates behaviour, not independent failure domains, partitions or a malicious validator; the deployed multi-host cluster remains NOT established.
[42.10.2] — 2026-09-21 (Cluster evidence: fault-injection on a disposable PBFT twin, a superseded unsupported claim, and an AI entry point)#
Status: Same-author evidence, one host, no third-party review. The deployed multi-host cluster is NOT established as Byzantine tolerant.
Start here: /proof/bundle/AI_VERIFY.json (on cainstudio.online and mcpgate.online; /evidence/AI_VERIFY.json on clawx.click). It lists each verification recipe with URLs on all three sites, the expected result, and what it does and does not prove.
What was tested and what happened (run 2, /proof/bundle/cluster-fault-test-2026-09-21-run2/)#
A disposable 4-node twin of the test cluster (same image, own network and keys; the live cluster was never touched), PBFT n=4, f=1, quorum 3:
- Commit with all 4 nodes: 4 signers. Crash 1 node, commit again: 3 signers verified independently, quorum met.
- Crash 2 nodes (beyond f=1): the request did not commit (
CONSENSUS_TIMEOUT, no quorum certificate), so safety held. The primary's sequence counter advanced from 2 to 3 without a commit; the state root did not change. - Unfavourable finding: a restarted node reported reachable but did not catch up on its own within 121 s (sequence 1 while the others were at 2). It converged only when the next request committed (all four at sequence 4, one state root). Recovery is therefore not passive.
- The standalone verifier (
verify_cluster_evidence.py, imports nothing from CAIN) re-checks every Ed25519 state proof, every quorum-certificate vote, that QC membership keys equal the keys the nodes report, and cross-node agreement: 33/33. Fifteen tamper tests show it rejects altered fields, forged and duplicate votes, sub-quorum certificates, and a certificate whose membership was swapped to attacker keys (caught only when keys are pinned to what the nodes report). - Corrected an earlier published note for run 1 that said "convergence observed"; in run 1 the restarted node had not converged when sampled.
A published claim was unsupported, and is now marked superseded#
cain_cluster_4node_bft_evidence.json asserted OPERATIONAL_AND_VERIFIED, quorum 3 and eight invariants ALL_VERIFIED. A read-only audit of the four endpoints it names (legacy-cluster-claim-audit.json, repeatable with the same GETs) found: three of four nodes report quorum 2, node2 reports 3, none exposes a PBFT endpoint, and none of the eight invariants carries any attached evidence. The file now says SUPERSEDED_CLAIMS_NOT_SUPPORTED; the original claims are kept inside it, labelled unverified, and the index hashes were updated. A separate public prober (/proof/bundle/byzantine-cluster-2026-09-21/) reaches the same NOT_ESTABLISHED verdict for the deployed cluster.
Not proven#
Independent failure domains (one host, one image, one Docker daemon for the twin), network partitions, a malicious equivocating validator on the live wire, long-duration behaviour, fault injection on the live cluster, and any third-party review. Bringing the remote nodes to the same build as the gateway node (quorum 3, signed state proofs, one version) is the step that would change the deployed-cluster verdict; it has not been done. The other legacy files under /proof/bundle/ have not been audited.
[42.11.0] — 2026-09-21 (CAIN-42 Frontier: trust primitives wired into the gateway; evidence bundle other AIs can validate)#
Status: PASS WITH LIMITATIONS. Release gate: NO_GO. Not production. Not a Byzantine cluster result. Enforcement is in SHADOW mode: the new gate records what it would block and blocks nothing in production today. This is self-generated evidence from one operator on one host; it has had no third-party review.
Verify it yourself (stdlib + cryptography, imports nothing from CAIN)#
curl -s https://clawx.click/evidence/frontier/verify_frontier_bundle.py.txt > verify.py && python3 verify.py
It fetches the bundle from cainstudio.online, mcpgate.online and clawx.click, checks every file has the same SHA-256 on all three, checks the published source against its manifest, verifies the transparency checkpoint signature and the RFC 6962 inclusion proof of the decision-log head, recomputes the AgentBench summary from its rows, and recomputes the release-gate decision from its own evidence. Pass --pin-key to pin the signer key yourself. Bundle index: /evidence/frontier/manifest.json · claims and what is NOT claimed: /evidence/frontier/claims.json · guide: /evidence/frontier/VALIDATION_GUIDE.txt · source: /evidence/frontier/source/manifest.json.
What shipped#
- A frontier gate (
cain/frontier): Ed25519 principal chain with attenuation-only, invocation-bound capabilities and optional proof-of-possession; world model that predicts and can deny but never grants authority; session-aware and trajectory containment; hard-ceiling budgets; memory-influence containment; structured errors so TIMEOUT / UNKNOWN / PARTIAL / UNVERIFIED are never success; hash-chained decision log anchored into a signed RFC 6962 transparency log. - Staged enforcement (operator-only): enforce per tool/agent/tenant with a deterministic canary percentage and per-category blocking, simulate a candidate rule against the real shadow log first, hot-reload with last-known-good on a bad edit, one-command panic and rollback. Production policy is currently empty (shadow everywhere). Reading the real shadow log already exposed a false positive (read-only tools such as
db_readandlookup_weatherwere classed as unknown), fixed with regression tests before any enforcement. - Real multi-process Byzantine experiments with raw signed messages (4 OS processes, own keys, one host, test harness): honest, wrong-commitment node, equivocating node, forged and relabeled votes, one crashed node, two crashed nodes. An independent verifier (
verify_bft_evidence.py.txt) re-derives signatures, quorum backing, safety and the Byzantine proofs from the exported messages, and tests show it rejects tampering. The experiment found a real liveness bug (one crashed node stalled the survivors); the failing run is preserved, the service is fixed, and the passing run is published. This does not establish f=1 for the live cluster, whose recorded probe verdict remains NOT_ESTABLISHED. - Measured on the real system: AgentBench 29 scenarios, all passed on task success and security-correctness, and mutation-tested (breaking a layer makes it fail); full suite 4291 passed, 0 failed, 42 skipped on the operator's host.
- Negative findings, published on purpose: the four cluster validators are not independent (epistemic independence 0.25, minimum collusion set 1: one image, one host); the release decision is NO_GO (independent verifier INCOMPLETE on day one, public-claims mapping unmeasured); the gate does not block production traffic yet.
Real defects found and fixed (each with a regression test)#
- A conformance run that executed zero tests reported CONFORMANT. Now UNKNOWN.
- Public evidence endpoints returned constants for reachability, sync and published-roots. Now measured.
- Tenant-filtered evidence exports could not be verified end to end. Now bridged with signed hash-only stubs.
- Journey-audit handlers called the gateway's own URL from inside its event loop and deadlocked it for 30 seconds until the watchdog killed the process. Fixed.
- A caller-supplied dual-custody flag was accepted by tests written before the hardening; the tests now require a real two-officer proposal.
Not tested / not implemented / not claimed#
Not tested: independent review; behaviour on more than one host; enforcement under real production traffic. Not implemented: real adapters for OpenClaw, Telegram, WhatsApp, Slack, Teams, email, browser and coding agents (only the signed-webhook adapter is complete); a learned world model. Byzantine fault tolerance across independent hosts is not claimed. Self-consistency only: the tests show the code and the verifier agree, not that the design is right. No third-party assessment exists.
[42.10.1] — 2026-09-21 (Frontier trust-engine hardening: 9 defects found and fixed, 11 attack handlers with positive controls, signed bundle)#
Status: Written and verified by the same author; no third party has reviewed it. The bundle and a standalone verifier are published live as static files on all three sites; implementation source is deliberately not published. The fixes themselves run in the gateway only after it is restarted; the /api/v1/frontier-trust/* routes exist in the repo and are not live before that.
Verify it yourself#
Needs Python 3 (and the cryptography package for the signature check). Same content on cainstudio.online, mcpgate.online and clawx.click (three profiles of one gateway process):
curl -s https://cainstudio.online/proof/bundle/v2/verify_frontier_trust_bundle.py -o v.py && python3 v.py --sites
The verifier imports nothing from CAIN. It checks that all three sites serve the same bytes, that the bundle hash and Ed25519 signature verify, that the counts equal what the per-attack list implies, that every claim is backed by tests that passed in the recorded run, and that no handler result is BLOCKED without a passing positive control. Artifacts: /proof/bundle/v2/CAIN42_FRONTIER_TRUST_ENGINE_BUNDLE.json and the verifier (on clawx.click: /evidence/..., verifier with a .txt suffix because that server serves no .py). The SHA-256 values of the implementation files are recorded in the bundle as commitments; the source itself is not published (it is proprietary), so a third party can check the recorded results, signature and consistency but cannot rebuild them without access. All three sites resolve to one host, so identical bytes across them show consistency, not independence.
Defects found and fixed (each has a regression test named in the bundle)#
- Trust computation could not run on the production schema.
compute_trust_deterministicselected a column no schema defines, joined a table from another database and read two tables nothing creates. Every call raised on a read-only copy of the production database observed in-session (not reproducible from the bundle). The adversarial engine's trust attacks had been reporting "blocked" partly because the verifier reads an exception as manipulation. Fixed; missing negative-evidence sources are now disclosed and cap the state below TRUSTED. - Stale trust cache. A cached state kept serving after 12 new violations. It is now recomputed when newer decisions exist.
trust_versionbumped on every write and a recompute returned a placeholder. It now bumps only on material change and returns the stored value.- Snapshot replay/forgery. There was no way to check a presented trust snapshot.
verify_trust_snapshotrejects a snapshot that is malformed, not derivable from evidence, stale or version-mismatched. - Attack routing. Three trust attack types were mapped to the identity runner and stayed INCONCLUSIVE whatever the handler did.
- Honesty-guard masking.
execute_attackdowngrades any ATTACK_SUCCEEDED withoutsimulated: true; the new handlers did not set it, so a real escape would have been reported as INCONCLUSIVE. Fixed and tested throughexecute_attack. - UNKNOWN became ALLOW. In the predictive recommendation a caller-supplied "trusted" + "low" overwrote the critical-unknowns denial. Critical unknowns are now a floor.
- cainbench accepted an uncompilable regex checker at registration and would then fail the run with a 500.
- Adversarial worker counted unsupported attacks as neither blocked nor escaped and could report RESILIENT; inconclusive attacks now yield INCOMPLETE.
What the numbers are#
On a fresh SQLite database whose schema the modules' own init functions create, the engine's 114 enumerated attack types were all BLOCKED, including the 11 handlers added here (each recorded with its positive control). That is not a security score: it reports which implemented attacks were blocked, and several older verifiers are weak (for example _verify_tenant_isolation returns valid on an exception). On an empty database with no schema the same 11 handlers return INCONCLUSIVE, by design.
Known limitations and open findings (also in the bundle)#
evidence_store.retrieve(id)takes no tenant argument; the cross-tenant evidence handler covers only the tenant-scoped fabric decision accessors.trust_graph.create_edgenever persistsvalid_until, so binding expiry cannot be created through the API.predictive_trust_decisionis defined twice intrust_graph.py.- 16 attack types are listed under one category but mapped to another; not audited.
- A wider regression run over every test file touching these modules was not re-run after the last edits; one older test file failed once and passed on four re-runs, cause unknown.
- Single host, single process. The signing key is held by the same author: the signature proves integrity since signing, not independent review. AGENTS.md figures (398/398, 120/120, 200 invariants) were not re-verified here.
[42.11.3] — 2026-09-21 (Verify-it-yourself on all three homepages; code stays private; an overstated number corrected)#
- Homepages: cainstudio.online, mcpgate.online and clawx.click now open with a "Verify it yourself" section linking the live-probe cluster bundle, the hardening bundle, the fault-test bundle, the frontier bundle and daily evidence, with a copy-paste command. Every link in it was crawled and resolves on each domain.
- Correction: the engine counted 19 of 114 attacks as blocked when their verifier had crashed or had no target. Errored attacks are now inconclusive; the published figure is 97 blocked, 17 inconclusive, 0 succeeded (see the note in 42.11.1).
- Code stays private: the public bundles publish outcomes, not implementation. Source files, file paths, module and class names, and raw error text were removed from them; an earlier bundle that included one source file was replaced. The checkers that remain import nothing from CAIN. Infrastructure paths (
.env,.git, keys,main.py, traversal attempts) were probed on all three domains and none is served. - Fixed link: the fault-test bundle's reproduce steps cite cainstudio.online, where it was not served; it is now mirrored (identical bytes) at
/proof/bundle/cluster-fault-test-2026-09-21-run2/. - Not yet live: the interactive public lab's API (
/api/v1/lab/*) is built but the running gateway has not been restarted onto it, so it is not linked from the homepages until it answers.
[42.11.2] — 2026-09-21 (Byzantine cluster: an independent prober, and an honest NOT_ESTABLISHED verdict)#
Status: Self-attested. We published a checker that any AI can run to probe the live cluster with no CAIN code. Its recorded verdict on 2026-09-21 is NOT_ESTABLISHED, and that is the point: it is derived from what the nodes return, not asserted.
curl -sO https://cainstudio.online/proof/bundle/byzantine-cluster-2026-09-21/verify_cluster_bundle.py.txt && mv verify_cluster_bundle.py.txt verify_cluster_bundle.py python3 verify_cluster_bundle.py https://cainstudio.online/proof/bundle/byzantine-cluster-2026-09-21/ --live
Bundle root 93c8a5f8fe0b3d5ae91e53581162bbf7…, served byte-identically from cainstudio.online, mcpgate.online and clawx.click.
- Established: four nodes (node2 plus three remote hosts) answer and agree on membership; node2's signed PBFT state proof verifies independently; the consensus logic passes 120 tests in 10 files (PBFT cluster, proof-carrying quorum certificates, independent verifier, multi-process BFT evidence, Byzantine swarm); the prober is itself tested against a fake cluster (healthy gives BFT_F1_ESTABLISHED, each injected fault is caught).
- Not established: the three remote nodes report a quorum of 2 where a Byzantine quorum for N=4 is 3 (a quorum of 2 lets two conflicting decisions both commit); only one node serves a signed state proof, so cross-node state agreement cannot be checked from outside; the other nodes report no software version and run a smaller, older API.
- Corrected claims:
llms.txtsaid "f = 1 proven resilience"; it now says designed-for, not established. Thebyzantine_f1_readiness: PROVENfield is computed from a membership count (N>=4, four trusted members), not from a fault-tolerance test, and the prober flags it as unsupported. The static 2026-09-16cain_cluster_4node_bft_evidence.jsonis a hand-authored snapshot, not a run output. Four unsourced files (competitive matrix, agent registry, exposure graph, trust-BOM) were withdrawn from the public CLAWX evidence. - Not done: no fault was injected into the live cluster; the remote nodes must be brought to the same build (quorum 3, state-proof endpoint, version reported) before the verdict can change. That needs access to those hosts.
[42.11.1] — 2026-09-21 (Hardening round: 16 defects found by attacking our own controls, fixed, and independently checkable)#
Status: Self-attested, single host. This round attacked our own verifiers instead of trusting them. Rebuilding the attack handlers on the real production code (not test-local stand-ins) is what exposed several of the bugs below.
Verify it yourself#
curl -sO https://cainstudio.online/proof/bundle/hardening-2026-09-21/verify_bundle.py.txt && mv verify_bundle.py.txt verify_bundle.py python3 verify_bundle.py https://cainstudio.online/proof/bundle/hardening-2026-09-21/
The same bundle (bundle root 26449783615937cc99931d9fd568db6a…) is served byte-identically from cainstudio.online, mcpgate.online and clawx.click. The checker imports nothing from CAIN; it verifies every file's SHA-256, the Ed25519 signature and that each number in the manifest can be recomputed from the bundle's own files. It does not re-run tests: the commands are in REPRODUCE.txt.
Found and fixed (each has a regression test, listed in defect-ledger.json)#
- Security-context verifier failed open: only 6 named checks could deny, so a context for read/report was allowed for delete/payroll_db under a different intent. Every failed check now denies.
audiencewas stored but never verified; now enforced. - MCP proxy: a principal could self-sign a capability for any tool it was never granted. Now checked against the identity's granted capabilities. New tool-poisoning / rug-pull guard for
tools/list(pins definitions, quarantines changes). - Kernel: a tampered call burned the token's nonce so the legitimate call was rejected as a replay; trust could be rebuilt from 0.20 to 0.91 in 100 s by repeating positive events (now time-paced).
- Evaluation fabric: the contamination scan never awaited its HTTP calls (the model was never queried; every scan passed), the report crashed, contaminated agents could pass, and the endpoint fetched caller-supplied URLs (SSRF).
- Honesty fixes: a benchmark reported a hard-coded "1-minute sustained load: completed, 0% errors"; the adversarial worker reported RESILIENT while attack types had no handler; the free-signup outage fallback implied a working key (now
registered:false). - A 26-route security-context API was dead code (import error, never mounted). Import fixed; deliberately not mounted pending an authentication review.
Evidence (all in the bundle)#
- 114 adversarial attack types exercised: 97 blocked by a real control, 17 inconclusive (no real control to attack yet), 0 succeeded, fresh run against a temporary database; each handler has an honest-path control and mutation tests showing it reports success when its defence is removed.
- 222/222 executable formal invariants; 585 tests passed, 0 failed in a serial, isolated run (
test-run.txt). - Measured performance on one host: p50 0.41 ms, p99 1.23 ms; a genuine 60.002 s sustained run of 106,650 iterations with 0 errors; 27/27 Byzantine vectors fail closed.
- Correction (same day): an earlier version of this entry said 114 of 114 attacks were blocked. 19 of those had been blocked only because their verifier crashed (a missing table or module, a dict-iteration bug) or had no target, and a crash is not a defence. The engine now reports such attacks as inconclusive; the two tool/MCP substitution attacks now run against a real control. The figures above are the corrected ones.
What this does NOT show#
- Not an independent audit: the same host runs, tests and signs it, and the three sites are one host (identical bytes, not independent trust domains).
- Attack coverage is limited to the attack types we defined; the count says nothing about ones we did not.
- Chaos was a 45 s run with simulated network faults, not a soak. No multi-host or Byzantine-cluster result.
- Not deployed by this round's author: production restarts are a separate operator step (
scripts/deploy_prod.sh).
[42.10.0] — 2026-09-20 (CAIN-42 Epoch 10 — Agentic Trust Fabric: audited, attacked, publicly verifiable)#
Status: PASS WITH LIMITATIONS. Not production. Not a Byzantine cluster result. Epoch 10 is a set of in-process Python modules (identity, invocation-bound authority, trust graph and path finder, negotiation, trajectory budgets, recovery, supply chain). The evidence below was produced by running those modules; it is published so that any reader, human or AI, can check it without trusting us.
Verify it yourself in about 30 seconds#
Run on any machine with Python 3 and the cryptography package. The same command works on all three sites because they are three profiles of one gateway process:
curl -s https://cainstudio.online/api/v1/epoch10/verify-sites.py | python3 -
It fetches the evidence from cainstudio.online, mcpgate.online and clawx.click, checks every artifact has the same SHA-256 on all three, downloads the bundle and two verifiers, runs them, then asks each site's running process to execute a fresh scenario and runs the clean-room verifier on the result. Every URL is listed in /api/v1/epoch10/manifest.json; the live self-test is at /api/v1/epoch10/selftest/run.
What was found and fixed (real defects, each with a regression test)#
- Handshake and claim challenge (audited last): the trust handshake accepted steps signed with a key supplied alongside them, let the requester issue its own authority, had no expiry and never revoked what it issued; claim challenges trusted an authority key nominated by the challenger. Fixed, 13 guard mutants killed.
- Ghost identity: an authority-signed token for an agent that was never registered was accepted (150/150 attempts). Fixed.
- Self-supplied verification keys: trust negotiation and Agent Cards verified signatures against a key the remote party supplied about itself. Now a locally held key is required.
- Execution proofs signed only two of their fields, so the resource, actor or action could be edited after the fact and the proof still verified. Now every field is signed.
- Identity revocation did not close trust-graph paths through the revoked agent. Fixed and covered by a red-team case.
- Evidence journal was not chained, so a deleted record went undetected. Now hash-chained and sequenced.
- Fail-open numeric handling (NaN or negative values disabled budgets, routing filters and blast-radius ratings), trust recovery that could be completed instantly, a research agent that stated results that had never been run, and identity records containing invented facts. All fixed; 20+ further items are listed in the commit history (
git log --oneline -- platform-gateway/*_10.py).
Evidence#
- Two evidence slices, reconciled. A second slice (trajectory, composition and common-mode attack matrices with its own clean-room verifier) lives beside this one;
status-matrix.jsongives every one of the 42 Epoch 10 items an assurance level (A0 none … A4 independent party) and lists what is not implemented (10.28, 10.34). Nothing reaches A4. - automated tests across Epoch 10 and the site registry, all passing at commit
a07a89b+. Guard mutations were applied to the fixes and to the verifier itself; a few redundant-check survivors are documented rather than hidden. - 17 machine-checked invariants (
INV-10-01..17), 17/17 passing after the fixes above. Before the fixes 3 of 17 failed. - Compromised-agent red team, 15 steps (
redteam.json): 11 blocked, 1 baseline, 2 not tested (memory poisoning and observation forgery are not covered by any Epoch 10 module), and 1 not blocked without a pinned head (removing the newest journal records is detectable only by a verifier that recorded the journal head earlier). - Decisions bundle (
bundle.json): 22 decisions (2 allowed, 20 denied), signed identities, the graph and revocations at each decision, a hash-chained journal and proofs. - Three verifiers, compared (
verification.json): A = the system's own check, B = clean-room re-derivation of every decision with zero CAIN imports, C = a ~40-line integrity check. Disagreement is reported asVERIFICATION_DISPUTE. B was also tested against bundles that are perfectly signed but semantically false (a compromised signer), and it caught them.
What this does NOT show#
- Nothing about consensus, node failure or network partitions: this is not evidence for the Byzantine cluster runtime.
- No MCP or A2A protocol-conformance suite, no benchmark, no long-duration chaos run, no third-party audit. The verifier was written by the same author as the fabric (independent implementation, not an independent party).
- Hardware (TEE) attestation is not implemented; privacy-preserving proofs are Merkle selective disclosure, not zero-knowledge.
- State is in one process's memory with ephemeral keys. The public key inside a bundle proves only self-consistency; pin a key and journal head obtained separately to prove more.
[42.1.0] — 2026-09-19 (CAIN-42 Epoch 6 — Autonomous World-State Integrity & Proof-Carrying Agency)#
The Autonomous World-State Integrity & Proof-Carrying Architecture#
CAIN-42 Epoch 6 evolves CAIN from a cognitively integrity-protected autonomous system into a self-verifying, world-state-aware, proof-carrying autonomous trust fabric. Built on the core governing doctrine: $$\text{COMPROMISED COGNITION} \ne \text{COMPROMISED AUTHORITY} \ne \text{COMPROMISED WORLD STATE}$$ $$\text{INTENT MUST NOT BECOME EFFECT WITHOUT CONTINUOUS PROOF}$$
- Observation ≠ Authority Segregation: Prevents replayed, stale, or forged observations from becoming execution authority. All sensory inputs require cryptographic provenance and quorum consensus.
- Closed-Loop Postcondition Reconciliation: Enforces that tool execution success is verified against external reality before world-state commit.
Tool return 0 is not success.Enforces the 4-stage lifecycle:ACCEPTED -> EXECUTED -> OBSERVED -> POSTCONDITION_VERIFIED. - ActionProofObject Subsystem: 24-field cryptographic action certificates binding cognition, intent, world-state versions, reversibility classification, and quorum signatures before MCPGate unblocks downstream tool sockets.
- Multi-Agent Epistemic Consensus: 6-tier epistemic state progression where independent witnesses overrule colluding Byzantine agent majorities.
- Long-Horizon Non-Escalating Governance: Formal mitigation against creeping drift and privilege accumulation across 1,000+ continuous execution steps.
- 22 Machine-Checkable Formal Invariants (
INV-E6-01toINV-E6-22): 22/22 evaluated and passed fail-closed. - 42 Adversarial Red-Team Attack Vectors Blocked: 42/42 vectors contained fail-closed with 0 physical tool executions on breach.
- Clean-Room Independent Verifier (
cain_verify_public.py): Standalone verifier with zero CAIN imports (proven via AST analysis), verifying RFC 8785 canonical JSON and RFC 6962 binary Merkle trees. Verified withFINAL VERDICT: VERIFIED. - High-Throughput Microsecond Performance: Full execution pipeline achieves 5,992.57 ops/sec (p50: 0.114 ms, p95: 0.181 ms).
- Public Evidence Package Released: 17 public evidence artifacts under
CAIN42_EPOCH6_PUBLIC_EVIDENCE/and master bundleCAIN42_EPOCH6_PUBLIC_EVIDENCE_BUNDLE.json.
[34.0.0] — 2026-09-17 (CAIN 34.0 — Production-Grade Byzantine CAIN Cluster Release)#
Live-Deployed Byzantine Fault Tolerant Cluster Runtime#
CAIN 34.0 transitions the Byzantine consensus substrate from an isolated engine module into a fully integrated, live-deployed, production-grade 4-node cluster with zero stubs, zero mocks, and zero unhandled failure modes.
- 100% Adversarial & Distributed Pass Rate: Rebuilt adversarial test suite (
tests/distributed/test_cain34_pbft_cluster.py) achieves 18/18 PASS in 4.30s; full distributed test suite achieves 108/108 PASS (100% pass rate). - 12/12 Baseline Defects Resolved: Eliminated all 12 defects identified in CAIN 33.0, including cryptographic message authentication, persistent SQLite WAL replay defense, idempotent request deduplication, view-change state preservation, cross-replica state root synchronization, and rate limiting.
- Production Docker Deployment: Released
cain-cluster-node:cain34-production-20260917(digestsha256:58ac30dbfacedf64a52932e7258528d0132bfbfe1befe023c9c5b85ab10457bf) rolled out to live containerscain-cluster-node-1..4with zero downtime ($Q=3$ quorum continuously preserved). Tested rollback safety on node 4. - Live Multi-Node State Convergence: Verified live consensus over HTTP (
/api/v1/cluster/pbft/request); all 4 independent containers converged to identical state root (f97ebffdc034de54a2c65e35b3a6629ec57c68a9f6afd3dd7781464730b7034a). - Cryptographic CLI Verification: Added
cain state prove,cain state verify, andcain state comparewith remote--endpointflags, verifying Ed25519 signatures and RFC 8785 canonical hashes against live running nodes. - Official Release Certification: Certified as
PRODUCTION_GRADEinevidence/releases/cain-34.0-production/CAIN_34_PRODUCTION_READINESS.jsonandCAIN_34_FINAL_FORENSIC_REPORT.md.
[3.0.0] — 2026-09-17 (CAIN Maximum Evolution — Phase 1, 2, 3: The $1B Enterprise Commercial & Developer Engine)#
The 32-Feature Monopoly & Dual-Channel Execution Governance#
CAIN establishes the first production execution-channel runtime for autonomous AI systems, overcoming the industry-wide Dual-Channel Control Problem. Governs actions over MCP, shell, database, cloud APIs, and financial rails through the canonical 7-Moat Trust Control System.
- 32/32 Formal Production Features Verified: Full A+++ compliance including RFC 8785 canonical action schemas, Z3 SMT formal semantic equivalence prover (
/verifygate), real enforcement proof tokens, offline court-admissible Merkle verifiers, and multi-tenant cryptographic isolation. - 100% Fail-Closed Security Doctrine: DENY, UNKNOWN, and ERROR strictly halt downstream execution with zero packets reaching unverified tools.
Phase 1: Rock-Solid Foundation & Subdomain Resilience#
- In-Process Billing Resilience: Fault-tolerant circuit breaker in
platform-gateway/routers/billing.pyguaranteeing 100% uptime (zero 502 Bad Gateway errors) during upstream payment processor degradation. - Live Mathematical Engines on Subdomains:
verifygate.mcpgate.online&/verifygate: Real Microsoft Z3 Theorem Prover verifying AST formal semantic equivalence and contract proofs.mcpsecurityscanner.mcpgate.online&/mcpsecurityscanner: Production static & semantic tool scanner flagging prompt injections, leaked credentials, and dangerous unconstrained parameters.analytics.mcpgate.online: High-throughput privacy-preserving telemetry beacon.- Smart Protocol Negotiation on
/mcp: Automated content negotiation serving interactive HTML protocol guides, SSE streams, or RFC JSON-RPC 2.0 based on client headers. - 100% Clean Link Audit: Site crawler verified 68/68 routes and links across
cainstudio.onlineandmcpgate.onlinereturn 200 OK.
Phase 2: The 10-Minute Adoption Loop (Developer Virality)#
- Universal CLI Interceptor (
cain mcp-wrap&cain proxy): Transparently wraps any downstream MCP server command with JIT capability token verification. - Automatic Desktop Client Protection (
cain guard --desktop): One-click injection into Claude Desktop (~/.config/Claude/claude_desktop_config.json) and Cursor (.cursor/mcp.json). Audited viacain guard --check(100% GUARDED). - Interactive Visual Terminal Firewall: ANSI terminal firewall rendering real-time risk alerts and blast radius bounds for high-risk tool proposals, requiring explicit operator authorization before execution.
- Downstream Result Attestation: Safe tool calls receive court-admissible
_cain_attestationcontaining Merkle evidence digests and latency metrics. - Universal Install Script: Hosted at
https://mcpgate.online/install.shfor one-line developer installation. - Zero-Dependency NPM Package (
@cain/guard): Published insdk/npm/guard/for frictionless Node.js / npx integration.
Phase 3: The Enterprise Commercial Wedge ($50k–$250k/yr)#
- Wedge 1 (MCPGate Sovereign Enterprise K8s Appliance):
- Production Helm chart
deploy/helm/mcpgate-appliance(Version 3.0.0) with local KMS root-of-trust, Traefik mTLS reverse-proxy sidecar, eBPF syscall filtering,seccompProfile: RuntimeDefault,drop: ALLLinux capabilities, read-only root FS, and fail-closed zero-trust network policies. - CLI lifecycle management:
cain appliance generate,cain appliance verify(7/7 invariants passed),cain appliance package. - Wedge 2 (Continuous EU AI Act Art. 72 WORM Notary & Discovery Bundle):
- One-click court-admissible evidence bundle export at
GET /compliance/bundle.zipand/api/v1/compliance/worm/bundle.zip. - Generates signed affidavits, Merkle inclusion proofs, JSONL ledgers, and a zero-dependency standalone offline verifier (
verify_offline.py) proving zero tampering under EU Regulation 2024/1689. - Wedge 3 (Actuarial Cyber Insurance Underwriting Protocol):
- Interactive Actuarial Portal launched at
GET /insuranceand/insurance/portal. - Dynamic Agent Volatility Index (AVI), Maximum Probable Loss (MPL), and Underwriting Credit Score (0–1000) under Lloyd's & Munich Re consortium standards.
- Real-time calculation unlocking up to 40% premium discounts (Score 944–948 AAA Premier Tier) and issuing signed Ed25519 underwriting certificates verified via
/api/v1/insurance/verify.
12 Monetization Channels Scaling to $1.13B+ Valuation#
- Full implementation of the 12 commercial revenue engines powering the 3-year financial model:
- Year 1 (2026): $4.55M ARR ($113M–$136M Series A)
- Year 2 (2027): $20.60M ARR ($412M–$515M Series B)
- Year 3 (2028): $113.10M ARR ($1.13B–$1.35B Enterprise Unicorn)
[2.4.1] — 2026-09-16 (CAIN 23.0 — Immutable Distributed Immune Consensus)#
Delegation-chain revocation cascade closed (VULN-001)#
- Revoking a delegation now transitively invalidates every delegation issued on its authority, not just the immediate delegator's identity. 10 new regression tests.
- Closed a divergence between the platform's two identity registries where a revocation applied through one store could leave the other still reporting an identity as valid.
Immune transition ledger is now tamper-evident#
- The trust-immune state-transition history is hash-chained so in-place row tampering or deletion after commit is detected, not merely disallowed by convention. 6 new regression tests.
Multi-process Byzantine consensus — real evidence, honestly scoped#
The existing BFT consensus primitive (real Ed25519 signing, real quorum math) previously ran all "nodes" as objects inside a single process, which proves the algorithm but not that independent processes can reach agreement over a real network with independently-verified signatures. This release adds that evidence:
- 4 genuinely separate OS processes, each independently generating its own Ed25519 keypair locally, communicating exclusively over real HTTP.
- A real Byzantine process (separate container, not an in-memory flag) broadcasts a genuinely altered commitment; the honest majority still converges correctly and independently identifies the Byzantine peer.
- An identity-spoofing probe — a genuinely-signed vote relabeled to claim another process's identity — is rejected by bootstrap-pinned public-key binding.
- Scope, stated precisely: this is multi-process/multi-container evidence on one shared host, not multi-independent-cloud-host evidence. It does not include the separately-hosted cluster nodes listed below — cross-host Byzantine fault tolerance across those specific machines is a tracked follow-up, not claimed here.
Full raw evidence and reproduction steps: CAIN_23_MULTIPROCESS_BFT_EVIDENCE/. Full claim-by-claim audit: CAIN_23_FINAL_FORENSIC_REPORT.md.
[2.4.2] — 2026-09-16 (CAIN 23.0 follow-up — real multi-independent-host Byzantine consensus)#
The [2.4.1] entry above proved Byzantine consensus across genuinely separate OS processes on one shared Docker host, and explicitly stated multi-independent-cloud-host evidence wasn't yet included. Same day, that gap was closed for real:
- 2 genuinely independent cloud VPS machines ran the consensus nodes, communicating over the real public internet — not a docker bridge, not localhost.
- Honest majority: all 4 nodes across both machines converge on the identical commitment. Real measured latency: ~17-42ms same-host, ~179-318ms cross-host — genuine internet round-trip time, not simulated.
- One real Byzantine process on the non-leader host: the 3 honest nodes, spanning both machines, still converge correctly and independently identify the Byzantine peer.
- Cross-host identity-spoofing probe: a vote genuinely signed on one host, relabeled to claim a node's identity on the other host, replayed over the real internet — rejected.
- A real bug was found and fixed, not hidden: the first cross-host attempt failed due to a cloud hairpin-NAT issue, causing a genuine false-positive Byzantine detection. Root-caused and fixed. Full account:
CAIN_23_MULTIPROCESS_BFT_EVIDENCE/README.md. - Still not covered: the cluster's other real hosts were not part of this test — a genuine 4-independent-host round remains a tracked follow-up. Existing live production containers on both hosts used were never stopped, restarted, or modified.
Raw evidence: CAIN_23_MULTIPROCESS_BFT_EVIDENCE/real_multihost_*.json.
[2.4.0] — 2026-09-16 (Past 96-Hour Maximum Platform Evolution)#
Unified 4-Node Byzantine Fault Tolerant (BFT) CAIN Cluster Architecture#
The entire CAIN platform has unified across all 4 cluster nodes into a single, identical CAIN Trust Runtime Kernel, providing mathematically proven Byzantine Fault Tolerance (f=1, N=4, Quorum Q=3):
- Single Unified Node Architecture: Every node in the cluster (
node1149.28.193.50,node245.76.60.231,node345.76.169.191,node4207.246.66.130, and local container meshcain-cluster-node-1..4) runs the exact same unified CAIN image, eliminating codebase drift. - 15 Distributed Cluster Endpoints: Mounted at
/api/v1/cluster/*across all nodes (/status,/health,/metrics,/nodes,/identity,/attestation,/invariants,/doctor,/health-score,/gossip,/envelope,/vote,/decision,/self-verify,/quarantine,/restore). - Byzantine Fault Tolerance f=1 Proven: 3f + 1 consensus guarantees that even if 1 node experiences network partition, Byzantine crash, or malicious compromise, the 3 remaining nodes reach quorum (Q=3) and maintain unbroken fail-closed consensus.
- Zero-Leakage Prometheus Observability:
/metricsscrubbed of all secrets/PII, emitting real-time cluster gauges (cain_up,cain_cluster_nodes_total,cain_cluster_healthy_nodes_total,cain_cluster_quorum_status,cain_cluster_byzantine_tolerant,cain_envelope_validations_total). - Multi-Agent Public Evidence Discovery: Real-time machine-readable manifests (
llms.txt,robots.txt,sitemap.xml,agent.json) and public evidence bundles published live across bothcainstudio.onlineandmcpgate.online.
[2.3.0] — 2026-09-15#
CAIN 17/18: Distributed Autonomy Constitution & JIT Capability Boundary#
- Autonomy Constitution Engine: Cryptographically hashed constitutional invariants governing autonomous agent authority boundaries and self-healing.
- Ephemeral JIT Action Capability Tokens: Sub-30s TTL single-use capability tokens verified at the MCPGate boundary with nonce-based anti-replay protection.
- 5-Layer Epistemological Fact Segregation: Explicit segregation of all evidence into
OBSERVED,VERIFIED,DERIVED,UNVERIFIED, andCOUNTERFACTUALlayers. - Transparent MCP Interception: Real-time streamable HTTP and SSE interception protecting Model Context Protocol tools from prompt injection and unauthorized state modification.
[2.2.0] — 2026-09-14#
CAIN 14.0: The Agentic Trust Intelligence Engine#
- 200 Machine-Checkable Formal Invariants: 100% verified across 16 formal invariant domains (Identity, Goal, Plan, World Model, Memory, Tools, MCP, A2A, Injection Defense, Blast Radius, Browser Control, Containment, Telemetry, Evaluation, Skills, Frontier Research).
- 120 Adversarial Red-Team Vectors: 100% blocked fail-closed across 12 attack categories (OWASP ASI01-10, OWASP AST10, tool hijacking, prompt injection, cross-tenant memory poisoning, Byzantine desync).
- Triple Verification System: Engine A (Production Evaluator), Engine B (Clean-Room Standalone Verifier with 0 CAIN imports), Engine C (Verifier of Verifiers Meta-Assurance with 8/8 corruption detection).
- Frontier AI Research Integration: Continuous loop incorporating MCP July 28 2026, A2A v1.0.0, OpenTelemetry gen_ai, AgentPRM, World Models, and Memory Sovereignty.
[2.0.0] — 2026-09-13 (48-Hour Major Platform Release)#
Trajectory Trust & The 16-Stage Dynamic Enforcement Loop#
CAIN has officially promoted Trajectory Trust from an internal invariant to a first-class, cryptographically verifiable, continuously enforced runtime primitive.
- Canonical 16-Stage Decision Pipeline: Fully wired and live across the platform runtime (
cain_canonical_pipeline.py), orchestrating the complete causal chain:
WHO → AUTHORITY → INTENT → SECURITY CONTEXT → POLICY → RISK → TRUST → TRAJECTORY → BLAST RADIUS → PREDICTION → DECISION → MCPGATE ENFORCEMENT → SYSTEM EXECUTION → EFFECT → EVIDENCE → ATTESTATION
- Unbroken Causal Hash Chaining: Every autonomous step in a trajectory cryptographically incorporates the prior step's cumulative hash, current action proposal hash, trust state snapshot, and Ed25519 signature (
cain_trajectory.py,cain_trajectory_enforcement.py). - Salami Slicing & Loop Trap Defenses: Active prevention of sub-threshold incremental attacks (
max_cumulative_delta) and circular loop traps across multi-agent handoffs. - Portable Trajectory Passports: Standardized Ed25519-signed trajectory credentials (
cain_passport.py) enabling cryptographically verified multi-agent custody transfer and inter-enterprise B2B trust settlement.
Mechanical Formal Verification via TLA+#
- Model-Checked WORM Immutability: Formal specification
CAINWormIntegrity.tlaverified via TLC model checker, mathematically guaranteeing append-only tamper evidence: no Byzantine actor or operator can mutate, delete, or rewrite historical execution evidence without breaking the cryptographic Merkle chain. - Fail-Closed State Machine Invariants: Formal specification
CAINTrustInvariants.tlaverified across all concurrent interleavings: - Non-escalation invariant: An autonomous agent cannot increase its own trust score or broaden its own authority.
- Fail-closed invariant: Under any network partition, parse failure, timeout, or ambiguity (
UNKNOWN/ERROR), the verdict strictly collapses to refusal.
BFT Multi-Validator Consensus & Confidential Computing Attestation#
- Byzantine Fault Tolerant (BFT) Quorum: Distributed multi-node consensus engine (
cain_bft_consensus.py) enforcing a 3f + 1 quorum requirement across independent validator nodes before finalizing Trajectory Passports. - Hardware Remote Attestation Notary: Hardware-rooted remote attestation engine (
cain_enclave_attestation.py) verifying execution inside confidential computing enclaves: - Intel SGX Quote verification and enclave signature validation.
- AMD SEV-SNP attestation report verification with firmware-signed measurement hashes.
- AWS Nitro Enclave PCR cryptographic measurement validation.
Universal SDK Release (cain-trust 2.0.0 on PyPI)#
- Standalone PyPI Package: Universal client library built and distributed (
dist/cain_trust-2.0.0-py3-none-any.whland.tar.gz). - One-Line Integration:
@cain.guarddecorator for securing any Python function or agent tool call.cain.wrapcontext manager for wrapping arbitrary agent frameworks (LangChain, AutoGen, CrewAI, LlamaIndex).- High-Throughput In-Memory Cache: Local cryptographic cache providing sub-15ms local decision validation with fail-closed offline fallback.
- Standalone CLI Verifier: Packaged
cain auditandcain verifycommands for offline verification of signed execution evidence bundles.
EU AI Act Statutory Pre-Conformity Portal#
- Pre-Conformity Engine: Statutory compliance framework (
cain.compliance) mapping runtime trust evidence directly to the European Union Artificial Intelligence Act (Regulation 2024/1689): - Article 9 (Risk Management System)
- Article 10 (Data & Governance Controls)
- Article 11 (Technical Documentation)
- Article 12 (Continuous Automated Record-Keeping & Logging)
- Article 13 (Transparency & Information Provision)
- Article 14 (Human-in-the-Loop & Fallback Authority Controls)
- Article 15 (Accuracy, Robustness & Cybersecurity)
- Article 72 (Post-Market Continuous Monitoring)
- Automated Discovery Portal: Live web-based statutory audit room (
/compliance/eu-ai-act) generating cryptographically signed, court-admissible WORM notary evidence packages insulating enterprises from €35M or 7% worldwide turnover fines.
Actuarial Cyber Insurance Consortium Protocol#
- Actuarial Risk Engine: Autonomous risk quantification module (
cain.insurance) calculating: - Actuarial Vulnerability Index (AVI, 0.000–1.000)
- Maximum Probable Loss (MPL) per autonomous workflow
- Underwriting Trust Credit Score (0–1000)
- Consortium Underwriting Data Rooms: Standardized data room generation for insurance syndicates (Lloyd's of London, Munich Re, Swiss Re, Beazley), unlocking 15% to 35% enterprise cyber premium discounts.
- CLI Underwriting Suite: Interactive CLI tools (
cain consortium status,cain consortium data-room) for real-time underwriting attestation.
Vertical Rego Policy & Threat Intelligence Marketplace#
- Commercial Rego Policy Packs: Production-grade Open Policy Agent (OPA) policy bundles (
cain.opa): FIN-REG: Financial controls enforcing GLBA, SOX-404, and SEC Rule 17a-4 transactional guardrails.HEALTH-REG: Healthcare security rules enforcing HIPAA, HITECH, and PHI de-identification boundaries.FED-REG: Public sector defense packs enforcing FedRAMP High and NIST SP 800-53 Rev 5 security controls.- Package Management CLI: Integrated commands (
cain marketplace list,cain marketplace install) for vertical policy lifecycle management.
Autonomous Swarm Fleet Quarantine & Emergency Circuit-Breaker#
- Sub-50ms Swarm Containment: Cross-process fleet isolation API (
cain.quarantine) backed by SQLite WAL sync, delivering sub-50 millisecond emergency kill-switch and quarantine capabilities across multi-agent swarms. - Immediate Certificate Revocation: Dynamic revocation of compromised agent credentials preventing cascade failures across distributed nodes.
Enterprise SIEM & SOC Connectors#
- Multi-Format Telemetry Streaming: Real-time log export engine (
cain.telemetry) delivering cryptographically notarized security telemetry formatted in: - Micro Focus ArcSight Common Event Format (CEF)
- IBM QRadar Log Event Extended Format (LEEF)
- Microsoft Azure Sentinel JSON Stream
- SecOps Alert Correlation: Live detection of prompt injection, policy violations, and trust drift integrated into existing enterprise Security Operations Centers.
[1.2.0] — 2026-09-10#
Strategic Identity & Validated Trust Runtime Baseline#
- Formalized CAIN strategic identity as the AI Infrastructure Validated Trust Runtime for Autonomous Systems.
- Established canonical 7-Moat Trust Control System architecture.
- Introduced
ValidatedTrustRuntimecore primitive with 19-category validation check graph. - Delivered
cain testindependent conformance and red-team testing suite.