Checkpoints & Fault Recovery
Two-phase commits, startup crash scanners, and 15 failure classifications.
M31A includes a fault-recovery subsystem designed around atomic two-phase checkpoints, differential task replanning, and fail-closed crash scanners.
Two-Phase Checkpoint Commit Protocol
Phase 1: External Artifact Staging
├── Write patches, test logs to staging/<checkpoint_id>/
├── Flush bytes with tokio::fs::File::sync_all
└── Verify content-addressed SHA-256 digests against manifest
↓
Phase 2: Atomic SQLite Transaction
├── BEGIN TRANSACTION; INSERT INTO checkpoints ...; COMMIT;
└── Promote staged artifacts to authoritative FsArtifactStore
Startup Crash Scanner
Upon startup, M31A scans for interrupted missions, categorizing state into 4 strict classifications:
- • SafeToResume: Manifest valid, artifacts intact, workspace clean. Safe to continue.
- • NeedsRepair: Checkpoint valid, but workspace files need restoration to baseline commit.
- • Ambiguous (Fail-Closed): Unexplained repo drift or conflicting jobs. Halts and prompts operator.
- • Corrupt (Fail-Closed): Missing artifacts or truncated SQLite state. Halts fail-closed.
15 Canonical Failure Classifications
Errors detected during task execution are deterministically classified across 15 structured categories (Transient, Timeout, Permission, Policy, Compilation, Test, RepositoryState, etc.). Security failures (Permission, Policy) strictly receive a retry budget of 0.