The audit is over, and it took the roadmap down with it.
For five months this section was a graveyard of half-remembered ambition. What climbed back out is narrower, and more honest. Sanctum’s core feature set is mechanically proven — so the remaining work is not to make the Living Force real. It is to make Sanctum reproducible, legible, and releaseable without depending on one Mac Mini and one operator’s memory.
One gate stands before that work resumes: a one-week Stability Window. If Sanctum cannot stay boring for seven days, it has not earned fresh ambition.
Sanctum is mechanically stable as a single-operator system. It is not yet productized in the stronger sense — portable, repeatable, supportable — without inheriting Bert’s machine shape.
Axis
Score
Where it stands
Stability
8/10
The runtime, audit wall, calibration checks, and live-system probes behave like real infrastructure instead of aspirational architecture.
Recoverability
7.5/10
Self-heal paths are real and tested, but some of that resilience still depends on bespoke restart scripts and machine-local supervision conventions.
Reproducibility
3.5/10
Too much of the system still assumes /Users/neo, local keychain state, ~/Projects/* repos, launchd behavior, and adjacent private runtime surfaces.
Productization readiness
4.5/10
Ready for continued hardening on the current machine, maybe a tightly controlled second-machine rollout — but not external beta users or a supportable install story.
The conclusion is simple: the feature surface is no longer the bottleneck. Turning this personal runtime into a portable operator product is.
Freeze the architecture and let it behave. Phase 2 work stays blocked until Sanctum survives a seven-day soak with the audit wall, runtime checks, and docs build still clean.
Turn sanctumctl from a strong operator surface into a genuinely reproducible install path. A fresh compatible machine should bootstrap the checked-in slice, render manifests, sync calibration artifacts, and run the audit wall without hidden shell lore.
Separate machine-specific state from the core product surface. Paths, host capabilities, LaunchAgent assumptions, and runtime conventions need a clearer boundary, so Sanctum can target more than one setup without pretending every machine is identical.
--dry-run previews the plan; --version vX.Y.Z overrides the stamp
The mechanism shipped. The first real run correctly blocked on 9 pre-existing test failures (xtts pin_deps, immune-system e2e, workspace/runtime/system audits) — the gate is real, not theater. Fixing those tests is separate work; the script’s job was to surface them, and it did.
sanctum-cli v0.8 — public install path (shipped 2026-05-21)
sanctum-cli v0.8.0 is tagged and the repo is public (council vote 4/4 APPROVE — brain/mlx/code/spacial). The Ogilthorp3/homebrew-sanctum tap repo exists; Formula/sanctum-cli.rb ships v0.8.0 (Apache-2.0, Language::Python::Virtualenv, 38-dep PyPI resolution verified end-to-end). Onboarding splash plus personalized completion celebration landed (commit 5b17c05).
The beta tester command is now real:
brew install ogilthorp3/sanctum/sanctum-cli && sanctum onboard --recipe family
Still open: Sigstore-signed releases (keyless OIDC via cosign in a GitHub Action), atomic-replace flow for cloud_backup, public demo recording, and the sanctum dashboard TUI.
Public Demo Path
Build the reveal layer. Sanctum needs a clean, intentional demo path that shows the parts worth believing: watchdog self-heal, Code Forge rollback, Jocasta context retrieval, Tech Lookout dispatch, and Force Flow delivery. The point is not theatrics. The point is making the real system legible.
Deliverables: scripted walkthrough, stable demo fixtures, screenshots or clips, one page tying the sequence together.
Cross-Repo Contract Cleanup
Reduce implicit coupling across sanctum, ~/.sanctum, openclaw-skills, and the support repos. Adjacent systems should declare what Sanctum expects from them instead of relying on shared memory and ambient convention.
The next work is content expansion: deeper L4 state-delta checks, broader scenario coverage, and tighter runtime attachment for real haushold automations and memory drift controls.
Why later: the mechanism exists and now runs on its own; the remaining work is coverage depth, not foundational productization.
Qwen3-TTS — Local Yoda Voice
Done 2026-04-19. Qwen3-TTS via mlx-audio replaced XTTS-v2 for Yoda’s voice agent — zero-shot voice cloning from reference audio, lower memory pressure than XTTS at equal quality.
com.sanctum.yoda-tts-worker on :8008 (workers.tts_server Python module)
XTTS-v2 plist → now .retired on disk
Council fallback — restored 2026-04-25
The dead council-guardian.sh::activate_fallback() automation was retired (script went 257 → 203 lines, three structural reasons documented). Its replacement lives in the Smart Router cathedral, where it always belonged. The point of a Python seat is stack independence: the cathedral (:1337) and devstral (:3301) seats share one binary file, so a corrupt or missing Rust binary kills both at once, and only a different-stack server survives it.
(an always-warm 35B would sit ~20 GB on top of the always-loaded Rust 35B
and push the 64 GB Mini deep into swap)
council-integrity-check promotes it on unrecoverable Rust-binary failure:
boots the broken Rust agent first (frees ~20 GB) → enables + kickstarts the Python 35B
wrapper ~/.sanctum/bin/mlx-server-with-caps.py caps the MLX cache on promotion
(mx.set_cache_limit + mx.clear_cache() before launch — same pattern as the
Rust binary's --metal-cache-limit-mb flag)
transient (non-binary) failover is separate, at the proxyd routing layer:
council-mlx → council-code (devstral :3301) → claude-cli-offline (Claude Max :3456)
Memory cost is zero at rest: the 35B loads only after council-integrity-check has booted the broken Rust agent and freed its ~20 GB. Holding both 35B seats resident is exactly the swap blowout the promote-only design avoids. End-to-end drill on the Mini (2026-04-25):
Stage
Latency
Rust primary, warm
0.51 s
Failover, cold-load (first request)
75 s
Failover, warm thereafter
1.07 s
Rust restored
0.51 s
With both seats on the same Qwen3.6-35B-A3B-4bit, the Python answer matches the Rust answer at full parity. Receipts: project_failover_drill_2026_04_25.md in agent memory.
Real product features, all of them — but downstream of making Sanctum easier to install, verify, and trust.
Monorepo or Workspace Consolidation
More of the Rust and support tooling could live in a tighter workspace someday — better release ergonomics, but secondary: first the system needs clearer contracts.
Why later: consolidation without stronger boundaries just relocates confusion.
Context-Aware Automation Expansion
Frigate-backed alert semantics and richer haushold automation remain appealing.
Why later: new features should follow product discipline, not precede it.
Firewalla Gold Pro Migration
Export the Firewalla Purple config and migrate to Gold Pro when the hardware arrives. Purple still works, so this waits.
SSH iptables fallback covers the current firmware's broken policy:delete API
— verify this is still needed on Gold Pro firmware
Full Local Voice Pipeline
Replace every cloud dependency in the voice agent. With Qwen3-TTS, that yields a zero-cloud pipeline: phone calls to Yoda with no data leaving the haus.
Deepgram STT → local Whisper via mlx-audio
Claude LLM → local Qwen3.6-35B-A3B-4bit via sanctum-mlx
Why later: each component must be fast enough on its own first, and Qwen3-TTS speed is the current bottleneck.
Gemma 4 (and other archs) in sanctum-mlx
Today sanctum-mlx::LoadedModel dispatches qwen3_5 and qwen3_5_moe only — anything else fails at config parse and falls back to Python mlx_lm.server. A gemma3.rs loader plus a third LoadedModel::Gemma3 variant would put Gemma on the same fast path Qwen rides — an estimated 1–2 days.
Gemma 3/4 share one Rust shape → a single gemma3.rs loader covers both
fused TurboQuant attention kernel needs a small adapter:
FullAttention::forward integration currently lives in qwen3_5.rs
Why later, not next: this is vendor diversification dressed as resilience. The shared substrate that actually fails — MLX runtime, Metal driver, custom kernels, launchd, mTLS — takes both models down at once, so a second loader adds nothing. The Python fallback already covers a broken Rust path for any model mlx-lm supports, Gemma included. Build it for a concrete workload — Gemma 4 LoRAs that want fast serving, or a seat needing Gemma’s tokenizer — not generic resilience.
SharePoint doc_type metadata + by-type views
The SharePoint bridge already stamps a doc_type tag on every upload — a controlled vocabulary (investment-memo, tearsheet, dd-summary, lp-update, …) from work-skills/sharepoint-structure.yaml. The goal is cross-cutting retrieval: every memo or LP update across all funds at once, not a tree walk. Two halves remain.
The column — holds the tag on the "Documents Triptyq" library
/sites/<work-site>, list 00000000-0000-0000-0000-000000000000
does not exist yet → the tag is dropped (metadata_applied: false)
provisioning is gated on a site-collection-admin grant:
and the operator is not an owner of that locked-down library
(createfieldasxml + Options:8 pins the exact internal name)
The consumer — a view or report that actually queries the tag — is unbuilt
Why later: the folder hierarchy already does the primary organizing and the bridge degrades gracefully; this pays off only once both halves ship.
Session Continuity Cache
A mid-session quota exhaustion on the wisdom-ladder chain (Max → Gemini → Grok → glm-52 → cathedral) hands the conversation to a different model family. Nothing is lost on the wire — the proxy is stateless, forwards the full transcript, and only rewrites the model field. But the new seat reasons differently, and a blinking client can drop its own transcript. A bounded, per-session cache at the proxy would inject a “here’s where we are” preamble on a seat change and pin a stable prefix for KV reuse.
Why later: session affinity plus a bounded cache in front of the router trades away the proxy’s statelessness — worth it only once switch frequency earns it. Design: Session Continuity Cache. Today’s mitigations (claude --resume, similar-model first rung) cover the common cases.
A native iPadOS app that turns any iPad into a full Sanctum control surface — not a web view in a frame, a real app that belongs on the device. Dock one at any satellite (chalet, office, cabin) and control the entire galaxy: voice input via on-device mic, camera feeds, climate zones, agent conversations, and a security dashboard, all through one interface that works offline when Tailscale drops and syncs when it reconnects.
Voice-first : wake word, local speech-to-text, response via AirPlay to HomePods
Dashboard : camera feeds (Blink, Ring), HVAC zones, alarm status, agent health
Agent chat : talk to any Jedi directly — routed to the right model via Smart Router
Offline : full functionality on the local network without internet
Multi-site : switch between Manoir, Chalet, and any future satellite with one tap
Kiosk mode : Guided Access lockdown for an always-on kitchen/hallway display
Tech stack : SwiftUI + local MLX inference + Home Assistant WebSocket + Sanctum API
Why Phase 3: the infrastructure — Smart Router, Model Tournament, 3-tier routing — must be solid before a consumer surface sits on top. Phase 2 made the foundation trustworthy. Phase 3 makes it beautiful.