Not more tools. Skills + purpose + capture. Upgrade to 0.19 first, compound procedures, one cloud agent brain, vault stays human. Docs + tapi X consensus, ranked for Dwad.
Open entry →Hard win → skill. Durable fact → OB1. Prose → vault. Cron keeps lights on. Don’t Hermes every nail.
Maximize Hermes by compounding skills + purpose + capture, not by bolting on more tools.
Get on 0.19 first — speed, delivery ledger, and smart approvals are free wins. You’re on 0.18.1 today.
v0.19.0 Quicksilver that matters on Telegram:
• ~80% faster first token (4.3s → ~0.9s)
• Smart approvals default — less approve-fatigue
• Delivery ledger — finished replies redeliver if gateway dies
• Live subagent transcripts — async work is watchable
• Bitwarden / 1Password secret sources
• Models you care about: grok-4.5, kimi-k3, etc.
Who: Mac Claude runs host image Plan A. Hermes verifies after. Plan file: out/hermes-0.19-seamless-upgrade-plan-2026-07-21.md
A. Purpose before tools. Lasting setups define the job (X consensus). You already have Donna identity + lifestyle cap. Don’t add 40 MCPs “because you can.”
B. Four memory types (CoALA map on X). Working = context. Semantic = facts/brain. Procedural = skills (Hermes’ real edge). Episodic = session lessons. Most people only ship working memory.
C. Skills = compound interest. Solve once → SKILL.md → load next time. After any 5+ tool win → save skill. Patch when wrong. /reload-skills. Curator later for agent-created skill hygiene.
D. Capture loop. X “game changer”: meetings/brain-dumps → agent knowledge. Your version: Telegram durable facts → capture; wrap/debrief → decisions; vault import later (OB1), not day-zero theater.
E. Right tool for the job. Don’t full-Hermes every PR/micro task (tonysimons_ tip). Small stuff → light surface. Multi-step life OS → Hermes.
F. Delegation that doesn’t go dark. 0.19 live transcripts + durable background results. delegate_task for parallel work; cron for anything that must outlive the chat.
G. Security without friction. Smart approvals, deny rules for never-dos, credential pools/vault secrets, pass-cli pattern stays.
H. Profiles only when jobs diverge. Keep default Donna + wuwei. No profile zoo.
Talk on Telegram → Hermes works → hard win becomes a skill → durable fact hits cloud brain (OB1) → human thinking stays in vault → cron keeps lights on → you don’t re-explain tomorrow.
That’s maximize. Not “more features.”
| # | Move | Why |
|---|---|---|
| 1 | 0.19 upgrade | Speed + no silent lost replies |
| 2 | Skills compounding | Free permanent IQ |
| 3 | OB1 one brain | Kill re-explain Mac↔Telegram |
| 4 | /goal contracts | Multi-step jobs finish for real |
| 5 | Smart approvals + secrets hygiene | Less nags, less leak risk |
| 6 | Ruthless non-use | Don’t Hermes every nail |
Desktop app obsession · Nous Portal upsell unless you want the bundle · extension zoo/dashboards before search works · parallel GBrain+OB1+Zep · running Hermes for every micro PR.
Proven maximizers: upgrade → skills → one memory door → goals/cron/delegation. Everything else is cosplay.
Install OB1. Don’t build “OB1 junior.”
Cloud Supabase = SSOT 24/7. Hermes + Mac Claude = clients on the same MCP URL. Vault = human notes only. Phone agent mem = Telegram.
• Mac not SSOT — laptop can sleep; brain stays up
• Multi-surface — one MCP for Hermes Telegram + Mac Claude
• Cheap — free Supabase tier + pennies embeds; fits $9 VPS frugality
• Solo — no staff, no second-brain theater, capture + search
• Vault import exists — keep Glass House; don’t migrate by hand
• Importers when needed — ChatGPT/Gmail/X later, not day one
• Lifestyle KPI — kills re-explain tax so presence stays on Kai/Elle
Skip for now: extension ladder, dashboards, Life Engine, wiki compiler, running GBrain in parallel.
capture + search working in one client first.1. Fact captured on Telegram is found on Mac Claude with source.
2. Mac powered off does not break recall on Telegram.
3. You’re not maintaining a second agent memory store.
No redesign of schema “for us.” No OB1+GBrain+Zep sandwich. No vault replacement. No dashboard before daily search hurts less.
out/one-human-door-one-agent-door-justified-2026-07-21.md · builds on Entries #003–#004 + OB1 SOTAYou do not need another second-brain product. You need one human thinking surface and one agent retrieval layer — then stop collecting memory systems.
Personal win: zero re-explain across Telegram, Mac Claude, and phone Obsidian, so presence stays on Kai and Elle, not on re-briefing Donna.
| Surface | Measured 2026-07-21 | Implication |
|---|---|---|
| Vault | 3,588 md · 247MB · VPS RO | Human SSOT already exists |
| MEMORY.md / USER.md | 816 B / 2.3 KB · caps 2.2/2.5K | Tiny always-on injection only |
| Memory provider | holographic | Hermes-local, not Mac-shared |
| MCP servers | agentmail, altfins, x, apify | No shared brain MCP |
| Write lane | /opt/data/out → Mac inbox | Delivery works; fusion does not |
| Candidate brains | OB1, GBrain, Zep, Mem0… | Research done; install not justified yet |
Plus overlapping agent memory already live: holographic facts, fabric, session_search, vault grep, skills, out artifacts. That is the silo problem Entry #004 named.
C1. No new second-brain product. Lifestyle cap + vault already is the human PKM. Nick/Cerebras: unused aesthetics are theater. Failure mode is retrieval silos, not missing Notion.
C2. Dual door, not mono store. Human door = vault (edit, phone sync, doctrine). Agent door = one MCP retrieval layer. Agents alone → slop. Humans alone → unqueryable across sessions.
C3. Zero re-explain is the product. Mac Claude ≠ Hermes holographic store. No brain in mcp_servers. session_search is Hermes-local. Every surface switch re-taxes LMS facts, constraints, open loops.
C4. KPI = re-explain minutes, not note count. More notes do not lower cognitive load. Capture-once + sourced retrieve + same agent door on Mac/VPS do.
C5. Presence tax. Priority stack is Kai/family first. Minutes recovered from hunting decisions are the lifestyle doctrine working, not soft wellness.
C6. Solo ops without staff. You will not manage people. Shared agent memory is how SOP/client state persists at lifestyle scale.
C7. Stop collecting brains. Seven overlapping stores. Each new layer without a kill criterion raises miss rate and token bleed (already flagged on interactive mega-sessions).
C8. P0 before OB1/GBrain. Fan-out + citations falsify “need a product” cheaply. Install-first creates an eighth silo. Infra law: durable, frugal, no compose roulette.
C9. GBrain lean if P0 fails. Prior analysis: Hermes-native markdown + hybrid/RRF. OB1 wins gentler setup/importers. Preference ≠ install order.
C10. Vault is not replaceable. Years of domain doctrine. OB1 community treats Obsidian as import source. Phone path already works.
| Pushback | Reply |
|---|---|
| “Just use Obsidian better” | Human door is fine. Agent door is the hole. |
| “Install OB1 tonight” | Seventh→eighth silo. P0 is cheaper falsification. |
| “Zep is SOTA graph” | Overkill before daily fusion proves value. Sequence after stable loops. |
| “Put everything in Postgres” | Breaks Glass House + phone sync. Extract, don’t migrate. |
| “One store only” | Agent-optimized slop vs human-optimized unqueryable. Dual-door is stable. |
You: Telegram (ops) · Mac Claude (build) · Obsidian (think/phone).
Agent door (ONE): shared retrieval over vault + sessions + captures, citations mandatory.
Not homes: ChatGPT memory, random Notion, third graph “for later.”
kb_search fan-out · citation law · 3 daily surfaces only.| Metric | Target |
|---|---|
| Mac↔Telegram re-explain | ≤1 / week |
| Memory answers with sources | ≥90% |
| New memory systems added | 0 unless P1 trigger |
| Vault remains human SSOT | yes |
Subtraction plus one fusion path — not acquisition. Human door: vault. Agent door: thin retrieval first, product only on failed P0. Kill criterion on every new memory idea.
Nick's open: almost every "knowledge base / second brain" has been hot air — until Cerebras shipped an internal system staff actually hammer (15,000+ queries/day in ~3 months), used by humans, automations, and agents.
The kill shot is not a prettier Obsidian theme. It is rejecting the single-source-of-truth migration fantasy. Meet data where it lives (Slack, GitHub, docs, Jira, custom DBs) → unify at the embeddings/query layer → hybrid retrieve → plan/execute tools in parallel → RRF + rerank → cited synthesis. Expose the same retrieval as MCP primitives so agents orchestrate without a hidden second brain LLM.
Evidence note: full captions blocked (Apify floor + SOCKS offline + YT bot wall). Delta grounded on indexed transcript open + full Cerebras architecture digest + live Donna memory audit. Artifact: out/cerebras-second-brain-nick-saraev-transcript-2026-07-21.md.
Ingest: one Postgres embeddings table; connectors via PR; Slack Socket Mode real-time; LLM distillation before embed (question/summary/resolution/systems); "bursting" for long threads; CocoIndex incremental code chunking (to 40GB+ repos).
Retrieve: four-signal Slack hybrid (full-text + embedding + IDF + age decay) because vectors alone promote filler. Query path: planner → parallel tools → RRF (k=60) → reranker top-10 → context expand → synthesis with citations.
Scope: Projects bundle channels/repos/docs per team so compiler engineers do not get DC runbooks. Agents: Web UI runs full pipeline; MCP exposes LLM-free search_slack / search_code / who_knows.
| Capability | We Have | Source Says | Gap |
|---|---|---|---|
| Meet data where it lives (no forced SSOT) | ⚠️ Multi-store by accident (memory/fact/fabric/vault/out) | Extract in place by design | Medium |
| Unified embeddings + common schema | ❌ Siloed stores, no shared embed table | One Postgres embeddings table | BIG |
| Continuous multi-source ingest | ⚠️ Manual delta-capture / ori_add / skill_manage | Slack/code/docs always-on | BIG |
| Hybrid retrieval (FTS+vector+IDF+age) | ⚠️ FTS (session_search/fabric) + ranked recall, no fusion | Four-signal hybrid + RRF | BIG |
| LLM distillation on ingest | ⚠️ Delta pipeline / wiki ingest pattern | Every Slack thread distilled | Medium |
| Planner → parallel retrieval tools | ✅ Parallel tools + delegate_task | Parallel fan-out executor | None |
| RRF + reranker + cited synthesis | ❌ Model synthesizes from whatever landed in context | Formal rank fusion + citations | BIG |
| MCP retrieval primitives for agents | ⚠️ MCP servers exist; no KB tool surface | search_slack/code/who_knows | Medium |
| Project-scoped knowledge packs | ⚠️ Profiles + vault domains, not query scopes | Projects at onboarding | Medium |
| Personal wiki / second brain discipline | ⚠️ llm-wiki skill + Obsidian vault RO + out/ | Nick: pure PKM theater fails at scale | Small* |
*Small for lifestyle-cap personal use — Dwad is not Cerebras. The transferable win is retrieval design, not cloning a 15k Q/day internal platform.
Do not rebuild Notion inside Hermes. Do not declare the vault "dead." Do adopt Cerebras's three laws at personal scale:
1. Extract, don't migrate — Telegram, Gmail, vault notes, session DB, X already exist. Ingest summaries into one queryable layer.
2. Hybrid rank, don't vibe-retrieve — stop hoping one store answers; fuse FTS + semantic + recency before the model speaks.
3. Agent-facing primitives — retrieval tools agents can call without a second hidden brain.
This upgrades Entry #003's Obsidian bonus from "browser save to vault" to "vault is one source among many in a fused index."
kb_search (or skill) that fans out to session_search + fabric_recall + fact_store + vault grep, returns a merged top-N with source tags. No new DB day one — fusion in code./opt/data/out, auto-write a normalised card (question/summary/resolution/tags) into a single kb/ markdown index the searcher prefers.Jack's claim is not "new Hermes concepts." It is that five already-shipped platform capabilities now compound into an agentic OS: swap brains, parallel tools, Firecrawl research, daily ops surface, and completion contracts — with Obsidian memory + skill-from-docs as the bonus layer.
For Donna this is a wiring audit, not a roadmap fantasy. Hermes 0.18.1 on this VPS already has Grok 4.5, parallel tool guidance, delegate_task, native /goal completion contracts, gws email/calendar, and skill writing. The gap is operational adoption + a few missing lanes (Kimi, Firecrawl budget, browser→vault).
Evidence note: full captions blocked this session (Apify usage floor + Mac SOCKS offline + YT bot wall). Delta grounded on full author description + chapter stamps + indexed transcript fragments + live config audit. Source artifact: out/hermes-10x-jack-roberts-transcript-2026-07-21.md.
L1 Swap Smarter Brains — GPT 5.6 via ChatGPT sub, Grok 4.5 cheap/fast + live X intel, Kimi K3 design at a fraction of Opus, one-prompt website test, price breakdown.
L2 Stop Waiting — native parallel tool calls; demo runs four tasks at once.
L3 60× Faster Research — Firecrawl as the research path; strips junk HTML; brand-identity extraction; model showdown (~60× faster / ~49× cheaper claim in fragments).
L4 Daily Edge — email + calendar connected; permission safety before acting; a brief that improves itself.
L5 Proof of Work — completion contracts: define verifiable DONE up front; agent keeps going until evidence matches.
Bonus — save anything to Obsidian from the browser, recall it later, turn docs into durable skills.
| Capability | We Have | Source Says | Gap |
|---|---|---|---|
| Grok 4.5 primary + X intel | ✅ Live (xai-oauth) | Fast/cheap brain + X | None |
| Multi-model swap (Kimi / GPT lane) | ⚠️ DeepSeek fallback only | Kimi K3 design + GPT 5.6 free lane | Medium |
| Parallel tool calls | ✅ Guidance on + native runtime | Four tasks at once | None |
| Subagent fan-out | ✅ delegate_task (max 10) | Parallel workstreams | None |
| Firecrawl / fast clean research | ⚠️ Key present, chronically 402 | 60× path, brand extract | BIG |
| Email + calendar connected | ✅ gws (draft-only mail) | Daily ops surface | None |
| Permission safety before external acts | ⚠️ Policy in SOUL, not a hard gate | Explicit permission rule | Small |
| Self-improving daily brief | ⚠️ Cron briefs exist | Brief that compounds | Medium |
| Completion contracts (/goal) | ⚠️ Native in 0.18.1, unused by Donna | Proof-of-work DONE checks | BIG |
| Browser → Obsidian real-time capture | ❌ Vault RO + out/ lane only | Save anything from browser | BIG |
| Compounding memory / recall | ⚠️ memory + fact_store + Ori | Wiki-style compounding memory | Medium |
| Build skills from docs | ✅ skill_manage + delta pipeline | Docs → durable skill | None |
Already done: Grok 4.5, X search, parallel tools, subagents, gws calendar/mail draft, skill writing.
Partial / unused: completion contracts exist but Donna does not run work through /goal; briefs cron but do not self-score; memory is multi-store not one compounding wiki.
Real holes: Kimi design lane, Firecrawl budget reliability, browser→Obsidian capture path.
/goal contract (outcome + verification + constraints) before execution. Stop accepting "looks done" without evidence checks. This is the same muscle Entry #001/#002 asked for — now native in Hermes./opt/data/out (syncs to Mac vault). Not full bidirectional Obsidian MCP day one — just the capture habit Jack demos.Karpathy's auto-research pattern is now a native Claude Code feature — /goal and /loop. The new unlock: evaluator cartridges — custom scoring algorithms that grade AI output autonomously. The combination — goal-driven agent + multi-cartridge evaluator + attempt budget — produces output that iteratively sharpens without you.
This is Entry #001's loop engineering with the inspector automated. The human graduates from checking every output to defining the scoring rubric once.
A cartridge is a self-contained scoring algorithm: output → number (1–10). Three shown:
1. Humanify Cartridge — voice, variety, predictability (AI slop detector). Only 9+ passes.
2. Hormozi AI Cartridge — scored against Alex Hormozi's public writing for marketing punch.
3. Open Rate Analytics Engine — predicted open rate from past email data + audience behavior.
Composite scoring across cartridges: newsletter required 27/30 total, no single cartridge below 8. Agent ran 5 attempts autonomously: 7.1 → 26.8 over ~40 minutes.
| Capability | Hermes Has | Source Says | Gap |
|---|---|---|---|
| Autonomous goal-driven loops | ⚠️ Manual pattern | /goal native | BIG |
| Evaluator cartridges (scoring) | ❌ None | Multi-cartridge | BIG |
| Condition-based wake-up | ⚠️ Cron (schedule) | /loop (condition) | Medium |
| Subagent swarm → single goal | ⚠️ Independent tasks | 650 experiments | Medium |
| Attempt tracking + scoring history | ⚠️ STOP AFTER only | Scoring table | Medium |
| Scoped file access for loops | ❌ Full FS | Scoped access | Small |
cronjob: "run every N minutes UNTIL condition met."Prompt engineering is dead. The new paradigm is loop engineering: AI tries → self-checks → fixes → retries → marks done. You move from being in the loop to being outside it. Four components: Goal, Checklist, Inspector, Budget.
Inner loop: Single task — try, inspect, fix, retry. AI self-corrects before delivering.
Outer loop: Recurring scheduled job running inner loops each iteration. Self-improves over time.
| Capability | Hermes Has | Loop Eng. | Gap |
|---|---|---|---|
| Scheduled recurring tasks | ✅ 17 crons | Outer loop | None |
| Self-review within task | ❌ Missing | Inner loop | BIG |
| Inspector/second pass | ❌ Missing | Grader model | BIG |
| Hard definition-of-done | ⚠️ Soft goals | Checklist | Medium |
| Skills auto-update from loops | ❌ Missing | Loop→skill | Medium |
| Per-task attempt budget | ⚠️ Repeat only | Stop after N | Small |
loop-engineer skill: GOAL + DONE WHEN + STOP AFTER → delegate → inspect → retry up to N.