DONNA

Universal output surface · Reports, deltas, plans, artifacts — all here
Maximize Hermes: skills + purpose + capture/0.19 first — speed + delivery ledger/Procedural memory = skills compound/OB1 one agent door · vault human only//goal contracts · smart approvals/Don't Hermes every micro-task/All output → delta111.pages.dev/ Maximize Hermes: skills + purpose + capture/0.19 first — speed + delivery ledger/Procedural memory = skills compound/OB1 one agent door · vault human only//goal contracts · smart approvals/Don't Hermes every micro-task/All output → delta111.pages.dev/
07
Entries
Published
06
Maximize
ROI Ranks
10
Action
Items
Latest Entry — Jul 21, 2026

Maximize Hermes — Proven Stack

Not more tools. Skills + purpose + capture. Upgrade to 0.19 first, compound procedures, one cloud agent brain, vault stays human. Docs + tapi X consensus, ranked for Dwad.

Open entry →
◆ Key Insight Jul 21
Procedural memory is the free IQ gain.

Hard win → skill. Durable fact → OB1. Prose → vault. Cron keeps lights on. Don’t Hermes every nail.

v0.19 docs · tapi X · hermes skill
// Active knowledge base

All Entries

◆ Entry #007 · Playbook Jul 21, 2026
Maximize Hermes — Skills, Purpose, Capture
Sources: Hermes v0.19.0 Quicksilver release · hermes-agent skill/docs · tapi X (NousResearch, HermesAgentTips, operators) · stack plan Entries #005–#006

Maximize Hermes by compounding skills + purpose + capture, not by bolting on more tools.

Get on 0.19 first — speed, delivery ledger, and smart approvals are free wins. You’re on 0.18.1 today.

v0.19.0 Quicksilver that matters on Telegram:

• ~80% faster first token (4.3s → ~0.9s)
• Smart approvals default — less approve-fatigue
• Delivery ledger — finished replies redeliver if gateway dies
• Live subagent transcripts — async work is watchable
• Bitwarden / 1Password secret sources
• Models you care about: grok-4.5, kimi-k3, etc.

Who: Mac Claude runs host image Plan A. Hermes verifies after. Plan file: out/hermes-0.19-seamless-upgrade-plan-2026-07-21.md

A. Purpose before tools. Lasting setups define the job (X consensus). You already have Donna identity + lifestyle cap. Don’t add 40 MCPs “because you can.”

B. Four memory types (CoALA map on X). Working = context. Semantic = facts/brain. Procedural = skills (Hermes’ real edge). Episodic = session lessons. Most people only ship working memory.

C. Skills = compound interest. Solve once → SKILL.md → load next time. After any 5+ tool win → save skill. Patch when wrong. /reload-skills. Curator later for agent-created skill hygiene.

D. Capture loop. X “game changer”: meetings/brain-dumps → agent knowledge. Your version: Telegram durable facts → capture; wrap/debrief → decisions; vault import later (OB1), not day-zero theater.

E. Right tool for the job. Don’t full-Hermes every PR/micro task (tonysimons_ tip). Small stuff → light surface. Multi-step life OS → Hermes.

F. Delegation that doesn’t go dark. 0.19 live transcripts + durable background results. delegate_task for parallel work; cron for anything that must outlive the chat.

G. Security without friction. Smart approvals, deny rules for never-dos, credential pools/vault secrets, pass-cli pattern stays.

H. Profiles only when jobs diverge. Keep default Donna + wuwei. No profile zoo.

Talk on Telegram → Hermes works → hard win becomes a skill → durable fact hits cloud brain (OB1) → human thinking stays in vault → cron keeps lights on → you don’t re-explain tomorrow.

That’s maximize. Not “more features.”

  1. P0: Mac Claude — 0.19 image upgrade (Plan A)
  2. P0: config migrate + doctor + Grok-only assert
  3. P0: Confirm smart approvals + delivery ledger
  4. P0: Skill discipline — every hard win → skill_manage
  5. P1: OB1 cloud SSOT — same MCP Hermes + Mac Claude
  6. P1: Capture habit — durable → OB1; prose → vault
  7. P2: /goal + completion contracts on multi-step jobs
  8. P2: Cron — no_agent watchdogs free; LLM crons only when needed
  9. P2: delegate_task + live probe on big research
  10. P2: Session export occasionally; prune junk
#MoveWhy
10.19 upgradeSpeed + no silent lost replies
2Skills compoundingFree permanent IQ
3OB1 one brainKill re-explain Mac↔Telegram
4/goal contractsMulti-step jobs finish for real
5Smart approvals + secrets hygieneLess nags, less leak risk
6Ruthless non-useDon’t Hermes every nail

Desktop app obsession · Nous Portal upsell unless you want the bundle · extension zoo/dashboards before search works · parallel GBrain+OB1+Zep · running Hermes for every micro PR.

Proven maximizers: upgrade → skills → one memory door → goals/cron/delegation. Everything else is cosplay.

◆ Entry #006 · Plan Jul 21, 2026
OB1 — Do This, Don’t Redesign
Repo: NateBJones-Projects/OB1 · Setup: docs/01-getting-started.md · Prior: Entries #004–#005

Install OB1. Don’t build “OB1 junior.”

Cloud Supabase = SSOT 24/7. Hermes + Mac Claude = clients on the same MCP URL. Vault = human notes only. Phone agent mem = Telegram.

Mac not SSOT — laptop can sleep; brain stays up
Multi-surface — one MCP for Hermes Telegram + Mac Claude
Cheap — free Supabase tier + pennies embeds; fits $9 VPS frugality
Solo — no staff, no second-brain theater, capture + search
Vault import exists — keep Glass House; don’t migrate by hand
Importers when needed — ChatGPT/Gmail/X later, not day one
Lifestyle KPI — kills re-explain tax so presence stays on Kai/Elle

Skip for now: extension ladder, dashboards, Life Engine, wiki compiler, running GBrain in parallel.

  1. Supabase project — free tier, pgvector, OB1 schema migration. Cloud only. No Mac DB.
  2. OB1 core — follow official getting-started (~45 min). Get capture + search working in one client first.
  3. Wire both mouths — same remote MCP URL on Hermes and Mac Claude Code. Smoke: remember on Telegram → recall on Mac (Mac asleep OK).
  4. Habit law — durable fact/decision → OB1 capture. Human prose → vault. Phone → Telegram Donna. No dual-write policy.
  5. Cutover — stop writing agent memory to holographic/fabric/MEMORY-as-brain. Keep vault. Optional vault import batch later.

1. Fact captured on Telegram is found on Mac Claude with source.
2. Mac powered off does not break recall on Telegram.
3. You’re not maintaining a second agent memory store.

No redesign of schema “for us.” No OB1+GBrain+Zep sandwich. No vault replacement. No dashboard before daily search hurts less.

◆ Entry #005 · Strategy Report Jul 21, 2026
One Human Door + One Agent Door — Justified
Full artifact: out/one-human-door-one-agent-door-justified-2026-07-21.md · builds on Entries #003–#004 + OB1 SOTA

You do not need another second-brain product. You need one human thinking surface and one agent retrieval layer — then stop collecting memory systems.

Personal win: zero re-explain across Telegram, Mac Claude, and phone Obsidian, so presence stays on Kai and Elle, not on re-briefing Donna.

SurfaceMeasured 2026-07-21Implication
Vault3,588 md · 247MB · VPS ROHuman SSOT already exists
MEMORY.md / USER.md816 B / 2.3 KB · caps 2.2/2.5KTiny always-on injection only
Memory providerholographicHermes-local, not Mac-shared
MCP serversagentmail, altfins, x, apifyNo shared brain MCP
Write lane/opt/data/out → Mac inboxDelivery works; fusion does not
Candidate brainsOB1, GBrain, Zep, Mem0…Research done; install not justified yet

Plus overlapping agent memory already live: holographic facts, fabric, session_search, vault grep, skills, out artifacts. That is the silo problem Entry #004 named.

C1. No new second-brain product. Lifestyle cap + vault already is the human PKM. Nick/Cerebras: unused aesthetics are theater. Failure mode is retrieval silos, not missing Notion.

C2. Dual door, not mono store. Human door = vault (edit, phone sync, doctrine). Agent door = one MCP retrieval layer. Agents alone → slop. Humans alone → unqueryable across sessions.

C3. Zero re-explain is the product. Mac Claude ≠ Hermes holographic store. No brain in mcp_servers. session_search is Hermes-local. Every surface switch re-taxes LMS facts, constraints, open loops.

C4. KPI = re-explain minutes, not note count. More notes do not lower cognitive load. Capture-once + sourced retrieve + same agent door on Mac/VPS do.

C5. Presence tax. Priority stack is Kai/family first. Minutes recovered from hunting decisions are the lifestyle doctrine working, not soft wellness.

C6. Solo ops without staff. You will not manage people. Shared agent memory is how SOP/client state persists at lifestyle scale.

C7. Stop collecting brains. Seven overlapping stores. Each new layer without a kill criterion raises miss rate and token bleed (already flagged on interactive mega-sessions).

C8. P0 before OB1/GBrain. Fan-out + citations falsify “need a product” cheaply. Install-first creates an eighth silo. Infra law: durable, frugal, no compose roulette.

C9. GBrain lean if P0 fails. Prior analysis: Hermes-native markdown + hybrid/RRF. OB1 wins gentler setup/importers. Preference ≠ install order.

C10. Vault is not replaceable. Years of domain doctrine. OB1 community treats Obsidian as import source. Phone path already works.

PushbackReply
“Just use Obsidian better”Human door is fine. Agent door is the hole.
“Install OB1 tonight”Seventh→eighth silo. P0 is cheaper falsification.
“Zep is SOTA graph”Overkill before daily fusion proves value. Sequence after stable loops.
“Put everything in Postgres”Breaks Glass House + phone sync. Extract, don’t migrate.
“One store only”Agent-optimized slop vs human-optimized unqueryable. Dual-door is stable.

You: Telegram (ops) · Mac Claude (build) · Obsidian (think/phone).

Agent door (ONE): shared retrieval over vault + sessions + captures, citations mandatory.

Not homes: ChatGPT memory, random Notion, third graph “for later.”

  1. P0 (7 days, no new product): one capture policy · kb_search fan-out · citation law · 3 daily surfaces only.
  2. P1 (trigger): same fact re-explained Mac↔Telegram twice in a week → install ONE of GBrain/OB1, vault import, both clients on same MCP.
  3. P2: demote holographic/fabric to feeds · new memory tool must kill an old one · no dashboard until felt daily.
MetricTarget
Mac↔Telegram re-explain≤1 / week
Memory answers with sources≥90%
New memory systems added0 unless P1 trigger
Vault remains human SSOTyes

Subtraction plus one fusion path — not acquisition. Human door: vault. Agent door: thin retrieval first, product only on failed P0. Kill criterion on every new memory idea.

◆ Entry #004 · Delta Jul 21, 2026
Second Brains Died. Retrieval Won.

Nick's open: almost every "knowledge base / second brain" has been hot air — until Cerebras shipped an internal system staff actually hammer (15,000+ queries/day in ~3 months), used by humans, automations, and agents.

The kill shot is not a prettier Obsidian theme. It is rejecting the single-source-of-truth migration fantasy. Meet data where it lives (Slack, GitHub, docs, Jira, custom DBs) → unify at the embeddings/query layer → hybrid retrieve → plan/execute tools in parallel → RRF + rerank → cited synthesis. Expose the same retrieval as MCP primitives so agents orchestrate without a hidden second brain LLM.

Evidence note: full captions blocked (Apify floor + SOCKS offline + YT bot wall). Delta grounded on indexed transcript open + full Cerebras architecture digest + live Donna memory audit. Artifact: out/cerebras-second-brain-nick-saraev-transcript-2026-07-21.md.

Ingest: one Postgres embeddings table; connectors via PR; Slack Socket Mode real-time; LLM distillation before embed (question/summary/resolution/systems); "bursting" for long threads; CocoIndex incremental code chunking (to 40GB+ repos).

Retrieve: four-signal Slack hybrid (full-text + embedding + IDF + age decay) because vectors alone promote filler. Query path: planner → parallel tools → RRF (k=60) → reranker top-10 → context expand → synthesis with citations.

Scope: Projects bundle channels/repos/docs per team so compiler engineers do not get DC runbooks. Agents: Web UI runs full pipeline; MCP exposes LLM-free search_slack / search_code / who_knows.

CapabilityWe HaveSource SaysGap
Meet data where it lives (no forced SSOT)⚠️ Multi-store by accident (memory/fact/fabric/vault/out)Extract in place by designMedium
Unified embeddings + common schema❌ Siloed stores, no shared embed tableOne Postgres embeddings tableBIG
Continuous multi-source ingest⚠️ Manual delta-capture / ori_add / skill_manageSlack/code/docs always-onBIG
Hybrid retrieval (FTS+vector+IDF+age)⚠️ FTS (session_search/fabric) + ranked recall, no fusionFour-signal hybrid + RRFBIG
LLM distillation on ingest⚠️ Delta pipeline / wiki ingest patternEvery Slack thread distilledMedium
Planner → parallel retrieval tools✅ Parallel tools + delegate_taskParallel fan-out executorNone
RRF + reranker + cited synthesis❌ Model synthesizes from whatever landed in contextFormal rank fusion + citationsBIG
MCP retrieval primitives for agents⚠️ MCP servers exist; no KB tool surfacesearch_slack/code/who_knowsMedium
Project-scoped knowledge packs⚠️ Profiles + vault domains, not query scopesProjects at onboardingMedium
Personal wiki / second brain discipline⚠️ llm-wiki skill + Obsidian vault RO + out/Nick: pure PKM theater fails at scaleSmall*

*Small for lifestyle-cap personal use — Dwad is not Cerebras. The transferable win is retrieval design, not cloning a 15k Q/day internal platform.

Do not rebuild Notion inside Hermes. Do not declare the vault "dead." Do adopt Cerebras's three laws at personal scale:

1. Extract, don't migrate — Telegram, Gmail, vault notes, session DB, X already exist. Ingest summaries into one queryable layer.
2. Hybrid rank, don't vibe-retrieve — stop hoping one store answers; fuse FTS + semantic + recency before the model speaks.
3. Agent-facing primitives — retrieval tools agents can call without a second hidden brain.

This upgrades Entry #003's Obsidian bonus from "browser save to vault" to "vault is one source among many in a fused index."

  1. P0 — One query surface over existing stores. Ship a kb_search (or skill) that fans out to session_search + fabric_recall + fact_store + vault grep, returns a merged top-N with source tags. No new DB day one — fusion in code.
  2. P0 — Citation law on memory answers. Any answer that used memory must list which store + which fact/file. Stops silent vibe-recall.
  3. P1 — Distill-on-ingest for high-value sources. When a delta/video/email lands in /opt/data/out, auto-write a normalised card (question/summary/resolution/tags) into a single kb/ markdown index the searcher prefers.
  4. P1 — Project scopes. Map Lane domains (Family, LHI, fCAIO, Trading deferred) to default source packs so queries do not cross-contaminate.
  5. P2 — True hybrid rank. Add recency decay + simple RRF across the fan-out results. Measure with 20 canned questions Dwad actually asks.
  6. P3 — MCP kb tools. Expose search primitives to other agents only after P0/P1 prove daily use. Do not build Cerebras-scale Slack Socket Mode for a single-user OS.
◆ Entry #003 · Delta Jul 21, 2026
Hermes 10X — Five Upgrades, Six Real Gaps

Jack's claim is not "new Hermes concepts." It is that five already-shipped platform capabilities now compound into an agentic OS: swap brains, parallel tools, Firecrawl research, daily ops surface, and completion contracts — with Obsidian memory + skill-from-docs as the bonus layer.

For Donna this is a wiring audit, not a roadmap fantasy. Hermes 0.18.1 on this VPS already has Grok 4.5, parallel tool guidance, delegate_task, native /goal completion contracts, gws email/calendar, and skill writing. The gap is operational adoption + a few missing lanes (Kimi, Firecrawl budget, browser→vault).

Evidence note: full captions blocked this session (Apify usage floor + Mac SOCKS offline + YT bot wall). Delta grounded on full author description + chapter stamps + indexed transcript fragments + live config audit. Source artifact: out/hermes-10x-jack-roberts-transcript-2026-07-21.md.

L1 Swap Smarter Brains — GPT 5.6 via ChatGPT sub, Grok 4.5 cheap/fast + live X intel, Kimi K3 design at a fraction of Opus, one-prompt website test, price breakdown.
L2 Stop Waiting — native parallel tool calls; demo runs four tasks at once.
L3 60× Faster Research — Firecrawl as the research path; strips junk HTML; brand-identity extraction; model showdown (~60× faster / ~49× cheaper claim in fragments).
L4 Daily Edge — email + calendar connected; permission safety before acting; a brief that improves itself.
L5 Proof of Work — completion contracts: define verifiable DONE up front; agent keeps going until evidence matches.
Bonus — save anything to Obsidian from the browser, recall it later, turn docs into durable skills.

CapabilityWe HaveSource SaysGap
Grok 4.5 primary + X intel✅ Live (xai-oauth)Fast/cheap brain + XNone
Multi-model swap (Kimi / GPT lane)⚠️ DeepSeek fallback onlyKimi K3 design + GPT 5.6 free laneMedium
Parallel tool calls✅ Guidance on + native runtimeFour tasks at onceNone
Subagent fan-out✅ delegate_task (max 10)Parallel workstreamsNone
Firecrawl / fast clean research⚠️ Key present, chronically 40260× path, brand extractBIG
Email + calendar connected✅ gws (draft-only mail)Daily ops surfaceNone
Permission safety before external acts⚠️ Policy in SOUL, not a hard gateExplicit permission ruleSmall
Self-improving daily brief⚠️ Cron briefs existBrief that compoundsMedium
Completion contracts (/goal)⚠️ Native in 0.18.1, unused by DonnaProof-of-work DONE checksBIG
Browser → Obsidian real-time capture❌ Vault RO + out/ lane onlySave anything from browserBIG
Compounding memory / recall⚠️ memory + fact_store + OriWiki-style compounding memoryMedium
Build skills from docs✅ skill_manage + delta pipelineDocs → durable skillNone

Already done: Grok 4.5, X search, parallel tools, subagents, gws calendar/mail draft, skill writing.
Partial / unused: completion contracts exist but Donna does not run work through /goal; briefs cron but do not self-score; memory is multi-store not one compounding wiki.
Real holes: Kimi design lane, Firecrawl budget reliability, browser→Obsidian capture path.

  1. P0 — Put completion contracts on the critical path. For any multi-step Donna job, draft a /goal contract (outcome + verification + constraints) before execution. Stop accepting "looks done" without evidence checks. This is the same muscle Entry #001/#002 asked for — now native in Hermes.
  2. P0 — Restore research path budget. Firecrawl is named as the 60× unlock and is the tool we keep dying on. Either top up credits, pin Exa as primary extract backend with a working key, or codify Jina/curl as the default so "research" never blocks on 402.
  3. P1 — Kimi K3 design lane (optional, cheap). Only if design/one-prompt site work becomes recurring. Wire via OpenRouter or Moonshot; keep Grok as default chat brain. Do not multi-model for sport.
  4. P1 — Self-improving brief loop. Nightly brief writes a scorecard against yesterday's brief (what was wrong/missing) into the next prompt. Closes L4 without new infra.
  5. P2 — Browser → vault capture. One command/path: selected page/text → normalized note in /opt/data/out (syncs to Mac vault). Not full bidirectional Obsidian MCP day one — just the capture habit Jack demos.
  6. P3 — Permission hard-gate checklist. Before any external side effect (send, post, pay, delete), force an explicit confirm step that cites the permission rule. Policy already exists; make it mechanical.
◆ Entry #002 · Delta Jul 18, 2026
Evaluator Cartridges & Autonomous Loops

Karpathy's auto-research pattern is now a native Claude Code feature/goal and /loop. The new unlock: evaluator cartridges — custom scoring algorithms that grade AI output autonomously. The combination — goal-driven agent + multi-cartridge evaluator + attempt budget — produces output that iteratively sharpens without you.

This is Entry #001's loop engineering with the inspector automated. The human graduates from checking every output to defining the scoring rubric once.

A cartridge is a self-contained scoring algorithm: output → number (1–10). Three shown:

1. Humanify Cartridge — voice, variety, predictability (AI slop detector). Only 9+ passes.
2. Hormozi AI Cartridge — scored against Alex Hormozi's public writing for marketing punch.
3. Open Rate Analytics Engine — predicted open rate from past email data + audience behavior.

Composite scoring across cartridges: newsletter required 27/30 total, no single cartridge below 8. Agent ran 5 attempts autonomously: 7.1 → 26.8 over ~40 minutes.

CapabilityHermes HasSource SaysGap
Autonomous goal-driven loops⚠️ Manual pattern/goal nativeBIG
Evaluator cartridges (scoring)❌ NoneMulti-cartridgeBIG
Condition-based wake-up⚠️ Cron (schedule)/loop (condition)Medium
Subagent swarm → single goal⚠️ Independent tasks650 experimentsMedium
Attempt tracking + scoring history⚠️ STOP AFTER onlyScoring tableMedium
Scoped file access for loops❌ Full FSScoped accessSmall
  1. P1 — Humanify cartridge. Standalone scorer script: text → 1-10 on voice, variety, predictability. Beachhead to prove the pattern.
  2. P1 — Autonomous retry in loop-engineer. Upgrade from manual inspection to subagent self-scoring against checklist, retry until pass or budget.
  3. P2 — Multi-cartridge composite. Run N cartridges, sum scores, enforce per-cartridge minimums.
  4. P2 — Scoring table in cron output. Per-attempt scoring history so Dwad sees improvement trajectory.
  5. P3 — Condition-based cron wake-up. Extend cronjob: "run every N minutes UNTIL condition met."
  6. P4 — Swarm coordinator. N subagents optimize toward shared best-score state.
◆ Entry #001 · Delta Jul 16, 2026
Loop Engineering — The End of One-Shot Prompts

Prompt engineering is dead. The new paradigm is loop engineering: AI tries → self-checks → fixes → retries → marks done. You move from being in the loop to being outside it. Four components: Goal, Checklist, Inspector, Budget.

Inner loop: Single task — try, inspect, fix, retry. AI self-corrects before delivering.
Outer loop: Recurring scheduled job running inner loops each iteration. Self-improves over time.

CapabilityHermes HasLoop Eng.Gap
Scheduled recurring tasks✅ 17 cronsOuter loopNone
Self-review within task❌ MissingInner loopBIG
Inspector/second pass❌ MissingGrader modelBIG
Hard definition-of-done⚠️ Soft goalsChecklistMedium
Skills auto-update from loops❌ MissingLoop→skillMedium
Per-task attempt budget⚠️ Repeat onlyStop after NSmall
  1. P0 — Inner loop skill. Created loop-engineer skill: GOAL + DONE WHEN + STOP AFTER → delegate → inspect → retry up to N.
  2. P0 — Inspector pass. Five-point checklist before subagent output reaches Dwad: exists, meets checklist, no hallucinations, correct format, sniff test.
  3. P1 — Upgrade crons. DONE WHEN checklists + STOP AFTER budgets on quality-critical cron jobs.
  4. P2 — Skills auto-update. LLM-driven cron prompt: "If you discover anything that improves the skill, patch it before delivering."
  5. P3 — Goal template library. Reusable DONE WHEN checklists for recurring task types.