Hands sketching a workflow diagram beside a laptop.
The operating layer: human oversight over memory, skills, todos, voice, channels, and the approval boundary. Photo via Unsplash (Unsplash License).
Hermes Journey #2

The Operating Layer: Memory, Skills, Todos, Voice, Channels, and the Approval Boundary

Second in the Hermes Journey series.

Hermes Journey · Operating Layer

A behind-the-scenes look at the plumbing that made Storm useful: memory, skills, task tracking, voice, channels, dashboards, and the approval boundary that keeps the system accountable.

Practical takeaway

An AI assistant becomes trustworthy when its memory, tasks, channels, and permissions are visible — and when the human approval line is designed into the workflow instead of patched on later.

A digital chief of staff sounds glamorous. The truth is, the useful 90% of it is plumbing — and it’s plumbing nobody posts about.

In the first post, I explained why I stopped chasing AI demos and started building a system I could actually work with. Now I want to take you inside the operating layer: the parts that made the difference between “impressive agent” and “assistant I’d genuinely miss.”

Because here’s what I’ve learned: the model isn’t the product. The system around the model is. Any capable agent can answer a clever question. Very few can remember your preferences, track your open threads, reach you on the channel you actually check, and know when to stop and ask before touching something important. That last piece — the stopping and asking — is the whole ballgame.

Memory

Memory: the thing everyone ignores

Memory is the single most underrated part of building an assistant that feels like yours.

A generic chatbot starts from zero every conversation. That’s fine for a search box. It’s useless for a colleague. So the first thing I built was a way for Storm to remember — not just facts, but how I like to work: how I prefer tasks structured, what standards I hold for the work I put my name on, and what’s happened recently that’s still relevant. When Storm starts a day already knowing the context, it stops feeling like a machine and starts feeling like a teammate who got the briefing.

The honest catch: memory only helps if it’s reliable. A system that thinks it remembers and gets it wrong is worse than one that admits it doesn’t know. So a big part of the work was teaching it to be honest about uncertainty — to check rather than guess.

Skills

Skills: teaching it to actually do things

Memory is what the assistant knows. Skills are what it can do.

A lot of agents come with impressive general abilities, but they’re not trained on your routines. So I started encoding the recurring things I do as repeatable skills — the steps, the checkpoints, the gotchas I’ve hit before, the way I want the result to look.

Here’s the pattern that changed everything: instead of giving the agent an instruction every time, I gave it a reusable procedure it could pull up when a task matched. Over time, the assistant gets better not because it’s a newer model, but because it has more experience written down and structured. That’s compounding improvement that has nothing to do with model releases and everything to do with how you use the system. It’s the difference between hiring a smart person and hiring a smart person who’s also been writing things down for years.

Todos + recovery

Todos and recovery loops: where it actually earns its keep

Now the part that made me a believer: task tracking with a recovery loop.

Here’s the reality of any working system, AI or human: things get missed. The value isn’t in never missing anything — it’s in catching the miss quickly, explaining it, and fixing the process so it doesn’t happen again.

So I built task tracking not as a static to-do list, but as a loop. The system keeps track of what’s open, what’s due, and what I said I’d follow up on. When something slips, the goal isn’t to pretend it didn’t. It’s to surface it, tell me what happened, and then — the part most people skip — improve the underlying process so the same miss is less likely next time.

That single mindset shift, from “tracking tasks” to “closing the loop on misses,” is worth more than any fancy feature. A to-do list tells you what’s left. A recovery loop makes the system smarter every time something goes wrong. It’s the difference between a ledger and a lesson.

Voice + channels

Voice and channels: meeting people where they are

A tool you have to remember to open is a tool you won’t use. So I made the system reachable the way I actually work — including voice, for the moments when I’m moving or my hands are busy, and across the communication channels I live in day to day. Make it meet you where you already are, and it becomes part of your rhythm instead of another tab.

Central Command

Central Command: one place to see the whole picture

This is where it started to feel real.

As I added memory, skills, todos, and channels, the risk was fragmentation — little islands of automation that didn’t talk to each other. So I built what I’ve come to call a Central Command dashboard: a single view of what the assistant is working on, what’s pending, what’s done, and where things stand across the whole system.

The dashboard does two jobs. First, situational awareness at a glance — I can see the shape of my workday without digging through a dozen logs. Second, and more importantly, it makes the system auditable: I can see what the assistant did, what it decided, and why. That visibility is what makes trust possible at all.

The approval boundary

The approval boundary: the part that makes it safe

And now the most important piece — the one that lets you sleep at night.

An autonomous assistant is powerful. An autonomous assistant with no boundaries is a liability. So the core design principle of everything I’ve built is a clear, enforced line between what the system may do on its own and what must always come back to a human first.

Think of it as the difference between an employee you’d trust with the coffee order and one you’d want to approve before they sign, send, or spend anything. The assistant handles routine, low-risk work freely. But anything with real consequence — anything that goes out under my name, anything that changes something significant, anything costly or embarrassing if wrong — stops and waits for me.

This boundary is what turns a demo into a tool you actually rely on, and it’s the part that gets skipped in most AI enthusiasm because it isn’t flashy. But it’s the difference between automation that works for you and automation that works on you.

That is the real job of the approval boundary: it keeps the system useful without asking me to surrender judgment.

The through-line

The through-line

None of these pieces is a silver bullet. Memory without recovery loops is trivia. Skills without a dashboard are hidden. Channels without boundaries are chaos. They only earn their keep as a system — working together, visible, accountable, and gated at the human line.

In the next post, I’ll take this outside the private system — and into what it means for people who aren’t building this for themselves, but need the result.

Ready to build your own operating layer?

If you’re trying to build your own operating layer and feel like you’re re-inventing the wheel, you’re not alone — and you don’t have to do it in the dark. I’m writing the Hermes Journey in the open because the practical, honest stuff is exactly what’s missing from the AI conversation. Follow along as I go public, and if you’re a consultant, owner, or team lead trying to figure out the right approval boundaries for an AI assistant, reach out — let’s talk about what should be automated and what should always come back to a human first.

Explore the Hermes Journey →

Continue the Hermes Journey

Start with the first entry, then return to the Hermes Journey landing page as the rest of the series goes live.