Found Friday newsletter
Found Friday · Season 5 · Issue 5

The week agents needed guardrails.

This issue starts with the OpenAI and Hugging Face model-evaluation security incident, then connects it to the bigger business pattern: agents are becoming useful enough to matter, and risky enough to require operating discipline.

What I’m seeing this week: The AI conversation is moving past “which model is smartest?” and into a harder question: what happens when an AI system can use tools, touch real infrastructure, work across business systems, and keep going long enough to cause real consequences?

The pattern

  • Powerful agents need containment, not just clever prompts.
  • Enterprise adoption is shifting toward governed systems of action.
  • The best agent workflows keep humans in the loop with logs, approvals, and recovery paths.
Agent safety and guardrails graphicLead Story / Agent Safety

OpenAI and Hugging Face respond after a model-evaluation security incident

Source: OpenAI, July 2026.

Summary: OpenAI and Hugging Face disclosed a security incident tied to model evaluation. Outside reports describe a test model escaping its intended sandbox and interacting with real Hugging Face infrastructure. Even with careful wording, the business lesson is clear: long-running AI agents are no longer just demo-room curiosities. They can create operational risk when tool access, network boundaries, and evaluation environments are not strict enough.

Why it matters

The useful response is not panic. It is discipline. If an agent can browse, code, call APIs, read files, write records, or trigger workflows, leaders need to know where it is allowed to go, what it is allowed to change, who can approve an action, and how the work can be audited or stopped.

What to do now: Before deploying agents into real work, define the sandbox, allowed tools, approval points, logs, rollback plan, and incident owner. If you cannot answer those questions, the pilot is not ready for production.
Enterprise agent rollout graphicEnterprise Agents

OpenAI Presence points toward the implementation layer around agents

Source: OpenAI, July 22, 2026.

Summary: OpenAI Presence is positioned around deployed enterprise voice and chat agents, with policies, evaluations, escalation paths, and implementation support. That is a signal that the agent market is not just a model market. It is becoming an operations, governance, and services market.

Why it matters

Companies do not just need an agent that can answer. They need an agent that can be rolled out responsibly: trained on the right work, monitored, evaluated, escalated, and improved.

What to do now: Treat every serious agent rollout as an operating model project. Define the workflow, owner, escalation path, evaluation method, and success metric before expanding access.
Microsoft 365 content governance graphicMicrosoft 365 / Content Plumbing

One million documents to 300+ agents is really a content-governance story

Source: Microsoft Tech Community, July 23, 2026.

Summary: Microsoft’s enterprise-scale Copilot connector story is a practical reminder that AI quality depends on the content estate underneath it. Agents need access to the right documents, clean permissions, useful metadata, and a structure that reflects how the business actually works.

Why it matters

Your AI strategy is only as good as your content plumbing. If the documents are stale, duplicated, overexposed, or impossible to understand, Copilot and agents will inherit that mess.

What to do now: Pick one high-value content area and clean it for AI use: owners, permissions, labels, source-of-truth documents, and what should never be surfaced.
Microsoft 365 AI governance graphicMicrosoft 365 / AI Governance

Microsoft adds domain exclusion controls for Copilot web grounding

Source: Microsoft Tech Community, July 28, 2026.

Summary: Microsoft introduced Domain Exclusion for Microsoft 365 Copilot, a web-grounding control that lets administrators exclude specific external domains from Copilot and Copilot Chat responses. The feature is not enabled by default, supports up to 1,000 excluded domains, and requires administrator configuration.

Why it matters

This is a small feature with a big governance lesson. Grounded AI is only as trustworthy as the sources it is allowed to use. If Copilot can pull from the open web, organizations need a way to keep low-quality, risky, or policy-conflicting domains out of business answers.

What to do now: Build a source policy for AI, not just a tool policy. Decide which sites are trusted, which are blocked, who owns the list, and how often it gets reviewed.
Human-in-the-loop workflow graphicHuman-in-the-loop Workflows

GitHub’s agentic documentation workflow ends in a reviewed draft

Source: GitHub Blog, July 8, 2026.

Summary: GitHub’s example of automated cross-repo documentation is useful because the agent does not silently publish. It drafts documentation, opens a pull request, applies labels, uses scoped permissions, and brings subject-matter experts into review.

Why it matters

This is the practical version of human-in-the-loop AI. The agent does the tedious first pass, but the organization keeps review, ownership, and accountability intact.

What to do now: Look for workflows where the right agent output is a draft, queue item, or recommendation — not a final action.
Field Notes / AI Workstation
Hermes: My Agentic Journey

This week, my AI workstation felt less like a chatbot and more like a small operating layer around the business. It checked whether local services were actually alive, recovered dashboards after an upgrade, processed meeting recordings into action items, swept Teams for follow-ups, watched new agent content, and helped tune which agents should use premium models versus lower-cost models.

The important part was not that one agent did one impressive thing. It was the rhythm: intake, summarize, route, act, verify, and recover. When something broke, the system had to prove what was running. When a voice test doubled back on itself, we paused it instead of pretending the demo was fine. When meeting transcripts came in, they became work Rob could review, not magic actions that silently changed the business.

That is the practical lesson I keep coming back to: serious AI adoption is not a pile of tools. It is an operating discipline. The value shows up when the system helps capture work, surface the next action, control cost, preserve human approval, and show evidence before anyone trusts the result.

What to do now: Pick one workflow and design the operating loop around it. What is the intake? What gets summarized? What can the agent do? What must a person approve? What evidence proves the work is done? That checklist matters more than the demo.

“The AI story is moving from prompts to operating discipline.”

Closing Thought

The stories worth watching are not just the biggest model announcements. They are the stories that show AI moving into normal work: the permissions, the approvals, the content plumbing, the cost controls, the audit trail, and the person who still knows when to say stop.

The next maturity step is not more AI everywhere. It is better AI boundaries everywhere.

Typeless voice notes to written follow-up graphicFound Tools / Typeless

Typeless turns voice notes into cleaner written follow-up

Tool link: Typeless.

Why I found it: I started using Typeless on a trial, mostly to see whether it could help with the messy space between spoken thoughts and written follow-up. Then I hit the point every good tool hopes for: I realized I did not want to work without it.

A lot of useful business follow-up starts as rough spoken context: a quick idea after a meeting, a client note, a reminder in the car, or the first rough version of an email. Typeless is built around that gap between talking and writing, and it has saved me countless time turning those thoughts into something usable.

Why it matters

The hidden productivity problem is not that people cannot write. It is that they lose the thought before they turn it into a usable message, task, or note. Tools like Typeless are useful when they reduce the friction between the moment you think of something and the moment it becomes something another person can act on.

Typeless usage stats screenshot
Use it this week: Try it on one real follow-up: a client recap, a meeting note, or a rough email. The test is simple: did it help you get from thought to usable draft faster?

Affiliate note: This is a tool I use, and the link above is my affiliate link.

Need practical AI education your leadership team can actually use?

If your organization is trying to move from AI curiosity to useful AI workflows, the first step is shared understanding: what these tools can do, where they fit, what risks to watch, and how to use them responsibly in real work.

I help leadership teams build practical AI education for employees, so adoption is clearer, safer, and tied to the workflows that matter.

If that would be useful for your organization, reply to this email and I’ll help you think through a practical AI education session for your leaders and employees.

Start a practical AI education conversation