Found Friday newsletter
Found Friday · Season 5 · Issue 4

The AI Horse Race Is Becoming an Operations Problem

OpenAI, Claude, Copilot, Gemini, and Grok all made moves this week. The business lesson is not hype. It is operational discipline.

Opening Note from Rob

A lot happened in AI this week.

OpenAI is pushing the GPT-5.6 line forward. Anthropic had a very public back-and-forth around Fable 5 access. Microsoft is making model choice part of Copilot. Google keeps putting Gemini directly into everyday Workspace tools. xAI is pushing Grok further into coding, agents, and voice.

That is exciting, but it is also where companies can get sloppy.

If AI is just a chat window, the blast radius is smaller. Once it starts touching files, email, calendars, code, spreadsheets, customer workflows, and internal knowledge, the conversation changes.

You need adoption. You also need boundaries.

This week’s issue is about the shift from “AI tool” to “AI operating layer.”

What I’m seeing this week

The AI platforms are all converging on the same idea: they do not want to be places where you ask questions. They want to be places where work gets done.

That creates opportunity for small and mid-sized businesses, but only if leaders slow down long enough to ask the practical questions:

  • Which tasks are safe to delegate?
  • Which require human approval?
  • Which model is right for which type of work?
  • What happens if the model changes, access is paused, or pricing shifts?
  • Where does the work get logged?
  • Who owns the outcome?

The companies that win with AI will not be the ones that chase every release. They will be the ones that turn these tools into governed, repeatable workflows.

Featured Stories

AI model routing and cost-control graphicOpenAI / Model Routing

OpenAI previews GPT-5.6 Sol as the next reasoning and agentic model line

Source: OpenAI and Neowin

Summary

OpenAI has previewed GPT-5.6 Sol as a next-generation model, with current reporting around a broader GPT-5.6 line that includes Sol, Terra, and Luna. The positioning is clear: higher-end reasoning and agentic work at the top, with lower-cost variants for more routine use cases.

Why it matters

This is where AI budgeting starts to get real. Businesses should not treat every task like it deserves the most expensive model. A high-stakes contract review, a multi-step research workflow, and a simple internal summary should not all be routed the same way.

The model race is becoming a routing problem.

What to do now:

Create a simple model-use map. Put your AI workflows into three buckets: low-risk routine work, medium-risk business work, and high-risk work that requires review. Then decide which model class belongs in each bucket.

AI access and model-lifecycle risk graphicAnthropic / Access & Fallback Risk

Anthropic’s Fable 5 access whiplash is a governance lesson, not just model news

Source: Anthropic

Summary

Anthropic says Claude Fable 5 and Mythos 5 were suspended after U.S. export controls were applied in June. Then Anthropic announced that export controls on Fable 5 had been lifted and that Fable 5 access would be restored globally starting July 1 across Claude Platform, Claude.ai, Claude Code, and Claude Cowork. Mythos 5 access was restored for a more limited set of U.S. organizations while broader access was still being coordinated.

Why it matters

This was the tennis match of the week: available, restricted, restored, still partly limited.

The business lesson is simple. Even the best model is not a stable operating dependency unless you have a fallback plan. Access can change because of policy, geography, vendor decisions, pricing, capacity, or compliance. If a workflow matters, you cannot build it on hope.

What to do now:

For every important AI workflow, write down the fallback. Which model would you switch to? Who approves the switch? What work pauses if access disappears? If you cannot answer that, the workflow is not operationally mature yet.

Claude agentic work graphicAnthropic / Agentic Workhorse

Claude Sonnet 5 is being positioned as the practical agentic workhorse

Source: Anthropic

Summary

Anthropic announced Claude Sonnet 5 as a model focused on coding, agents, and professional work at scale. Anthropic says Sonnet 5 is available across all plans, is the default model for Free and Pro plans, and is available in Claude Code and through the Claude Platform with introductory API pricing through August 31, 2026.

Why it matters

Most businesses do not need a leaderboard champion for every task. They need a dependable workhorse model that can handle documents, coding, structured work, and tool use without blowing up the budget or requiring a specialist at every step.

That is why Sonnet 5 matters. It is part of the move from “impressive demo” to “reliable daily worker.”

What to do now:

Pick one real workflow and test it against your current AI setup. Score the result on output quality, review time, cost, and how often a person had to rescue the process. That is more useful than arguing about model rankings.

Microsoft 365 governance graphicMicrosoft 365 / Model Governance

Microsoft 365 Copilot is turning model choice into a tenant-level decision

Source: Microsoft 365 Roadmap

Summary

Microsoft’s current Microsoft 365 roadmap lists “Available today: Anthropic’s Claude Sonnet 5 in Microsoft 365 Copilot” dated July 2. The same roadmap continues to frame Copilot Cowork as long-running, multi-step work inside Microsoft 365.

Why it matters

This is a big Microsoft 365 story. Copilot is no longer just “Microsoft’s AI assistant.” It is becoming a governed work surface where multiple models may be used inside the tenant.

That means model selection becomes an IT, security, compliance, and training issue. If people can choose different models inside Copilot, leaders need to understand what those models are for, where data goes, and how usage is controlled.

What to do now:

Review your Copilot governance plan. Add a section for model selection: who can choose alternate models, what types of work they are approved for, and what sensitive content should stay out of experimental workflows.

Agentic cowork graphicAgentic Platforms / Cowork

Copilot Cowork and Claude Cowork show where agentic work is headed

Source: Microsoft 365 Blog and Claude

Summary

Microsoft has made Copilot Cowork generally available, and Anthropic is expanding Claude Cowork beyond a single desktop workflow. Both moves point in the same direction: AI systems that can carry work forward across sessions, files, tools, and business context.

Why it matters

This is where AI gets useful and risky at the same time.

A chatbot helps you think. A coworking agent starts to act. That means approvals, logging, permissions, recovery, and human review matter a lot more than prompt style.

What to do now:

Choose one safe delegated workflow before trying ten risky ones. Define the task, the data it can touch, the approval point, and the recovery step if the output is wrong. Start with boring controls before scaling the exciting part.

Gemini workflow and model-lifecycle graphicGoogle / Gemini Workflows

Google keeps putting Gemini into the normal places people already work

Source: Google Workspace Updates and Gemini API Release Notes

Summary

Google continues to expand Gemini inside Workspace and developer surfaces. Recent updates include Fill with Gemini in Sheets expanding to additional languages, Ask Gemini and AI Overviews moving further into Drive and mobile workflows, and Gemini API updates such as Gemini Omni Flash in public preview.

Why it matters

Google’s strongest business story may not be a single giant model headline. It is that AI keeps showing up inside spreadsheets, files, presentations, mobile apps, and everyday work surfaces.

That is how adoption really happens. Not through a big AI strategy deck, but through someone cleaning up a spreadsheet, searching Drive, drafting a slide deck, or summarizing work from their phone.

What to do now:

Pick one Workspace workflow and define the guardrails. What content can Gemini use? What must be reviewed? What admin settings need to be checked? Practical adoption starts with one controlled workflow, not a company-wide “everyone go use AI” announcement.

Voice agent workflow guardrails graphicVoice Agents / Grok

xAI pushes Grok further into coding, agents, and voice

Source: xAI and xAI Grok Voices

Summary

xAI announced Grok 4.5 on July 8, positioning it around coding, agentic tasks, and knowledge work. xAI also announced 21 new flagship Grok voices on July 6, expanding the voice-agent side of its platform.

Why it matters

Voice agents are becoming a real front door for work, not just a novelty feature. The more natural the interface becomes, the easier it is for people to ask the system to do things without thinking through the operational consequences.

That is where governance has to follow the interface.

What to do now:

If you test voice agents, test the handoff, not just the voice. Where does the transcript go? Can a person interrupt? Is the action logged? What happens if the agent misunderstands the request? Those questions matter more than whether the demo sounds impressive.

AI workstation field notes graphicField Notes / AI Workstation

Field Notes from My AI Workstation

I’m learning that the real value in my AI workstation is not just having more models, agents, or automations. The useful part is when the system starts behaving more like an operating rhythm for the work itself.

This week, that showed up in a few practical ways. Daily accomplishment summaries now capture what actually happened so the work does not disappear into chat history. A failed 5am maintenance job was not just ignored. Storm traced the failure to an oversized backup step, changed the routine, verified the smoke test, and documented the fix. Victor finds potential companies, Nikki now picks up five a day for deeper research, and Plaud conversations are being turned into routed notes and follow-up tasks.

That is the part of agentic work I think matters for businesses.

One more operating lesson is becoming clear: completion gates matter. If different groups, agents, or workstreams each have different tasks, they need their own checklist. The system should not move forward just because a general task list looks active. It should wait until the required item for that group is actually checked off.

The headline is not “AI can do tasks.” The headline is that useful AI systems need memory, handoffs, recovery loops, review points, and a clear human approval boundary.

A workflow that runs once is interesting. A workflow that can notice what happened, summarize it, pick up the next step, and recover when something breaks is much closer to a business system.

That is where I think companies should focus. Not on collecting more tools, but on building operating discipline around the tools they already have.

“The model race is loud. The operational work is quieter. That is probably why it matters more.”

Closing Thought

The next phase of AI will not be won by the company with the longest list of tools.

It will be won by the company that knows which work to delegate, which work to review, which tools to trust, and what to do when the system changes.

The model race is loud. The operational work is quieter. That is probably why it matters more.

Need practical AI education your leadership team can actually use?

If your organization is trying to move from AI curiosity to useful AI workflows, the first step is shared understanding: what these tools can do, where they fit, what risks to watch, and how to use them responsibly in real work.

I help leadership teams build practical AI education for employees, so adoption is clearer, safer, and tied to the workflows that matter.

If that would be useful for your organization, reply to this email and I’ll help you think through a practical AI education session for your leaders and employees.

Start a practical AI education conversation

More Found Friday