← ALL THOUGHTS

SEP 18, 2026 · AI

AI Agents Are Picking Up Real Work. Here’s How I’m Making Sense of It.

Written by Jeremy McKellar

The AI conversation is moving beyond what a chatbot can tell us. I’m paying attention to what it can actually help us finish: a report, a working page, a recurring task, or the follow-through that usually gets lost between apps.

That raises a different set of questions. Where does the work happen? What can the agent access? What should it do on its own, and where should it come back to us?

I put together this interactive field guide to make those differences easier to explore. It is my read on the landscape, based on the products’ published capabilities—not a claim that I’ve tested every platform head-to-head. Choose a starting point below and use the source links to look closer.

FIELD NOTES / SEPTEMBER 18, 2026From answering a question
to carrying the work forward.

The useful question: what can I responsibly hand off?

ContextToolsAgentActionsHuman review

01 / THE LANDSCAPE

Seven approaches to getting work done.

These are different starting points, not a ranked leaderboard. The fit depends on your task, your tools, and the access you are comfortable granting.

01

A team of persistent agents

Grok Bot

Grok Bot gives bots a persistent cloud computer, tools, and routines. Multiple bots can work in parallel and hand work between themselves. That team metaphor is interesting when a project has distinct jobs.

What I would check

Bots belonging to one user share files, browser sessions, and logins. Separate bot names do not create separate security boundaries.

Read the official details ↗
02

Control the setup yourself

OpenClaw

OpenClaw is a self-hosted gateway connecting messaging channels to agents and tools. You choose the machine or server and configure the model, integrations, and workflows.

What I would check

You also own the upkeep. Self-hosting the gateway does not keep data local if your chosen model or tool sends it to a cloud service.

Read the official details ↗
03

Start with the finished result

ChatGPT Work

Work brings apps, files, and tools into tasks that can produce documents, spreadsheets, presentations, reports, and sites. It supports local and cloud work, with recurring tasks also part of the experience.

What I would check

Choose the environment and connected apps deliberately. Review the finished artifact and any consequential actions.

Read the official details ↗
04

An agent that keeps working

Gemini Spark

Google positions Spark as a persistent personal agent that can work in the background even when your devices are off. The appeal is ongoing help with goals rather than a fresh conversation for every step.

What I would check

Check current access and supported connections. A product announcement is not a promise that every integration is available to your account.

Read the official details ↗
05

A familiar messaging doorway

Muse

Meta’s Muse pairs a personal agent with a dedicated secure virtual machine. You can interact through the Muse app or WhatsApp, bringing longer-running tasks into a familiar messaging flow.

What I would check

A familiar chat interface can still initiate real actions. Be clear about which accounts and tasks you are handing over.

Read the official details ↗
06

Cowork, now with cloud tasks

Claude

Claude’s task experience has expanded beyond the original desktop-only Cowork setup. Cloud tasks are in beta, and the shared experience brings chats, skills, connectors, and task context together across devices.

What I would check

Older comparisons can miss this change. Verify which cloud and local capabilities are enabled for your account before designing a workflow around them.

Read the official details ↗
07

Work inside company boundaries

Microsoft Scout

Scout combines OpenClaw with Microsoft Work IQ and Microsoft 365 context. It is designed for work spanning tools such as Teams, Outlook, OneDrive, and SharePoint, with organizational controls.

What I would check

Scout is an experimental Frontier offering. Treat access, approved resources, identity, and organizational policy as part of the evaluation.

Read the official details ↗

02 / YOUR STARTING POINT

What do you want help with?

Choose the closest fit. I’ll point you toward a starting comparison.

MY STARTING SUGGESTION

Start with the ecosystem you already use.

Look at Gemini Spark if Google is your starting point, or Muse if messaging is the natural doorway. Check actual availability and connections before committing. My first test would be one small recurring task.

Compare the shape of the work.

A dated snapshot, not a live product feed. Access and capabilities can change; use the official links before choosing.

Agent environments and possible uses
AgentWhere it runsWorth exploring for
Grok Bot ↗Shared cloud computerCoordinating recurring work
OpenClaw ↗Your machine or serverA configurable personal system
ChatGPT Work ↗Local or cloudProjects and finished deliverables
Gemini Spark ↗CloudOngoing goals and assistance
Muse ↗Dedicated cloud virtual machinePersonal tasks through messaging
Claude ↗Cloud tasks and desktop capabilitiesMulti-step work with context
Microsoft Scout ↗Cloud, with desktop extensionsGoverned Microsoft 365 workflows

03 / WHAT I’M WATCHING

The connections matter as much as the model.

My read: the next useful comparison will be less about who writes the best paragraph and more about who can carry context through the tools we actually use. A strong model still needs the right information, a workable handoff, and a clear stopping point.

Tools that can connect

MCP ↗ standardizes connections between AI applications and external tools or data. I’m watching whether that makes useful workflows easier to move and maintain.

Agents that can coordinate

A2A ↗ addresses communication between agents. That is a different layer from connecting an agent to a tool. Neither protocol guarantees that every product works with every other one.

Control that stays visible

I want to see what an agent accessed, what it changed, and where it needs me. Messaging may make delegation easier, but the underlying permissions and review points still matter.

Cloud, local, and hybrid setups each involve tradeoffs. Local execution can give more control over a machine; cloud execution can keep work running when that machine is off. In either case, I would look at the complete route the data takes, including model providers and connected services.

04 / MY PRACTICAL PLAYBOOK

Give one agent one useful job.

  1. Pick a task you already understand.A weekly research brief or a draft report is easier to evaluate than an open-ended instruction to “run everything.”
  2. Define the finish line.Specify the format, sources, audience, and what a useful result looks like.
  3. Connect only what the task needs.Start with limited access and expand it for a concrete reason.
  4. Keep consequential actions reviewable.Be explicit about approval before sending messages, spending money, publishing, or deleting information.
  5. Measure the cleanup, too.Time saved is useful only after accounting for corrections, checking, and maintenance.
  6. Keep what works. Adjust what doesn’t.A small reliable routine is a better starting point than a collection of unfinished automations.

That is the shift I’m interested in: less time managing the handoffs, more time deciding what is worth doing. I’ll keep watching how these products develop—and where the promise turns into something useful in everyday work.

What would you hand off? Let’s talk ↗
OLDER → I’m Building Form Forward for Fitness That Fits Real Life