Six stages every "AI employee" needs, and the one question no analytics platform can answer
͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­
Forwarded this email? Subscribe here for more

Managing AI Employees

Six stages every "AI employee" needs, and the one question no analytics platform can answer

Matteo Cellini
Sep 23
 
READ IN APP
 
AI Employee Searches Trend - what happened in August this year??

C-levels today are more focused on saying they have agents, and many of them, than on knowing what those agents actually do: what they cost once you count human supervision, or whose performance review suffers when they fail. The “AI Employee” narrative has already curdled into something bigger: agent teams, interconnected systems meant to replace whole groups of humans, sold with the same clickbait certainty as everything before it. But if agent teams are going to be treated as teams, they need accountability, and right now, nobody’s job is to notice when they don’t have it.

The signals

We have already argued that adoption intensity is not structural transformation, and that tracking usage of tools tells you nothing about value. Four more signals are coming to the surface that we need to keep in mind:

Work3 - The Future of Work is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Upgrade to paid

Agent sprawl and abandonware: agents get built in a rush, without clear goals or real integration into the ecosystem. They work in demos and small pilots, then fail at scale. Some keep running long after the process, owner, or data they depended on has changed. Gartner recently forecast that over 40% of agentic AI projects will be cancelled over cost, unclear value, and weak risk controls. That’s a staggering amount of budget going to waste.

Cognitive exhaustion: one of our earliest pieces (Buried by Bots) foresaw this, and it was just picked up by McKinsey: the heaviest AI users are often the most exhausted, because automation removes the routine work but leaves the judgment calls, the decisions, the difficult conversations. This is what success looks like when nobody redesigns the job around it: the cost has simply moved into people’s heads.

Economics arrive after the demo: prototypes run on fifty cases, while production runs on fifty thousand. Token costs, latency and review queues appear only at volume, which is why most agent business cases are missing a line for supervision. Even more importantly, many of those tasks never needed a frontier model, and some never needed a model at all.

Domain knowledge is the scarce input: I recently read a line that stuck with me: “It can be easier to teach a metallurgist AI than to teach an AI specialist metallurgy.” In a recent study of 5,172 customer-support agents, AI assistance raised productivity by 15% on average - less experienced workers improved in both speed and quality, whilethe most skilled saw small speed gains and small declines in quality. The valuable input was top performers' know-how built into the workflow, that is why domain leaders should own agent outcomes, and IT should own the standards.

The Agent Lifecycle

The simplest way to stop treating agents as magic software is to start giving them a lifecycle. As they increasingly behave less like static software and more like delegated workers, they are assigned tasks, given access, trained on context, measured through outputs, adapted over time and embedded into workflows that other people come to depend on.

A useful lifecycle should have six stages:

Hire: Before an agent is built, the domain leader should have to define the job. What decision or workflow is it supposed to improve? What would count as success? Why is an agent the right form, rather than conventional software, automation, search, a dashboard or a better process? A surprising amount of agent work fails at this first question. The agent exists because the model could do something impressive in a demo, not because the organization had a durable role for it.

Onboarding: This is where the agent receives context, instructions, permissions and data access. It is also where risk usually enters quietly. The fastest way to make an agent useful is to give it broad access “just for now.” But “just for now” often becomes production architecture. Onboarding should therefore be treated as a joint responsibility between the domain leader and IT: the business defines the work; IT defines the access pattern, logging, identity and control surface.

Performance: This is where most organizations substitute usage for value. A dashboard that says people used the agent 4,000 times this month tells you almost nothing. Did it reduce cycle time? Did it improve accuracy? Did it create more review work? Did it shift costs from one team to another? Did customers notice? An agent’s performance review should be tied to the business result it was hired to produce, not to the enthusiasm of early adopters.

Development: Human workers learn through coaching, feedback and changed responsibilities. Agents improve through prompt revisions, workflow redesign, additional tools, better eval sets, fine-tuning or model swaps. The danger is that this learning often accumulates outside the firm. The best version of the workflow may live inside a vendor interface, a prompt chain controlled by a consultant, or the private habits of one power user. If the organization cannot explain what the agent has learned, where that learning lives and how it can be transferred, then the capability is not fully owned.

Restriction: This is the missing management move: when a person is underperforming, managers can narrow their scope, add review points, remove privileges, change responsibilities or put them on a performance plan. Agents need the same logic. Some should not be turned off; they should be demoted. A customer-facing agent might be moved back into draft-only mode. A finance agent might lose write access. A research agent might require citation checks before its output can enter a client document. But this only happens if someone has the authority to say: the agent is useful, but not safe enough for its current scope.

Retirement: This is where software thinking is most dangerous. Old software usually sits there until it breaks or gets replaced. Old agents can keep acting. They can continue sending messages, drafting decisions, summarizing outdated policies, triggering workflows and consuming budget long after the process around them has changed. Retirement means revoking access, archiving useful context, preserving audit trails and making explicit what should not be reused. An agent should not be able to become organizational abandonware simply because nobody wants to own the cleanup.

Agent Analytics and Scorecards

Microsoft tried to answer this by launching Agent 365, a registry for every agent in the org, access control through Entra identity, and analytics that promise to show ROI. Salesforce also launched Agentforce Observability - which looks far more granular with metrics like deflection rate (sessions closed without human escalation), escalation rate, abandonment rate, task resolution rate, session count and duration, cost per agent in consumption credits, step-level latency, error rate, a 1–5 quality score plus answer-faithfulness and context-relevance scores, and user sentiment/toxicity scoring per session.

This is a genuinely new practice, and it’s likely already generating new titles, like Head of Agent Analytics, AI Analytics Engineer, AI Agent Analyst, or AI Orchestrator - jobs that definitely didn’t exist two years ago.

Microsoft Agent 365: The control plane for AI agents | Microsoft 365 Blog

But an ownerless agent and an unaccountable one are not the same problem. Agent 365 can tell you a name isn’t attached to an agent’s identity record and it can’t tell you whether a business leader would put their name on what that agent is producing.

It has nothing to say about whether the agent is worth what it costs once you count the review time, or whether it’s helping a team or quietly becoming their second job. Those are management questions and no platform vendor is going to answer them for you, because the answer is specific to what the agent was hired to do.

That’s the gap the scorecard we built below is built for, which aims at giving operating questions a business owner should answer regardless of what the platform reports:

Two Critical Implications

The first implication is for organization design: management will become the management of mixed production systems.

Most companies still draw the org chart as if work is done by people using toolsm but increasingly, work is done by combinations of people, models, workflows, automations, vendors and review systems. A finance process may involve accountants, an ERP, a reconciliation agent, a policy checker, a human approver and a risk dashboard. A sales process may involve account executives, enrichment agents, writing agents, CRM automations and human managers reviewing exceptions. A customer-service process may involve a chatbot, a routing model, a knowledge-base agent, frontline staff and escalation rules.

The unit of management is no longer just the team, it's going to be the production system.

That means domain leaders have to own blended teams: finance owns both the reconciliation humans and the reconciliation agents, legal owns both the lawyers and the contract-review agents, customer support owns both the service representatives and the resolution agents and IT provides the infrastructure and guardrails, but it cannot be the universal manager of every agent any more than it can be the universal manager of every spreadsheet.

As for middle management: someone has to translate business goals into agent instructions, to decide what gets escalated, inspect output quality, know when the workflow has drifted and explain to leadership why a supposedly automated process still consumes so much human attention.

Span of control should therefore count more than direct reports: a manager responsible for eight people and twenty agents may have a heavier coordination burden than a manager responsible for fifteen people and no agents. The relevant question is not just how many humans report to you, but how many semi-autonomous systems depend on your judgment, review capacity and escalation decisions.

This is also why workflow redesign has to come before tool rollout. Adding an agent to a broken process is not transformation; it is a faster way to produce ambiguity, as we have argued in the Agent Graveyard a while back. Decision rights, review points, escalation paths and retirement triggers need to be redesigned first. The sequence should be: measure the work, govern the system, then build the agent - not the other way around.

The second implication is for the workforce: the valuable work shifts toward judgment, teaching and the “why.”

When agents absorb routine tasks, they do not remove human work so much as concentrate it because people are left with framing, validation, redirection and exception handling. They decide what the agent should attempt, whether the output is good enough, when the edge case matters and how the work should change after repeated failures. That is higher-cognitive-load work which is also less visible than the work it replaces.

For example: a support agent may reduce the number of tickets a person writes from scratch, but increase the number of ambiguous escalations they handle. A research agent may reduce search time, but increase the burden of source judgment. A coding agent may generate more code, but increase the need for architectural review. A manager may see output rising and miss the fact that people are spending the day in a continuous stream of micro-decisions.

That is why agent management needs a cognitive-load budget. If every automation leaves behind only exceptions, then people’s jobs become denser, not easier and the productivity gain is real, but so is the fatigue. Without measuring review minutes, escalation quality and rework, the organization will mistake hidden strain for efficiency.

Final Thoughts

In a few years, the search term will fade, the way they all do.

"AI Employee" will be replaced by whatever the next wave of interconnected systems gets branded, and in a year someone will chart that curve cooling too. But the org chart will still be missing the same box it's missing right now: not who built the agent, not who uses it, but who's on the hook when it's wrong.

That box doesn't get filled by better analytics. Microsoft and Salesforce can tell you an agent has no owner; they cannot make a domain leader stand up and say, in a performance review, this one is mine, I own what it produces, and I will answer for it the way I'd answer for a direct report.

Most organizations are not short on dashboards, they are short on someone willing to say that sentence and take responsibility and effort. Until one does, the next platform launch will just be a better way to measure a job nobody has agreed to do.

Until next week!
Matteo

Work3 - The Future of Work is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Upgrade to paid

You're currently a free subscriber to Work3 - The Future of Work. For the full experience, upgrade your subscription.

Upgrade to paid

 
Like
Comment
Restack
 

© 2026 Matteo Cellini
From Rome with ❤️️
Unsubscribe

Get the appStart writing