What separates an AI experiment from an AI System
͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­͏     ­
Forwarded this email? Subscribe here for more

Dear Reader, it is very hard to discuss the future of work, without getting into the future of society and politics.

We hope you enjoyed our article last week about intergenerational politics, where we tried to steer away from the usual generational clichés, to look at why our policies need to change. This week you can read the 2nd part on how we might be able to respond positively. As always, we are interested in your views, and examples.

Andy and Matteo


Beyond the Agent Graveyard: AI Production Lifecycles

What separates an AI experiment from an AI System

Matteo Cellini
May 12
 
READ IN APP
 

Welcome to this week’s edition! We are continuing a series on Enterprise AI adoption in the workplace, this week looking at how we can move from AI experiments to AI Systems.

Let’s dive in!

Work3 - The Future of Work is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Upgrade to paid

When Agents leave the Demo

There is a strange moment that happens after a good AI demo: the agent works, it finds the right document, drafts the right answer, pulls together the right context, and does in a few seconds what would normally take someone twenty minutes of searching, copying, checking, and rewriting. Everyone in the room has the same reaction: this is useful, we should have more of this.

In 2026, most large organizations have moved past the question of whether they should experiment with AI agents. Some have dozens, others have hundreds. McKinsey has talked about running 20,000 AI agents internally. Agents exist across HR, sales, customer support, finance, IT, and knowledge management. Some are built centrally by technical teams, but many emerge from the edges of the business, created by people trying to solve the specific, visible friction in their daily work.

This is genuinely a good thing: for years, enterprise software has been something done to employees rather than with them: bought centrally, implemented slowly, and pushed into workflows that didn’t always match how people actually worked. Agents change some of that energy when the people closest to the work can now shape the tools around it.

But once the agent leaves the controlled environment of the demo, it enters the much messier reality of the company it’s supposed to serve, and that’s where things tend to get complicated.

Enter the agent graveyard: every organization has a graveyard of internal tools. Dashboards nobody opens anymore, some automations built by someone who has since moved teams, knowledge bases that were useful for six months, then slowly drifted out of date as the processes they documented evolved. Some are workflows that technically still run, even though nobody can quite remember why they were created or whether they still reflect how the work is done. Most of these failures are not dramatic - the tools didn’t announce themselves as failures. They were useful once, then gradually underused, forgotten, or broken. The more common pattern is quieter: the tool keeps existing, but nobody is quite sure whether it still reflects how the business actually works.

I suspect many agents will follow exactly the same path. They won’t stop working in any obvious way. The underlying business process will shift and the documents they rely on will be updated, duplicated, or contradicted somewhere else in the organization. Six months later, someone will ask whether the agent is still accurate, and the honest answer will be that nobody has been watching closely enough to know.

What is Shadow IT | With Examples, Risk, Benefits & Policy

This is what AI sprawl (uncontrolled proliferation of AI models, tools, and agents across an organization, creating security, governance, and efficiency risks) actually looks like in practice: a gradual accumulation of promising tools that never quite became reliable systems. It’s a pattern (or vicious circle) enterprises have lived through before. Useful tools appear at the edge of the business, solve real problems, spread faster than the governance structures around them, and eventually become another layer of organizational complexity to manage. Shadow IT wasn’t born from bad intentions, it came from good intentions meeting the wrong infrastructure, and agents risk following the same trajectory, except with considerably more surface area.

The CIO’s problem is not “can we build agents?” It is “can we have hundreds or thousands of agents operating across the enterprise without creating unmanaged risk, duplication, shadow AI and another layer of digital clutter?”

Five ways an agent fails

The failure usually doesn’t happen at the model level, it happens at the five levels underneath, and in my experience they tend to chain rather than appearing separately: an agent built with poor context gets used less, which means nobody notices the problems accumulating, which means nobody owns fixing them, which means governance gaps go unaddressed, which means when someone finally asks whether the program is working there is nothing to measure it against. Each failure makes the next one harder to catch.

  1. The first is context. Most agents are built and tested against a simplified version of the business. In the demo, the documents are clean, the use case is narrow, and the permissions work. In production, the agent meets the actual company: siloed data, inconsistent metadata, outdated policies, undocumented processes, and knowledge that lives partly in systems and partly in the memory of whoever has been around longest. Without a shared context architecture, every new agent starts from zero, with the same data sources getting reconnected, the same permissions getting negotiated, and the same business logic getting reconstructed by a different builder in a different team.

  2. The second is adoption. Many agent rollouts are still treated as a distribution problem: announce the tool, show the demo, share the link. The pattern I see most often is this: the agent gets announced in a team meeting, four or five people try it in the first week, usage drifts back to two by week three, and by month two it’s one person occasionally checking whether it still works. The announcement created awareness but it didn’t create habit. Without structured enablement, feedback loops, and visible integration into the workflows people already use, most agents end up bypassed by everyone except the people who built them.

  3. The third is ownership. The builder moves on, the workflow changes, the underlying source gets updated, and the agent keeps running, still available, still answering, but no longer fully aligned with the process or the data it was built around. A particularly common version of this: the organization goes through a reorganization, a process changes, and the agent keeps confidently answering questions about the old workflow because nobody updated its knowledge sources. Nobody notices immediately because the people who knew the original process have moved on too.

  4. The fourth is governance. Even when IT and security are involved, they’re often trying to govern an expanding surface area with limited visibility. If every team makes its own decisions about permissions, access, and risk, governance becomes inconsistent at exactly the moment the organization needs more consistency, with no org-wide guardrails and every team making its own security decisions at speed and in isolation.

  5. The fifth is measurement, and this is the one I find hardest to explain to people because it seems so obviously fixable and yet it almost never gets fixed in time. Agents get launched without a clear definition of what good looks like, and a few months later, when someone asks whether the program is working, there’s no baseline, no measurement framework, and no business case to interrogate. The agent either feels useful or it doesn’t, and that’s not a sufficient basis for deciding whether to maintain it, improve it, or scale it.

What makes these five problems particularly hard to solve is that they don’t belong to the same person: the builder owns context and adoption, nobody clearly owns maintenance. IT owns governance, when asked. The CIO owns measurement, when someone thinks to involve them early enough. Each group does their job reasonably well within their own lane, but none of them do it in sequence with the others. This is less a collection of operational failures and more a single coordination failure that shows up in five different places.

Not that this is a new kind of scenario: Software has had a development lifecycle for decades: define the problem, design the solution, build it, test it, ship it with governance, monitor it in production, and improve it over time. You don’t deploy code into a business-critical environment and then walk away. But agents feel lighter than traditional software: the interface is conversational, the build is fast, and the use case seems obvious, so organizations sometimes treat them more like experiments than production systems. That informality is useful in the early stages of discovery and becomes a problem when agents begin to touch real workflows, real data, real users, and real decisions. Enterprise software has tried to solve versions of this before. Agile addressed fast iteration. DevOps connected development and operations. MLOps emerged to govern model training and monitoring. Security frameworks cover access controls and risk management. But none of these were designed for the specific challenge agents create: dozens or hundreds of them being built by non-technical business teams, operating inside complex permissioned environments, touching real data and decisions. Agents are too easy to build to warrant major software project treatment, but too consequential to be treated as casual experiments. They currently slip through all of these frameworks simultaneously.

The question then is not whether organizations should build agents more carefully, but where that care actually has to start, and the answer is usually much earlier than anyone involved in the build expects.

Running the sequence backwards

The order most enterprises follow: build something, launch it, and figure out whether it worked. After all, you need something to exist before you can govern it or measure it, so the natural instinct is to build first and sort out the rest afterwards.

The problem is that this ordering, however logical it feels in the moment, is also why so many agent programs end up with tools nobody can honestly evaluate. By the time the CIO asks whether the agent is delivering value, the builder has moved on and IT is still finding out the agent existed.

The order that works is the reverse: Measure, Govern, Build. Before the builder writes a single prompt, the CIO’s office agrees on what success looks like and how it will be tracked, and that baseline becomes the brief. Before the agent launches, IT and security define the access policies, the system boundaries, and the failure protocols, not as a review gate at the end but as constraints that shape the design from the beginning. Then the builder starts, with a clear problem definition, guardrails to work within, and a measurement framework to build toward.

What I find interesting about this sequence is that it doesn’t actually require new tools or new roles to implement, it requires the most senior stakeholders to do their work before there’s a polished demo to react to, and that’s a much harder cultural shift than any technical one. The CIO naturally wants to see the agent before committing to a definition of success, IT naturally waits to be asked. The builder, who has the clearest mandate and the most momentum, fills the vacuum that this creates, which is how organizations end up in the situation they were already trying to avoid.

Reversing the sequence is fundamentally a question of discipline rather than technology, and it requires a shared operating model that works across teams, vendors, and business functions, much like a software development cycle.

Context as infrastructure

There’s a reason we have started talking about context engineering rather than prompt engineering: the prompt is a relatively small part of what determines whether an agent works in production. The larger part is everything around it: what company data the agent can see, what it knows about the process it’s operating inside, what it remembers from prior runs, and what it’s blocked from touching. Getting that right is an information design problem, not a language one.

Good context engineering has two qualities: it’s dynamic, meaning the agent fetches what it needs automatically rather than relying on a human to pass it every time. And it’s multi-player, meaning it’s designed for the whole organization rather than one person’s carefully configured setup.

Most enterprise AI efforts struggle with both. What passes for context engineering in most teams is one builder spending a weekend connecting data sources, writing system prompts that encode their own understanding of the business, and configuring an environment that works well for them and nobody else. When that builder moves on, the context architecture they built tends to go with them, because it was never documented or shared in a way that others could maintain or build on.

Making context engineering genuinely multi-player requires agreement across teams on what the agent should know, what it should have access to, and what the standards are for how it gets built. That agreement doesn’t happen at the builder level. It requires the same coordination across the three roles described earlier, which is why the sequence matters as much as the tooling.

The Enterprise Agent Development Lifecycle

This is the problem Glean, whose Enterprise Graph we covered earlier in this series, has formalized into what they call the Enterprise Agent Development Lifecycle. I found it interesting not because it invents lifecycle management, but because it applies that discipline to the specific problem AI agents create: fast, distributed creation inside complex, permissioned, constantly changing organisations. This also matters because it provides the shared knowledge infrastructure foundation, and this is the operating model that runs on top of it and turns agent building into something repeatable.

The framework spans seven stages.

  1. Opportunity – Start by spelling out the business problem the agent is intended to solve. Use plain language: who is affected today, how work happens without the agent, and what tangible change you expect if the agent succeeds. This anchors everything that follows to outcomes the business already cares about, not just an interesting demo.

  2. Design – Describe what the agent is actually responsible for. Define the unit of work (a ticket, an incident, a call, a workflow), when it should run, what information it needs, and what it should produce every time. Be explicit about what’s in scope, what’s out of scope, and which assumptions need to be tested early.

  3. Performance – Turn that intent into a small, concrete set of success metrics: the business KPIs you expect to move, plus agent‑centric quality and safety signals. Agree on baselines and target ranges, and decide up front what would cause you to expand, pause, or roll back the agent.

  4. Context – Identify the minimum set of permission‑aware data sources, tools, examples, and feedback signals the agent needs to do its job and to be evaluated fairly. This includes which systems it can read from, which actions it is allowed to take, and what telemetry you will use to understand how it’s performing.

  5. Develop – Turn the design into a reliable agent. Choose the right execution model (for example, a structured workflow versus more flexible auto‑mode), then test against golden examples, run it in parallel with the existing process, and pilot with design partners until the agent behaves predictably, not just on a hand‑picked demo set.

  6. Launch – Treat rollout as a change‑management exercise, not a switch flip. Decide who will see the agent first, how it will show up in their workflow, what training and communication they need, and which guardrails, SLOs, and kill switches must be in place before you broaden access.

  7. Monitor & Improve – Operate the agent like any other critical system. Use dashboards, alerts, and runbooks to track impact and quality over time, handle incidents, and feed real‑world signals like user feedback, overrides, drift in key metrics, back into earlier stages of the lifecycle so the agent keeps improving.

The practical question for most people reading this is how to actually introduce this sequence in an organization where the builder already has momentum and the CIO hasn’t been asked. The honest answer is that you rarely reverse the sequence across the whole organization at once. What tends to work is picking one upcoming agent build and treating it as a pilot for the new process: asking the business owner to define the success metric before the configuration starts, getting IT in the room during the design rather than the review, and documenting the context architecture in a way that survives the builder moving on. One agent built this way is more persuasive than any argument about methodology, because it gives you something concrete to point to when the next build starts and someone asks why the process is different this time.

What makes this more than a framework is that each stage requires concrete tooling behind it. The platform must be designed to lower the barrier for non-technical teams without sacrificing the context quality that makes agents useful in production. Builders should be able to describe what they want an agent to do in plain language and have the platform handle execution across company data without predefined workflows. The key differentiator for any platform is its foundation: a robust, permission-aware enterprise graph ensures the agent isn’t searching a curated subset of documents but is accessing a continuously updated map of how the company actually works across all connected systems. When something fails, the platform must offer tools that surface what happened at every step: inputs, tool calls, model decisions, and output, giving teams a basis for diagnosis rather than guesswork. For instance, Glean provides features like Debug and Trace Views to achieve this visibility. Furthermore, sophisticated architectures should support coordination at runtime, allowing agents to react to enterprise events automatically rather than waiting to be manually invoked.

The launch stage is built around the premise that governance designed after deployment isn’t really governance so much as an attempt to retrofit control onto something that was already built without it. An Agent Library gives IT teams the controls to verify agents, organize them by department, and manage their lifecycle, turning informal distribution into something the organization can actually see and track. Agent Access Policies sit above that, applying org-wide guardrails consistently regardless of who built the agent or which team owns it, so that the same controls that apply to the tenth agent apply to the thousandth.

The Monitor & Improve stage closes the loop that most agent programs leave open. An Agent Insights dashboard tracks adoption rates, top use cases, estimated hours saved, and ROI trends over time. The business value isn’t just tracking a single agent’s performance but being able to compare impact across the portfolio, build a defensible case for continued investment, and identify which use cases deserve more development and which should be retired. That’s what turns monitoring from a reporting exercise into an operational decision.

The next divide

So far, the conversation about enterprise AI has been dominated, understandably, by what agents can do.

But questions shift quickly: early on, the question is whether the agent can work at all. Then it’s whether it works reliably. Then it’s whether anyone is actually using it. Then it’s whether it’s still accurate six months later. Then it’s whether it can be trusted with something more consequential. Each of those transitions requires a different kind of organizational response, and most enterprises are still working through the first few while the more mature programs are already navigating the later ones.

The organizations that handle those transitions well probably won’t be distinguished by how many agents they’ve created, they’ll be the ones where someone, at some point, decided to run one build properly before scaling the approach. That’s usually how the sequence gets reversed in practice: not through a top-down mandate but through one team demonstrating that the slower, more disciplined approach produces something the organization can actually rely on.

The next divide will not be between companies that use AI and companies that do not, it will be between organizations that can turn AI experiments into trusted systems, and those that simply accumulate another layer of digital clutter.

Work3 - The Future of Work is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Upgrade to paid

You're currently a free subscriber to Work3 - The Future of Work. For the full experience, upgrade your subscription.

Upgrade to paid

 
Like
Comment
Restack
 

© 2026 Matteo Cellini
From Rome with ❤️️
Unsubscribe

Get the appStart writing