Build with live web data

The Outcome Era: Why Enterprise AI Agents Need an Execution Factory

Sudheesh Nair
enterprise ai agents, enterprise ai agent infrastructure

AI agents have already proved they can do remarkable work once. The next question is whether they can do it again and again, at enterprise scale, without a human standing nearby. The outcome era begins when that work becomes repeatable, observable, and economically predictable.

Intelligence was act one for AI agents

Within three weeks, two of the largest AI platforms showed us where this industry is heading. Meta launched Muse on September 8, a personal agent that runs on its own virtual machine and keeps working after you close the app. On September 29th, OpenAI took the DevDay stage to launch dots, always-on agents powered by GPT-6 Astra, each with its own cloud computer and browser and access to more than 4,000 apps.

Enterprise AI agents are moving toward persistent execution on cloud computers. Give an agent its own computer in the cloud, let it keep working while the human goes somewhere else, and judge it by whether the job gets finished.

That shift is more interesting than another generation of model intelligence. For three years the industry kept score with one question: how smart is the model? We compared benchmarks, context windows, reasoning scores, coding ability, multimodality, latency and price, and we were right to, because the first problem really was intelligence and the progress was extraordinary.

Then success did what success does and moved the bottleneck. Intelligence is becoming abundant and dramatically cheaper, and while better models will keep mattering, model quality alone no longer defines the product. The model now sits inside a larger system whose job is to get something done.

The first era of AI produced answers. The next one produces outcomes, and in that era the finished task is the product and the model is one of its parts.

Why one successful agent task is not a production system

There is something close to magic in watching an agent finish a task end to end for the first time. You ask it to research a trip, compare a few flights, find the right hotel and fill in the details, and ten minutes later most of the work is done. If it hesitates, asks a question, picks the wrong fare class, stalls on a CAPTCHA or needs approval before it touches an account, the experience survives, because a human still sits on the other side of the task.

That human does more work than we give them credit for. They notice when something looks strange, forgive a bad run, answer the unexpected question, correct a wrong assumption, approve the sensitive step and, most importantly, decide whether the result is good enough. OpenAI's own launch makes the point. Its showcase story is a dot that noticed a tester had forgotten to invoice a publication, prepared the invoice, and sent it once he approved, and the launch post tells users to "always review consequential work." The most capable agents shipped this month still assume a person in the loop.

For a single task, that arrangement is a bargain. Trade thirty seconds of supervision for forty-five minutes of tab switching and most of us call it a win.

Now remove the person. Instead of one traveler booking a flight, picture an AI-native company running 40,000 supplier lookups overnight, a bank reconciling records across hundreds of portals, or a developer embedding an agent inside a product used by thousands of other companies. Nobody watches each run anymore, and nobody is there to forgive the strange one. The output of one system becomes the input to the next, and the next system assumes the first did its job. That is where autonomy stops being enough.

A production system needs something much harder: predictability. No amount of engineering will make every run deterministic in a world of probabilistic models, changing websites, expired sessions, network errors and CAPTCHA challenges.

Predictability means the system knows what happened when things go wrong. It knows where the task failed, what state it reached, whether the operation is safe to retry, whether it can resume from the last good step, and whether the result deserves trust at all. Failure is inevitable. Unobservable failure is what kills production systems.

Then there is cost. A consumer on a monthly subscription has no reason to care whether one task used three model calls or thirty. A company running that task hundreds of thousands of times cares a great deal, because a workflow that costs forty cents today and four dollars tomorrow is a CFO problem.

At scale, the unit of work moves past the token, the search and the browser session to the outcome itself. The questions get brutally practical. Did it work? Will it work again tomorrow? If it fails, can we recover? What does one completed outcome cost? What happens when we run it ten thousand times?

From AI agent demos to enterprise-scale execution

In 1785, French gunsmith Honoré Blanc demonstrated interchangeable musket-lock parts by mixing components and assembling working locks. Thomas Jefferson, then the American minister in Paris, described the demonstration in correspondence. Both historical claims require inline links to reliable primary or scholarly sources.

Master gunsmiths had been making excellent weapons for generations, with extraordinary skill and precision. Blanc was solving a different problem. An army needed thousands of muskets whose parts could be replaced, repaired and reproduced without depending on the particular hands that made the original, and one magnificent musket from one magnificent craftsman did nothing for it. The breakthrough was not craftsmanship, but repeatability.

It was also far harder to industrialize than to demonstrate. Blanc's approach did not become an established French production system. In the 1820s, John Hall advanced mechanized rifle production with interchangeable parts at Harpers Ferry Armory, then in Virginia. These historical claims require inline links to reliable sources. The decades between demonstration and production show how much engineering repeatability required.

We have watched models write software, reason across enormous bodies of information, operate browsers, research markets and complete tasks that would have sounded absurd a few years ago. The interesting question is no longer whether one extraordinary run is possible, but whether those capabilities can be industrialized into systems that produce the same class of outcome again and again, at enormous scale, without depending on a human standing nearby.

Give one enough time, enough context, enough retries and a human supervising it, and it can produce something astonishing. An enterprise needs that result on the ten-thousandth run, on a Tuesday night, with nobody watching. The hard problem now is building the factory around these probabilistic craftsmen, one that manufactures reliable outcomes from models, browsers, websites and networks that never stop changing.

Where search fits in AI agent infrastructure

Search was one of the first agent infrastructure businesses to take off, and the reason is simple. Before an agent can do something, it needs to find something, so the first generation of agent infrastructure monetized search by the query. That made sense while search was the product.

It makes much less sense once the product is the outcome. A single completed task can trigger dozens of searches, page fetches, browser loads, retries and follow-up queries before the agent knows enough to act. Charging for every lookup at that point stops pricing the value and starts putting a toll booth between the agent and the work it is trying to finish.

Worse, the meter creates the wrong incentive. When every search costs money, developers, routing systems and optimization layers are naturally pushed toward fewer searches, fewer fetches and less inspection, even when one more lookup would materially improve the final answer. That protects query margin and works directly against the best possible outcome.

So TinyFish made search free. Search is the front door to an outcome, and the value sits past it: getting through the login, holding the right identity, operating the page, surviving the weirdness of the modern web, completing the workflow and returning something another system can trust. A company that only sells search has to charge admission at the door. We want to make money on the work that gets done inside.

What we built, and why

When we started TinyFish, our core bet was that agents would become some of the web's most important users, and that when they did, they would need infrastructure built for machine execution rather than human browsing. The architecture follows from that bet.

Repeatable execution starts with persistent identity. Cookies, credentials, login state, and the browser profile must persist across pauses so an authorized agent can return with the same account context. Websites extend trust to the users they recognize, and continuity is how an agent stays recognizable.

Recoverability depends on the session holding enough state for an agent to resume instead of starting over. We built fast hydration so a browser can freeze and restore its state, which spares the agent from replaying twenty previous steps and hoping the website behaves the same way twice.

For isolation at scale, every browser has to run independently without paying for a heavyweight machine per session. TinyFish runs its browser infrastructure on unikernels, which are minimal virtual machines configured with the components each browser session needs. This design isolates sessions from one another. Any claim that it also improves startup speed should be supported with product measurements.

An outcome also demands more than a pile of web pages. A web agent must operate page controls, authenticate with permission, complete multi-step workflows, and return structured results with an execution record.

Predictable cost at enormous scale eventually means controlling more of the intelligence that performs those operations rather than handing every decision back to an expensive frontier model. That is why we built Mako, our own mixture-of-experts model trained for web execution, so we can keep moving more of the execution stack onto intelligence whose cost, latency and behavior we control.

A browser stack, a credential vault, a model and an execution layer are very different pieces of technology, and every one of them answers the same question. Together, these components help an agent produce observable and recoverable results on a changing web.

Enterprise AI agents are moving into production

For the last two years, selling agent infrastructure to an enterprise meant starting the conversation one level too early. Before anyone would discuss reliability, architecture or economics, you first had to convince them that agents would do meaningful work at all.

That part of the conversation is disappearing fast. Muse and dots make the future visible in a way no slide deck could, because people can now watch software take a goal, operate a computer, navigate the web and come back later with the work done.

The debate has moved. Serious companies are no longer asking only whether agents can do real work; they are already building around the assumption that they will. Now we get to argue about the harder problem. Will the same workflow work tomorrow? Will it still work when the website changes? If it fails halfway through, can it recover? Can another system trust the output without a human checking it? Can you run it ten thousand times without the economics falling apart?

Models are probabilistic. The web is nondeterministic. Networks fail, websites change, sessions expire, and agents will keep making the occasional strange choice. Businesses still need predictable outcomes, and the gap between probabilistic intelligence and predictable execution is where I believe the next great infrastructure companies in AI will be built.

Frontier models are becoming extraordinary craftsmen.

The enterprise needs the factory.

AI disclosure

Content on this website may be created or refined with the assistance of AI tools and is subject to human editorial review.

Get started

Start building.

No credit card. No setup. Run your first operation in under a minute.

Get $8 in Wallet fundsRead the docs