The web is getting its second user

For thirty years, the web has had one kind of user. A person, with a browser, reading a page. Every business model on the web was built for that person: the ad next to the article, the paywall, the seat license, the free tier that turns into a sales call. Every technical convention, too - robots.txt, the referral, the click. All of it assumes someone is looking at a screen.
That assumption broke quietly, and recently. Automated traffic passed human traffic on the web in 2024 and reached 53 percent of all traffic in 2025, according to Imperva. In June, Cloudflare put machines at more than 57% of the web’s HTML requests. And the fastest-growing slice of that traffic isn’t the old kind of crawler indexing pages for search. It’s agents: software acting on behalf of a specific person, on a specific task, right now. HUMAN Security measured traffic from AI agents and agentic browsers growing more than seventy-fold in a year.
The web is getting its second user. And that second user doesn’t behave like the first.
What a second user changes
A person browses. An agent queries. A person reads a page; an agent needs one field on it. A person visits a site a few times a month; an agent might hit it four hundred times in an hour, then never again. A person arrives through a link and can be shown an ad or offered a subscription. An agent arrives, takes what it needs, and leaves nothing behind - not a click, not a session, not a reader.
Cloudflare’s own measurements make the shift concrete. Classic search crawlers fetched a page or two for every visitor they sent back to a site. AI operators now fetch hundreds or thousands of pages per referral. That isn’t a variation on the old pattern. It’s a different user, with a different relationship to what it consumes.
I run product and everything around it at TinyFish, and we see this from the inside. Our infrastructure handles more than forty million web operations a month for agents doing research, enrichment, monitoring, diligence, and many other critical workflows that involve web automation. We watch what those agents reach for. And the thing they reach for most, and get least, is data.
Three problems, one cause
For the people building agents, the best data on the web is unreachable. Not technically unreachable - commercially. Market data, private-company records, business identity, court filings, verified contacts, scholarly literature: it all exists, cleanly structured, behind an API. But it also sits behind a license, a sales cycle, a minimum commitment, and a form that might say “talk to sales.” An agent can’t navigate all that at once, and a developer building one on a Tuesday afternoon doesn’t want to. So they do the only thing available. They point their own browser at the site and scrape what should have been queried - slowly, unreliably, without a citation, and often without the right to. Or they give up, and the agent gets built with a worse substitute nobody can trust. We know, because they do it with our tools.
For the companies that hold the data, demand is arriving from users they can’t see. Their distribution was built for people: a dashboard, a seat, an account manager. When an agent reads their data through a scraped page, they don’t see the usage, can’t serve it well, and aren’t credited for it. The most valuable inputs to the agentic web are owned by companies with no way to reach the agents that most need them.
For the web itself, this breaks its original contract. You cite the source, the source gets the traffic, the traffic fuels the work. When data reaches an agent through a scraped page, the answer is unsourced, possibly stale, and the people who assembled it are nowhere in the loop. Multiply that across a web where machines are the majority and you’ve broken the mechanism that made the web worth building on.
So, three problems, one cause: a second user arriving on infrastructure built for the first.
What we believe
We didn’t start with a spec or alignment doc. We started with a set of positions we were willing to be held to, listening to what the market demands.
Every one of them comes back to trust, and trust here runs in two directions. The person whose agent pulls a court record needs to trust that it’s current, correct, and from the source accountable for it. The company that assembled that record needs to trust where it goes: that it’s credited, that the usage is visible, that reaching an agent’s builder doesn’t mean losing the relationship with the customer. Neither side gets that today. The scraped page fails the builder, who can’t know what they got, and fails the owner, who can’t see that it was taken. An alliance is a structure where both sides can trust the exchange, because the terms are known before the first query and the source is never hidden. The principles below are how we do it.
The source is credited, every time. Whatever shape the plumbing takes, the answer knows where it came from. It’s the condition for the whole thing being worth doing - what turns a scraped fact into a sourced one.
The owner sees the demand. The company that built the data should know how agents are using it. Today they can’t; they see a scraper, or nothing. Making agent demand visible and attributable to the source is the first step toward it being served properly, and toward the value it creates flowing back to where the data came from.
Curation over catalog. The instinct in a new market is to list everything and let the market sort it out. We think the point of specialized data is that someone is accountable for it being right. So the alliance is curated: leading providers, organized by category, each there because their data is what an agent should reach for when a plausible answer isn’t good enough.
Commit on evidence, not intentions. We launched a partnership before we launched an integration, on purpose. Nobody has actually measured how agents consume specialized data yet, and it doesn’t look like human usage. We’d rather learn what agents really pull, with fifteen partners watching the same numbers, and build from there.
Why we started with a curated room, not a marketplace
Today we announced the Data Partners Alliance: fifteen founding data providers - Alpha Vantage, Baselayer, Crunchbase, Databento, Enigma, Faraday, Fiber AI, Intellizence, Jinko, OpenAlex, Particle, RocketReach, Similarweb, Tracxn, and Trellis Law - launching together for the builders and enterprises creating AI agents on TinyFish.
The mechanics at launch are deliberately simple. Builders discover the partners on our site and connect with each one directly; the account and the terms stay between them. What matters more than the mechanics is what the fifteen represent: the companies holding the leading data in their categories, agreeing that the second user is real, and that they’d rather meet it together than wait for it to arrive through a scraper.
I spent the summer talking with the people who run these companies. What struck me was how many had already reached the same conclusion on their own. Their buyers are changing. Their pricing pages assume a human. And every one of them had a story about watching demand show up in a form they couldn’t serve. The alliance didn’t need convincing. It needed a place to stand.
What comes next
The founding fifteen are data providers, and that’s where we started because the problem is sharpest there: structured data, clear ownership, an API already waiting. But the same questions are arriving for everyone who makes the web worth reading. Publishers are asking how agents should identify themselves and what they owe the sites they read. Content owners are asking whether the referral economy that funded the open web survives a user that doesn’t click. We don’t have all the answers, and we’re suspicious of anyone who says they do. But the principles above weren’t written for data vendors alone. They’re what we think a web with two users needs, and we intend to hold ourselves to them as this grows.
The direction of travel is clear enough to say out loud. Agents on TinyFish already read the live web. The next step is data on tap: the right record, filing, or price arriving inside the work, from the source, with the source credited. The exact shape gets built with each partner as real demand from real agents makes the case, category by category.
The web is getting its second user. This is where it starts learning to serve them, with trust at the center of it all - running both ways, from the first query.
Two ways in
If you build agents: register for early access and tell us which partners you’d want first. That’s how we decide what to build, in what order.
If you hold data that belongs where agents work: tell us about it. The founding fifteen are where the alliance starts, not where it ends.
Homer Wang is Head of Product at TinyFish. Explore the Data Partners Alliance and register for early access at tinyfish.ai/partners.
AI disclosure
Content on this website may be created or refined with the assistance of AI tools and is subject to human editorial review.



