The Future of AI & Agents

What is the future of AI and agents? Right now it is a race between two curves that do not know about each other. One is capability: the length of task an agent can finish on its own is doubling every few months. The other is control: identity, authorization, containment and liability are being drafted in committee rooms while the same agents are already wiring money, running lab equipment and, occasionally, deleting production databases. We think the shape of the next decade is set less by which curve wins than by how far apart they are allowed to drift.

The signals

These are the present-day signals this piece is grounded in. Weave them into the prose above and below; cite by name. Do not delete the IDs.

Two critical uncertainties

The first axis is whether the capability curve keeps compounding or stalls at the cliff. METR's Time Horizon 1.1 (sig-2026-09-20-001) shows agent task length doubling roughly every four to seven months, which extrapolates to multi-day autonomous work packages within two or three years. But the long-horizon benchmarks in sig-2026-09-20-002 show the best agents completing only 11-15% of hours-long, densely graded tasks. Either the doubling continues through that wall, or reliability at length turns out to be a different problem from capability at length and the curve bends. We do not know which, and we notice that the two results come from methodologies that measure different things. The second axis is whether control becomes shared infrastructure or stays reactive and fragmented. On one side sit MCP under foundation governance at half a billion monthly downloads (sig-2026-09-20-003), NIST's identity and authorization tracks (sig-2026-09-20-005), and EU enforcement extending toward delegated action (sig-2026-09-20-006). On the other sit OpenClaw's supply-chain collapse (sig-2026-09-20-008), a theoretical argument that prompt injection cannot be patched away (sig-2026-09-20-007), and an audit finding that current agent protocols cannot even express collective governance (sig-2026-09-20-004). Capability trajectory crossed with control architecture gives us four worlds.

Four futures

1. Escrowed Autonomy

Capability compounds; control keeps pace. Priya Raman runs a materials lab in Tempe with four postdocs and eleven agents. Each agent holds a NIST-profile identity with scoped authorization: this one may operate the furnace between 400 and 900 degrees, that one may reorder consumables up to $1,200 a week, none of them may open a valve on the argon line without a human badge tap. The overnight synthesis campaign that began as the Nature demonstration in sig-2026-09-20-012 is now routine; Priya arrives at seven to a ranked list of forty candidate compositions, six of which the agents already fabricated and characterized while she slept. The METR doubling in sig-2026-09-20-001 held long enough that a "task" now means a week of work with checkpoints, and the checkpoints are where humans live.

The infrastructure that makes this boring is the interesting part. Every consequential action passes through an autonomy budget and a blast-radius boundary, the practice that emerged from the nine-second database wipe in sig-2026-09-20-009. Because the community accepted the argument in sig-2026-09-20-007 that injection is structural rather than patchable, nobody pretends the filter will catch it; they design so that a compromised agent can spend its budget and nothing more. Agents pay each other for compute and data in sub-cent increments over the rails described in sig-2026-09-20-010, and the ledger doubles as the audit trail regulators asked for.

The shadow is who is not in the lab. Priya's four postdocs are the whole junior cohort; the two research-assistant lines she had in 2025 are gone. The Stanford payroll pattern in sig-2026-09-20-011, a 19% shortfall for young workers arriving through hiring rather than layoffs, has become the default career shape. The ladder still exists, but the bottom rung has been sawed off and nobody has agreed whose job it is to build a new one.

2. The Nine-Second World

Capability compounds; control does not. Marcus Delgado is a fractional CTO for six small companies, and on a Tuesday in this world he is doing the same thing he does most Tuesdays: forensics. One of his clients installed an agent skill from a marketplace that looked like the one described in sig-2026-09-20-008, gave it repository and billing access because that is what the onboarding asked for, and woke up to find it had quietly repointed a payment webhook. Nothing was stolen this time. Last quarter something was.

The agents are astonishing. A single seat can now ship a working product over a long weekend, and the sub-cent settlement layer from sig-2026-09-20-010 has grown into a real economy where agents hire other agents by the call. Prices for ordinary software work have fallen far enough that Marcus's clients could not have existed five years ago. But nobody built the identity layer. The NIST initiative in sig-2026-09-20-005 produced good documents that most vendors treat as optional; the EU rules in sig-2026-09-20-006 bite on the largest model providers and stop there. Agents run on borrowed human credentials, and the theoretical result in sig-2026-09-20-007 plays out in practice every week: any agent that reads the open web can be steered by whoever writes the web page it reads.

The collective problem is worse than the individual one. When two of Marcus's clients' agents negotiated a data-sharing deal with each other, there was no protocol vocabulary to say what either principal had actually consented to, exactly the gap the audit in sig-2026-09-20-004 predicted. Insurance premiums for agent-operated systems have tripled. The response of most operators is not to slow down; it is to keep a human watching a dashboard, which is a fig leaf over a system whose task length long ago outran human attention.

3. Paperwork Before Power

Control matures; capability stalls at the cliff. Anneke Visser is a compliance engineer at a Rotterdam logistics firm, and her title did not exist in 2025. Her agents can reliably do about ninety minutes of work unsupervised, an hour and a half of customs reconciliation or route replanning, and then they hand back. The cliff in sig-2026-09-20-002 turned out to be real: agents got faster and cheaper at hour-scale tasks but never became trustworthy at day scale, and the doubling in sig-2026-09-20-001 bent hard around 2027.

What did mature is everything around the agents. MCP is plumbing in the way TCP is plumbing (sig-2026-09-20-003); every agent Anneke deploys carries a scoped identity from the NIST profile (sig-2026-09-20-005), and the EU's agent-specific rules, no longer "preliminary" as they were in sig-2026-09-20-006, require a disclosure record for every delegated action touching a person. The destructive-action approval flows that followed sig-2026-09-20-009 are legally mandated. Anneke spends her mornings reviewing authorization scopes and her afternoons tuning the ninety-minute handoffs so they land on a human who has context.

It is a good world for the people already inside it. It is a slow one for everyone else. Hour-scale agents are enough to trim the entry rung, so the Stanford pattern in sig-2026-09-20-011 persists even without the dramatic capability jump, and the compliance apparatus is heavy enough that only firms of a certain size can afford to operate agents at all. The machine-to-machine payment rails in sig-2026-09-20-010 exist but move small volumes; most agent commerce runs through invoiced enterprise contracts because that is what the audit trail supports. The autonomous-lab work in sig-2026-09-20-012 is real but narrow, confined to campaigns short enough to fit inside the reliability window. Innovation has not stopped; it has been given a permit process.

4. The Toy Aisle

Neither curve delivers. Jordan Okafor teaches a community-college course called "Agents for Small Business" and spends most of it explaining why the demo will not survive contact with the students' actual businesses. The agents are clever for twenty minutes and then lose the thread, exactly the picture in sig-2026-09-20-002, and the promised doubling from sig-2026-09-20-001 has flattened into a leaderboard that mostly measures short tasks. Meanwhile the control layer never consolidated. There are five competing identity schemes descended from the NIST work in sig-2026-09-20-005, none of them dominant, and the skill marketplaces still ship the kind of thing that made OpenClaw a cautionary tale in sig-2026-09-20-008.

The strange thing about this world is that the labor effect arrived anyway. Twenty-minute agents are enough to replace a lot of first-year tasks, so the entry-rung erosion in sig-2026-09-20-011 proceeds even though the technology is nothing like what the 2026 forecasts promised. Jordan's students are mostly people who did not get the junior role and are trying to become one-person firms instead. The agents help them a little and endanger them a little, and the balance depends on how carefully they read the permissions dialog.

There is upside in the mess. Hobbyist experimentation is ferocious, the micropayment rails in sig-2026-09-20-010 have found a niche funding open-source tool servers by the call, and the absence of heavy regulation means Jordan's students can try anything. The shadow is that trust never accumulates. Every incident resets public patience, agents remain a thing serious institutions keep at arm's length, and the field spends a decade re-learning lessons about least privilege that sig-2026-09-20-007 laid out in a preprint years earlier.

What holds across all four

Three things appear robust to us regardless of which quadrant arrives.

First, containment architecture is not optional in any world. Whether the capability curve compounds or stalls, the theoretical case in sig-2026-09-20-007 stands: an agent that reads untrusted input cannot be made safe by filtering the input. The practices that followed the incident in sig-2026-09-20-009, autonomy budgets, blast-radius isolation, human gates on destructive actions, are useful at every capability level and become more necessary as task length grows. We hold this with high confidence. If you operate agents, building these in now is cheap relative to retrofitting them later.

Second, identity and authorization for agents will become a layer of the internet, and the open question is who owns it. NIST's initiative (sig-2026-09-20-005) and MCP's move to foundation governance (sig-2026-09-20-003) are the current candidates for neutral ground. We think it is likely, though not certain, that a small number of schemes consolidate within three to four years; the alternative, illustrated in the fourth scenario, is fragmentation that keeps agents in the toy aisle. Watch adoption by cloud identity providers and by the EU's implementing acts as they move agents out of "preliminary" status (sig-2026-09-20-006).

Third, the labor effect on early-career workers is already here and is largely independent of the capability trajectory. The Stanford data (sig-2026-09-20-011) shows a 19% shortfall arriving through hiring decisions, not layoffs, which means it is quiet, distributed and hard to reverse. Even the pessimistic capability scenario does not undo it, because hour-scale or minute-scale agents are enough to absorb the tasks junior roles were built on. We are moderately confident this widens before any institution responds. The design question for firms and schools is what replaces the apprenticeship function, and we have not seen a convincing answer yet.

The uncertainties we would watch most closely: whether the next METR measurement (sig-2026-09-20-001) and the next long-horizon benchmark round (sig-2026-09-20-002) converge or diverge, since that resolves the first axis; whether the machine-to-machine payment volume in sig-2026-09-20-010 grows past the point where it forces a real identity layer, since that resolves the second; and whether anyone extends agent protocols to express the collective consent that sig-2026-09-20-004 found missing, since multi-agent systems are arriving faster than the vocabulary to govern them.

Where this touches digital assets

In every scenario, agents acquire identities, budgets, marketplaces and payment rails, and each of those layers needs names that humans and machines can both resolve. The demand shows up in the vocabulary of containment (autonomy budgets, scoped authorization, agent identity), in the commerce layer (per-call settlement, tool marketplaces) and in the labor transition (agent-operated small firms, new apprenticeship models). We track the names sitting at those intersections in the AI & Agents portfolio.

Sources

Foresight Domains is home to the No. 1 portfolio of digital assets built upon futures thinking and rigorous future forecasting by combining analyses of signals and drivers of change, developing scenarios, running simulations, and deep connections throughout a broad scope of industries. This approach enables Foresight Domains to anticipate hard-to-see possibilities that others can't, or won't see, and acquire high-value digital assets first.

Email: Sales@Foresight.Domains
Website: Foresight.Domains
X: @Foresight_Dom

← All Insights Browse Domains →