Essay

The AI Skeptics Are Right, but They're Missing the Point

The autonomous agent economy the largest companies are quietly building, and the one property it cannot open without.

Author: Victor Davidenko ยท ORCID 0009-0001-9073-8628

License: All Rights Reserved

The public conversation about AI has narrowed to a few questions. Whether it is a bubble. How much power the data centers will draw. Whether the frontier labs can hold a lead now that open models are closing the gap almost as fast as the closed ones can extend it. The skeptics in that conversation are not wrong. Adoption inside most companies is running well behind what was promised, the returns are hard to find, and some of the people best placed to judge, the ones shipping real products on these tools every day, will tell you plainly they are not betting on it. If you think the current wave of AI has been oversold, the evidence is on your side.

All of that is an argument about one thing: who sells the intelligence. It is loud because it is where the money and the reputations sit today. Underneath it, drawing far less attention, a different thing is being built, and it has nothing to do with selling models. It is about giving AI a way to act in the world. Humanoid robots and the actuators that drive them are moving AI off the screen and onto the factory floor and the loading dock. Autonomous vehicles are putting it on the road. AI systems are now working at the frontier of science, predicting how proteins fold well enough to earn a Nobel Prize and proposing materials no chemist had drawn. None of these are chatbots. They are AI given hands, wheels, and instruments, and every one of them is real, funded, and shipping while the bubble argument runs on the surface.

Follow those movements to where they lead and they converge. A warehouse robot has to work in a building its owner does not own. A car has to draw power from a network its maker does not run. A research agent has to reach data that lives inside another institution's walls. The value shows up the moment one owner's AI has to deal with another owner's AI, and the sharpest form of dealing is commerce: negotiating terms and committing to them. The largest players in payments already see this. Over the past year Visa, Mastercard, and a Google-led coalition that includes PayPal and American Express have shipped rails for agent-initiated payments, and Mastercard and Banco Santander ran a live one on production infrastructure in March 2026. When companies of that size build for a market before it visibly exists, they are telling you where they think this goes.

But every one of those roads hits the same wall just short of the destination, and the payment rails show exactly where it is. The rails prove an agent was allowed to pay and they move the money, within limits the account holder set in advance. What they do not reach is the step before the payment, where two companies' agents negotiate a commitment and each has to come away holding its own record of what was agreed and of the authority the other side was working within, a record it can prove later to a counterparty, a regulator, or a court without depending on anyone else. Until that record exists, no company that manages its risk will let its agent close a real deal with an outside agent, and the autonomous economy stays a set of demos. The skeptics are right that today's AI stops at the edge of the company. The people building the rails are right that something much larger is coming. Both are looking past the single property that decides whether the larger thing arrives: a deal each side can prove.

What has to be true

Start with why a company would want this at all. Every deal it closes today runs through a person who negotiates and signs, and that person's hours are the tightest limit the business has. Only the deals large enough to be worth those hours get done, and the smaller ones, and the ones that would have to happen faster than a person can act, do not happen. An agent that can negotiate and close removes that limit, and a whole tier of deals that were never worth a person's time comes into reach.

For a company to allow it, two things have to be true, and neither is true today. The company is bound by what its agent commits to, the same way it is bound by what an employee does. So first it needs to bound what the agent may commit to, so that a runaway or a manipulated agent cannot obligate it past set limits. Second, it needs to come away from every deal holding a record of what was agreed and of the authority its agent carried, provable later without anyone else's cooperation.

That second requirement is not academic, because these negotiations are adversarial. This year MIT researchers ran more than 180,000 agent-against-agent negotiations through a single competition, and the agents negotiated with real strategy. One of them won by feeding its counterparts what looked like a system instruction, drawing out their private limits. When the agent across the table may be working to manipulate yours, what was actually agreed is something you will have to prove.

Neither requirement exists in usable form today, so for now a person is put back on the signature line and the deal runs at human speed. That is not a small tax. It removes the reason to have agents transact at all. An assistant that can negotiate your supplier contracts but still needs you to review and sign each one is a demonstration of a skill, and the business runs exactly as it did before.

Why the obvious answer is the wrong one

The obvious place to keep that record is the platform both agents run on. That is the one place it cannot go, and the reason has nothing to do with anyone acting in bad faith.

Contoso.com and Tech.co each connect their agents to the same vendor's platform, and the two agents negotiate and agree the terms of a deal: the price, the volume, the conditions each side will hold the other to. Before they can negotiate at all, each agent needs a verifiable identity binding it to the company that authorized it, carrying the scope of authority it is allowed to exercise. Without that, neither side knows whether the agent across the table is who it claims to be, or whether it has the authority to commit to anything. Assume both have that credential. The moment they agree, both companies are committed. And right then, before anything has gone wrong, the problem is already present. Neither company walks away holding its own record of what its agent agreed to, or of the authority its agent was working within. The only account of the deal lives inside the platform, produced by the platform's software, held on the platform's infrastructure, and where it is signed at all, signed with the platform's keys.

So the first time anyone asks Contoso.com to show what its agent committed it to, whether a counterparty or a regulator, Contoso.com has to go back to the platform and ask for its own agent's word, in the platform's format, on the platform's schedule. And the platform has an interest in the answer. If the record it produced were to show that the platform's own software mishandled the limits Contoso.com had set, that would put the platform's product at fault, with exposure to Contoso.com and to every other customer on the same system. The only account of the deal sits with the party that has the most to lose from what it says.

For evidence, this is the worst possible arrangement: one interested party holding the only record of a dispute between two others, about that party's own product. And it does not depend on the platform ever altering a record. A challenger does not have to prove the record was changed. It only has to point out that the party holding it had both the means and the motive, and that nothing in the arrangement rules it out. More logging does not close that gap. Every additional record is authored and held by the same party whose conduct is in question.

Who audits the auditor

Every public company faces a version of this and settled it a century ago. A company does not audit its own books. The auditor's independence was made a structural requirement, written into securities law, after enough decades of watching what a self-checked opinion is worth when something goes wrong. Independence has to be structural, which means the auditor's incentives, access, and tools sit outside the company being audited. A careful internal process is not a substitute, because the question is never only whether the work was honest. It is whether an outside party can rely on it.

Agent governance has arrived at the same problem without the hundred years of hindsight. When a single vendor supplies the agent, supplies the layer that governs the agent, and supplies the system that records what that layer decided, the evidence and the conduct it describes share one root of trust. Picture the hearing after an incident. A general counsel stands in front of a regulator holding a clean chain of records from request to decision to action. The regulator asks who outside the vendor has ever seen this evidence, and who could have changed it before today. The honest answers are no one, and the vendor. That is enough to move the records from proof to assertion, without anyone alleging a thing.

This has nothing to do with the vendor's integrity. It is about where the record lives, and it applies to any first-party arrangement, including an independent governance vendor that held the signing keys itself. If the party that runs the governance also holds the only provable copy of the evidence, the same question lands on it. The categories shipping governance today, from the platforms that authenticate agents to the clouds that host them, differ in the name on the door and share this one structural spot.

What the record has to be

The fix is not a more trusted middleman. A trusted party in the middle holding the record is the same arrangement with a friendlier name. The record has to leave the middle entirely.

Think about how two companies close a deal on paper. Each signs, and each keeps a copy carrying both signatures. The moment the pens lift, both are bound, and neither has to take the other's word for what was agreed, because each holds the other's signature on it. That is the standard to reach for. Two agents closing a deal should end the same way: each company holding a record of what was agreed, signed by the other side's agent and bounded to the authority its own agent carried, verifiable by itself without the platform's cooperation and without trusting the other side's systems.

Strip away the branding and that requirement reduces to four properties. All four have to hold at once, or the independence falls apart.

First, the evidence is created at the trust boundary, the point where a call crosses an organizational or vendor line, and not somewhere downstream inside a system one party controls. A record written after the fact by the same system that took the action is a self-report whatever format it wears.

Second, the evidence is signed with keys the customer holds in their own key infrastructure, never the governance vendor. A signature proves non-repudiation only when the party being governed does not also hold the key that could sign a revised version. If the vendor holds the key, the vendor can in principle sign whatever account is convenient.

Third, the evidence is stored in the customer's own infrastructure, with no vendor copy of record. This closes the retention and access gap that leaves most platform audit logs reachable only through the vendor, on the vendor's timeline, in the vendor's format. If the authoritative copy lives with the vendor, the vendor decides what the record means.

Fourth, and this is the property most designs skip, the evidence is verifiable by any third party with an open, inspectable tool and the public key, offline, with no one to contact. A regulator, an opposing counsel, or an insurance adjuster should be able to take the signed record and the public key and confirm it on their own machine, with no call to make and no dependence on the vendor's goodwill or survival. If verification requires trusting a party to the dispute, the independence was never there.

The result carries the weight of a paper contract both companies signed. What was agreed is a single instrument that both parties hold and either can prove, offline, to anyone, and there is nothing to argue about afterward. Set against a platform-held log that no outside party can examine, that is far stronger ground under the tests the law already applies to evidence.

What governance has to do before the record can exist

Independent evidence is the proof. What it proves is that governance was running. Signing a record is worthless if the governance machinery behind it is not doing real work.

Governance at this level means every agent action is classified before it proceeds, because a system that does not know whether a request carries sensitive data, or triggers a regulated workflow, or exceeds a cost threshold cannot decide whether to allow it. It means access control that determines which agents, models, and tools may interact at all, and under what conditions. It means kill switches at graduated scopes, from shutting down a single tool to halting all agent activity across an organization, because an incident in production does not wait for a committee. It means budget enforcement that prevents a runaway agent from consuming resources past the limits its owner set. It means detecting credentials embedded in prompts before they leak into a model's context. It means catching adversarial prompt patterns before they manipulate an agent into exceeding its authority. It means tracking every hop in a multi-step chain, so that when an agent delegates to a second agent that calls a third-party tool, the full path from request to action is accounted for and none of it escalates beyond the authority the first caller carried.

All of that has to run at the speed the agents move, which is the speed of a network call. And every decision the governance layer makes, every classification, every access check, every kill switch state, every policy evaluation, has to produce a signed evidence record meeting the four properties above, or the governance is no more provable than the self-attested logs it was meant to replace. Evidence without governance is an empty signature on a blank page. Governance without independent evidence is unverifiable.

This is the reason the problem cannot be solved by adding a signing step to an existing platform. The governance layer and the evidence architecture have to be designed together, from the ground up, as a single system that runs inside the customer's own infrastructure, signs with the customer's own keys, stores evidence in the customer's own storage, and never holds a copy of any of it.

What you can check today

None of this asks you to take my word for it. An open-source verifier called scarp-verify is published on GitHub, and it runs the offline check the fourth property describes. Take a signed evidence record and the public key, run the tool on your own machine, and confirm the signature and the contents hold. Change one field and run it again, and it fails. There is no server to reach and no account to create.

The evidence records available alongside the verifier are produced by Scarp Governance Gateway, a governance infrastructure product that implements all of the properties described above. It deploys inside the customer's own infrastructure. It signs with keys held in the customer's own key infrastructure, keys that the gateway itself cannot access or export. It stores evidence in the customer's own storage with no vendor copy. It classifies every request, enforces access control and budget limits, provides graduated kill switches, tracks multi-hop delegation chains, detects credential leakage and adversarial prompts, and produces a signed evidence record for every governance decision. The records are what the verifier checks.

That answers the one question the whole architecture turns on, whether an outside party can confirm a record without trusting whoever produced it. The rest is a matter of building to the four properties.

The companion paper describes the broader architecture in which this governance layer sits: four foundational layers (Identity, Cooperation, Governance, and Settlement) that together provide the infrastructure for AI systems to identify themselves, cooperate across organizational boundaries, operate under independently verifiable governance, and settle value with full provenance.

Where the regulation is heading

The specific deadlines keep moving. The direction does not.

The EU AI Act's obligations for high-risk systems were first set to apply in August 2026. The EU's Digital Omnibus pushed the standalone high-risk obligations to December 2, 2027. The obligations themselves did not change. The reason given for the delay was that the technical standards and compliance tooling companies need to meet them were not ready. That gap is the subject of this essay, named in the text of the deferral. The requirement still stands, and it still expects documented, auditable evidence of how a system's decisions were governed. A description of the policy in the abstract will not satisfy it.

State activity in the United States has been less stable. Colorado's early AI law was repealed and replaced inside of two years, with enforcement paused in the meantime, which is a reminder that a headline statute is a poor foundation for a compliance posture. The durable expectations are older and sector-specific. HIPAA's audit-control requirements, Canada's OSFI B-13 guideline for technology and cyber risk, and PIPEDA's accountability principle already expect regulated entities to produce compliance evidence that stands up to independent scrutiny, rather than evidence their own vendor generated on their behalf.

None of these frameworks were drafted with autonomous agents in mind. All of them carry an assumption that becomes load-bearing the moment agents take consequential actions: the entity under scrutiny cannot be the sole author of the evidence used to judge it.

The AI Kill Switch Act, introduced in the United States in July 2026, adds a further signal. It gives the Department of Homeland Security shutdown authority over frontier AI models meeting certain revenue and compute thresholds. The question it does not answer is how that shutdown is executed at the infrastructure level and how an organization proves compliance with it. A governance layer with graduated kill switches and independently verifiable evidence of their activation is one way to answer both.

What this opens

Once a deal between two agents can be bounded and proven without either side trusting the other or the platform between them, the movements from the opening stop being demonstrations and start to compound. The warehouse robot contracts for space by the hour with a building it has never dealt with before. The car settles charging and parking and road priority with networks its maker never partnered with, its agent striking terms and keeping the record. Your own assistant negotiates your utility rates, your insurance renewal, and the recurring purchases you would rather never think about again, closing each within limits you set and returning with proof of exactly what it committed you to. A company's procurement agent closes with a supplier it found and vetted the same morning, and the supply chain renegotiates itself as conditions move. Research combines data and models no single institution owns, on terms the parties' agents settle between themselves.

Past everything that makes existing work faster, a further layer opens where the buyers and sellers are themselves agents and the goods are compute, energy, bandwidth, and data, priced and traded between machines in real time. None of it runs on evidence authored by an interested party, because the moment any of it is disputed, the record has to convince someone who was not in the room and does not take the platform's word for anything.

The roadmaps skip this part, and so does most of the public argument about AI. That argument is a contest over who captures the market that already exists, the business of selling the intelligence itself and the chips and power beneath it. Those are real businesses and their winners will be large, but they are a small set of companies splitting one prize. The expansion behind the missing property is a larger and stranger thing. When a deal between two organizations' agents can be struck and proven in seconds at a cost near zero, the deals worth doing are no longer limited to the ones a person had time for, and their number climbs by orders of magnitude. The gains from that reach past the firms selling the models and land across every business that runs on agreements, which is very nearly all of them, and that is the part the loud conversation is talking over.

The whole branch of the economy where agents transact and commit across the lines that separate one owner from another, whether the owner is a person, a company, or a machine, sits behind a single missing property: a record of the deal each side can hold and prove without trusting the other or the platform between them. That branch reaches nearly every industry there is, and today it is closed.

Where this goes from here

The architecture described here rests on patent filings on file and pending. The governance product that implements it, Scarp Governance Gateway, is described in detail on the product page. The broader substrate architecture spanning identity, cooperation, governance, and settlement is laid out in the companion paper and the vision document. The independently verifiable evidence records are published with full source and verification tooling in the evidence repository.

The readers this essay is written for are enterprises deploying autonomous agents, platform teams deciding whether to build governance as a first-party feature or to partner with an independent layer, and the standards bodies working on agent identity and audit frameworks.

If you are evaluating agent governance for a deployment, or you are on a platform team weighing whether self-attested governance will hold up over the next few years of regulatory tightening, I am glad to talk specifics. Serious inquiries: info@scarpprotocol.com


Victor Davidenko is a senior enterprise architect. ORCID 0009-0001-9073-8628. Patent filings on file and pending cover independent, cryptographically verifiable AI agent governance evidence.