Second-order implications of on-device, consent-gated agentic execution — once an agent can act on the software it does not control: the legacy line-of-business systems, government and banking portals, and third-party consoles and SaaS dashboards the rest of the world's work runs on — software built for human hands that was never given an API and never will be, alongside newer systems where the GUI still outpaces what any API exposes. What changes downstream when it can operate them, verify the result, and prove a human said yes.
The agentic stack became very good at thinking and stayed unable to act on the software it does not control. Software you own is already automatable — expose an API, or add a test seam, and drive it in CI. The durable execution gap is everyone else's software, plus your own legacy systems that can no longer be changed: the consoles, dashboards, portals, and old line-of-business interfaces you can only operate through their screens, exactly as they ship. We treat the closure of that loop — an on-device supervised operator that works the visible interface on your behalf, checks the result on the real screen, and pauses for a real human approval before any consequential step — as a given, and ask the research question that follows: what changes once the loop is closed?
We argue four things, then turn to the risks. The agent-addressable surface inverts — from software you control to software you don't, which is the substance of regular white-collar work. The bottleneck moves rather than vanishes — from human actuation to human judgment — so one person can stand behind a fleet of agents instead of doing the clicking. Trust becomes the scarce resource, and proof of who authorized an action becomes the unit of governance. And an execution market emerges, in which agents hire operators to act on the software they cannot otherwise reach. We then turn to the broader risks — economic, social, and security — of equipping agents with this at scale, and close with predictions and the limits of the claim.
First, what the loop is for. The execution gap is not simply "software without an API" — it is software you do not control. Software you own you can already automate: expose an endpoint, add a test seam, drive it headless. What does not yield to that is everyone else's software — and your own legacy systems with no source to hand, no vendor still answering, and no API that was ever written — which you can only operate through their interfaces, exactly as shipped. Every claim in this paper is about that software.
By closing the loop we mean a specific, operational sequence, not a metaphor:
Three properties make this closure meaningful and distinguish it from adjacent approaches:
Two contrasts sharpen the picture. It is not a cloud bot working on a copy of the software — a copy has none of your real accounts, and nothing you would trust it to act on. And it is not handing an agent free run of your machine — that trades the problem for a worse one. It is an operator working your real software, on your own machine, with you approving the steps that matter. The rest of this paper is about what that picture implies.
Tool-call agents address software that exposes an API; developers automate software they own by adding one, or a test seam. Both escape hatches require control of the software, and together they cover a minority of the surface. The dashboards, OAuth screens, settings panels, SaaS portals, government and banking interfaces, and legacy line-of-business systems that constitute the majority of real-world workflows by surface area are operated, not owned by the people using them; API coverage is expanding but remains structurally incomplete for the most regulated, legacy, and fragmented contexts — and for these the GUI is the only path.
Closing the loop changes the addressing scheme from software you control — software you own, or have an API into — to software you only operate. Work that was out-of-distribution for an agent — anything reachable only through someone else's GUI — becomes in-distribution. This is not an efficiency gain on existing agent work; it is a category expansion of what agent work is. The interesting consequence is compositional: a single task often spans both worlds (call an API, then click through a console that has none). Before closure, any such task degrades to a human handoff at the GUI step; after closure, the whole task stays inside one loop.
A developer reading this has a ready objection: I don't hand tasks to a GUI — when my agent hits a login, I add a test seam or a backdoor and let it through. That works, and for the inner build-and-test loop it is the right move. But it generalizes only to software you own. You can add a bypass to your own app because you hold its source; you cannot add one to the software the rest of the work runs on. And that software, not your codebase, is the shape of regular white-collar work.
The median office day is spent operating systems nobody in the building can modify: the HR and payroll suite, the CRM, the expense and procurement portals, the bank's web app, an insurer's claims screen, a government filing site — and the legacy line-of-business tools with no source anyone can find, no vendor still answering, and no API that was ever written. None of them expose a backdoor you can add, and most expose an API that covers a fraction of what the interface actually does. The escape hatch that makes a developer's own login a non-problem does not exist for the bulk of knowledge work, which is operating someone else's software — or your own frozen software — through its screen.
The legacy line-of-business system makes the line permanent — no source, no vendor, no API ever written, no escape hatch to add. The AWS and Google Cloud consoles illustrate the near-term version: the same engineer who drops a ?test_auth=1 into their own app cannot do the equivalent to the GCP or AWS console today — there is no seam to add, and the console is the system of record for the click that has to happen. Those modern consoles will eventually get deeper APIs and MCP servers; the legacy long tail will not. At every step the agent either hands it back to a human (the open loop) or a supervised operator performs it on the real screen, with the real account, behind the consent gate (the closed loop). A test seam closes the loop for the code you wrote; the screen — the universal port — is the only path for the software you merely use, especially the software that will never cooperate any other way.
Before closure, throughput is bounded by human actuation: a person is the cursor, and work scales with how fast that person can click. After closure, throughput is bounded by human judgment: the person sets goals and approves consequential steps, and work scales with how much they can meaningfully attend to and authorize.
This yields a new scaling relation. One supervisor can stand behind many operators — a fleet — bounded not by manual labor but by attention and the cost of each approval. The design goal therefore shifts to asking for as few human approvals per task as possible without weakening them: interrupting only for the steps that matter, batching, and limited standing permissions that still need a real approval to set up. A system that closes the loop but demands an approval per click has simply relocated the bottleneck without widening it.
Restatement. "Human-in-the-loop" framed the person as a step to route around. Once execution is solved, the person is not a step — they are the governor of a fleet. The scarce quantity is their judgment, and the engineering problem is spending it well.
When acting on real software is no longer the hard part, the scarce resource becomes trust. The question shifts from "can it do this?" to "do I trust it to — and can I prove I allowed it?" The guarantee that matters is not just that the agent acted, but that a human authorized the consequential step and there is evidence of it. That turns "the agent did it" into "a human authorized it, and the agent did it, on the record" — which holds only insofar as people actually attend to what they approve rather than rubber-stamp it (see §6).
This makes authorization, not model behavior alone, the unit of governance. Every consequential action carries proof that a human said yes, on the record. For regulated and enterprise settings, the difference between an agent that can act and one that is allowed to act is whether it can show who authorized what, and when. Proof of consent becomes the thing institutions actually buy.
Once execution is a verifiable service — submit a goal, receive evidence of the verified outcome — it becomes tradeable. A third-party agent can pay for execution on real software it has no other way to reach, and the human's consented local sessions become a capability that can be offered (always behind the gate). Pricing could shift from paying for the model's thinking to paying per verified action — you are billed for results that actually happened on the real software, not for the reasoning behind them — though per-seat, per-task, and mixed models could equally emerge.
We flag this as the most forward-looking implication and the least settled: the path that would let any outside agent hire an operator is a direction we are heading, not something that ships today. The claim is structural, not a product announcement — once execution is reliable and kept behind the approval gate, a market for it is the natural layer above.
The implications above are about capability and value. Equipping agents with reliable execution on real software also carries side effects — economic, social, and security — that grow with adoption. Honest research names them, including the ones this approach itself creates or worsens. We separate risks that come with the capability itself from risks that come from how it is deployed.
The automatable frontier moves from software you control to software you only operate. That captures a vast band of GUI knowledge work — back-office operations, procurement, claims, onboarding, compliance, account administration, data entry — that resisted automation precisely because it had no API. The displacement surface potentially expands into GUI-mediated administrative work that resisted prior API-driven automation; whether the net scope exceeds prior automation waves is an open empirical question that depends on adoption pace and the rate at which new roles absorb displaced labor.
Human-at-the-helm softens this without erasing it: one supervisor behind a fleet produces more output per person, which at the margin means fewer people per unit of output. New roles appear — fleet supervisors, procedure curators, consent auditors — but nothing guarantees they absorb the displaced one-for-one or at comparable wages. Value capture shifts from doing to judging: returns accrue to those who set goals, exercise taste, and can be held accountable, concentrating gains among capital and the highly skilled. Metering by verified-operation also creates a new rent that whoever owns the execution-and-skill rail can extract.
The exposure is uneven and fairly predictable: it tracks how much of a role is operating software you don't control — forms, portals, ledgers, consoles — versus judgment, relationships, or physical presence. The most exposed are the desk professions whose day is largely that operating: bookkeepers and accountants, accounts-payable and -receivable and payroll clerks, tax preparers, claims processors and insurance/underwriting administrators, loan and mortgage processors, medical billing and coding staff, procurement and expense administrators, KYC and compliance analysts, data-entry and back-office operations, and paralegal and e-filing work. Their core tools — the accounting suites, bank and tax portals, claims and HR systems, and (for the IT side) the cloud consoles — are exactly the software no one in the building can modify. Roles weighted toward judgment, persuasion, care, negotiation, or physical work are far less exposed by this mechanism, though not untouched. This is exposure by task composition, not a forecast of elimination — the caveats above (new roles, adoption pace, the judgment bottleneck) still apply.
The biggest risk to the approval gate is that people stop reading it. A gate that fires constantly trains them to approve on reflex — rubber-stamping. The very thing that makes delegation safe weakens under volume, and the obvious fixes (interrupting less often, batching, standing permissions) each risk approving too much at once. This is an ongoing failure mode, not a solved problem.
The approval proves a person was present, not which person, and not that they truly agreed — presence can be borrowed or coerced. Responsibility can blur in both directions — "the agent did it / I only approved it" — and the record helps trace what happened, but does not settle who is morally or legally to blame. And the capability is not evenly available: those with capable agents plus execution pull steadily ahead of those without.
The capability cuts both ways. The same operator that does legitimate work also raises the ceiling for abuse on systems that assume a person is doing the clicking — account takeover, fraud, and mass actions at machine speed. Running on the user's own machine, a human approval on the steps that matter, and keeping the agent from touching the controls directly are the brakes; without them, the capability is dangerous by default.
One side effect deserves naming directly: "are you human" checks lose their meaning. If an operator clears those challenges by riding a real session behind a single approval, the web's basic assumption — that passing a human-check means a human did the work — weakens. Modern defenses watch a whole session, not just one check; but once a real person starts the session, the agent's later actions ride on that person's good standing. The hard part is the trust a real session lends to everything that follows. This is a genuine downside of the whole category, our own approach included; the honest response is tying actions to real approvals and limiting how much and how fast — not pretending it does not happen.
Two more failure modes grow with adoption. An agent that gets hijacked — say, by a malicious instruction hidden in a web page it reads — can now reach real money and real accounts; the supervision and the approval gate exist for exactly this, but if the gate fails the cost is far higher than a mistake in a sandbox. And lightly-supervised fleets could act in concert across real accounts — fake-grassroots campaigns with authentic-looking accounts, mass submissions, market manipulation. The approval on each action is the brake; reflexive approving and over-broad standing permissions are how it wears down.
When one agent hires another, what exactly did the human approve? This is the sharpest version of the problem, and the market in §5 forces it. A person approves the first agent's broad goal — but the risky action happens several steps later, chosen by the agent, not the human. A hijacked first agent could turn one honest "yes" into something the person never meant. The open question: should every risky action ask for its own approval, no matter which agent set it in motion? This paper names the problem; it does not solve it.
As more real-account actions are taken by agents — visibly and with consent — it gets harder to tell what online activity a person actually chose. A real account doing a real thing no longer means a human decided each step. "No human was clicking," spread across the whole web, complicates reputation systems, fraud detection, and anything that reads account activity as a sign a human did it. The detail that makes one demo impressive becomes, at scale, a problem of knowing what is real.
A trusted execution layer, a shared library of how to operate things, and the approval system together become a huge point of concentration — able to see across many users' screens and holding the record of what was approved. Keeping the work on each person's own machine, with their data local, is a real safeguard; a market where one provider dispatches operators pulls it back to the center, making that provider both a bottleneck and a high-value target. This concentration is built into the "agents hire execution" endgame, and is the part most worth designing against early.
Regulation cuts both ways. A clear record of who approved each action is a gift to oversight — but standardizing it, deciding who is liable for an agent's actions, and defining what "an agent acted on my behalf" means in law are all unsettled. Without rules, the capability outruns accountability.
These are not arguments against closing the loop. The loop is closing regardless of any one product. They are the reason the architecture matters: the same choices that make execution safe — running on the user's own machine, a single supervised operator, and a real human approval on consequential steps — are the levers that decide which of these risks are contained and which are amplified. The capability is coming; the open question is whether it arrives with the brakes attached.
Implications are only research if they can be wrong. Each prediction below states what would falsify it.
A claim that overreaches is not research. Closure is a precondition, not a panacea, and its boundaries are part of the result:
Closing the loop is not the destination. It is the precondition that makes the rest of the agentic era legible. Once execution on the software you do not control — the consoles, portals, and legacy interfaces the world's work runs on — is reliable and consent is provable, the binding constraints relocate — to judgment, to trust, to who is allowed to act — and the shape of the work changes from "an agent that can think" to "a fleet of agents that can act, governed by a human who no longer holds the cursor."
The questions worth researching are no longer "can an AI use a computer." They are: can your agent hire one, can it prove a human said yes, and can it compound skill across sessions? Those are the implications of closing the loop.