Logan Kelly

Human-in-the-Loop Is a SHOULD in the MCP Spec, Not a MUST — and That Changes Who Is Accountable

Human-in-the-Loop Is a SHOULD in the MCP Spec, Not a MUST — and That Changes Who Is Accountable

MCP's spec says a human SHOULD be able to deny tool calls — never MUST. Oversight that lives in the client is a liability transfer, not a control.

Waxell blog cover: human-in-the-loop approval is a SHOULD in the MCP spec, not a MUST

The Model Context Protocol's specification states its own limitation in its security section, in a sentence most implementers skim past: "While MCP itself cannot enforce these security principles at the protocol level, implementors SHOULD: Build robust consent and authorization flows into their applications." The principles it cannot enforce include "Hosts must obtain explicit user consent before invoking any tool." That line is lowercase, and the same page declares that the RFC 2119 keywords in the document bind "when, and only when, they appear in all capitals." Where the capitals do appear, human oversight is a SHOULD. The Tools page of the 2026-07-28 revision: "there SHOULD always be a human in the loop with the ability to deny tool invocations" — mirrored for sampling requests on the Sampling page, while the Elicitation page has clients SHOULD "allow users to decline elicitation requests at any time."

Human-in-the-loop, in agent governance, is the requirement that a person be able to inspect and deny an action before the agent takes it. In protocol terms it is a recommendation addressed to the client application, not a property of the connection. Whether a human is actually in the loop is therefore a fact about which client software happens to be running, not a fact about the system on the governance diagram.

That distinction is not pedantry. It decides who is accountable when the approval turns out to have been wrong.

Why is the approval prompt written by the thing being approved?

Read the two halves of the tools specification together and the problem turns out to be structural rather than incidental.

Half one: a tool definition is supplied by the server. The name, the title ("optional human-readable name of the tool for display purposes"), the description ("human-readable description of functionality") and the input schema all originate upstream. The elicitation flow works the same way — its message field is documented as "a human-readable message explaining why the interaction is needed," and the server writes it.

Half two, from the same document: "For trust & safety and security, clients MUST consider tool annotations to be untrusted unless they come from trusted servers." The specification's root page puts it in plainer words: "descriptions of tool behavior such as annotations should be considered untrusted, unless obtained from a trusted server."

So the specification instructs clients to treat server-supplied metadata as untrusted, and separately recommends that clients render that metadata to a human who will decide whether to permit the call. The reviewer is being asked to adjudicate a request described to them by the party whose behaviour is in question. This is the tool-poisoning shape moved up one layer: the injection target is no longer the model's context window but the reviewer's screen.

The spec's own security guidance points at the mitigation — clients SHOULD "show tool inputs to the user before calling the server, to avoid malicious or accidental data exfiltration." Arguments are the part an attacker cannot pre-write, because they are produced at call time by the agent. A prompt that shows a friendly description and hides the arguments is worse than no prompt at all, because it manufactures a signed approval with no evidentiary content behind it.

What the regulator actually asks for

The EU AI Act puts the burden on design rather than on the diligence of the person clicking. Article 14(1) requires that high-risk systems "be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use."

Article 14(4) then enumerates what the people assigned oversight must be enabled to do, and the second item is the one that bites here: to "remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias), in particular for high-risk AI systems used to provide information or recommendations for decisions to be taken by natural persons."

Automation bias is named in the statute as a failure mode the interface is supposed to counteract. An approval queue that an agent fills faster than a person can read is not a counteraction — it is a generator of exactly the condition the Article describes. On the Act's own logic, approvals produced that way are evidence of oversight failing rather than oversight working.

"Human on the hook"

Practitioners have already renamed the pattern. On a Hacker News thread about approval and review fatigue that drew 318 points and about 200 comments, one commenter described the term their workplace had adopted: "We've started calling it 'Human on the hook' instead of human in the loop in work. It is more accurate, in terms of, it only matters when something goes wrong." A reply compressed it further: "Employee Role: Blameable Component."

That is a description of an architecture, not a complaint about morale. When the approval step exists but the approver cannot reconstruct what they were shown, the step has stopped working as a control and started working as an accountability sink. It moves liability from the system onto a person without moving any information in the opposite direction.

The architectural test has nothing to do with how conscientious the reviewer is. Three questions:

Where does the gate run — inside the client, or somewhere the client cannot route around? Oversight implemented as client UI is defeated by changing clients, and the client config file is the easiest thing in an agent stack to edit.

What is on the screen at decision time — a server-written summary, or the arguments about to be sent?

What survives the decision — the verdict alone, or the verdict plus the evidence, the identity, the rule that demanded review, and what happened when nobody answered?

A stack that answers "client / summary / verdict" has a human on the hook. A stack that answers "outside the client / arguments / the full record" has a human in the loop. Both get described the same way in a governance review, because the description tracks the presence of the step rather than its content. Our earlier piece on agent approval workflows works through the adjacent gap between what an approval grants and what it was asked about.

How Waxell handles this

The Waxell MCP Gateway moves the gate out of the client. Agents point at one governed MCP endpoint per tenant, and a tools/call is evaluated against tenant policy rules before dispatch, with the result evaluated again before it returns to the agent. Because the rule lives at the gateway rather than in the agent's config file, swapping Claude Code for Cursor does not swap out the oversight.

The require_approval action answers those three questions directly. When a rule fires, the call parks and a reviewer decides with what the documentation describes as "full context — who, which agent, which tool, the exact arguments — with approve/deny." The arguments are the artifact, not a server-written précis of them.

Rules scope by upstream, tool glob, user email, role, team and agent profile — so that, as the docs put it, "the same human gets different rules depending on which agent is acting for them." That is the right granularity for the automation-bias problem, because it makes queue volume a policy decision. What reaches a person can be tuned down to the calls that warrant reading.

Silence is handled explicitly rather than by default. Each rule sets a timeout and an auto_decision_on_timeout, documented as typically deny; Waxell's starter rule set pairs require_approval on payment and outbound-communication tools with a fifteen-minute timeout and denial on expiry. An unanswered approval resolves to whatever that rule says, instead of ageing quietly into a pass. Meanwhile the agent waits rather than breaking — the gateway holds the call until it is decided or times out, then proceeds or returns a structured denial, which removes both the retry choreography and the incentive to route around the gate.

The record is the decision plus its context. Waxell's documentation states that approval lifecycle events — requested, granted, denied, timed out — flow into incident notifications, and that every decision, along with the rule that fired, is recorded in the audit log. That is the difference between an approval that is auditable six months later and one that is merely asserted, when the question has moved from "did someone approve this" to "on what basis."

One further piece of the record matters for accountability specifically. Every call through the gateway is tied to a real user identity on the way in — the caller authenticates to Waxell, so who called, and through which agent, is in the gateway's audit record whatever the upstream auth mode. And when an upstream is connected with per-user OAuth — the zero-credential default for spec-compliant servers in the catalog — that upstream's own audit log shows the person too. An approval that cannot be tied to a named human is not much of an approval.

FAQ

Does the MCP specification require human approval of tool calls?

No. The specification states that "MCP itself cannot enforce these security principles at the protocol level," and its consent principles are written in lowercase prose, which the document's own RFC 2119 convention excludes from normative force. Where the all-capital keywords do appear, on the Tools, Sampling and Elicitation pages, human oversight is a SHOULD. Human oversight is therefore a property of the client application you happen to be running, not a guarantee of the protocol.

What is the difference between human-in-the-loop and human-on-the-loop?

In-the-loop means the action cannot proceed until a person decides. On-the-loop means the action proceeds while a person supervises and intervenes if something looks wrong. The distinction collapses in practice when an agent issues requests faster than a person can evaluate them: nominal in-the-loop review degrades into on-the-loop supervision, then into rubber-stamping, with no change to the architecture or to how it is described in a compliance document.

Why is showing the tool description not enough for an approval decision?

Because the description is written by the upstream server, and the MCP specification instructs clients to treat server-supplied tool annotations as untrusted unless the server is trusted. A reviewer shown only the description is reading the account of the party being reviewed. The arguments are generated at call time by the agent, which is why the spec's security guidance asks clients to show tool inputs before the call goes out.

What does the EU AI Act say about automation bias?

Article 14(4)(b) requires that people assigned human oversight of a high-risk AI system be enabled "to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)." Article 14(1) places the obligation on system design, including the human-machine interface, rather than on an individual reviewer's vigilance. Whether the Article applies to a given system depends on whether that system is high-risk under the Act's own criteria.

How do you stop an approval queue from becoming a rubber stamp?

Reduce what reaches it, and make what does reach it decidable. Scope approval rules narrowly — by tool, upstream, team and which agent is acting — so the queue carries calls whose blast radius justifies a person's attention, and put the actual arguments in front of the reviewer rather than a summary. Then define what happens when nobody answers: a timeout that denies is a decision, and a request that expires into silence is not.

Does an approval step satisfy an auditor?

Only if the record shows the basis of the decision. An audit trail that stores "approved by X at 14:02" cannot answer what X was looking at, which rule demanded the review, or what would have happened had X not responded. Retaining the identity, the rule that fired, the decision and the lifecycle events around it is what turns an approval into evidence rather than an assertion.

Sources

Model Context Protocol, "Specification 2026-07-28", 28 July 2026.

Model Context Protocol, "Tools" — Specification 2026-07-28, 28 July 2026.

Model Context Protocol, "Elicitation" — Specification 2026-07-28, 28 July 2026.

Model Context Protocol, "Sampling" — Specification 2026-07-28, 28 July 2026.

Official Journal of the European Union, "Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 14 — Human oversight", 12 July 2024.

Pydantic (Laura Summers), "The Human-in-the-Loop is Tired", 18 February 2026.

Hacker News, discussion of "The human-in-the-loop is tired".

Waxell Docs, "Policies & Approvals".

Your agents are already calling tools, and something is already deciding which of those calls a person sees. Book a 30-minute demo of the Waxell MCP Gateway and bring whoever signs off on your approvals. The useful part is watching a require_approval rule fire on a real destructive tool, and seeing exactly what lands in front of the reviewer.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.