MCP Server Governance: What Actually Gets Enforced, Not Just Written Down
Published:
Last Updated:
In March 2026, one endpoint in nginx-ui’s MCP integration sat behind an IP allowlist and auth middleware. The endpoint next to it, the one that actually executed tool calls, including configuration writes and server restarts, had none. Shodan found more than 2,600 instances running that configuration in production. Nobody had decided the second endpoint needed the same protection as the first. Nobody had decided it didn’t, either. It just shipped that way.
That’s MCP server governance failing in the most common form it actually takes: not an absent policy, a policy that stopped at the point where someone assumed the next component would handle it.
MCP Governance and Project Governance Are Not the Same Thing
Search for MCP governance, and one of the first results is the Model Context Protocol project’s own governance page, which has nothing to do with securing a deployment. It documents how the specification itself gets decided: a Lead Maintainer structure inherited from projects like Python and PyTorch, Core Maintainers who steer the spec, Specification Enhancement Proposals as the mechanism for changing it, all operating under LF Projects, LLC.
That’s project governance. It answers who gets to change the protocol. It says nothing about who gets to grant an agent a database credential, and conflating the two is how some early “MCP governance” content ends up thin on anything a security team can actually use. What this article means by MCP server governance is the enterprise-side question: access controls, audit logs, and policy enforcement applied to the MCP servers an organization actually runs, independent of how the protocol’s maintainers vote on the next spec revision.
Why MCP Isn’t API Governance With Extra Steps
Treating MCP servers like REST APIs with a different transport is the single assumption behind most of what shows up downstream as excessive access. It’s worth naming precisely why that assumption breaks.
Model context protocol MCP’s own authorization specification treats authorization as optional, a position that hasn’t moved through the July 2026 revision either. More specifically, it directs implementations using STDIO transport away from its OAuth 2.1 flow entirely, toward pulling credentials from environment variables instead. That’s a documented design choice, made because STDIO was assumed to run locally, under one user, with no delegation involved. It is not a design choice that survives the moment a local mcp server gets wrapped in a proxy and exposed to a team. Transport layer security at the connection level, TLS on the wire, solves a different problem entirely and doesn’t touch this gap at all, since the credential model is the issue, not whether the traffic itself is encrypted.
A traditional API has a fixed, coded call pattern: a developer writes GET /orders/{id}, and that’s the entire request surface for that endpoint. An MCP client acting on an agent’s behalf doesn’t call a fixed pattern. It interprets intent and decides which tool to invoke, with what parameters, in a sequence nobody scripted in advance, sometimes across multi-agent systems where one agent’s output becomes the input to a second agent’s tool call. A tool call that’s syntactically valid and semantically wrong, an agent invoking delete_record because it misread ambiguous intent, is not a category of failure traditional API security was ever built to catch, because traditional API security assumes the caller’s intent, not just its syntax. Tool descriptions, the metadata an MCP server publishes describing what a given capability does, become an attack surface of their own here: a server that responds with an accurate-looking description at approval time can be modified afterward, and nothing in the base protocol re-verifies that description before the next call. The July 2026 spec does flag one adjacent field, tool annotations, as untrusted by default unless the server is already trusted. It draws no equivalent line around the description field itself, the one users actually read before approving a call.
Risk One: Excessive Permissions
An October 2025 credential audit of more than 5,200 mcp server implementations, conducted by Astrix, found that 88% require credentials to operate. Only 8.5% use OAuth. 53% rely on static API keys or personal access tokens, and 79% pass those credentials through environment variables, precisely the pattern the specification still assumes for local, undelegated use, unchanged through its July 2026 revision. Most production MCP deployments are running the local-trust model in a context that stopped being local months ago.
Broad access granted once and never revisited is what turns an ordinary prompt injection into a serious incident. Security risks at this layer rarely start as a dramatic exploit; they start as a token that could access data far beyond what its actual task required. In May 2025, Invariant Labs disclosed a vulnerability in the official GitHub MCP server, at the time carrying over 14,000 GitHub stars: an attacker plants a prompt injection payload inside a public repository’s issue. When a connected agent reviews that issue as part of a routine task, the injected instruction directs it to read and post content from a private repository the same token has access to. The MCP server wasn’t misbehaving. It was faithfully honoring a request from an agent that had no reason to distrust it. Invariant Labs named this pattern toxic agent flow: legitimate MCP server access, redirected by untrusted external data the agent had no way to distinguish from a real instruction. Prompt injection attacks of this kind don’t require credential theft or a flaw in the server’s code, which is precisely why security controls aimed only at the server’s own logic miss the actual mechanism.
The control here isn’t more scanning. It’s scope minimization: a token that can only read orders shouldn’t also carry write access to every other remote mcp server the user happens to have. Short-lived access tokens, issued per session rather than standing indefinitely, shrink the actual blast radius of exactly this kind of injection, because a token that expires in an hour limits how long a compromised agent’s borrowed authority actually lasts.
Risk Two: No Audit Trail Worth the Name
“We have logs” and “we have an audit trail” are different claims, and the gap between them is where most incident reviews stall. A usable audit trail for MCP tools needs to answer five things for any given call: which agent identity made it, which tool, with what parameters, what came back, and whether a human approved it if approval was required. A log that captures connection events but not tool-call parameters answers none of those on its own, and it won’t support regulatory compliance work when an auditor asks for a specific record rather than a general statement that logging exists.
Anthropic’s own documentation for Claude Cowork is a useful, concrete illustration of this gap and its fix, even outside the MCP context specifically. Cowork activity isn’t captured in Anthropic’s Compliance API. Team and Enterprise plan owners can export Cowork events to a SIEM through OpenTelemetry instead, a real, documented path, but one that has to be explicitly configured before an incident, not discovered as an option after one. The same distinction applies to MCP-mediated tool calls generally: the capability existing and the capability being wired into something a security team actually watches are two different operational states, and treating the first as sufficient is how organizations discover the gap mid-investigation instead of before, often while trying to determine whether sensitive data actually left the environment or just passed through it.
Shadow AI in the MCP context has a specific mechanism, not just a label. A local MCP server bundled inside a plugin or developer tool runs with exactly the same OS-level permissions as any other program a user installs on their own machine. Nothing in the protocol requires that connection to be reviewed, approved, or even logged centrally before it starts handling requests, and nothing enforces role based access control on a connection nobody centrally knows exists.
That mechanism scales badly in exactly the environments where MCP adoption is happening fastest. Community-built MCP servers pulled from GitHub repositories, connected to codebases containing proprietary source or enterprise data, with no data visibility into what the server reads or transmits and no identity record tying its activity back to whoever installed it, are not a hypothetical. They describe the default state of a fast-moving developer ecosystem, the same pattern this site has documented in the context of Claude Code’s own plugin ecosystem, where practitioners routinely compose agent workflows spanning several MCP servers without any centralized review of what’s actually connected to which external systems. Where one of those unreviewed servers is compromised or simply misbehaves, malicious instructions riding along in whatever content it processes have exactly the same reach that server’s original, unreviewed grant gave it.
Multiple MCP servers accumulating this way isn’t a policy failure exactly. It’s the absence of any point where a policy could have been applied, because nothing in the default workflow creates one, and no amount of security expertise on a central team compensates for a connection that team never knew to review. Security threats introduced this way don’t announce themselves at connection time. They surface later, when something the server was never supposed to reach turns out to be reachable.
A Governance Framework for MCP Servers: Registry, Isolation, and Trust Boundaries
The three risks above share a fix that’s less about any single control and more about treating MCP servers as inventory, not ambient infrastructure that exists once someone stands it up.
Centralized inventory and allow-listing starts with an MCP registry: an authoritative list of approved MCP servers, tracking tool metadata, who’s responsible for each one, and which version is currently sanctioned. Without that registry, individual MCP servers accumulate the same way any unowned system does, functional, unreviewed, and impossible to assign ownership to after the fact. Scoped authentication follows the same least-privilege logic already covered for tokens, but applied at connection time: an agent’s network access to a given server should be granted for a specific purpose, not inherited wholesale because a service account already had broad reach. Where a server genuinely needs a standing identity, service account tokens should carry only what that server’s function requires, nothing closer to admin-level reach “in case it’s needed later.”
Isolating runtimes in hardened containers matters specifically because it changes what a compromised tool can actually do next. A tool call that executes inside a properly isolated runtime, rather than a shared process with access to the host’s file systems, can be compromised without that compromise automatically becoming a path to everything else running alongside it, the same privilege-escalation logic behind the SharedRoot sandbox-escape disclosed against Claude Cowork’s local execution mode earlier in 2026. Trust boundary enforcement extends that isolation logic organizationally: an MCP environment handling internal databases or regulated data needs a harder boundary around it than one serving low-risk, read-only lookups, and collapsing both into the same trust zone because they happen to run the same protocol is how a low-risk connection ends up with an unearned path to sensitive systems.
Data leakage prevention circles back to context minimization from the mechanism section above: a tool call that returns only the fields a task actually needs limits what’s exposed if that specific response is ever intercepted or logged somewhere it shouldn’t be. This matters most for data agents, the tools whose entire function is querying data sources, since their normal operation already involves handling exactly the content a leak would expose. Taken together, these aren’t four unrelated best practices. They’re what compliance controls and security operations actually look like once “governance” stops being a document and starts being architecture: a registry that says what exists, scoped credentials that say what it can reach, isolation that contains what happens if it’s compromised anyway, and a trust boundary that keeps a low-risk mistake from becoming a high-risk one. None of this is about model quality or training data. It’s about whether the infrastructure around the model can enforce security policies and protect data regardless of what any single call decides to do.
The Controls: Access Controls, Centralized Monitoring, and MCP Authentication Standards
Each risk above has a specific answer, not a generic “improve security posture” gesture.
Access controls need to operate at the tool level, not just the server level. A security team that can say “this identity can call orders:read but not orders:write” has a working access policy. One that can only say “this identity is connected to the orders server” does not, because everything that server exposes is implicitly in scope. Fine grained access control at this granularity is the difference between least privilege as a stated principle and least privilege as something actually enforced at the point a tool executes. Governance policies written at the server level rather than the tool level are, functionally, no policy at all once a server exposes more than one capability.
Centralized monitoring means one place, not per-team visibility scattered across however many teams independently stood up their own MCP access. An organization that can’t produce, on request, a current inventory of every MCP server in use and what each one can reach doesn’t have monitoring. It has logging from whichever teams happened to set it up. Security events, a spike in tool-call volume from one identity, a call pattern that doesn’t match a server’s normal usage, only become detectable once they’re visible in one place instead of scattered across multiple systems nobody’s correlating.
Authentication and Authorization for AI Agents
Authentication and authorization solve different problems, and MCP governance has to treat them as two steps, not one. Authentication (MCP authentication) confirms which agent identity is making a call. Authorization confirms what that identity is allowed to do once confirmed, evaluated per tool call rather than granted once at connection time. Standardized authentication means OAuth 2.1 wherever the transport supports it, full stop, and an explicit, written policy for the STDIO-transport servers the spec itself excludes from that flow, rather than silently inheriting whatever the default happened to be. That policy has to say something specific: where environment-variable credentials are acceptable, where they aren’t, and who’s responsible for rotating them.
Getting this right for AI agents specifically, rather than treating them like any other service account, matters because an agent’s calling pattern isn’t fixed the way a service account’s usually is. Users authenticate once and their session persists; an agent’s authority needs re-evaluating call by call, because the sequence of tools it invokes in a given session isn’t scripted the way a traditional integration’s is.
Incident Response for MCP Security Threats
Responding to a compromised MCP server isn’t the same exercise as responding to a compromised API, and treating it as one misses the part that actually matters.
The confused deputy dynamic behind the GitHub MCP incident means the agent itself may keep acting with legitimate-looking, still-valid credentials after the initial compromise. Revoking access to “the server” isn’t sufficient if the actual exposure is a specific token scope that’s still valid against every other system it was ever granted. A working incident response runbook for this category needs three things in sequence: revoke the specific token and scope involved, not just disable the server; pull the tool-call-level audit trail to determine actual blast radius, which tools were invoked, with what parameters, in the window of compromise; and re-vet the server explicitly before reconnecting it, rather than assuming the same approval that covered it originally still applies.
That middle step is where the audit-trail gap from earlier becomes an operational cost, not just a compliance one. An organization that can only say a server was compromised, without being able to say which specific calls it made while compromised, is investigating blind.
The MCP Gateway as the Agent Governance Control Point
None of the four controls above are enforceable as a written policy. They’re enforceable as infrastructure that sits between an agent and the servers it calls, whether or not anyone remembers to check, and that infrastructure is what actually delivers centralized governance across an organization’s MCP deployment rather than leaving it to whichever team got there first.
Obot MCP Gateway is open-source under the MIT license, self-hostable on Kubernetes or Docker, and available as a managed service running the identical codebase. It enforces tool-level access policies rather than server-level ones, the same distinction described above, and recent releases have extended that enforcement further: tool call enforcement that evaluates individual calls against policy rather than trusting a session-level grant, MCP Tunnels for routing traffic to internal servers without exposing them directly, and Agent Auth Scopes that tie a specific identity’s permissions to what it’s actually allowed to invoke. It maintains audit logs at the tool-call level, the granularity the incident response section above depends on, and integrates with identity providers an organization already runs rather than introducing a parallel identity system for agent workloads alone. This is what agent governance looks like enforced rather than described: a gateway that knows, for every call, which identity made it and whether that call fell inside what governance policy actually permits, not a document describing what should happen.
The four risks in this article aren’t separate problems needing four separate tools. They’re symptoms of the same missing layer: a control point between agents and MCP servers that enforces what a policy document only describes.
FAQ
What is MCP server governance?
MCP server governance is the set of access controls, audit logging, and policy enforcement applied to the MCP servers an organization runs in production, distinct from the Model Context Protocol project’s own governance of the specification itself. It covers who an agent identity is, what it’s permitted to call, and what gets recorded when it does.
Is MCP server governance different from general API governance?
Yes, structurally. Traditional API security assumes a fixed, coded call pattern and benign client intent. MCP traffic is intent-based: an agent interprets a goal and decides which tool to invoke and with what parameters, which means governance has to account for agent behavior and injected intent, not just endpoint access.
What is a toxic agent flow?
A term coined by Invariant Labs after their May 2025 disclosure against the GitHub MCP server. It describes an attack where legitimate, authorized tool access is redirected by a prompt injection hidden in untrusted content the agent processes, rather than by any flaw in the MCP server or credential theft.
Does the MCP specification require authentication?
No. As of the July 2026 revision, the specification still treats authorization as optional and specifically directs STDIO-transport implementations away from its OAuth 2.1 flow, toward environment-variable credentials instead, a pattern assumed safe for local use that often persists after a server is exposed more broadly.
What does a real audit trail need to capture for MCP traffic?
Five things for every call: the agent identity that made it, which tool was invoked, what parameters were passed, what was returned, and whether human approval was required and obtained. Connection-level logging without tool-call parameters doesn’t meet this bar.
How is shadow MCP different from shadow IT?
The mechanism is the same, unsanctioned connections outside central visibility, but the consequence is different. Shadow IT typically exposes data. A shadow MCP server can take actions, since it’s granting an agent tool access, not just a login to a SaaS product.
What should incident response look like for a compromised MCP server?
Revoke the specific token and scope involved, not just the server connection. Pull the tool-call-level audit trail to determine what was actually invoked during the compromise window. Re-vet the server explicitly before reconnecting rather than assuming the original approval still holds.
Does an MCP gateway replace the need for these policies?
No, it’s the enforcement layer for them. The policies (who can access what, what requires approval, what gets logged) still have to be defined by the organization. A gateway is what makes those definitions operate on every call instead of depending on individual teams to implement them consistently.