Most enterprise AI governance frameworks—the EU AI Act, NIST AI RMF, ISO/IEC 42001—were built to govern models: their training data, outputs, documented risk classification, and the core principles that help ensure AI operates transparently, accountably, fairly, securely, and in compliance with legal and organizational standards. None of them were designed to answer for an agent’s runtime decision to call an MCP tool, leaving the entire tool layer effectively ungoverned even in organizations with mature AI governance programs.
An organization can have its EU AI Act risk tiers mapped, ISO/IEC 42001 certification underway, and a model review board that meets every month, and still have no answer to a basic operational question: which of the forty-odd MCP servers its agents can reach were ever vetted for anything beyond whether the team that added them remembered to write documentation. The governance program is real. It was built to answer questions about models. It was never built to answer questions about tools.
AI governance, in the form most organizations have actually implemented, is a system of policies and review processes aimed at a specific target: the model, its training data, its outputs, and its documented risk tier. For enterprise IT and security leaders, AI governance teams, and organizations managing multi-component AI systems, that definition has now transformed. Once an agent can discover a tool at runtime and call it, the governance program has to answer for an action an AI agent took, not just an output a model produced. That gap creates unmanaged security risk, weakens auditability, and raises compliance exposure as AI adoption scales.
This article examines where general AI governance frameworks stop at the model boundary, which principles of AI governance still apply when MCP tools enter the picture, and what an MCP-specific governance layer needs to include: policy, risk classification, registration-time vetting, runtime monitoring, audit trails, and a practical maturity model for managing the tool layer.
Why General AI Governance Frameworks Stop at the Model Boundary
Effective AI governance exists because the risks it manages are real and already documented, not hypothetical. AI models trained on skewed data can perpetuate existing biases into hiring, lending, and other high-stakes decisions, producing discriminatory outcomes an organization is legally and reputationally exposed to. A model that infers sensitive information from indirect signals, inferring health status from purchase history, for instance, raises data protection questions regardless of whether that inference was the system’s intended purpose. Yet confidence and readiness haven’t kept pace with each other: per IBM’s AI governance overview, 80% of business leaders cite explainability, ethics, bias, or trust as a major roadblock to generative AI adoption, and 80% of organizations have a dedicated AI risk function. Formal structure alone hasn’t closed the gap between innovation and regulation. AI governance matters because none of these risks resolve themselves as adoption scales: they compound.
The major AI governance framework options in production today are genuinely well built for what they cover, including regulatory compliance, and it’s worth being precise about what that is before pointing at the gap. What is AI governance framework design supposed to answer, at its core: who is accountable, what gets reviewed, and on what evidence. Only 19% of organizations have full visibility into AI usage. The frameworks below answer that question thoroughly for one specific unit, the model.
The EU AI Act, the world’s first comprehensive AI regulation, classifies systems into four risk management tiers, unacceptable, high-risk, limited-risk, and minimal-risk, and attaches conformity assessments, transparency obligations, and human oversight mandates to high-risk AI applications. Its high-risk AI systems provisions were originally set to take effect August 2, 2026, though the Digital Omnibus agreement reached in May 2026 pushed the deadline for most standalone high-risk use cases to December 2, 2027, according to Gibson Dunn’s analysis of the Omnibus agreement. The deferral moved a date. It didn’t change what the Act actually regulates: the model’s classification, its training data, and its documented risk profile.
The NIST AI Risk Management Framework, published January 26, 2023, organizes work into four functions, Govern, Map, Measure, and Manage, each broken into categories an organization tailors to its own environment, per NIST’s own AI Risk Management Framework resource center. It’s voluntary, sector-agnostic, and widely referenced as a starting point for responsible AI programs that operationalize governance principles and ethical principles. It’s also, like the EU AI Act, built around evaluating an AI system’s design, training, and deployment, not around what happens when that system reaches outside itself to call a tool mid-task.
ISO/IEC 42001, published December 2023 as the world’s first certifiable AI management system standard, gives auditors a formal way to evaluate whether an organization’s AI governance practices actually function day to day, not just whether they exist on paper, per ISO’s own standard listing. It covers data quality, continuous improvement, lifecycle management for the AI system as a whole, and the data governance needed to support compliant oversight.
Beyond these three, the OECD AI Principles, first adopted in 2019 and updated on May 3, 2024 to address risks specific to generative and general-purpose AI, are the closest thing to a global baseline for trustworthy ai, endorsed by 47 governments and built around five values including transparency and explainability, reflecting broader economic co-operation on responsible standards. They also help frame ai ethics, ethical development, and responsible ai development in policy terms organizations can apply. And regional data protection law shapes what an organization’s governance program has to enforce even when it says nothing about AI directly: India’s Digital Personal Data Protection Act, enacted August 11, 2023, governs how sensitive data used in AI systems can be collected and processed within India, with implementing rules notified as recently as November 2025. Outside the EU, Canada’s Directive on Automated Decision-Making is a government example of ai governance regulations.
All of these frameworks share the same structural assumption: the unit being governed is the model, evaluated at points along its entire AI lifecycle, training through retirement. None of them were designed around a model that discovers a tool it’s never seen before, decides to call it based on its own reasoning, and takes an action with consequences outside the conversation. That’s not a criticism of the frameworks. It’s a description of what they were built to regulate, and MCP wasn’t a mainstream enterprise AI governance concern when any of them were drafted.And the gap between how confident organizations are and how ready they actually are is well documented: research from the IBM Institute for Business Value found that 80% of business leaders see AI explainability, ethics, bias, or trust as a major roadblock to generative AI adoption, according to IBM’s own AI governance overview. AI governance matters because none of these risks resolve themselves as AI adoption scales. They compound.
The major AI governance framework options in production today are genuinely well built for what they cover, and it’s worth being precise about what that is before pointing at the gap. What is AI governance framework design supposed to answer, at its core: who is accountable, what gets reviewed, and on what evidence. The frameworks below answer that question thoroughly for one specific unit, the model.
The EU AI Act, the world’s first comprehensive AI regulation, classifies systems into four risk management tiers, unacceptable, high-risk, limited-risk, and minimal-risk, and attaches conformity assessments, transparency obligations, and human oversight mandates to the high-risk category. Its high-risk AI systems provisions were originally set to take effect August 2, 2026, though the Digital Omnibus agreement reached in May 2026 pushed the deadline for most standalone high-risk use cases to December 2, 2027, according to Gibson Dunn’s analysis of the Omnibus agreement. The deferral moved a date. It didn’t change what the Act actually regulates: the model’s classification, its training data, and its documented risk profile.
The NIST AI Risk Management Framework, published January 26, 2023, organizes work into four functions, Govern, Map, Measure, and Manage, each broken into categories an organization tailors to its own environment, per NIST’s own AI Risk Management Framework resource center. It’s voluntary, sector-agnostic, and widely referenced as a starting point for responsible AI programs. It’s also, like the EU AI Act, built around evaluating an AI system’s design, training, and deployment, not around what happens when that system reaches outside itself to call a tool mid-task.
ISO/IEC 42001, published December 2023 as the world’s first certifiable AI management system standard, gives auditors a formal way to evaluate whether an organization’s AI governance practices actually function day to day, not just whether they exist on paper, per ISO’s own standard listing. It covers data quality, continuous improvement, and lifecycle management for the AI system as a whole.
Beyond these three, the OECD AI Principles, first adopted in 2019 and updated on May 3, 2024 to address risks specific to generative and general-purpose AI, are the closest thing to a global baseline for trustworthy ai, endorsed by 47 governments and built around five values including transparency and explainability. And regional data protection law shapes what an organization’s governance program has to enforce even when it says nothing about AI directly: India’s Digital Personal Data Protection Act, enacted August 11, 2023, governs how sensitive data used in AI systems can be collected and processed within India, with implementing rules notified as recently as November 2025.
All of these frameworks share the same structural assumption: the unit being governed is the model, evaluated at points along its entire AI lifecycle, training through retirement. None of them were designed around a model that discovers a tool it’s never seen before, decides to call it based on its own reasoning, and takes an action with consequences outside the conversation. That’s not a criticism of the frameworks. It’s a description of what they were built to regulate, and MCP wasn’t a mainstream enterprise AI governance concern when any of them were drafted.
An MCP Governance Framework, Applied at Each Layer
An effective AI governance framework for organizations running MCP in production needs five components, each one a direct answer to a question the model-level frameworks above don’t ask, and together helping support responsible AI governance across the tool layer.
1. Policy and Standards: What “Approved Server” Actually Means
A policy that says “only use approved MCP servers” is not a policy until “approved” has a concrete definition aligned with ethical guidelines and business and regulatory expectations: who can register a server, what its tool descriptions have to disclose before approval, and what happens when a previously approved server’s tool definitions change. Most organizations that think they have this covered actually have an informal convention, not a governance policy anyone could point to during an audit, and consistent governance requires the same approval standard across teams and servers.
2. Risk Classification Scoped to What a Server Can Reach
Risk assessment at the MCP layer isn’t a generic AI-risk category. It’s a direct function of what a server is authorized to touch: a read-only connection to an internal wiki carries a different risk tier than write access to a production database or a payments system, regardless of which model is calling it, and classification should account for the likely AI outcomes of that access, not just the connection type. A risk management framework built for MCP has to tier by reach, not by model provider or use case, which distinguishes this analysis from traditional model risk management.
3. Registration-Time Vetting, Before a Server Is Discoverable
The AWS-backed, open-source MCP Gateway and Registry project (a separate initiative from Obot’s own gateway) runs three Cisco AI Defense scanners against every asset at the moment it’s registered, with these checks acting as governance controls embedded before discovery: one analyzing tool definitions for malicious patterns, one checking agent-to-agent specification compliance, one inspecting instruction files for injected content and exposed credentials, disabling anything that fails before an agent can ever connect to it. That’s a governance control most AI governance programs don’t have an equivalent for at the model level, because a model doesn’t get “registered” the same way a tool does, and it supports responsible AI practices by preventing unvetted tool exposure early.
4. Runtime Monitoring at the Tool-Call Level
Ongoing monitoring for a model usually means output quality, drift, and bias metrics, but when AI-driven decisions trigger tool calls rather than just model outputs, runtime monitoring becomes necessary. Monitoring an MCP deployment means something structurally different: the OWASP Top 10 for Agentic Applications, finalized in December 2025 with input from over 100 security researchers and practitioners including contributors from NIST and the European Commission, names tool misuse and agentic supply chain manipulation as distinct risk categories precisely because they only become visible once a tool is live and in use, not at a one-time review. Continuous monitoring here means watching whether a tool’s behavior has drifted from what it was approved to do, not whether a model’s accuracy has drifted, and it is also how organizations verify transparent AI systems in practice at the tool-call level.
5. Audit Trails and Lifecycle Management That Survive an Agent’s Retirement
Audit trails at the tool-call level, and credential revocation when an agent or server is retired, close the loop the other four components open by letting organizations trace AI applications and agent actions after deployment and retirement. A registration policy without an audit trail is a one-time approval with no record of what happened after. A risk tier without lifecycle management is a classification nobody revisits when the server it describes changes hands or gets deprecated, which helps limit unintended consequences.
👉 Obot helps enterprises operationalize MCP governance with centralized control, observability, and secure connector management. Try Obot today.
Where CVE-2026-32211 Fits Into a Governance Framework, Not Just a Postmortem
CVE-2026-32211, a missing-authentication flaw Microsoft rated CVSS 9.1 in the Azure MCP Server, is useful here specifically because it shows what a model-level governance program would have missed entirely. A risk-tiering process built around model classification, with governance also covering how AI operates through tools, a review board evaluating training data and output quality, an EU AI Act conformity assessment focused on the model’s risk category: none of these processes would have had a reason to inventory the Azure MCP Server at all, because it isn’t a model. It’s a tool a governed model was allowed to call, and the vulnerability lived in a layer the governance program was never built to see.
That’s the concrete version of the abstract argument in the previous section. A framework can be fully implemented, audited, and compliant with every applicable regulation, and still have zero coverage of the specific layer where this vulnerability sat. Even when the model program is compliant, governance gaps at the tool layer can still undermine trustworthy AI systems.
A Maturity Model for MCP Governance Specifically
General AI governance maturity models typically run from ad hoc to fully integrated, evaluating how consistently policies apply across an organization as AI systems evolve and AI initiatives multiply, with responsible AI adoption as the broader outcome those models are usually meant to track. An MCP-specific version of that same idea needs different criteria, because “consistent policy” means nothing if the policy never accounted for tools in the first place.
Stage
What it looks like at the MCP layer
No registry
Developers connect to MCP servers directly, no central record of what’s running or who approved it
Informal allowlists
A spreadsheet or wiki page lists approved servers, updated inconsistently, no enforcement mechanism tied to it
Per-tool policy
A gateway enforces access control at the individual tool level, not just the server level, with identity-scoped permissions
Full lifecycle coverage
Registration-time vetting, runtime monitoring, and tool-call audit trails operate together, with credential revocation built into agent and server retirement
Most organizations with a mature general AI governance program are still somewhere between the first two stages specifically for MCP, because their governance maturity was measured against a framework that never asked this question, and true maturity also depends on applying governance principles consistently to MCP, not just to models.
What This Forces IT and Security Leaders to Decide
Extending an existing AI governance program to cover MCP isn’t a matter of adding a paragraph to an existing policy document. It forces real ownership decisions, the kind a RACI matrix exists to force into the open rather than leaving assumed: who is responsible for the MCP registry day to day, who is accountable when a server turns out to have been mis-scoped, who gets consulted before a new risk tier ships, and who simply needs to be informed after the fact. Establishing an AI governance committee with explicit authority over MCP-specific decisions, distinct from whoever chairs the model review board, is what turns that matrix from a diagram into an enforced structure. What risk tier requires a human sign-off before a server goes live, and who has the authority to grant it, has to be decided by that structure, not inferred after an incident.
How MCP audit events get folded into whatever compliance reporting already exists for the broader AI governance programs an organization runs is the other open question. A security review that asks for complete audit trails and gets model-level logs with a gap where tool activity should be is not going to satisfy anyone auditing against ISO/IEC 42001 or a regulator applying regulatory requirements under the EU AI Act.AI Governance Best Practices Applied to the MCP Layer
What is enterprise AI governance framework best practices at the model level usually translates into some version of: inventory what exists, assign ownership, classify by risk, embed review into existing workflows, monitor continuously. Those same five moves hold up when implementing AI governance at the MCP layer, helping organizations apply AI ethics at the MCP layer, not only at the model layer. What changes is what each one actually has to inventory, own, and classify.
Start with an inventory that includes every MCP server already connected, not just the ones a formal request went through. Assign an owner for the registry itself, separate from whoever owns model governance, since the two roles require different technical context. Classify by what a server can reach rather than by a generic risk label borrowed from model governance. Build registration-time and runtime checks into the same pipeline developers already use to add a new server, rather than a separate approval process people route around under deadline pressure. Monitor tool-call behavior continuously, the way the OWASP Agentic Top 10 framework above describes, not just at the point a server was first approved.
None of this replaces human review for the decisions that warrant it. Ethical considerations and business risk don’t disappear just because a check is automated, and a robust AI governance framework still routes the genuinely ambiguous cases, a server requesting broader access than its stated purpose implies or access that could conflict with societal values, to a person with the authority to say no. Ai governance tools that automate registration-time scanning and runtime monitoring exist specifically to make that human review possible at scale, by filtering out the routine approvals so the ambiguous ones actually get attention instead of being rubber-stamped alongside everything else.
Organizations that already run a mature general AI governance program have most of the organizational muscle this requires. They’ve already done the harder work of aligning organizational values with business objectives and building oversight processes that hold up under audit. What they’re usually missing isn’t governance instinct. It’s a technical layer that makes MCP-specific enforcement possible in the first place, since a policy document alone doesn’t stop an unvetted server from being reachable, no matter how clearly AI governance policies describe what should have happened instead.
AI Governance Best Practices Applied to the MCP Layer
What is enterprise AI governance framework best practices at the model level usually translates into some version of: inventory what exists, assign ownership, classify by risk, embed review into existing workflows, monitor continuously. Those same five moves hold up when implementing AI governance at the MCP layer, helping organizations apply AI ethics at the MCP layer, not only at the model layer. What changes is what each one actually has to inventory, own, and classify.
Start with an inventory that includes every MCP server already connected, not just the ones a formal request went through. Assign an owner for the registry itself, separate from whoever owns model governance, since the two roles require different technical context. Classify by what a server can reach rather than by a generic risk label borrowed from model governance. Build registration-time and runtime checks into the same pipeline developers already use to add a new server, rather than a separate approval process people route around under deadline pressure. Monitor tool-call behavior continuously, the way the OWASP Agentic Top 10 framework above describes, not just at the point a server was first approved.
None of this replaces human review for the decisions that warrant it. Ethical considerations and business risk don’t disappear just because a check is automated, and a robust AI governance framework still routes the genuinely ambiguous cases, a server requesting broader access than its stated purpose implies or access that could conflict with societal values, to a person with the authority to say no. Ai governance tools that automate registration-time scanning and runtime monitoring exist specifically to make that human review possible at scale, by filtering out the routine approvals so the ambiguous ones actually get attention instead of being rubber-stamped alongside everything else.
Organizations that already run a mature general AI governance program have most of the organizational muscle this requires. They’ve already done the harder work of aligning organizational values with business objectives and building oversight processes that hold up under audit. What they’re usually missing isn’t governance instinct. It’s a technical layer that makes MCP-specific enforcement possible in the first place, since a policy document alone doesn’t stop an unvetted server from being reachable, no matter how clearly AI governance policies describe what should have happened instead.
Where Obot Fits
Obot MCP Gateway is not a replacement for an organization’s broader AI governance program, and it doesn’t try to be. It doesn’t do bias testing, model cards, or EU AI Act conformity assessments. What it does is make the five components in this framework enforceable at the infrastructure level rather than aspirational on a policy document: a curated MCP catalog for the registration and policy layer, tool-level RBAC scoped to individual tools rather than whole servers, OAuth 2.1 enforcement with server-side credential brokering so raw credentials never reach an MCP server directly, and an audit trail that operates at the tool-call level with exportable logs.
Because Obot sits in front of MCP servers as a gateway rather than living inside any one client, the policy and audit layer is client-agnostic, and gateway-layer filtering can screen PII in tool responses before it ever reaches the agent. Obot is open-source under the MIT license, self-hostable on Kubernetes or Docker, or available as a managed service, same product either way. It connects to an organization’s existing identity providers, including Okta, Microsoft Entra, Google, and GitHub, with Okta and Entra connections gated to the Enterprise edition. For organizations that already have a functioning AI governance program covering models and data, Obot is the layer beneath it specifically for MCP, not a second governance program running in parallel.
Where This Leaves You
General AI governance frameworks are maturing fast, and the ones built by regulators and standards bodies are genuinely well constructed for what they regulate. Governance in AI terms usually still means governance of the model, built to reflect societal expectations about how AI technologies should be developed and used AI responsibly. What they regulate stops at the model boundary. The tool layer an agent reaches through MCP needs the same structural rigor, accountable AI governance applied specifically to what a tool can do and who approved it, not inherited by assumption from a framework that was never built to see it.