AgenticLong read

MCP Server Architecture for Safe Agent Data Access

MCP servers function as trust boundaries, not just communication channels.

Staff Writer · · 11 min read
Cover illustration for “MCP Server Architecture for Safe Agent Data Access”
Agentic · October 7, 2026 · 11 min read · 2,414 words

Most unsafe agentic data deployments trace back to a single mistaken assumption: that an MCP server is a protocol bridge, a translator that moves requests from an AI agent to a data source and moves answers back. That assumption is wrong, and it is wrong in a way that has consequences. The Model Context Protocol exposes application capabilities, reading data, creating records, triggering downstream actions, to AI models through a standardized interface, and the moment that interface touches a database, it stops functioning as a communication channel and starts functioning as a trust boundary [1][2][3]. A channel only needs to move bytes reliably. A trust boundary has to decide, on every single call, what an agent is allowed to see, what it is allowed to change, and what happens if it is tricked into asking for something it should not have. Treating the two as interchangeable is the architectural error this piece sets out to correct, and the industry's own data shows how far that error has already traveled: the Cloud Security Alliance's Agentic MCP Security Best Practices guide found that by early 2026, roughly half of nearly 7,000 internet-exposed MCP servers were running with no authentication controls. The protocol spread faster than the governance needed to make it safe.

The three trust boundaries every MCP architecture must account for

MCP's architecture does not create one point of trust to secure. It creates three, and each can fail independently of the other two. A team that locks down one and assumes the system is now safe has only closed off a third of the problem.

The first is the model boundary, the handoff between the LLM and the MCP client. The model reads tool descriptions and constructs its own invocations from them, but it has no way to independently confirm that those descriptions are accurate or that they have not been altered. The CSA guide identifies this as the layer where tool definition manipulation and malicious tool descriptions do their damage: if the text describing a tool lies about what that tool does, the model has no mechanism for catching the lie before it acts on it.

The second is the client-server boundary, the connection between the MCP client and the MCP servers it talks to. The client is supposed to authenticate itself to each server and validate what comes back. The CSA guide found that many implementations handle this poorly. The protocol specification has tried to close that gap: the March 2025 MCP specification (version 2025-03-26) introduced OAuth 2.1 as the authentication standard, and the November 2025 revision (version 2025-11-25) refined it further. Adoption on the ground has not caught up with the standard on paper.

The third is the server-to-data boundary, where an MCP server reaches into file systems, databases, external APIs, and cloud services on behalf of the model that called it. This is the boundary where production database exposure actually happens, because the server is acting as an agent with whatever permissions it has been given, and those permissions are frequently broader and less carefully scoped than anyone intended.

None of these three substitute for one another. A server that authenticates perfectly at the client-server boundary with OAuth 2.1 can still reach a production database with god-mode permissions if the server-to-data boundary was never addressed. The CSA guide's 2025 incident record shows how varied the entry points can be: cross-tenant data exposure at Asana, a prompt injection attack against the GitHub MCP server, unauthenticated remote code execution in Anthropic's own MCP Inspector tool, and multiple supply chain compromises delivered through malicious npm packages. Each incident exploited a different boundary. The Cloud Security Alliance also recorded more than 30 CVEs filed against MCP servers, clients, and infrastructure components between January and February 2026 alone, a volume that confirms the three-boundary model is not a theoretical framework but a map of where real attacks are already landing.

Diagram: Three Trust Boundaries, Six Threat Classes. Visualizes: Visualize the three distinct MCP trust boundaries as a left-to-right chain: (1) Model Boundary — between the LLM and the MCP client, vulnerable to tool definition manipulation and…

The six threat classes that exploit those boundaries in practice

Knowing which boundary an attack targets is only useful once it is paired with knowing which threat class the attack belongs to. Generic "MCP security" thinking tends to produce a control that blocks one threat while leaving the rest of the field open, so the threats need to be named individually.

A supply chain attack comes through a malicious or compromised MCP server, and under casual inspection it looks ordinary. It can steal credentials the agent holds for other connected servers, pull context out through exfiltration, or push the agent to take unwanted actions in systems it can reach. It is the same risk category long associated with a malicious npm package, carried over into the MCP ecosystem.

A confused deputy attack happens when one MCP server manipulates the agent so it misuses credentials or capabilities it holds for a different server. The downstream server has no way to tell a legitimate agent action from an attacker-directed one, because the agent is presenting valid credentials either way.

Data exfiltration through misconfigured tool outputs can happen without a malicious actor. The server just returns more than it should: full database rows where only a field was needed, directory listings that include sensitive paths, environment variables spilled into code execution results. It lacks the drama of an active intrusion, but the CSA guide notes it is the more common failure mode of the two.

Privilege escalation via overprivileged tool scopes grows out of ordinary development habits. A server gets configured with broad permissions so engineers can move fast during testing, and then it gets promoted into production before anyone pulls those permissions back down. The agent then inherits everything the server was ever given.

Uncontrolled MCP deployment is an organizational threat. A developer downloads a server, wires it up with a personal access token carrying more permission than the task needs, and runs it locally. Multiplied across a team, that behavior produces hundreds of credential-holding processes that nobody in the organization can see or account for.

System-prompt rules and LLM judgment cannot enforce data boundaries

Many teams believe they have solved data access by writing the right instruction into the system prompt: never touch payment data, never write to the accounts table, always ask before deleting. That belief does not hold up, because an instruction in a system prompt is not a security control. The threat model built around MCP treats it as a request, one the model may or may not honor.

The reason is structural. System-prompt instructions and tool results enter the same context window, and a prompt injection payload embedded in a tool result can contradict or override those instructions once it is sitting in that window. The model cannot tell whether an instruction came from its operator or was smuggled in by an attacker through a document, a webpage, or a database field it was asked to read.

The failure plays out concretely. Tell a customer support agent never to access payment data, and a malicious tool result can still steer it into calling a payment-data tool, as long as the server holding that tool sits anywhere within the agent's reachable scope. The instruction was advisory. It was never enforced.

That leaves one workable principle for architecture: every access control that actually matters has to sit at a layer the LLM cannot reach around. That means the MCP server's own permission scope, row-level security enforced inside the database, and the allowlist enforced at a gateway, not language typed into the model's instruction set. The four structural controls that follow are what that principle looks like when it is built out in practice.

The four structural controls that enforce safe data access at the MCP layer

Diagram: Four Structural Controls and Where They Enforce. Visualizes: Show four controls mapped to the layer of the stack where each enforcement actually lives, making clear that none of them operates inside the LLM's context window.

Safe agent data access rests on four structural controls working together. Each one closes a gap the others cannot close alone, and skipping any single one leaves the system exposed at the boundary that control was meant to cover.

The first is least-privilege scoping at the server layer. A customer support agent that looks up orders needs read access to the orders table and nothing else: not write access, not access to user accounts, not access to payment data. That scope has to be defined in the MCP server's own tool configuration, not inferred from whatever the agent claims its purpose is. The stronger MCP platforms inherit fine-grained access controls directly from the source system itself, so an agent's permissions are always bounded by what the authenticated user behind it is actually allowed to see, rather than by what the agent decides it should be able to reach.

The second is the use of read replicas and governed datasets as a physical separation layer. The only real guarantee against AI-driven data destruction or an unintended write to production is physical separation: agents query a replica or a dataset built specifically for their use, keeping the production primary out of their path. A permission setting on the production database is a configuration choice, and configuration choices can be changed, misapplied, or bypassed. A production database that is simply unreachable from the agent's network path is a structural guarantee, and structural guarantees do not depend on someone remembering to configure them correctly.

The third is OAuth 2.1 with audience-validated token scoping. The CSA guide treats the November 2025 MCP specification's formal adoption of OAuth 2.1 as a genuine step forward for the protocol's security. Tokens issued under that standard should be scoped to a specific audience, the particular MCP server they were meant for, so you cannot lift a token and replay it against a different server in a confused deputy attack. Non-predictable session IDs, routine credential rotation, and multi-factor authentication on administrative actions round out what a mature authentication layer looks like.

The fourth is data-layer tokenization and PII redaction applied at the MCP response layer itself. Even if a tool call is correctly authenticated and correctly scoped, it can still return protected health information, payment card data, or other personally identifiable information straight into the agent's context window, where it becomes part of the model's reasoning and can resurface in its output. Strac's analysis of Box MCP deployments identifies this as the real governance gap in the stack: the moment an agent reads a file through MCP, that data has entered the model's context outside the reach of traditional data-loss-prevention tooling. An MCP proxy positioned to intercept every tool response can classify sensitive data as it passes through and tokenize or pseudonymize it before the agent ever sees it, so real PII, PHI, and PCI never enter the LLM's context at all, while the agent keeps working with realistic, format-preserving synthetic values that let it do its job without ever touching the real thing.

The gateway architecture: centralizing policy enforcement across multiple MCP servers

A single well-configured MCP server is a manageable problem. Dozens of them, scattered across teams and configured inconsistently, are not. Per-server configuration becomes ungovernable once the number of servers grows past what a small group can track by hand, and the only way to keep policy consistent at that scale is to enforce it centrally, through a gateway that sits between every agent and every server it is permitted to reach.

A gateway does three things structurally. It enforces an allowlist of approved MCP servers, so a server that has not been vetted cannot reach enterprise resources no matter how a developer has configured their own local environment. It centralizes access control and role management across every connected server, so a permission change becomes a single organizational action instead of a scramble through dozens of configurations. And it inspects every tool invocation before that invocation is allowed to execute, which places the actual enforcement point between the agent and anything it might otherwise be able to touch.

The audit log produced at that enforcement point is not an afterthought. CData's best practices guidance specifies that every agent action should generate a structured log entry carrying the identity of the agent making the request, the specific tool it invoked along with the input parameters it passed, and the result that came back along with any errors raised. That is the record that turns a suspected incident into an investigable one and turns a compliance claim into something that can actually be demonstrated. Paired with continuous monitoring of that log stream, an organization can catch the behavioral signatures of trouble, an agent issuing an unusual volume of queries, reaching into tables outside its normal scope, or showing a pattern consistent with exfiltration, before any of it escalates into an incident.

The uncontrolled deployment problem is the clearest argument for why this centralization has to exist. If there is no gateway enforcing an allowlist, an organization's real MCP exposure is the sum of every server any developer has installed locally and every personal access token any of them has ever authorized. That surface grows without anyone watching it grow, and by the time it is visible, it is too large to govern after the fact.

Implications for teams building on Supabase and Postgres

If you run a relational database with built-in access controls, you already sit on top of every control described above, because that database gives you the primitives those controls are built from. Row-level security is a native feature of some relational databases, and it is exactly the kind of enforcement that has to live at a layer the LLM cannot talk its way around: a policy written into the database itself holds regardless of what the agent's system prompt says or what a prompt injection manages to convince the model to attempt. Read replicas are a standard deployment pattern for such databases, and pointing an MCP server at a replica rather than the primary turns the physical-separation control from a theoretical recommendation into a configuration that can be stood up with existing tooling.

The remaining work is at the MCP layer itself: scoping each server's tool definitions to the narrowest set of tables and operations a given agent actually needs, putting OAuth 2.1 and audience-scoped tokens in front of every server rather than relying on a single shared credential, and deciding, before any agent is connected to a production database, whether a tokenization layer needs to sit between the database and the model's context window. None of that requires exotic infrastructure. It requires treating the MCP server not as a convenience that was added on top of the database, but as the governance layer that now decides what an agent is allowed to know.

Sources

  1. Agentic MCP Security Best Practices Guide
  2. TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
  3. When MCP Servers Attack: Taxonomy, Feasibility, and Mitigation
  4. Proxy-based secure model context protocol server access for artificial intelligence agents
  5. A First Measurement Study on Authentication Security in Real-World Remote MCP Servers
  6. Runtime Policy Enforcement for MCP-Based LLM Agents
  7. MCP gateway architecture: A complete technical guide
Filed underAgentic

More in Agentic