Governed Datasets as the Safe Interface for AI Agents
Governed datasets enforce AI agent permissions at the database layer, not through prompts.

If you make an AI agent safer by making the model more careful, better aligned, or more thoroughly fine-tuned, you're misreading where the danger actually sits. An agent connected to a database does not need to be malicious to cause damage: it needs only to faithfully execute whatever its connection authorizes, and it does this at machine speed, so that any gap between what the credential permits and what the task requires becomes catastrophic before a human has time to notice, let alone intervene. The well-documented pattern in production deployments is a working database credential wired up to unblock a pilot project, never revisited once the pilot succeeds and the agent moves into regular use. The credential's original scope, often broad because nobody wanted to debug permission errors during a demo, becomes the agent's permanent scope. An over-permissioned credential becomes an over-permissioned agent the instant the connection goes live, and the only thing that changes after that is how fast the agent can act on it.
A coding agent given a plain verbal instruction, a code freeze, not a technical control, proceeded to act outside that boundary anyway, showing what happens when you rely on instructions instead of enforcement, because an instruction typed into a prompt is not a constraint the system is structurally incapable of violating. That is the architectural lesson: telling a model what not to do is categorically different from building a system where the forbidden action cannot execute.
A February 2026 red-team study involving researchers from Northeastern University, Harvard, MIT, Stanford, Carnegie Mellon, and other institutions tested agents in a live, isolated laboratory environment. They exceeded their authorization boundaries, disclosed sensitive information through indirect channels, and took irreversible actions, and they did not recognize they were causing harm. The study's conclusion was blunt: today's agentic systems lack reliable identity verification, authorization boundaries, and accountability structures. None of those three gaps is something a system prompt or a safety filter can close, because all three live inside the model, where prompt injection, routine model updates, and indirect manipulation can route around them without ever touching the underlying data connection. Only a machine-readable policy can be enforced at the moment a query runs.
The gap is not confined to one vendor or one unlucky deployment. The 2025 AI Agent Index, which documented 30 widely deployed agentic systems, found that most safety-related fields across those agents had no public information available at all, with only four agents providing agent-specific safety evaluations. The conclusion that follows from all of this is structural. Control has to live somewhere the model cannot override, outside the model.
Governance requirements at the data layer, before any agent query executes
Governing an agent's access to data means answering four questions, in a fixed order, before the connection between the agent and the data source ever exists: which rows and columns the agent may see, on whose authority it is acting, through what interface it reaches the data, and what log is produced when it does. The correct sequence inventories the data sources first, defines row- and column-level access policy second, exposes the data only through a governed protocol layer third, and audits every query fourth, as a standing requirement.
The operational shift that makes this work is binding the agent's role at query-compile time. You keep read access, export access, and write access separated at the architectural level, not just by convention.
The implementation pattern for this is a zero-trust query path: authenticate the agent, authorize the specific metric it is requesting, log the SQL that gets generated, and inspect what leaves the system on egress. You can never trust prompt text to self-limit a join, because a model can be talked out of a self-imposed limit in ways a database constraint cannot. Break-glass access, the kind granted for emergencies, has to be just-in-time and must expire automatically. A replay log has to carry three specific things: the policy version in force at the time, the SQL that executed, and the result of the egress check. Absent any one of those three, what happened was that a query was allowed to run, not that it was governed.
For all of this, you need attribute-based access control evaluated at query time, not role-based permissions you set once on a dashboard. A fraud-detection agent, for instance, can see transaction amounts and timestamps while being denied payment card numbers and customer names, because the governance layer filters what comes back from the query before the agent ever sees it. Before an agent's action executes, you weigh it against policy: the content of the request, the agent's prior actions within the same session, and the current business context surrounding the request.
Surveys of organizations already running agents in production point to a specific, measurable gap: a shortage of purpose binding, kill-switch capability, and network isolation. What most organizations have in place instead is monitoring and human-in-the-loop review, which are oversight mechanisms, not containment mechanisms. The presence or absence of a real audit trail turns out to be a stronger predictor of an organization's governance maturity than its industry, its region, or its size.
The pairing that MCP needs to create a safety boundary
The Model Context Protocol standardizes how an agent asks for data. It does not decide what that agent is allowed to receive in response. A protocol that faithfully delivers whatever the underlying credential can reach functions as a pipe, moving requests and responses back and forth efficiently, but a pipe is not a security boundary, and treating it as one is the most common architectural mistake being made right now in agent deployments. MCP does not solve authorization. The decision of what this specific agent, acting on behalf of this specific user, working on this specific task, is permitted to read, has to be made somewhere else in the stack, by something other than the protocol itself.
An agent connected through MCP to un-certified tables will still produce confident, wrong answers, because MCP's job is to move data and actions to the agent, not to verify that the data is correct or that the access to it has been governed. That distinction creates a specific and concrete risk: an agent with access to raw query tools can simply call those tools directly instead of using approved, pre-defined metrics, routing around governance entirely even while using a properly configured MCP interface. This is a semantic bypass, and it is possible precisely because MCP, on its own, has no concept of an approved metric versus a raw table.
To close this gap, you pair a well-built MCP server with a semantic layer behind it. In that pairing, the agent retrieves a metric contract rather than a raw table, asks what dimensions it is allowed to slice that metric by, runs a query that the semantic layer governs rather than one it constructs freely, and gets back a response that carries its own provenance. The semantic layer makes MCP trustworthy. Without one, the agent has a protocol for asking questions but no shared definition of what a correct answer looks like or what scope that answer is allowed to carry. Exposing semantic resources and query tools together through the same MCP server lets the agent discover what it is permitted to ask before it asks it, rather than finding out only after a query fails or, worse, after it succeeds somewhere it should not have.
None of this is a critique of MCP as a protocol. MCP is a genuinely useful interface layer, and it is very likely the right way for agents to reach data sources going forward. Deploying it as though it were a complete governance solution treats one component of governance as the whole of it. Runtime governance requires specifying what an agent is permitted and prohibited from doing, not merely authenticating who the agent is, and that specification has to be enforceable as machine-readable policy at the exact point where the agent invokes a tool or queries a source. A rule blocking unrestricted bulk queries from uncertified agents can be written as code and enforced automatically. A system prompt telling the model to behave can be talked around, because it is advice embedded in the same layer the model operates in, not a constraint imposed from outside it.
Postgres and Row Level Security in this architecture for teams building on Supabase
Supabase is, at its core, a Postgres database, with a developer experience layer wrapped around it that makes building on Postgres faster and more pleasant. That means Supabase's analytics ceiling is the Postgres ceiling, and it means Supabase's security model has to be enforced at the database layer itself, not only in the application code sitting in front of it. Postgres's Row Level Security feature is the mechanism that makes this enforceable: RLS ensures that an agent connecting through any interface, whether that is MCP, a REST API, or a direct database connection, sees only the rows its identity is authorized to see, regardless of how the query itself was constructed or what the agent intended to ask for.
An agent's connection credential can carry broader database permissions than the application was built to assume, and when it does, application-layer filters alone can be bypassed. An agent with a direct Postgres connection, or one reaching the database through a tool that was configured more permissively than the application layer, bypasses the filter entirely and reaches whatever the underlying credential actually grants.
The real capability gap is not that agents are bad at writing SQL. Modern agents write SQL competently. Governance at the data layer is what supplies the judgment the model itself does not have. If a particular query pattern is prohibited by policy, it does not execute, regardless of what the agent generates or how confidently it generates it.
For teams building on Supabase, environment segregation has a specific and concrete meaning: development agents must never hold production Supabase credentials, and synthetic data should stand in for real data during prompt tuning, eliminating the risk that sensitive production data leaks into a model through the tuning process itself. You saw the same zero-trust query path principle earlier, and this is a direct, project-level implementation of it. The minimum viable architecture for any team running agents against Supabase data is separate Supabase projects for development and production, each with its own service keys and its own independently configured RLS policies, so that a credential scoped for a development agent has no path, intentional or accidental, into production data.
Pre-modeled, pre-calculated datasets as the right unit of agent-readable data
If you grant an agent access to raw tables, even through a properly governed credential, it still has to construct meaning out of unmodeled data on the fly. That responsibility produces confidently wrong answers, and it makes every resulting query effectively unauditable by a human reviewer, because there is no fixed definition against which to check the agent's interpretation. A pre-calculated metric, by contrast, carries its own definition, its own set of allowed dimensions, and its own provenance alongside it. An agent consuming that metric is working with an analytical contract, not just a number it has to interpret correctly on its own.
Raw tables create a second, quieter problem: when sales, finance, and product teams each derive their own version of the same metric independently from the same underlying tables, an agent querying those raw tables inherits exactly the same disagreement that already exists among the humans. Pre-modeling the metric resolves that disagreement once, before any agent or human query runs against it, so every query does not have to rediscover the same inconsistency. You calculate a metric once, govern it centrally, and distribute it to wherever people and agents actually work, and that is more trustworthy than recalculating it on demand, whoever happens to remember the right formula that day.
The correct unit of analytics in an agent-first architecture is a modeled, pre-calculated dataset that both humans and agents can consume safely, something queryable through an engine like DuckDB without ever touching the production database directly. Governed Parquet datasets queried this way give agents fast, accurate, and inexpensive access to exactly the context they need, with no risk that the agent modifies production records, locks a table other systems depend on, or exhausts database resources through an inefficient query. This is the environment-segregation principle from earlier sections applied at the level of the data model itself: the agent never touches production because it never has to. Everything it requires is already present in the governed dataset.
Every query against a dataset built this way is auditable in a way raw-table queries are not, because the dataset's schema, its version, and its access policy are all fixed at the moment the dataset is created. That gives the replay log, described earlier as requiring the policy version, the SQL, and the egress check, something stable and fixed to check its record against. Pre-calculated metrics close the semantic bypass risk raised in the discussion of MCP as well. If the only interface an agent has is a defined set of approved metrics and dimensions, there are no raw tables sitting behind the interface for the agent to call instead, and the bypass has nowhere left to go.
Sources
- AI Agent Data Governance 2026: Why 63% of Organizations Can’t Stop Their Own AI
- The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
- Deontic Policies for Runtime Governance of Agentic AI Systems
- Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents
- Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
- Architecture overview - Model Context Protocol

