Why organizations need a kill switch for their MCP servers

An MCP server your team approved six weeks ago starts making tool calls nobody recognizes. Maybe a dependency changed underneath it. Maybe a token is being used from somewhere it shouldn’t be. Maybe it is behaving exactly as configured, and the configuration was wrong. Your audit log is filling with calls you can’t explain, and the person who has to decide what happens next is whoever owns security, not whoever installed the server.

The question in the room is not architectural. It is: what do I stop, right now, and how do I know it stopped?

Most enterprises cannot answer that about their MCP connections. Closing that gap is worth doing before the incident rather than during one, and it is worth being precise about what closing it actually takes.


Can an administrator immediately stop the gateway for all affected users?

Only if every agent-to-tool call already routes through a control point that you operate.

If it does, stopping access is an action your platform and security teams already know how to take: terminate the workload, revoke the group in your identity provider, or stop the control point itself. Your audit log then confirms the traffic went quiet, which is the part that matters.

If it doesn’t, there is no button anywhere that helps you, because nobody can stop traffic they never see. A kill switch is not a feature you buy late in an evaluation. It is a property of where your agent traffic goes.


Why this became urgent so quickly

MCP arrived in enterprises from the bottom up. Developers adopted assistants, wired up servers because it made them faster, and governance showed up afterward as a retrofit. Platform teams are now putting controls around tools their organization already depends on, and security teams inherited a blast-radius problem they did not design: mishandled tokens, servers nobody registered, and no consistent place to see or stop any of it.

That ordering is the problem itself. Controls that arrive after adoption have to be retrofitted onto traffic that was never designed to pass through them — and the traffic that didn’t get routed through a control point is exactly the traffic you can’t stop.


What a real kill switch requires

Three things, and none of them is a button.

A control point the traffic actually traverses. This is the precondition for everything else. If agent-to-tool calls reach your internal systems by several different paths, there is no single place to intervene and you are chasing processes across a fleet while the incident runs. One governance model for how agents reach internal and external connectors is what makes a single action possible at all.

Identity, so you can stop a who and not only a what. Plenty of incidents are scoped to a team, a token, or one service account rather than to a server. If your control point is wired into the identity provider you already run (EntraID, Okta, etc.), then revocation follows the same path your security team uses for every other system, and you don’t need a bespoke process invented under pressure.

An audit trail, so you can prove it stopped. An action you can’t verify is a guess. Audit logging and OpenTelemetry traces flowing into the observability stack you already operate — Splunk, Datadog, Grafana, New Relic,  are what turn “I think we stopped it” into something you can show an auditor afterward. Don’t trust the action; trust the log.

Notice that two of the three are things a mature platform team already has. The missing piece is usually the first one.


Why a SaaS control plane can’t give you this

AI access cannot depend on a black-box control layer that runs in someone else’s environment. If that’s where your control point lives, your emergency lever is a support ticket, and your incident timeline now includes a vendor’s response time.

Self-hosted changes the shape of the answer. When the control point runs in the Kubernetes cluster you already operate, the levers are ones your team has practiced on every other workload: stop a workload, change an authorization policy, revoke access. Your data stays in your environment, and the controls are yours on your schedule. For a security architect who can block a deployment, this is usually the deciding property — not a feature list.


The version of this question nobody asks out loud

Teams rarely ask “do you have a kill switch” because they want a demo of a button. They ask because they are trying to work out whether they would be able to act at all. The honest framing is more useful than a yes: a vendor who answers “yes, there’s a button” without asking how your agent traffic is routed has told you they don’t understand the question.

So ask it the other way around. If a connected MCP server started behaving badly this afternoon, which single action covers the blast radius, who is authorized to take it, and what would you look at to confirm it worked? If any of those three has no answer, that’s the quarter’s work.


What to do this quarter

  1. Inventory your MCP connections. List every server in circulation and note how traffic reaches it. Most teams have never written this down, and the list alone changes the conversation with security.
  2. Route everything that touches internal systems through a control point you own. A server that reaches your ticketing system, data warehouse, code host, or internal APIs has no good reason to be reachable by an unmanaged path.
  3. Wire that control point to your identity provider. Revocation should use the mechanism your organization already trusts, not a vendor-specific side channel.
  4. Send audit events to the observability stack you already run. You need the confirmation step to exist before you need it.
  5. Decide who pulls the lever, and rehearse it once. The organizations that handle an MCP incident well are not the ones with a special button. They are the ones where somebody has already decided who acts and what it stops.

Each of those is available on infrastructure you already run. None requires a new plane to manage.


Where to start

Write your own first-five-minutes plan against the MCP connections you have today, and see how far you get before you hit a path you cannot stop. That exercise is free, it takes an afternoon, and it tells you exactly how much of your agent surface is governed rather than merely observed.

If you want help mapping it to your own cluster and identity provider, Stacklok’s forward-deployed engineers do this with platform teams directly on the Kubernetes infrastructure you already run, built by the people who created it.

October 01, 2026

How-To

Skailar Hage

Senior Product Marketing Manager

More by Skailar Hage