If an MCP tool executes code chosen by a model that reads untrusted content, you have remote code execution with extra steps. The question is not whether isolation is worth it — it is how much isolation to buy.
Three levels, three costs
Standard container. Namespaces and cgroups. Cheap, fast, and sharing the host kernel. Good for separating tools from each other and capping resources, but a kernel flaw goes straight through. With restrictive seccomp, a non-root user, a read-only filesystem and no capabilities, it covers most internal cases.
gVisor or syscall sandboxing. Interposes a layer between the process and the real kernel, sharply reducing exposed syscall surface. It costs I/O latency and some compatibility. This is the defensible middle ground when a tool processes files supplied by third parties.
microVM. Its own kernel, hardware isolation, boot in tens of milliseconds. It is the standard for anyone executing arbitrary customer code. If your tool runs model-generated code, that is the correct level — and it costs less than the incident.
Choosing without over-engineering
- Read-only tool over internal data → a hardened container is enough.
- Tool that processes external files (convert, extract, parse) → gVisor. Parsers have historically been the origin of half the RCEs.
- Tool that executes code or commands → microVM, discarded after each task.
The isolation level should follow the origin of the input, not the criticality of the output.
What a sandbox does not solve
Isolation contains execution, not authorisation. A perfect sandbox with a broad write credential inside it still lets the agent drop the table — legitimately, through the API, without touching the kernel. That is why sandboxing and credential least-privilege are complementary controls, and the second usually pays more.
And remember the client-tooling lesson: researchers demonstrated RCE in an MCP inspection tool, meaning that examining a suspicious server compromised the examiner's machine. Sandboxing applies to the developer workstation too.