Blog

How Much Latency Does an MCP Gateway Add? We Measured Ours: 52 Milliseconds, Then 29.

We timed 10,800 tool calls, direct to a server and through Gatana, from inside the same cluster. In the morning the gateway added 52 ms per call; after a fix the same afternoon, 29 ms. The setup, the raw numbers, where the time goes and what we make of it.

· Erik Jonsson Thorén, Founder, Gatana

Three app-style icon tiles on a soft background: a blue stopwatch, the Gatana mark on a dark tile, and a teal lightning bolt

Short answer: on the morning of 5 October 2026 a tool call through a Gatana gateway took 53.6 ms at the median, against 1.4 ms for the same call made directly. 45 ms of that was the gateway’s own work: authentication, authorization, tool resolution and the audit record. We profiled that path, changed it, and measured again the same afternoon: 30.3 ms through the gateway, 22 ms of it the gateway’s own work. Against the 1.0 second a real tool call takes at the median in production, the overhead went from 5% to 3%.

What we measured

  • Target. A minimal MCP server with one tool, echo, that returns its input. Every millisecond it shows belongs to the network and the gateway.
  • Position. Probe and echo server ran inside the same cluster as the gateway. A user on a laptop adds their own round trip on top, on both paths alike.
  • Paths. Direct to the echo server, by its internal address and by its public hostname through the ingress. Through the gateway, which authenticates the caller, resolves the tool, calls the echo server by its public hostname, records the call and returns the result.
  • Protocol. MCP over streamable HTTP, one kept-alive connection per run, 20 warm-up calls, then 300 to 1,000 timed tools/call round trips. Direct and gateway runs were interleaved.
  • Twice. The full set ran in the morning, and again in the afternoon after a change to the gateway’s per-call path. Staging environment, so without the extra proxy hop of production.

The raw per-run figures are in gateway-latency-2026-10-05.json. No call failed.

Results

Median round trip per tool call, millisecondsLower is betterProbe and echo server inside the same cluster as the gateway. A range is the spread of the session's runs; hover for every run and the 95th percentile.
Direct to the echo server, public hostname
Morning1.2 to 1.4
Afternoon1.2 to 1.4
Through the gateway to the echo tool
Morning53.4 to 53.8
Afternoon30.3 to 31.2
Through the gateway to a built-in tool, no upstream server
Morning45.6
Afternoon21.9

The exact figures. Medians, with the 95th percentile in brackets. Gateway figures are the range over the session’s runs.

Path Morning Afternoon, after the change
Direct, internal address 0.6 ms (1.5) not repeated
Direct, public hostname through the ingress 1.2 to 1.4 ms (1.7 to 1.8) 1.2 to 1.4 ms (1.6 to 1.7)
Through the gateway to the echo tool 53.4 to 53.8 ms (64.8 to 66.9) 30.3 to 31.2 ms (37.4 to 41.9)
Through the gateway, built-in tool with no upstream server 45.6 ms (57.0) 21.9 ms (27.3)

The direct path did not move between the sessions, so the afternoon difference is the gateway, not the cluster. Medians are stable to within a millisecond across runs. The 99th percentile through the gateway is not: 61 to 520 ms depending on the run, a handful of calls per thousand that waited on something.

Where the 52 milliseconds go

The built-in tool run is the control: the full gateway path, but no upstream server. Morning, 45.6 ms against 53.6 ms for the echo run. Afternoon, 21.9 against 30.3. So the gateway’s own processing was about 45 ms and is now about 22 ms. The upstream MCP exchange costs 8 to 9 ms either way, of which the network itself is 1.4 ms.

Where the milliseconds goLower is betterGateway work is the built-in tool run; the upstream exchange is the echo-tool run minus it.
Gateway's own work per call
Morning45.6
Afternoon21.9
Upstream MCP exchange, gateway to echo server and back
Morning7.8 to 8.2
Afternoon8.4 to 9.3
Network round trip alone, the direct path
Morning1.2 to 1.4
Afternoon1.2 to 1.4

45 ms for a handful of lookups and one audit write was more than it needed to be. About half of it was avoidable work on the per-call path. 22 ms is not the floor; this page will be updated, with the date, when it moves again.

What it means in practice

The audit log of one customer organization puts the median real tool call at 1.0 second and the 90th percentile at 2.8 seconds. The time is in the upstream service. Against that, 29 ms is 3% at the median and 1% at the 90th percentile.

Gateway overhead against a real tool call, millisecondsThe audit log of one customer organization, analysed separately.
A real tool call, from one customer organization's audit log
Median1,000
90th pct.2,800
What the gateway adds per call, after the change
Overhead29

What the 29 ms buys: the call ran under the caller’s own identity, the credential never left the gateway, the tool was one the caller was allowed to use, and there is a record of it. For a team connecting production systems to agents, we think that trade is right. For a latency-critical loop that calls a local tool a hundred times a second, a gateway is the wrong layer, ours or anyone’s.

This run is listed with our other measurements, each with its date and raw data, on the measurements page.

All posts