Measured 5 October 2026
OpenAPI and Swagger specifications to MCP tools
How many public OpenAPI specifications convert into MCP tools?
Gatana converted 299 of 300 randomly selected public OpenAPI and Swagger specifications sampled from apis.guru into MCP tools (99.7%), and every operation became a tool. The one failure is a webhook specification with no paths. The median input schema per tool is 1,424 bytes. Measured on 5 October 2026.
| First run | After the fix | |
|---|---|---|
| Specifications converted | 262 of 300 (87.3%) | 299 of 300 (99.7%) |
| Swagger 2.0 | 121 of 122 | 122 of 122 |
| OpenAPI 3.x | 141 of 178 | 177 of 178 |
| Tools generated | 6,964 | 10,834 |
| Largest specification | NetBox, 386 tools | GitHub REST API, 845 tools |
| Conversion time per specification, median (maximum) | 0.2 s (8 s) | 0.15 s (3.7 s) |
| Input schema per tool, median | 11,276 bytes | 1,424 bytes |
| Input schema per tool, 90th percentile | 94,668 bytes | 4,808 bytes |
| Tool names over 64 characters | not counted | 84 of 10,834 (0.8%) |
How we measured it
A sample of 300 specifications from apis.guru (122 Swagger 2.0, 178 OpenAPI 3.x, seed 20261005), each created as an OpenAPI server in one Gatana organization through the public API, the tools refreshed and read back. The first run converted 262. All but one of the 38 failures were defects in our converter, fixed the same day; the table shows both runs.
Full write-up: We Converted 300 Public OpenAPI Specs to MCP Tools. Here Is What Broke.
Raw data: CSV, 300 rows: the sample, both result sets, schema sizes and the error texts
Measured 5 October 2026
Latency added by the gateway
How much latency does the Gatana gateway add to a tool call?
A tool call through the Gatana gateway took 30.3 ms at the median against 1.4 ms for the same call made directly, so the gateway adds about 29 ms for authentication, authorization, tool resolution and the audit record. A real tool call takes 1.0 s at the median, so the overhead is about 3%. Measured on 5 October 2026.
| Path | Morning, median (95th pct.) | Afternoon after the change, median (95th pct.) |
|---|---|---|
| Direct to the server, public hostname | 1.2 to 1.4 ms (1.7 to 1.8 ms) | 1.2 to 1.4 ms (1.6 to 1.7 ms) |
| Through the gateway to the same server | 53.4 to 53.8 ms (64.8 to 66.9 ms) | 30.3 to 31.2 ms (37.4 to 41.9 ms) |
| Through the gateway, built-in tool with no upstream server | 45.6 ms (57.0 ms) | 21.9 ms (27.3 ms) |
| Gateway overhead per call | about 52 ms | about 29 ms |
| Share of a real tool call (1.0 s median, 2.8 s 90th pct.) | 5% at the median | 3% at the median, 1% at the 90th percentile |
How we measured it
10,800 tool calls from a probe pod to an echo MCP server in the same cluster as a staging gateway, over one kept-alive connection after 20 warm-up calls, direct and through the gateway. Measured in the morning and again after a change to the gateway's per-call path the same afternoon. The real-call baseline is the audit log of one customer organization, used with its consent.
Full write-up: How Much Latency Does an MCP Gateway Add? We Measured Ours: 52 Milliseconds, Then 29.
Raw data: JSON: every run with its percentiles, morning and afternoon
Measured 5 October 2026
Code mode against direct tool calls
How many tokens does code mode save against direct tool calls?
On two tasks with more data than a model should hold in context, code mode used 4 to 44 times fewer prompt tokens than direct tool calls (30,458 against 1.33 million on the largest task) and 17 to 51 times less wall time. The data stays in the sandbox; only the result reaches the model. Measured on 5 October 2026 with DeepSeek-V4.1-Flash.
| Task | Mode | Correct | Prompt tokens (median) | Seconds (median) | USD per run |
|---|---|---|---|---|---|
| Revenue by region, 10,000 orders | Direct | 2 of 5 | 1,330,460 | 617 | 0.389 |
| Code mode | 5 of 5 | 30,458 | 12 | 0.003 | |
| Word counts, 40 documents | Direct | 0 of 5 | 242,630 | 318 | 0.103 |
| Code mode | 5 of 5 | 59,153 | 19 | 0.006 | |
| Five largest of 500 files | Direct | 5 of 5 | 31,110 | 6 | 0.006 |
| Code mode | 5 of 5 | 17,670 | 9 | 0.003 |
How we measured it
A synthetic server with four data tools (10,000 orders, 500 files, 40 documents of 65,000 words) and three tasks with answers known in advance. DeepSeek-V4.1-Flash at temperature 0 ran each task five times with the four tools exposed directly and five times through Gatana's two code mode tools, 30 runs in all. Correctness was checked against ground truth computed outside the model; cost is at DeepSeek list prices.
Full write-up: Code Mode vs. Direct Tool Calls: The Same Three Tasks, Five Times Each
Raw data: JSON: all 30 runs with tokens, tool calls, time and the answer given
How we measure
- Every figure has a date. Software changes, and a number without a date is a claim. The gateway latency figure changed within one day.
- The raw data is published next to the figure, as CSV or JSON, so anyone can check the arithmetic or repeat the run.
- We measure Gatana, not competitors. The direct-call figures are the baseline, not a rival product.
- Customer data appears only as aggregates, with the organization's consent, never per record.
- A repeated run does not overwrite the old one. The old figure stays on this page, marked as superseded, with the date of the new run.
Reading this with a script or an agent?
The same figures are published as Markdown at gatana.ai/measurements.md.