Blog
Code Mode vs. Direct Tool Calls: The Same Three Tasks, Five Times Each
One model, three data tasks, five runs each: tools called directly, then through Gatana's code mode. On 10,000 orders, direct calls were right 2 times out of 5 at 1.3 million prompt tokens per run. Code mode was right 5 of 5 at 30,000 tokens, in 12 seconds instead of 10 minutes.
· Erik Jonsson Thorén, Founder, Gatana
Short answer: when a task needs more data than a model should hold in its context, let the model write code that calls the tools. On 5 October 2026 we ran three tasks five times each with DeepSeek-V4.1-Flash, once with direct tool calls and once through Gatana’s code mode. On the two large tasks, direct calls were right 2 times out of 10. Code mode was right 10 of 10, with 4 to 44 times fewer prompt tokens and 17 to 51 times less wall time. On the small task both were right, and direct was a little faster.
What code mode is
In the usual MCP setup, every tool result goes into the model’s context. A 500-row page of orders is 15,000 tokens the model reads and re-sends on every later turn. Code mode replaces the tool list with two tools: one to search the available tools, one to run code. The model writes a short program. The program runs in a sandbox inside the gateway, calls the same tools, does the loop and the arithmetic, and returns only the result. The data never passes through the model.
What we measured
- Data. A synthetic server with four tools: list orders (10,000 rows, at most 500 per page), list files (500 rows), list and read documents (40 documents, 65,000 words). A fixed seed, so the right answers are known in advance.
- Tasks. T1: total revenue per region over all orders, within 0.5%. T2: the five largest files. T3: the three documents in which the word “audit” occurs most often, plus the total word count, within 1%.
- Two modes, one model. Direct: the four data tools exposed to the model as functions, with their real schemas. Code mode: Gatana’s two code mode tools, nothing else. DeepSeek-V4.1-Flash, temperature 0, at most 120 turns, the same system prompt apart from one sentence.
- Five runs per task and mode, 30 in total. We report medians and the count of correct answers. Correctness was checked against ground truth computed outside the model.
All 30 runs are in code-mode-benchmark-2026-10-05.json.
Results
The exact medians. Cost at DeepSeek list prices, see the note below.
| Task | Mode | Correct | Prompt tokens | Completion tokens | Tool calls | Seconds | USD per run |
|---|---|---|---|---|---|---|---|
| T1: revenue by region, 10,000 orders | Direct | 2 of 5 | 1,330,460 | 127,900 | 20 | 617 | 0.389 |
| Code mode | 5 of 5 | 30,458 | 1,274 | 6 | 12 | 0.003 | |
| T3: word counts, 40 documents | Direct | 0 of 5 | 242,630 | 67,535 | 41 | 318 | 0.103 |
| Code mode | 5 of 5 | 59,153 | 1,802 | 6 | 19 | 0.006 | |
| T2: five largest of 500 files | Direct | 5 of 5 | 31,110 | 905 | 2 | 6 | 0.006 |
| Code mode | 5 of 5 | 17,670 | 1,134 | 5 | 9 | 0.003 |
How direct mode fails
The failures are not timeouts. The model fetched the data. It could not do the arithmetic on it.
- T1. The two correct direct runs paged through all 10,000 orders and summed them in the open: 247,000 and 283,000 completion tokens of running totals, in 12 and 14 minutes. Two runs stopped early or lost precision. The fifth asked for 100 rows per page, hit the turn cap after 124 tool calls, and consumed 28.9 million prompt tokens, because every call re-sends the whole conversation.
- T3. All five direct runs read all 40 documents. None counted 65,000 words correctly. Models do not count; they estimate.
- Code mode wrote a loop each time, ran it in the sandbox, and returned the figures. Its tokens are the tool search, the code, and the result.
Where direct mode is fine
T2 is the control. One call returns all 500 files in 10,000 tokens, and picking the five largest is easy for the model. Direct was right every time and a little faster, because code mode spends two turns on tool discovery and code first. Below a few thousand tokens of data, there is nothing for code mode to save.
A note on cost
DeepSeek charges 0.30 USD per million input tokens on a cache miss and 0.006 on a hit, and most of direct mode’s 1.3 million tokens were hits. At a provider without that discount, the same T1 run at 3 USD per million input tokens costs about 4 USD, against 0.09 USD for code mode. The token figures transfer between providers; price them with your own rates.
What this means
- Size the task, not the tool. If one tool result fits in a few thousand tokens, call it directly. If the task walks through thousands of rows or dozens of documents, let the model write the loop.
- Correctness, not only cost. The point is not 44 times fewer prompt tokens. It is that direct mode was wrong in 8 of the 10 large-task runs, and code mode in none.
- Watch the page size. A model that picks a small page on a large dataset multiplies its own cost on every turn. The 28.9 million token run was one such choice.
We will repeat the run with a second model and publish both, with the date.
This run is listed with our other measurements, each with its date and raw data, on the measurements page.