Skip to main content

How Code Mode turns MCP tools into a TypeScript API

Code Mode gives the model a TypeScript API instead of a list of tools, then runs the program it writes in a sandbox. How the pattern works, built from scratch on MCP, QuickJS and Gemini, and where it gets fiddly.

7 min read
Read with ClaudeRead with ChatGPTMarkdown

Every tool result an agent receives passes through the model. The model asks for a call, the result lands in its context, and it reads that result before deciding what to do next. Chain 5 calls and the model has read 5 results, including the ones it only needed to copy into the next call’s arguments.

Cloudflare’s Code Mode: the better way to use MCP, published in September 2025, rearranges this. Convert the MCP tools into a TypeScript API, ask the model to write a program against it, and run the program in a sandbox. Their reasoning fits in two sentences:

LLMs have seen a lot of code. They have not seen a lot of “tool calls”.

The name comes from their Agents SDK, where code mode is a way of connecting to an MCP server. Connect in code mode and the agent gets an API to write code against instead of tools to call.

I built the pattern from scratch to see what it takes. It came to about 670 lines across five modules, with no framework. Gemini writes the programs, a QuickJS sandbox compiled to WASM runs them, and a host process connects the sandbox to the MCP servers.

QuickJS is only there because I needed somewhere safe to run code the model writes. Cloudflare runs it in V8 isolates on their own Workers platform, and I don’t have that. QuickJS was the easiest sandbox to get going inside Node, since it’s one npm package with nothing to compile or install. Any sandbox that can run JavaScript and expose a few functions to it would do the same job.

Two shapes of agent

Plain tool calling looks like this:

LLM -> tool -> LLM -> tool -> LLM -> tool -> LLM -> answer

Every intermediate result travels back through the model, and every hop is a chance to paraphrase a number slightly wrong.

Code Mode looks like this:

LLM writes program -> [tool, tool, reduce, filter] -> LLM reads one result

The model still decides what happens. Writing the program is where all the judgement goes. What moves out of the model is everything between the tool calls: the loops, the filtering, the arithmetic, the plumbing of one result into the next call. Cloudflare’s line is that the model “can skip all that, and only read back the final results it needs”.

How the pieces fit

The model gets exactly one tool, execute_code, which takes a string of JavaScript. The agent loop hands that string to a sandbox, the sandbox runs it, and whatever the program logs goes back to the model as the tool’s result. The MCP servers themselves stay on the host, out of the sandbox’s reach.

The host connects to each MCP server and calls listTools(). It owns the transports and any credentials, and the sandbox never sees either.

A generator then turns each tool’s JSON Schema into TypeScript. My toy weather server has a list_readings tool that returns a week of hourly records for a city. Here’s the generator’s output for it, trimmed:

interface ListReadingsInput {
  /** City name, e.g. 'Lisbon' */
  city: string
}

interface ListReadingsOutput {
  city: string
  readings: {
    hour: number
    tempC: number
    humidity: number
  }[]
}

declare const codemode: {
  /** Hourly weather readings for the past 168 hours for a city. Returns one record per hour. */
  list_readings: (input: ListReadingsInput) => Promise<ListReadingsOutput>;
  // ...one method per tool
};

That block goes into the system instruction, with a few lines of guidance: chain whatever calls you need in one execute_code, keep intermediate data inside the code, and console.log() only the result you need.

Ask which of Lisbon, Tokyo, London and Singapore had the warmest week, and the program the model writes looks like this:

const cities = ['Lisbon', 'Tokyo', 'London', 'Singapore'];
const averages = await Promise.all(
  cities.map(async (city) => {
    const { readings } = await codemode.list_readings({ city });
    const avgC = readings.reduce((sum, r) => sum + r.tempC, 0) / readings.length;
    const { value } = await codemode.convert_temperature({
      value: avgC, from: 'celsius', to: 'fahrenheit'
    });
    return { city, avgF: value };
  })
);
console.log(JSON.stringify(averages.sort((a, b) => b.avgF - a.avgF)));

The sandbox runs it in a fresh QuickJS runtime and throws the runtime away afterwards. When the guest calls codemode.list_readings(), the call crosses to the host as JSON, the host forwards it to the MCP server, and the result goes back into the sandbox as a resolved promise. Each call is a promise, so Promise.all over 4 cities runs them concurrently.

Whatever the program logs becomes the tool result the model sees. The 672 hourly records exist inside the sandbox and never enter the model’s context. 4 averages do.

What the model stops doing

Two jobs leave the model.

The first is carrying data. With plain tool calling, all 672 records go through the context window, because that’s the only place the model can look at them. In Code Mode they stay in the sandbox and the model reads 4 numbers.

The second is arithmetic. Without Code Mode, the model has to average 168 numbers by reading them, one token at a time, and it can come out slightly off. Here reduce() does the sum.

There’s a cost on the other side. The generated TypeScript goes into the system instruction, so it’s resent on every request. When the tool results are small, there’s nothing for the sandbox to keep out of the context, and the API is overhead you pay on every turn.

The sandbox decides what code can reach

QuickJS starts empty. There’s no fetch, no process, no require, no timers and no filesystem. The sandbox injects a console that collects output and a single bridge function out to the host, and the guest builds the codemode object on top of that bridge. The program can log and call the connected tools, and nothing else. Cloudflare does the same with bindings, granting capabilities instead of filtering the network.

API keys and server URLs stay on the host. Code the model writes can’t leak a key it never receives.

Each run also gets 64 MB of memory and a 30 second deadline. QuickJS’s interrupt handler enforces the deadline while guest code is running, and the host enforces it while the program waits on tool calls.

Getting the TypeScript right

The generated TypeScript is the only description of the tools the model ever sees, so the generator matters more than it looks.

Tool names are the first snag. MCP is happy with names like brave-search or notion.query, but neither is a valid TypeScript property key, so a naive generator writes a declaration block that doesn’t parse. I quote the key and add a note to the JSDoc telling the model to use brackets, as in codemode["list-cities"]({}).

Return types matter more. With plain tool calling, the model sees each result before it does anything with it. In Code Mode it writes r.tempC before any result exists, so the field names have to come from the types. MCP tools can declare an outputSchema, and if the generator ignores it, every method returns Promise<unknown> and the model is left guessing whether the field is tempC or temperature. I compile the schema into a real return type with json-schema-to-typescript, and if I built this again it’s the part I’d write first.

When I’d reach for it

For a filesystem agent doing 3 small reads, I wouldn’t. The TypeScript API costs prompt tokens on every turn and there’s nothing big to keep out of the context.

For an agent that queries, filters or aggregates real data, I would, and mostly so that code does the maths instead of the model.

Credit where it’s due: the idea, the TypeScript API framing and the sandbox-per-execution design all come from Cloudflare’s Code Mode post. It’s worth reading in full, particularly the section on why V8 isolates make the sandboxing cheap enough to do on every call.