Skip to content

Performance

These ratings are a rough guide to the cost of one call. A Slow entry can still be 2–4× faster than postMessage, depending on the payload and the work being done.

  • Best — almost no per-call overhead
  • Fast — inexpensive for most workloads
  • Good — suitable for normal workloads
  • Fair — fine occasionally; watch hot paths
  • Slow — consider another representation in tight loops

The thresholds below are intentionally broad. They describe the approximate cost of one call, not the total time spent running the task:

  • Best: < 1 µs
  • Fast: < 2 µs
  • Good: < 4 µs
  • Fair: < 9 µs
  • Slow: > 9 µs

The figures on this page were measured on an Apple M3 Ultra running Node 24.12.0 on arm64-darwin, at roughly 3.86 GHz. Treat them as useful comparisons, not promises: your runtime, CPU, payload shape, and task itself will change the result.


Most performance differences come down to how much data has to cross the worker boundary and whether that data is copied.

  • Small values. Numbers, booleans, short strings, and similar values fit in the call header, so encoding and decoding are very cheap.
  • Small binary payloads. Typed arrays and other small values use the transport’s preallocated space and usually stay fast.
  • Larger payloads. The transport may need to allocate more space and copy the data. This is still efficient, but the cost becomes visible in a hot loop.
  • Shared memory. SharedArrayBuffer and ProcessSharedBuffer avoid copying the bytes. Knitting passes a handle instead, so the cost of sending a 1 KiB buffer and a 64 MiB buffer is roughly the same. BufferReference (knitting/unsafe) is the thread-only move variant: it detaches the source and hands the bytes to the worker.

These ratings describe one value passed in a single call. They are useful for choosing a representation, but measure your real workload before tuning around them.

ValueTypical cost
Primitives: boolean, undefined, nullBest
Numbers: numberBest
Time/IDs: DateBest
Strings: small stringBest
Symbols: Symbol.forFast
BigInt: small bigintBest
BigInt: large bigintFast
Binary: typed arraysBest
Views: DataViewGood
Structured: JSON objectGood
Structured: JSON arrayGood
Errors: ErrorSlow

Thread count depends on what the host thread must still do. For a mixed HTTP service, start with one worker: the host still accepts connections, routes, encodes requests, and sends responses. Add workers only when CPU-heavy calls queue and lower tail latency is worth the extra coordination. For independent batch compute, os.availableParallelism() - 1 is a reasonable first trial.

Every worker also consumes memory for payload buffers, shared locks, and cancellation state. Adding threads beyond the available cores usually stops helping. See Multi-threading for a server-focused selection process and timer configuration.

Compatible multi-worker pools use native work stealing by default. The host publishes calls to one shared submit region, and workers claim available tasks as they become free. Responses still use private return lanes, so the task API and promise behavior do not change.

This helps most when many CPU-bound calls compete for workers and task durations are uneven. It is unlikely to be the main lever for tasks that mostly wait on databases, networks, or other external I/O. Set host.steal: false for a private-lane baseline, or tune host.stealRegionLanes for the workload:

  • Wider regions reduce arbitration overhead for many cheap, similarly sized calls.
  • Narrower regions expose more independent work for expensive or uneven calls. 1 is a useful starting point when one task can take much longer than another.

The host doorbell is a separate completion optimization. Node.js and Bun thread pools can use Atomics.waitAsync instead of repeatedly polling for responses; Deno, process, compiled, and browser pools use the polling fallback. Compare like with like when benchmarking: keep the worker topology fixed and change one of host.steal or host.doorbell at a time. See Work stealing for the full option and compatibility details.

The inliner runs eligible calls without sending them through a worker, skipping transport encoding and decoding. It is most useful for very small tasks, such as simple arithmetic. For example, inliner: { position: "last", batchSize: 64 } can improve throughput. See the Inliner guide.

Strict permissions add a small amount of startup work for each worker, such as generating flags and resolving lock files. They do not add measurable overhead to individual calls once the workers are running.

payload.payloadInitialBytes, payload.payloadMaxByteLength, and payload.maxPayloadBytes control the transport buffer used by each worker. Increasing the initial size avoids growth later, but uses more memory up front. For consistently small payloads such as primitives and short strings, the defaults—4 MiB initially, 64 MiB maximum length, and an 8 MiB dynamic payload cap—are usually enough.


Threads and processes have different memory boundaries, so the best zero-copy option depends on which one you use:

RuntimeIsolationZero-copy tools
thread (default)shares the host address spaceSharedArrayBuffer, BufferReference (move)
processseparate memory and permissionsProcessSharedBuffer (OS shared memory)

A SharedArrayBuffer or BufferReference cannot cross a process boundary. Use ProcessSharedBuffer when the worker runs in a separate process. See Shared memory and Buffer reference.

Because a process worker is a child process, you can start it through another tool using worker.processCommandPrefix. This is useful for a sandbox such as bwrap or a container such as Docker:

worker: {
runtime: "process",
processCommandPrefix: ["bwrap", "--unshare-all", "--ro-bind", "/", "/"],
}

See Process workers for the full wrapper recipes.


call.*() accepts promises for supported inputs. In an HTTP handler, you can pass the request body’s promise straight to a task instead of awaiting it on the request thread:

app.post("/jwt", async (c) => {
const responseJson = await handlers.call.issueJwt(c.req.arrayBuffer());
return c.body(responseJson ?? "Bad request", responseJson ? 200 : 400, {
"content-type": "application/json; charset=utf-8",
});
});

The promise itself is not faster. The benefit is that the request thread can hand the body to Knitting immediately:

  • c.req.arrayBuffer() already returns a promise, so forwarding it skips an await in the handler.
  • UTF-8 decoding and JSON parsing happen in the worker, not on the request thread.
  • ArrayBuffer stays on the binary fast path.

When a task needs both request metadata and the raw body, put them in an Envelope. The header carries the metadata and the payload carries the bytes. Use .then(...) to build the envelope from the body promise without awaiting the body on the request thread:

import { Envelope } from "knitting";
app.post("/upload", async (c) => {
const result = await handlers.call.storeUpload(
c.req.arrayBuffer().then(
(body) =>
new Envelope(
{ contentType: c.req.header("content-type") ?? "application/octet-stream" },
body,
),
),
);
return c.json(result);
});

For a large binary body sent to a thread worker, use a BufferReference instead of an ArrayBuffer. The bytes move without a copy; only the body line changes:

import { BufferReference } from "knitting/unsafe";
new Envelope(
{ contentType: c.req.header("content-type") ?? "application/octet-stream" },
new BufferReference(body), // moves the body bytes without copying
);

This works best when the route is mainly forwarding data and the worker does the parsing, such as SSR or JWT issuance. If the main thread needs to inspect or validate the body first, await it there instead.

Returning data has the same costs as sending it, just in the other direction. Choose the return type based on the size of the result and how the data is already represented:

  • JSON object / array — serialized on the worker and parsed again on the host, so it makes two passes over the data. This is fine for small results but expensive for large ones.
  • SharedArrayBuffer / ProcessSharedBuffer — shared memory is usually the cheapest way to return bytes when the result can use it. ProcessSharedBuffer also works with process workers.
  • BufferReference — useful for large binary results, especially when a library gives you an ordinary ArrayBuffer that you cannot turn into shared memory. Returned references can stay zero-copy on Node with native ownership; the safe default makes one copy on Deno, Bun, and some Node builds. See the Buffer reference guide for the explicit borrow option and its lifetime rules.

See Payloads and Buffer reference for the full type list.