# Knitting — full documentation > Knitting is a zero-dependency, shared-memory concurrency runtime for Node.js, Deno, and Bun. Move typed JavaScript work to threads, separate processes, or browser workers and call it like an async function. Use Knitting when CPU-heavy, bursty, or isolation-sensitive work should leave the main thread without becoming a separate service. Its compact API combines typed calls with shared-memory IPC, work stealing, timeouts, cancellation, worker permissions, and zero-copy paths for large binary payloads. Its scheduling defaults are built to keep the CPU cost of threading low rather than to maximize a benchmark number: idle workers park instead of spinning, and on supported runtimes the host waits on a doorbell instead of polling, so a pool that is not saturated costs close to nothing while it waits. Knitting is Apache-2.0 open source on [GitHub](https://github.com/mimiMonads/knitting). Its [test suite](https://github.com/mimiMonads/knitting/tree/main/test) covers runtime behavior, shared-memory transport, process workers, work stealing, permissions, package output, browser execution, and compiled workers. [Continuous integration](https://github.com/mimiMonads/knitting/actions/workflows/test.yml) exercises Node.js, Deno, and Bun across a multi-OS matrix, with a [90% Node line-coverage gate](https://github.com/mimiMonads/knitting/actions/workflows/coverage.yml). **Essentials** - Install: `npm install knitting` (the npm package is `knitting`; it is also on JSR as `@vixeny/knitting`). Requires Node 22+, Deno 2+, or Bun 1+. - A task is an exported function at module scope. Wrap it with `task({ f })` only when you want options like a timeout or an abort signal. - Tasks take ONE argument. Use a tuple or object for multiple values: `([a, b]) => a + b`. - Guard host-only code with `isMain` — workers re-import the module. - Module loading: each worker re-imports the module that DEFINES your tasks, and its top-level `import`s run in every worker (they are hoisted — `isMain` does NOT gate them). Keep tasks in a lean module separate from your server/framework code. Tasks must be `export`ed or the loader can't find them and the call silently hangs. `importTask` targets must be plain functions, not `task()` wrappers. - Create a pool with `createPool(options)({ taskA, taskB })`, then call `await pool.call.taskA(args)`. - Scheduling: compatible multi-worker pools use native work stealing by default. Workers claim tasks from a shared submit region while keeping private return lanes; control it with `host.steal`, `host.stealRegionLanes`, and `host.doorbell`. The task API does not change, and unsupported runtimes fall back to private lanes or polling. - Idle cost is a design goal: waiting threads are not allowed to burn CPU. A single worker spins 50us before parking because it is on the request's critical path; multi-worker pools park immediately, since a peer is already awake to take the work. The host doorbell replaces polling wake-ups on Node and Bun thread pools. Expect a bigger pool to raise CPU per request without raising throughput when the host is the only producer (a server), so size `threads` from measurements, not from core count — see the Multi-threading guide. - Cleanup: `using pool = createPool(...)` disposes the pool at scope exit. `await pool.shutdown()` still exists to close it earlier or to await teardown. - Isolation: `importTask({ href, name })` keeps a task's code off the host (only the worker imports it). Set `worker.runtime: "process"` to run each worker as a separate process — including inside a bwrap sandbox or a container. - Security: `importTask` prevents the task module from being imported or evaluated at host scope, but it is not a sandbox. For genuinely untrusted code, use process workers with an OS sandbox or container and restrictive permissions; runtime permissions are guardrails, not a complete security boundary. - Zero-copy IN: `ProcessSharedBuffer` (`knitting/shared-memory`) shares bytes across processes; `SharedArrayBuffer` and `BufferReference` (`knitting/unsafe`) move bytes to thread workers without copying. Pick by boundary — process vs thread. - Binary results: for large results from a thread worker, RETURN a `BufferReference`; owning Node addons can move them back zero-copy, while the safe default may take one copy on Deno/Bun (use the explicit borrow mode only when its lifetime rules fit). `knitting/utils` converts string/JSON/number ↔ `SharedArrayBuffer`. - Optimized for HTTP: `call.*()` accepts `Promise` inputs, so forward `request.arrayBuffer()` (e.g. Hono `c.req.arrayBuffer()`) straight into a task without awaiting it on the request thread — UTF-8 decode / JSON parse then happens in the worker. Ideal for SSR, JWT, and upload routes. - Workers are quiet by default: in strict mode worker `console.*` does NOT reach the host — set `permission: { console: true }` to surface it. Common direct exit calls (`process.exit`, `process.kill`, `process.abort`, and `Deno.exit`) are blocked, but this is not a complete security boundary; resource exhaustion and runtime or native-code vulnerabilities still require OS-level isolation. - Debugging goes to STDERR: pass `debug: true` to `createPool` (or set the `KNITTING_DEBUG=*` env var) to stream diagnostics, each line tagged with the worker (`host`, `w0`, `w1`, …), the runtime, and a per-worker ms timer. Select namespaces instead of all — `host` (pool/task setup), `imports` (which modules each worker loaded), `lifecycle` (worker ready / process events), `signals` (per-dispatch traffic, very chatty), `globals` (`globalThis` pollution per load phase) — via `debug: { host: true, imports: true }` or `KNITTING_DEBUG=host,imports`. The option and the env var merge; either can enable a namespace. Zero-cost when off: the logger module isn't even imported. - Payload size: dynamic payloads are hard-capped at ~8 MiB by default (over-cap calls reject with `KNT_ERROR_3`). Raise it with `payload: { maxPayloadBytes, payloadMaxByteLength }` — `maxPayloadBytes` must be `<= payloadMaxByteLength >> 3`; the buffer growth cap defaults to 64 MiB. - Cancellation & timeouts: `task({ f, timeout: { time: 100 } })` bounds a call, `task({ f, abortSignal: true })` injects an abort toolkit (`signal.hasAborted()`, `signal.now()`) as the task's second argument — it is NOT a DOM `AbortSignal` (no `.aborted`, no `addEventListener`, cannot be passed to `fetch`) — and `worker.hardTimeoutMs` is a hard wall-clock kill for runaway CPU. - Browser: `knitting/browser` runs the same pool API on web workers. Two hard requirements: the page must be cross-origin isolated (`Cross-Origin-Opener-Policy: same-origin` plus `Cross-Origin-Embedder-Policy: require-corp`, or `createPool` throws), and every task module must call `setModuleUrl(import.meta.url)` before defining tasks, because stack-based module discovery needs V8's `Error.prepareStackTrace`, which Firefox and Safari do not have. Not available in a page: process workers, compiled/Porffor workers, `BufferReference`, `ProcessSharedBuffer`, and passing a `SharedArrayBuffer` as a task argument. `permission: {...}` is accepted but IGNORED — a web worker holds the full privileges of the page that started it. - Errors are real: thrown errors and rejected promises return to the host as `Error` objects with `name`, `message`, `stack`, and the full `cause` chain. ```ts import { createPool, isMain } from "knitting"; export const square = (n: number) => n * n; export const greet = (name: string) => `hello ${name}`; if (isMain) { // `using` shuts the pool down when this block ends. using pool = createPool({ threads: 2 })({ square, greet }); const [n, msg] = await Promise.all([ pool.call.square(8), pool.call.greet("knitting"), ]); console.log({ n, msg }); // { n: 64, msg: "hello knitting" } } ``` --- # Installation URL: https://knittingdocs.vercel.app/start/installation/ Install Knitting from npm for Node.js, Deno, and Bun, and verify runtime requirements for shared-memory worker IPC. Use this page to check runtime requirements and install Knitting correctly for your JavaScript runtime. ## Requirements - Node.js 22+ - Deno 2+ - Bun 1.0+ ## Install Knitting is published on npm as `knitting`. Latest release: `0.1.62`. > Note: Knitting is also published on JSR as `@vixeny/knitting` — the same code under a > scoped name. The npm package is the simplest path on every runtime, so prefer it > unless you specifically need the JSR entry. --- # Quick Start URL: https://knittingdocs.vercel.app/start/quick-start/ Build your first Knitting pool: define tasks, call them like async functions, run them in parallel, and shut down cleanly. Define some tasks, create a pool, call the tasks like async functions, let the pool close itself. That is the whole loop, and it takes a few minutes. ## Introduction If you have used workers before: Knitting is workers with a function-call API and much less overhead. If you haven’t, the terms below are all the vocabulary you need. ### Quick definitions - **Task**: a function the workers can run. An exported function already counts as one; wrap it in `task({ f })` when you want options like timeouts or aborts. - **`call.*()`**: runs a task on the pool and returns a `Promise` with the result. - **`isMain`**: `true` on the host, `false` inside a worker. Workers re-import your module, so anything that should happen once — creating the pool, starting a server — belongs behind this check. - **`createPool()`**: starts the workers and hands back a typed `call` object and a `shutdown`. It is disposable too, which is what lets `using` close it. - **Host ↔ Worker**: the host is the process that creates the pool. The workers are the threads (or processes) that run the tasks. ### Examples Four snippets, roughly in the order you will need them. `hello_world.ts` ```ts import { createPool, isMain } from "knitting"; // A task is just an exported function the workers can run. export const greet = (name: string) => `hello ${name}`; if (isMain) { // `using` shuts the pool down automatically when this block ends. using pool = createPool({ threads: 1 })({ greet }); console.log(await pool.call.greet("knitting")); // hello knitting } ``` `multi_thread.ts` ```ts import { createPool, isMain } from "knitting"; export const hello = () => "hello "; export const world = (prefix: string) => `${prefix}world!`; if (isMain) { using pool = createPool({ threads: 2 })({ hello, world }); // call.hello() returns a promise; Knitting resolves it before world runs. const lines = await Promise.all( Array.from({ length: 3 }, () => pool.call.world(pool.call.hello())), ); console.log(lines.join(" ")); // hello world! hello world! hello world! } ``` `multi_task.ts` ```ts import { createPool, isMain } from "knitting"; // Several tasks share one pool. Calls are promises, so you can chain them. export const double = (n: number) => n * 2; export const square = (n: number) => n * n; if (isMain) { using pool = createPool({ threads: 2 })({ double, square }); const results = await Promise.all( [1, 2, 3, 4, 5].map(async (n) => pool.call.square(await pool.call.double(n))), ); console.log(results); // [4, 16, 36, 64, 100] } ``` `task_options.ts` ```ts import { createPool, isMain, task } from "knitting"; // Wrap a function with task() when you want options like a timeout. // This call is too slow, so it falls back to the default instead of hanging. export const slow = task({ timeout: { time: 100, default: "timed out" }, f: async (name: string) => { await new Promise((resolve) => setTimeout(resolve, 1_000)); return `hello ${name}`; }, }); if (isMain) { using pool = createPool({ threads: 1 })({ slow }); console.log(await pool.call.slow("knitting")); // timed out } ``` ## Build it step by step 1. Import what you need: ```ts import { createPool, isMain } from "knitting"; ``` 2. Export your tasks at module scope. Workers find them by name, so they have to be reachable from the top level of the file: ```ts export const square = (n: number) => n * n; export const greet = (name: string) => `hello ${name}`; ``` 3. Create the pool behind `isMain`. Workers re-import this module, and without the guard every one of them would try to start a pool of its own: ```ts if (isMain) { using pool = createPool({ threads: 2 })({ square, greet }); } ``` `using` closes the pool when the block ends, so there is nothing to clean up. 4. Call the tasks. They hand back ordinary promises, so `Promise.all` batches them: ```ts if (isMain) { using pool = createPool({ threads: 2 })({ square, greet }); const [n, message] = await Promise.all([ pool.call.square(8), pool.call.greet("knitting"), ]); console.log({ n, message }); // { n: 64, message: "hello knitting" } } ``` 5. Shut down when you are done. With `using` that already happened at the end of the block: the workers stop and the process can exit. Call `shutdown()` yourself when you want to pick the moment, or when your runtime has no `using`: ```ts const pool = createPool({ threads: 2 })({ square, greet }); try { console.log(await pool.call.square(8)); } finally { await pool.shutdown(); } ``` ### A task can make its own pool One task in a short script? Skip the separate `createPool` and let the task carry its own: ```ts import { isMain, task } from "knitting"; export const double = task({ f: (n: number) => n * 2, }).createPool({ threads: 2 }); if (isMain) { try { console.log(await double.call(21)); // 42 } finally { await double.shutdown(); } } ``` ## Good habits None of this is required. All of it saves you an afternoon later. ### Keep tasks in their own module(s) Workers import that module and nothing around it, so they start fast and stay out of the way of your app code. - package.json - deno.json - src - knitting - database.ts - img_parsing.ts - jwt.ts - app/ - pages/ ### Promise inputs are awaited on the host `call.*()` accepts `Promise` inputs. Knitting resolves them on the host before dispatch, so unresolved promise state never crosses the thread boundary. Request handlers get this for free — hand the call a body you have not read yet: ```ts app.post("/validate", async (c) => { const result = await pool.call.validate(c.req.text()); return c.json(result); }); ``` If that promise rejects, the call rejects on the host and the worker never runs. ### Pick the cheapest payload that fits Smaller and flatter travels faster. Numbers, booleans and short strings first, then typed arrays and `Buffer`, then compact JSON. When you have bytes plus some metadata to describe them, use `Envelope`. See [Supported payloads](/guides/payloads/). ### Start strict, open up later Worker permissions start restricted. Sensitive paths like `.env`, `.git`, `~/.ssh` and `/etc` are blocked, and `node_modules` is deny-write. Open up only what your tasks actually touch. See [Permissions](/guides/permissions/). ## Footguns Almost everything here follows from one fact: **workers re-import the module that defines your tasks.** - **One argument per task.** `pool.call.add(a, b)` won't work — pass a tuple or object: `pool.call.add([a, b])` for `([a, b]) => a + b`. - **Guard host code with `isMain`.** Without it, pool creation (and any other host-only code) re-runs inside every worker. - **Top-level imports run in every worker.** `import` is hoisted, so it executes before any `isMain` check. Keep tasks in their own lean module so workers don't load your whole server framework. - **Export your tasks.** An unexported `task()` / `importTask()` is invisible to the worker loader, so the call just hangs — no handler is ever registered. - **`importTask` targets are plain functions**, not `task()` wrappers (that throws a `TypeError`). Put `timeout` / `abortSignal` options on the `importTask` call instead. - **Worker `console.*` is silent by default** in strict mode. Pass `permission: { console: true }` to surface worker logs. - **Can't tell what the pool is doing?** Pass `debug: true` to `createPool`, or set `KNITTING_DEBUG=*`. Setup, import and lifecycle diagnostics go to **stderr**, each line tagged with its worker and a millisecond timer. If that is too much, name the parts you care about: `debug: { host: true, imports: true }`. Turned off, it costs nothing. - **Only supported payloads cross the boundary.** `Map`, `Set`, class instances, and functions are rejected — see [Payloads](/guides/payloads/). - **Dynamic payloads cap at ~8 MiB** by default; raise `payload.maxPayloadBytes` (and `payload.payloadMaxByteLength`) for larger ones. > Tip: Great question — and you're absolutely right to be thinking about context. 🚀 > > Knitting ships machine-readable documentation, so let's dive in: > > - **[`/llms.txt`](/llms.txt)** — the essentials, distilled for maximum signal. > - **[`/llms-full.txt`](/llms-full.txt)** — every page, inlined into a single file. > > It's not just documentation — it's a force multiplier. By leveraging these > resources, your agent can unlock a deeper understanding of the footguns above > and avoid them far more often. ✨ > > Would you like me to break any of this down further? ## Where to go next - [Defining tasks](/guides/defining-tasks/) — `task()`, `importTask()`, timeouts, and aborts. - [Creating pools](/guides/creating-pools/) — threads, balancers, and shutdown. - [Payloads](/guides/payloads/) — what crosses the boundary, and `Envelope`. - [Performance](/guides/performance/) and the [inliner](/guides/inliner/) — when to let the host run some work too. Two things worth knowing about before you need them. Workers can each run as a [separate process](/guides/process-workers/), inside a `bwrap` sandbox or a container, when threads are not enough isolation. And large buffers can skip the copy entirely with [`ProcessSharedBuffer`](/guides/shared-memory/), which puts the bytes in shared memory instead of sending them. --- # Defining tasks URL: https://knittingdocs.vercel.app/guides/defining-tasks/ A task is a function your workers run. Use a plain function for the simple case, or task() when you want timeouts, aborts, or imported worker code. A **task** is a function your workers run. An exported function already counts as one; wrap it in `task({ f })` when you want options like timeouts or abort signals. ## The rules - Define tasks at module scope — no conditional or dynamic exports. - Export them from the module where they are defined. - One argument in, one value out. Use a tuple or object for multiple values. - Keep tasks in their own file(s) so workers load only what they need. ## Module loading Each worker **re-imports the module that defines your tasks** — the file that calls `task()` / `importTask()` and hands them to `createPool`. Two things follow from that: - **Top-level `import`s run in every worker.** `import` statements are hoisted, so they execute *before* any `if (isMain)` guard — `isMain` gates your executable code, not your imports. If you define tasks in the same file as a web framework, every worker loads that framework too. Keep tasks in their own lean module and import it from your server, so workers only load what they run. - **Tasks must be exported.** The worker discovers tasks by scanning the module's exports. An unexported `const myTask = task(...)` (or `importTask`) is invisible to the loader, so calling it just **hangs** — no handler is ever registered. Always `export` your tasks and `importTask` wrappers. For full isolation — keeping a task's *own* code off the host entirely — use [`importTask`](#importing-worker-side-code-with-importtask): only the worker imports the target module, and that target must be a **plain exported function**, not a `task()` wrapper. ## A plain function When a task needs no options, a bare exported function is enough: ```ts import { createPool, isMain } from "knitting"; export const greet = (name: string) => `hello ${name}`; if (isMain) { using pool = createPool({ threads: 1 })({ greet }); console.log(await pool.call.greet("knitting")); // hello knitting } ``` `greet` has to be a real exported binding. Workers look tasks up by name in the module they re-import, so an inline `{ greet: (name) => ... }` handed straight to `createPool` leaves them nothing to find. Wrapping it in `task()` does not change that — `task()` adds options, not discoverability. ## Wrapping with `task()` Use `task({ f })` when you want options — a timeout, an abort signal, or explicit types: ```ts import { task } from "knitting"; export const add = task({ f: ([a, b]: [number, number]) => a + b, }); ``` Return types are inferred, but argument types are not — annotate the parameter, or pin both with generics: ```ts export const add = task<[number, number], number>({ f: ([a, b]) => a + b, }); ``` ## Arguments and return values Each task receives **one** argument and returns **one** value. For multiple inputs, pass a tuple or an object: ```ts type ResizeInput = { width: number; height: number }; export const pixels = task({ f: ({ width, height }) => width * height, }); ``` See [Payloads](/guides/payloads/) for everything that can cross the boundary. ### Promise inputs are awaited on the host `call.*()` also accepts a `Promise` as input. Knitting awaits it **on the host** before dispatch, so only plain values ever reach the worker: - Fulfilled input → the worker runs with the resolved value. - Rejected input → the host call rejects and the worker never runs. - Only a native `Promise` is awaited; thenables are not. That's why chaining works — `call.hello()` returns a promise, and Knitting resolves it before `world` runs: ```ts const lines = await pool.call.world(pool.call.hello()); ``` ## Options Beyond `f`, there are two: | Option | Purpose | | --- | --- | | `timeout` | Bound how long a call may run. | | `abortSignal` | Make a task cancellable and abort-aware. | ### Timeouts Use a timeout when a call should not wait forever: ```ts export const maybeSlow = task({ timeout: { time: 100, default: "timed out" }, f: async (value) => value, }); ``` The shape of `timeout` decides what happens when the budget runs out: | Form | Outcome | | --- | --- | | `number` (ms) | Rejects with `Error("Task timeout")`. | | `{ time, default }` | Resolves with `default`. | | `{ time, maybe: true }` | Resolves with `undefined`. | | `{ time, error }` | Rejects with `error`. | A missing or negative `time` disables the timeout. > Note: The timeout races the task on the worker using whatever budget is left (wall > time minus dispatch latency). The work itself keeps going, so it can still > finish after its promise has already resolved or rejected. ### Abort signals Abort signals opt a task into cooperative cancellation. `abortSignal: true` is all it takes: the task becomes abort-aware, so its in-flight calls reject with `"Thread closed"` when `shutdown()` runs instead of hanging, and it receives an **abort toolkit** as a second argument. ```ts export const cpuWork = task({ abortSignal: true, f: (items: number[], signal) => { let sum = 0; for (const item of items) { if (signal.hasAborted()) throw new Error("Task aborted"); sum += item; } return sum; }, }); ``` The toolkit is not a DOM `AbortSignal`. There is no `.aborted`, no `addEventListener`, and you cannot hand it to `fetch`. It carries two methods: | Method | What it gives you | | --- | --- | | `signal.hasAborted()` | `true` once the call has been cancelled. Poll it in a loop and bail out. | | `signal.now()` | A monotonic millisecond clock, for measuring elapsed time inside the task. | Monotonic means a clock adjustment can't make a duration come out negative, which is what makes `now()` safe for giving a loop its own budget: ```ts export const budgeted = task({ abortSignal: true, f: (items: number[], signal) => { const started = signal.now(); let sum = 0; for (const item of items) { if (signal.hasAborted() || signal.now() - started > 50) break; sum += item; } return sum; }, }); ``` `abortSignal: { hasAborted: true }` is accepted as well, and does exactly the same thing. Reach for the plain `true`. The promise returned by an abort-aware call also exposes `.reject()`, so the **host** can cancel without touching the worker: ```ts import { createPool, isMain, task } from "knitting"; export const slow = task({ abortSignal: true, f: async () => { await new Promise((r) => setTimeout(r, 10_000)); return "done"; }, }); if (isMain) { using pool = createPool({ threads: 1 })({ slow }); const promise = pool.call.slow(); setTimeout(() => promise.reject?.("cancelled by host"), 100); try { await promise; } catch (e) { console.log(e); // "cancelled by host" } } ``` > Note: Abort signals use a shared-memory bitset pool, sized by `abortSignalCapacity` > in `createPool` (default `258`). It applies whenever at least one task declares > `abortSignal`. When every slot is in use, further abort-aware calls reject with > the `AbortSignalPoolExhausted` symbol. ## Importing worker-side code with `importTask` `importTask({ href, name?, timeout?, abortSignal? })` points at a function in another module. The host gets a typed task wrapper but **never imports or evaluates that module itself** — only the worker does. That is what makes it the right tool for [process workers](/guides/process-workers/) and sandboxing: keep the code you want isolated in its own file, and the worker's permissions are what it runs under. ```ts // worker-tasks.ts — only the worker imports this. export const add = ([a, b]: [number, number]) => a + b; ``` ```ts // main.ts import { createPool, importTask, isMain } from "knitting"; export const add = importTask<[number, number], number>({ href: "./worker-tasks.ts", name: "add", }); if (isMain) { using pool = createPool({ threads: 2 })({ add }); console.log(await pool.call.add([2, 3])); // 5 } ``` `href` can be a relative path (resolved from the calling module), an absolute path, or a URL. `name` is the export to call and defaults to `"default"`. Worker permission policy applies to the import, so a strict pool can still load task modules but limit what they read, write, or reach. See [Permissions](/guides/permissions/). > Caution: The target export must be a **plain function** — pointing `importTask` at a > `task()` wrapper throws `TypeError: importTask expected export "…" to be a function`. Put task options on the host side instead: `importTask` accepts the > same `timeout` and `abortSignal` options as `task()`. ### Importing from a URL `href` can be remote, which is handy for shared task bundles: ```ts import { createPool, importTask, isMain } from "knitting"; const REMOTE = "https://knittingdocs.netlify.app/example-task.mjs"; export const addFromWeb = importTask<[number, number], number>({ href: REMOTE, name: "add", }); if (isMain) { using pool = createPool({ threads: 2 })({ addFromWeb }); console.log(await pool.call.addFromWeb([8, 5])); // 13 } ``` > Note: On Deno, keep `deno.lock` current for remote imports — update it with > `deno cache --lock=deno.lock --frozen=false --reload`, then run frozen with > `--lock=deno.lock --frozen=true`. Re-cache and commit the lockfile whenever the > remote module changes. ## A task can make its own pool One task in a short script? Chain `.createPool()` onto it and skip the separate call: ```ts import { isMain, task } from "knitting"; export const double = task({ f: (n: number) => n * 2, }).createPool({ threads: 2 }); if (isMain) { try { console.log(await double.call(21)); // 42 } finally { await double.shutdown(); } } ``` The single-task pool is created where it's defined (module scope), so close it with `shutdown()` rather than `using`. ## Advanced: overriding `href` on `task()` By default `task()` records the URL of the module it was called in, and workers import from there. Passing `href` points them at a different module. > Caution: The `href` override on `task()` is not a stable public contract and may be > removed in a future major release. Prefer the default caller resolution; if you > must use it, pass an absolute `file://` URL to a module that exports the task at > top level, and pin your version. --- # Creating pools URL: https://knittingdocs.vercel.app/guides/creating-pools/ Every createPool option in one place: threads, balancers, work stealing, the inliner lane, payload buffers, permissions, debug namespaces, and shutdown. `createPool(options)(tasks)` starts the worker threads and hands back three things: - `call.(args)` enqueues a task and returns a promise. - `shutdown(delayMs?)` stops the workers, either now or after a delay. - `[Symbol.dispose]`, which is what lets a `using` declaration close the pool when its scope ends. Reach for `using pool = createPool(...)({ ... })` and let the pool close itself. `await pool.shutdown()` is for the cases `using` cannot cover: closing **before** the scope ends, awaiting teardown, or running where `using` does not exist. Arguments may be promises as well as plain values. Knitting resolves them on the host before dispatch, so a rejected input rejects the call and the worker never runs it. See [Promise inputs are awaited on the host](/guides/defining-tasks/#promise-inputs-are-awaited-on-the-host). > Note: `createPool(...)` only works on the main thread. Guard any entrypoint that > creates one with `isMain`. ## Batching pattern Enqueued calls dispatch on their own, so the way to get a batch moving is to create every call first and await them together afterwards. ```ts const jobs = Array.from({ length: 1_000 }, () => call.hello()); const results = await Promise.all(jobs); ``` ## Options ```ts createPool({ threads?: number, inliner?: { position?: "first" | "last", batchSize?: number, dispatchThreshold?: number, }, balancer?: { strategy?: | "roundRobin" | "robinRound" | "firstIdle" | "randomLane" | "firstIdleOrRandom" } | "roundRobin" | "robinRound" | "firstIdle" | "randomLane" | "firstIdleOrRandom", worker?: { runtime?: "thread" | "process" | "compiled", processRuntime?: "node" | "deno" | "bun" | "porffor", processCommandPrefix?: string[], processSharedMemory?: "inherit" | "named" | { mode?: "inherit" | "named", namePrefix?: string, unlinkOnShutdown?: boolean, }, bootstrap?: { href: string, name?: string, data?: unknown }, resolveAfterFinishingAll?: true, timers?: { spinMicroseconds?: number, parkMs?: number, pauseNanoseconds?: number, }, hardTimeoutMs?: number, resourceLimits?: { maxOldGenerationSizeMb?: number, maxYoungGenerationSizeMb?: number, codeRangeSizeMb?: number, stackSizeMb?: number, }, }, payload?: { mode?: "growable" | "fixed", payloadInitialBytes?: number, payloadMaxByteLength?: number, maxPayloadBytes?: number, }, abortSignalCapacity?: number, host?: { steal?: boolean, stealRegionLanes?: number, doorbell?: boolean, stallFreeLoops?: number, maxBackoffMs?: number, dispatcher?: "per-thread" | "serial-channel", }, workerExecArgv?: string[], permission?: "strict" | "unsafe" | PermissionProtocol, dispatcher?: DispatcherSettings, // deprecated alias of host debug?: boolean | { host?: boolean, globals?: boolean, signals?: boolean, imports?: boolean, lifecycle?: boolean, }, source?: string, }) ``` Deprecated payload aliases are still accepted at the top level: - `payloadInitialBytes` -> `payload.payloadInitialBytes` - `payloadMaxBytes` -> `payload.payloadMaxByteLength` - `bufferMode` -> `payload.mode` - `maxPayloadBytes` -> `payload.maxPayloadBytes` ## threads How many worker threads to spawn (default `1`). Lanes are counted as `threads + (inliner ? 1 : 0)`. See [Multi-threading](/guides/multi-threading/) for choosing a worker count and configuring idle-worker timers in a server. ## payload These options tune the shared buffers that carry arguments out and results back. ### mode Picks how the shared buffer is allocated: - `"growable"`: starts at `payloadInitialBytes` and grows on demand, up to `payloadMaxByteLength`. - `"fixed"`: allocates `payloadMaxByteLength` at startup and stays that size. Growable is the default wherever growable `SharedArrayBuffer`s exist, fixed everywhere else. Asking for `"growable"` on a runtime without them gets you `"fixed"` regardless — the request is downgraded, not refused. ### payloadMaxByteLength How large a single payload buffer may grow, in bytes. Default `64 MiB`. ### payloadInitialBytes The size a buffer starts at, in bytes, default `4 MiB`. Growable mode clamps it to `payloadMaxByteLength`; fixed mode ignores it and allocates the full `payloadMaxByteLength` up front. ### maxPayloadBytes A hard ceiling on any one dynamically encoded payload. Must be `> 0` and `<= payloadMaxByteLength >> 3`, which is also the default (`8 MiB` with everything else left alone). Go over it and the call is rejected with `KNT_ERROR_3`, before a slot is even reserved. In `"fixed"` mode a payload can sit under the cap and still not fit what is left of the buffer. There is no room to grow into, so that call is rejected with an encoder error. ### How the limits are applied Every limit above is per worker and per direction. Each worker allocates two buffers, one for arguments and one for results, and each buffer gets the full allowance. ## abortSignalCapacity How many abort-aware calls the pool can track at once, default `258`. It only matters if a task declares `abortSignal` in the first place. Only tasks defined with `abortSignal: true` or `abortSignal: { hasAborted: true }` count against the limit. ```ts const pool = createPool({ threads: 4, abortSignalCapacity: 1024, })({ myAbortableTask }); ``` ## Example ```ts import { createPool, isMain, task } from "knitting"; export const add = task<[number, number], number>({ f: async ([a, b]) => a + b, }); if (isMain) { using pool = createPool({ threads: 2 })({ add }); const results = await Promise.all([ pool.call.add([1, 2]), pool.call.add([3, 4]), ]); console.log(results); // [3, 7] } ``` > Danger: After `shutdown()` runs, in-flight and future `call.*()` promises reject > with `"Thread closed"`. ## balancer Controls how calls are routed across lanes (threads, plus optional inliner). Pass a string or an object with a `strategy` key. - `roundRobin` (default): round-robin rotation through all lanes. - `robinRound`: legacy alias of `roundRobin`. - `firstIdle`: pick the first idle lane, else fall back to round-robin. - `randomLane`: pick a random lane. - `firstIdleOrRandom`: pick the first idle lane, else random. With one thread and no inliner there is nothing to balance, so calls skip the balancer and go straight to that worker. ## inliner Adds an extra lane that runs tasks on the main thread. - `position`: whether the inline lane appears before (`"first"`) or after (`"last"`) the worker lanes for balancing. - `batchSize`: max tasks processed per event-loop tick (default `1` when enabled). - `dispatchThreshold`: minimum in-flight calls per invoker before inline lane is eligible (default `1`). See [Inliner guide](/guides/inliner) for detail. ## worker ### runtime `"thread"` (default) runs workers as runtime-local threads — the lowest-overhead option. `"process"` runs each worker as a separate OS process for stronger isolation, and unlocks `processRuntime`, `processCommandPrefix`, and `processSharedMemory`. See [Process workers](/guides/process-workers) for the full story — sandboxes, containers, and the stdin / fd-0 handshake. `"compiled"` builds the task module into a native executable with Porffor and runs that as a child process. It is experimental, and supports a smaller feature set than the other two — see [Compiled workers](/guides/compiled-workers/). ### bootstrap A privileged module — `{ href, name?, data? }` — that every worker imports and awaits once, before any task module loads. Use it to install runtime guards, strip environment variables, or set up worker-only globals. It is worker-only, so it cannot be combined with the inline lane. ### resolveAfterFinishingAll Set this to `true` and workers wait for every pending promise to settle before they exit. ### timers What a worker does while it has nothing to run: - `spinMicroseconds`: busy-spin budget before parking. - `parkMs`: `Atomics.wait` timeout while parked. - `pauseNanoseconds`: `Atomics.pause` duration while spinning. Set `0` to disable. ### hardTimeoutMs A wall-clock timeout on every task call. Tripping it shuts the whole pool down, which is the only way to stop a worker already burning CPU in a tight loop. ### resourceLimits Memory and stack limits for Node.js workers: - `maxOldGenerationSizeMb` - `maxYoungGenerationSizeMb` - `codeRangeSizeMb` - `stackSizeMb` ## workerExecArgv Extra Node.js `execArgv` flags passed to workers, for example `["--expose-gc", "--max-old-space-size=4096"]`. When `permission` is set to `"unsafe"`, inherited Node permission flags (`--allow-fs-read`, `--allow-fs-write`, etc.) are stripped. ## permission Which permission flags the workers start with. - Leave `permission` out: strict defaults, plus `allowImport: true` so web imports still work. - `"strict"` (what you get when you pass an object): conservative defaults, worked out per runtime. - `"unsafe"`: no permission flags at all, and any inherited Node ones are stripped. - In object mode, `console` is `false` under strict and `true` under unsafe. See [Permissions guide](/guides/permissions) for runtime-specific mapping and strict defaults. ## Timing note Each worker takes one high-resolution `performance.now()` reading at startup and measures everything against it. Scheduling and timeouts stay precise that way, and global `performance` is left alone for your own code to use. ## Safety hardening defaults - The guards go in once, before the worker loop starts, so nothing extra runs inside the hot task loop. - Task code cannot take the process down: `process.exit`, `process.kill` and `process.abort` are blocked, along with `Deno.exit` where it exists. - Permissions are enforced by the runtime itself — Node's worker permission flags, Deno's worker permissions — rather than by monkey-patching FS, network or env from inside the worker. ## host Controls host-side scheduling and completion handling. The defaults select native work stealing for compatible multi-worker pools and use the best completion waiter available on the runtime. These options affect the host dispatcher, not task arguments or worker code. See [Work stealing](/guides/work-stealing/) for the topology, runtime support, and tuning guidance. ```ts using pool = createPool({ threads: 4, host: { steal: true, doorbell: true, }, })({ task }); ``` ### steal When enabled, workers claim tasks from one shared submit region instead of waiting behind private request lanes. Each worker keeps a private return lane, so the worker that claims a task also owns its response. The pool's pending registry still resolves the correct promise. For compatible pools with more than one worker, native work stealing is selected automatically. A one-worker pool has nothing to steal from. An explicit `balancer`, private-lane dispatcher, inliner, compiled worker, or an unsupported worker count can change that compatibility decision. Set `host.steal: false` to measure or use private request lanes. The task API, payload types, and completion semantics stay the same in either topology. ### stealRegionLanes Controls how many submit slots one stealing handshake claims. It must be a positive power of two; the default is the widest valid region for the worker count. Wider regions reduce arbitration overhead for many cheap, similarly sized calls. Smaller regions expose more independent work for expensive or uneven tasks. Start with `1` when task durations vary substantially, then benchmark the real workload. ### doorbell Requests an asynchronous host completion waiter instead of repeated response mailbox polling. It is enabled by default when Node.js or Bun thread workers provide `Atomics.waitAsync`. It is disabled for Deno, process workers, compiled workers, and browsers, where polling is used instead. Setting it to `true` cannot override a runtime limitation; set it to `false` for a polling baseline. ### stallFreeLoops How many immediate dispatcher turns run before escalation. The default is `1` when the doorbell is active and `128` when the dispatcher must poll. ### maxBackoffMs The longest the dispatcher will wait between polls once it starts stalling, in milliseconds (default `10`). `maxBackoffMs` affects only the polling fallback; it does not change the doorbell's asynchronous wait. ## dispatcher Deprecated alias of `host`. ## debug Streams diagnostics to **stderr**, each line tagged with the worker (`host`, or `w0`, `w1`, … for the thread/process workers), the runtime, and a millisecond timer relative to when that worker's debug initialised. Pass `true` to enable everything, or turn on individual namespaces: - `host`: host-side pool setup — cwd and caller, each registered task, runtime / workers / lanes / inliner, the module list, permission mode, and worker bootstrap. - `imports`: how many tasks each worker loaded, and from which modules. - `lifecycle`: the worker "ready" line and process-worker lifecycle events. - `signals`: per-dispatch worker traffic (work / result / run / idle). Very chatty. - `globals`: `globalThis` changes across the worker's bootstrap and task phases, so you can see which loader injected which global. Enable the same namespaces without touching code through the `KNITTING_DEBUG` environment variable — a comma-separated list (`KNITTING_DEBUG=host,imports`) or `*` for all. The option and the env var merge; either one can turn a namespace on. Debug is zero-cost when off: with no namespace active, the logger module is never even imported. ## source Point workers at a specific entry module instead of the one Knitting resolves for you. ## Limits One pool holds up to **65,536 tasks** — function IDs are `Uint16`, so the range is `0..0xFFFF`. Hand it more than that and you get a `RangeError`. --- # Payloads URL: https://knittingdocs.vercel.app/guides/payloads/ The data types you can pass into a task and return from one. The transport supports the following payloads: - `number` (including `NaN`, `Infinity`, and `-Infinity`) - `string` - `boolean` - `bigint` - `undefined` and `null` - plain JSON-like `Object` and `Array` - `Envelope` where `H` is JSON-like and `B` is `ArrayBuffer` (default), `SharedArrayBuffer`, `ProcessSharedBuffer`, or `BufferReference` - `Buffer` (Node.js), `ArrayBuffer` - `Uint8Array`, `Int32Array`, `Float64Array`, `BigInt64Array`, `BigUint64Array` - `DataView` - `ProcessSharedBuffer` for zero-copy shared memory (see [Shared memory](/guides/shared-memory/)) - `BufferReference` from `knitting/unsafe` for zero-copy buffers to thread workers (see [Buffer reference](/guides/buffer-reference/)) - `Error` (name, message, stack, and cause chain) - `Date` - `symbol` from `Symbol.for(...)` only - native `Promise` values at the host call boundary If you need multiple values, pass a tuple or object as the single argument. These types are not supported directly: - `Map`, `Set`, `WeakMap`, and custom class instances (except `Envelope` and subclasses) - non-global symbols - `Blob` - functions Promise values are accepted at the `call.*()` boundary, but promises themselves are runtime state and are not transferred through IPC as payloads. Only resolved values are serialized and sent to workers. If a promise input rejects, the host call rejects and the worker task is not executed. Only native `Promise` is accepted; thenables are treated as regular values. See [Promise inputs are awaited on the host](/guides/defining-tasks/#promise-inputs-are-awaited-on-the-host) for details. ```ts type JSONValue = | string | number | boolean | null | JSONArray | JSONObject; interface JSONObject { [key: string]: JSONValue; } interface JSONArray extends Array {} type ValidInput = | bigint | void | JSONValue | symbol | Envelope | Uint8Array | Int32Array | Float64Array | BigInt64Array | BigUint64Array | DataView | ArrayBuffer | Error | Date; type Args = ValidInput | Serializable; type TaskInput = NoBlob | Promise>; ``` ## Picking the right payload type Different types take different code paths and have very different costs. Use the cheapest type that fits your data. **Header-only (fastest):** `number`, `boolean`, `undefined`, `null`, `Date`, small `string`, small `bigint`. These fit in the call header with near-zero overhead. Prefer these whenever possible. **Static payload (fast):** `Symbol.for(...)`, large `bigint`, typed arrays (`Uint8Array`, `Int32Array`, etc.). Reuses a small buffer region alongside the header. **Dynamic payload (allocator path):** `Object`, `Array` (JSON-serialized), `Error`, `Date`, and larger strings. These need allocation, serialization, and copying. Still faster than `postMessage`, but measurably heavier in hot loops. See the [Performance guide](/guides/performance) for per-type benchmarks and batching guidance. ## Envelope `Envelope` pairs a JSON-serializable header with a binary body. Use it when a call needs both structured metadata and raw bytes — the transport carries one special binary value per call, so an envelope is how you attach a header to one. ```ts import { Envelope } from "knitting"; const message = new Envelope( { route: "/upload", contentType: "application/octet-stream" }, new Uint8Array([1, 2, 3]).buffer, ); ``` The second type parameter `B` sets the body type and defaults to `ArrayBuffer`. The supported body types are: | Body | Copy? | Works with | Notes | |------|-------|------------|-------| | `ArrayBuffer` | copied | thread + process | Default. Works everywhere. | | `SharedArrayBuffer` | zero-copy, shared | thread only | Shared by reference; process workers reject it. | | `ProcessSharedBuffer` | zero-copy, shared | thread + process | Cross-process shared memory. | | `BufferReference` | zero-copy, moved | thread only | From `knitting/unsafe`. Source is detached on construction. | `Envelope` is disposable — `[Symbol.dispose]()` disposes a disposable body (a `BufferReference`) and is a no-op for `ArrayBuffer` and `SharedArrayBuffer`. Use `using` to dispose automatically when the envelope goes out of scope: ```ts using result = await pool.call.processImage( new Envelope({ format: "png" }, buffer), ); console.log(result.header); // result is disposed here when the `using` scope exits ``` For `BufferReference` bodies, see [Buffer reference](/guides/buffer-reference/). For `ProcessSharedBuffer` bodies, see [Shared memory](/guides/shared-memory/). ## Errors `Error` payloads preserve `name`, `message`, `stack`, and recursive `cause` chains. Both sync throws and async rejections inside tasks propagate back to the host as `Error` objects. See also: [knitting/utils](/guides/utils/) for buffer serialization helpers and [Buffer reference](/guides/buffer-reference/) for zero-copy thread payloads. --- # Buffer utilities URL: https://knittingdocs.vercel.app/guides/utils/ knitting/utils — helpers for serializing strings, JSON, and numbers into SharedArrayBuffer and back. `knitting/utils` is a set of helpers for converting between JavaScript values and raw buffer bytes. Import from the subpath: ```ts import { bufferToBytes, bytesToBuffer, bufferToString, stringToBuffer, bufferToJson, jsonToBuffer, numbersToBuffer, bufferToNumbers, } from "knitting/utils"; ``` All functions accept `ArrayBuffer`, `SharedArrayBuffer`, or any typed-array view as input (`BufferLike`). Encode functions return `SharedArrayBuffer` so the result can be handed to a worker payload or stored in shared memory without an extra copy. --- ## Strings ```ts const sab = stringToBuffer("hello"); // UTF-8 → SharedArrayBuffer const text = bufferToString(sab); // SharedArrayBuffer → string ``` `stringToBuffer` uses Node's `Buffer.from` when available (avoids a TextEncoder allocation) and falls back to `TextEncoder` on Deno and Bun. --- ## JSON ```ts const sab = jsonToBuffer({ status: "ok" }); // JSON.stringify → SharedArrayBuffer const obj = bufferToJson(sab); // SharedArrayBuffer → parsed value ``` `bufferToJson` is `JSON.parse(bufferToString(source))`. If the value is not JSON-serializable, `jsonToBuffer` throws. --- ## Raw bytes ```ts const bytes: Uint8Array = bufferToBytes(source); // any BufferLike → Uint8Array const sab: SharedArrayBuffer = bytesToBuffer(source); // any BufferLike → SharedArrayBuffer ``` `bufferToBytes` does not copy if the source is already a `Uint8Array`. `bytesToBuffer` always copies into a fresh `SharedArrayBuffer`. --- ## Numbers ```ts type NumberFormat = "f64" | "f32" | "i32"; // default: "f64" const sab = numbersToBuffer([1.1, 2.2, 3.3], { format: "f64" }); const arr: Float64Array = bufferToNumbers(sab, { format: "f64" }); const isab = numbersToBuffer([1, 2, 3], { format: "i32" }); const iarr: Int32Array = bufferToNumbers(isab, { format: "i32" }); ``` `bufferToNumbers` returns a typed-array view (zero-copy when the byte offset is aligned). If the buffer's byte length is not a multiple of the element size for the chosen format, it throws a `RangeError`. --- ## Typical use Passing a string to a worker through shared memory so the bytes are not copied on every call: ```ts import { createPool, isMain, task } from "knitting"; import { getDefaultProcessSharedBufferPrimitives, ProcessSharedBuffer, } from "knitting/shared-memory"; import { stringToBuffer, bufferToString } from "knitting/utils"; export const shout = task({ f: (buf) => bufferToString(buf.view(Uint8Array)).toUpperCase(), }); if (isMain) { using pool = createPool({ threads: 1 })({ shout }); const primitives = getDefaultProcessSharedBufferPrimitives(); const encoded = stringToBuffer("hello knitting"); // SharedArrayBuffer const shared = ProcessSharedBuffer.create(encoded.byteLength, primitives); new Uint8Array(shared.view(Uint8Array)).set(new Uint8Array(encoded)); console.log(await pool.call.shout(shared)); // HELLO KNITTING shared.descriptor.mapping?.close?.(); } ``` See [Payloads](/guides/payloads/) for the full list of types that can cross the worker boundary. --- # Permissions URL: https://knittingdocs.vercel.app/guides/permissions/ Control what worker tasks can read, write, import, and run. Workers run your task code, and the `permission` option controls what that code can access at runtime. You can restrict file access, network connections, environment variables, imports, and subprocesses. Set the policy when you create the pool: ```ts import { createPool, isMain, task } from "knitting"; export const work = task({ f: async (x: number) => x * 2, }); if (isMain) { using pool = createPool({ threads: 2, permission: { mode: "strict" }, })({ work }); console.log(await pool.call.work(21)); // 42 } ``` ## Choose a mode There are three ways to configure permissions: - **Leave `permission` out.** Knitting uses the strict defaults and allows imports, including web imports. - **Use `permission: {}` or `permission: { mode: "strict" }`.** Knitting uses the conservative strict defaults and lets you add only the access your tasks need. - **Use `permission: "unsafe"`.** Knitting disables runtime permission flags and removes inherited Node permission flags from the worker. For most applications, keep strict mode and add a small allow-list. Use `"unsafe"` only when a dependency needs access that the runtime cannot express with the strict policy. The `console` option controls whether worker `console.*` calls are forwarded to the host. It defaults to `false` in strict mode and `true` in unsafe mode. ```ts createPool({ permission: { mode: "strict", console: true } })({ work }); createPool({ permission: "unsafe" })({ work }); ``` ## Grant only what the task needs Object mode lets you grant access one capability at a time. Anything you do not list stays denied: ```ts createPool({ permission: { mode: "strict", allowImport: true, // allow task-module imports read: ["./data"], // path allow-list (or `true` for all) write: ["./out"], net: ["api.example.com"], // host allow-list (or `true` for all) env: { allow: ["NODE_ENV"] }, run: ["git"], // subprocess allow-list console: true, }, })({ work }); ``` | Field | Controls | | --- | --- | | `read` / `write` | Filesystem allow-lists. `true` means unrestricted access. | | `denyRead` / `denyWrite` | Explicit denials applied after the allow-list. | | `net` / `denyNet` | Network host allow- and deny-lists. | | `allowImport` | Modules the worker may import. `true` allows all imports. | | `env` | Environment access through `{ allow, deny, files }`. | | `run` / `denyRun` | Subprocess execution allow- and deny-lists. | | `console` | Whether worker `console.*` output reaches the host. | Knitting uses each runtime's own permission mechanism, so the exact coverage varies by runtime (see [how each runtime enforces it](#how-the-runtimes-enforce-it)). Permissions are a guardrail, not a complete boundary for untrusted code. For stronger isolation, combine them with [process workers](/guides/process-workers/). ## What strict mode allows Strict mode starts with a conservative policy: - reads and writes are limited to the current `cwd`; - writes to `node_modules` are denied; - sensitive files and directories are denied, including `.env`, `.git`, `.npmrc`, `.docker`, `.secrets`, `~/.ssh`, `~/.gnupg`, `~/.aws`, `~/.azure`, `~/.config/gcloud`, and `~/.kube`; - sensitive POSIX paths are denied, including `/proc`, `/sys`, `/dev`, and `/etc`; - `deno.lock` and `bun.lock*` can still be read. `permission: "unsafe"` turns off runtime permission flags and removes inherited Node permission flags from the worker's `execArgv`. ## How the runtimes enforce it Knitting translates the same policy into each runtime's native permission system. The result is slightly different on Node.js, Deno, and Bun. ### Node.js workers Node workers receive `--permission` or `--experimental-permission`, along with the relevant allow flags: - `--allow-fs-read` - `--allow-fs-write` - `--allow-worker` - `--allow-child-process` - `--allow-addons` - `--allow-wasi` Node's worker flags are allow-list based. That means Knitting cannot represent every protocol-level deny-list rule as a native Node flag. ### Deno workers When enabled, Deno workers receive a `Worker.deno.permissions` policy. Knitting applies it only when one of these is true: - it detects `--unstable-worker-options` (using a Linux `/proc` check); or - `KNITTING_DENO_WORKER_PERMISSIONS=1` is set. ### Bun workers Bun does not currently provide worker permission flags. Knitting accepts the permission values for API compatibility, but Bun cannot enforce them through runtime flags yet. ## Controlling subprocesses If a task needs to start another process, object mode also supports runtime-specific overrides: - `node.allowChildProcess?: boolean` - `deno.allowRun?: boolean` — a legacy option, superseded by the top-level `run` allow-list. Both default to `false` in strict mode. Prefer the top-level `run` allow-list when you need to permit specific commands. --- # Process workers URL: https://knittingdocs.vercel.app/guides/process-workers/ Run each worker as a separate OS process for stronger isolation — through a sandbox like bubblewrap or a container like Docker. By default Knitting runs workers as **threads** — the lowest-overhead option. When you need stronger isolation, run each worker as a **separate process**. A process worker has its own memory and permissions, and — because it's just a child process — you can launch it inside a sandbox like `bwrap` or a container like Docker. ```ts const pool = createPool({ threads: 2, worker: { runtime: "process", processRuntime: "node", // "node" | "deno" | "bun" (default "deno") }, })({ add }); ``` That's the whole switch: the same `call` API and the same tasks, now running in another process. Everything below is about wrapping that process. ## Keep isolated code out of the host Isolation only helps if the code you want contained never runs on the host. Use [`importTask`](/guides/defining-tasks/#importing-worker-side-code-with-importtask) so the host holds a typed wrapper while only the worker imports and evaluates the module: ```ts // worker-tasks.ts — only the sandboxed worker imports this. export const add = ([a, b]: [number, number]) => a + b; ``` ```ts // main.ts export const add = importTask<[number, number], number>({ href: "./worker-tasks.ts", name: "add", }); ``` ## The fd-0 handshake Process workers receive their shared-memory handle on **stdin — file descriptor 0**. That one detail decides how a wrapper has to behave: - Wrappers that leave stdin alone (most sandboxes) work as-is, inheriting the fd. - Wrappers that replace, close, or proxy stdin (most containers) break the handshake. For those, switch to **named** shared memory so the worker reopens the mapping by name instead of inheriting an fd: ```ts worker: { runtime: "process", processSharedMemory: "named", // or { mode: "named", namePrefix: "knit" } } ``` Named memory needs both sides in the **same OS IPC namespace** (for Docker, `--ipc=host`). > Note: On Windows, Knitting always uses named shared memory for process workers, so you > don't need to set `processSharedMemory` yourself. ## `processCommandPrefix` `processCommandPrefix` is the wrapper command placed **before** Knitting's own worker launch. Knitting appends the real command — runtime plus the worker file — after your prefix, so the prefix is just "how to start a process that will then run the worker." ### Bubblewrap (keeps fd 0) `bwrap` preserves stdin, so the inherited-fd path works and no named memory is needed. This runs Bun workers with new namespaces (no network) and a read-only filesystem: ```ts import { createPool, importTask, isMain } from "knitting"; export const add = importTask<[number, number], number>({ href: "./worker-tasks.ts", name: "add", }); if (isMain) { using pool = createPool({ worker: { runtime: "process", processRuntime: "bun", processCommandPrefix: [ "bwrap", "--unshare-all", // new namespaces, no network "--ro-bind", "/", "/", // read-only filesystem "--dev-bind", "/dev", "/dev", "--proc", "/proc", "--tmpfs", "/tmp", // writable scratch "--die-with-parent", ], }, })({ add }); console.log(await pool.call.add([1, 2])); // 3 } ``` > Caution: `--tmpfs ` masks whatever is at that path. Don't put your project — the > files the worker imports — under a tmpfs-masked path, or the worker can't load > them. And remember it's isolation, not a full sandbox: a read-only bind is still > readable. ### Docker (named shared memory) Containers replace stdin, so use `processSharedMemory: "named"` and share the IPC namespace with `--ipc=host`. The container also needs to see the same files at the same path (a volume mount) and receive Knitting's two boot env vars. ```ts // docker-worker-tasks.ts — imported only inside the container. import { isMain } from "knitting"; export const addOne = (n: number) => n + 1; export const reportIsMain = () => isMain; // false: this runs off the host ``` ```ts // main.ts import { spawn } from "node:child_process"; import { mkdtemp, readFile, rm } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; import { createPool, importTask, isMain } from "knitting"; const docker = process.env.DOCKER_BINARY ?? "docker"; const image = process.env.KNITTING_DOCKER_IMAGE ?? "node:24-trixie-slim"; const cwd = process.cwd(); export const addOne = importTask({ href: "./docker-worker-tasks.ts", name: "addOne", }); export const reportIsMain = importTask({ href: "./docker-worker-tasks.ts", name: "reportIsMain", }); // Record the container id so we can always clean it up. const dockerPrefix = (cidfile: string): string[] => [ docker, "run", "--ipc=host", // share the shared-memory namespace "--cidfile", cidfile, "-v", `${cwd}:${cwd}`, // same files, same path "-w", cwd, "-e", "KNITTING_PROCESS_WORKER", // forward Knitting's boot payload "-e", "KNITTING_PROCESS_WORKER_BOOT", image, ]; const removeContainer = async (cidfile: string) => { const id = await readFile(cidfile, "utf8").then((s) => s.trim()).catch(() => ""); if (!id) return; await new Promise((resolve) => { spawn(docker, ["rm", "-f", id], { stdio: "ignore" }) .once("error", () => resolve()) .once("exit", () => resolve()); }); }; if (isMain) { const dir = await mkdtemp(join(tmpdir(), "knitting-docker-")); const cidfile = join(dir, "worker.cid"); const pool = createPool({ threads: 1, worker: { runtime: "process", processRuntime: "node", processSharedMemory: "named", processCommandPrefix: dockerPrefix(cidfile), }, permission: "unsafe", // the container is the boundary here })({ addOne, reportIsMain }); try { const [value, workerIsMain] = await Promise.all([ pool.call.addOne(41), pool.call.reportIsMain(), ]); console.log({ value, workerIsMain, ok: value === 42 && workerIsMain === false }); } finally { await removeContainer(cidfile); await pool.shutdown().catch(() => undefined); await rm(dir, { recursive: true, force: true }); } } ``` Why `permission: "unsafe"` here? The container — not Knitting's per-worker flags — is the isolation boundary, so the in-container worker runs without extra permission flags. Keep the image and mounts tight instead. See [Permissions](/guides/permissions/). ## Choosing a wrapper | Wrapper | stdin (fd 0) | Shared memory | | --- | --- | --- | | None (plain process) | inherited | inherited (default) | | `bwrap` / sandbox | preserved | inherited (default) | | Docker / container | replaced | `"named"` + `--ipc=host` | This is same-host communication. Named shared memory is fast because both sides map the same bytes, but it's not a network transport. For the lower-level building block behind all of this, see [Shared memory](/guides/shared-memory/). --- # Shared memory URL: https://knittingdocs.vercel.app/guides/shared-memory/ ProcessSharedBuffer — the lower-level shared-memory channel for passing bytes between workers and processes without copying. `ProcessSharedBuffer` is the building block under process workers: a block of shared memory two processes can read and write **without copying** the payload on every call. Reach for it when workers or processes need to see the *same bytes* — counters, ring buffers, large frames — instead of message-passing copies. It lives on a subpath: ```ts import { getDefaultProcessSharedBufferPrimitives, ProcessSharedBuffer, } from "knitting/shared-memory"; ``` The `primitives` are the platform's shared-memory functions. Grab the defaults once and reuse them. ## Anonymous buffers (parent ↔ child) The default is anonymous: a private handle passed intentionally through Knitting's transport. It's the safest option and needs no name. ```ts import { createPool, isMain, task } from "knitting"; import { getDefaultProcessSharedBufferPrimitives, ProcessSharedBuffer, } from "knitting/shared-memory"; export const readFirstCell = task({ f: (buffer) => Atomics.load(buffer.view(Int32Array), 0), }); if (isMain) { using pool = createPool({ threads: 1 })({ readFirstCell }); const primitives = getDefaultProcessSharedBufferPrimitives(); const shared = ProcessSharedBuffer.create(64, primitives); try { Atomics.store(shared.view(Int32Array), 0, 42); console.log(await pool.call.readFirstCell(shared)); // 42 } finally { shared.descriptor.mapping?.close?.(); } } ``` A `ProcessSharedBuffer` is a supported [payload](/guides/payloads/), so you pass it straight to a task. `view(Int32Array)` returns a typed-array view over the same memory — pair it with `Atomics` for safe cross-process reads and writes. ## Named channels (independent processes) When two processes don't share a parent — so there's no fd to inherit — use a **named** channel. One side creates the name, the other opens it. ```ts const name = "knitting-demo-channel"; const primitives = getDefaultProcessSharedBufferPrimitives(); const owner = ProcessSharedBuffer.create( { name, size: 64, mode: "create" }, primitives, ); try { Atomics.store(owner.view(Int32Array), 0, 7); const peer = ProcessSharedBuffer.create( { name, size: 64, mode: "open" }, primitives, ); try { console.log(Atomics.load(peer.view(Int32Array), 0)); // 7 } finally { peer.descriptor.mapping?.close?.(); } } finally { owner.descriptor.mapping?.close?.(); primitives.unlinkSharedMemory?.(name); } ``` Use `"create"` on the owner and `"open"` on the peer. **The name is the capability** — anyone who knows it can map the memory — so generate a hard-to-guess name, keep it private, and `unlinkSharedMemory` it when you're done. ## Sending one to a container Docker process workers can receive a `ProcessSharedBuffer`, but it must be **named** — the default anonymous form is fd-backed and private to the parent/child path, which a container can't reopen. Create the payload with `mode: "create"` and a name, run the pool with `processSharedMemory: "named"`, and add `--ipc=host` so the container shares the namespace. See [Process workers](/guides/process-workers/) for the pool side. ## Cleaning up Shared memory is not garbage-collected for you: - Close every mapping you open with `descriptor.mapping?.close?.()`. - For named channels, the owner also calls `primitives.unlinkSharedMemory?.(name)` once nobody needs the name anymore. This is a same-host, fast path — both sides map the same bytes — not a network transport. Anonymous is the safe default; reach for named only when processes can't inherit a handle. For thread-only zero-copy transfers within the same process, see [Buffer reference](/guides/buffer-reference/). --- # Buffer reference URL: https://knittingdocs.vercel.app/guides/buffer-reference/ Move large ArrayBuffers to thread workers without copying. `BufferReference` is useful when you need to send a large `ArrayBuffer` to a **thread** worker without paying for a copy. It moves the bytes instead of sharing them. Import it from the `knitting/unsafe` subpath: ```ts import { BufferReference } from "knitting/unsafe"; ``` This API is experimental, so it may still change. Normal use is safe on all supported runtimes. The optional `borrow` mode can avoid a copy for returned buffers, but it comes with stricter lifetime rules; see [Copying and borrowing returned buffers](#copying-and-borrowing-returned-buffers). --- ## Moving the buffer Creating a `BufferReference` **detaches the source immediately**. The bytes now belong to the reference, so the original view can no longer access them. For a typed-array view, `byteLength` and `length` become zero; APIs that require an attached `ArrayBuffer` may throw. ```ts const pixels = new Uint8Array([0, 64, 128, 192, 255]); const ref = new BufferReference(pixels); // pixels.buffer is now detached console.log(pixels.byteLength); // 0 — the source was moved console.log(ref.byteLength); // 5 ``` That is the trade-off that makes the transfer zero-copy: after the move, there is only one owner of the bytes. --- ## Send it to a worker Wrap the buffer and pass the reference as a task argument: ```ts import { createPool, isMain, task } from "knitting"; import { BufferReference } from "knitting/unsafe"; export const invert = task({ f: (ref) => { const pixels = ref.toUint8Array(); const out = new Uint8Array(pixels.length); for (let i = 0; i < pixels.length; i++) out[i] = 255 - pixels[i]; return new BufferReference(out); }, }); if (isMain) { const pixels = new Uint8Array([0, 64, 128, 192, 255]); using pool = createPool({ threads: 1 })({ invert }); const result = await pool.call.invert(new BufferReference(pixels)); console.log([...result.toUint8Array()]); // [255, 191, 127, 63, 0] } ``` --- ## Reading the bytes Use either accessor to read the bytes: | Method | Returns | Notes | |--------|---------|-------| | `toUint8Array()` | `Uint8Array` | A view over the bytes. | | `toArrayBuffer()` | `ArrayBuffer` | An `ArrayBuffer` containing the bytes; a subview may require a copy. | Both methods can be called more than once while the reference is active. Keep the reference alive while you use the returned view or buffer, and call `release()` when you are done. Releasing the reference detaches any views it created first, so an escaped view becomes empty or throws instead of reading freed memory. --- ## Releasing the reference `BufferReference` implements `Symbol.dispose`, so `using` is usually the easiest way to clean it up: ```ts { using result = await pool.call.invert(new BufferReference(pixels)); const out = result.toUint8Array(); console.log([...out]); } // result is released here ``` If you are not using `using`, call `release()` yourself. After `release()`, stop using any view you took from the reference. On runtimes where the host and worker cannot safely keep the same backing store alive, Knitting detaches those views before releasing the worker's memory. If detaching fails, it keeps the memory alive rather than risk a use-after-free. --- ## Important constraints - **Thread workers only.** The handle refers to memory in the current process. Sending it to a process worker throws. For cross-process sharing, use `ProcessSharedBuffer` (see [Shared memory](/guides/shared-memory/)). - **`ArrayBuffer`-backed views only.** `SharedArrayBuffer` cannot be detached and is rejected. SAB-backed typed-array views are also rejected. - **The move is one-way.** A reference may be read more than once while it is active, but it cannot be used after `release()`. Do not hand its view to a timer, stream, or other work that continues after the task. `BufferReference` is intended for trusted, same-process code. It is not a security boundary, so do not accept raw metadata or native pointers from untrusted code. --- ## Copying and borrowing returned buffers When a worker returns a `BufferReference`, Knitting must make those bytes readable on the host. By default, it chooses the safe option. You can make that choice explicit with `BufferReferenceReturn: "copy"`: ```ts using pool = createPool({ threads: 1, unsafe: { BufferReferenceReturn: "copy" }, })({ invert }); ``` On Node 22/24 with the current native support, the host can share the returned memory, so no copy is needed. On Deno, Bun, and Node builds without that support, Knitting makes one safe copy. Choose `"borrow"` when you want to skip that copy and can keep the result's lifetime under control: ```ts import { BufferReferenceReturn } from "knitting/unsafe"; using pool = createPool({ threads: 1, unsafe: { BufferReferenceReturn: BufferReferenceReturn.Borrow }, })({ invert }); const input = new Uint8Array([0, 64, 128, 192, 255]); { using result = await pool.call.invert(new BufferReference(input)); const out = result.toUint8Array(); // valid while result is alive console.log([...out]); } // result is released here; do not read out after this point ``` On Deno, Bun, and Node builds without native ownership support, `"borrow"` reads directly from worker memory. Release the result before its worker shuts down. If the bytes need to outlive the result—for example in an HTTP response, stream, timer, or cache—make a normal copy while the reference is still active, for example with `new Uint8Array(result.toUint8Array())`. Worker shutdown revokes outstanding borrows, but you should not rely on a view surviving it. If you prefer named constants, `BufferReferenceReturn.Copy` (`"copy"`) and `BufferReferenceReturn.Borrow` (`"borrow"`) are exported from `knitting/unsafe`. --- ## Use it in an `Envelope` An `Envelope` can carry a `BufferReference` body when you need a JSON header alongside binary data: ```ts import { Envelope, task } from "knitting"; import { BufferReference } from "knitting/unsafe"; export const processImage = task< Envelope<{ op: string }, BufferReference>, Envelope<{ done: boolean }, BufferReference> >({ f: (env) => { const pixels = env.payload.toUint8Array(); const out = new Uint8Array(pixels.length); for (let i = 0; i < pixels.length; i++) out[i] = 255 - pixels[i]; return new Envelope({ done: true }, new BufferReference(out)); }, }); ``` Disposing the envelope also disposes a `BufferReference` body. An `ArrayBuffer` or `SharedArrayBuffer` body has nothing to dispose. See [Payloads — Envelope](/guides/payloads#envelope) for the full body type table. --- ## When to use it For smaller buffers, the setup cost can outweigh the time saved by avoiding a copy. Use `BufferReference` when profiling shows that copying large buffers is actually a bottleneck; this is usually more relevant for buffers that are hundreds of kilobytes or several megabytes in size. For process workers, use `ProcessSharedBuffer` instead. For smaller payloads, a plain `ArrayBuffer` or typed array is simpler and works with both worker types. --- # Performance URL: https://knittingdocs.vercel.app/guides/performance/ How payloads, worker counts, and runtime choices affect performance. ## How to read the ratings These ratings are a rough guide to the cost of one call. A Slow entry can still be **2–4× faster** than `postMessage`, depending on the payload and the work being done. - Best — almost no per-call overhead - Fast — inexpensive for most workloads - Good — suitable for normal workloads - Fair — fine occasionally; watch hot paths - Slow — consider another representation in tight loops ### Thresholds The thresholds below are intentionally broad. They describe the approximate cost of one call, not the total time spent running the task: - Best: < 1 µs - Fast: < 2 µs - Good: < 4 µs - Fair: < 9 µs - Slow: > 9 µs --- ## Benchmark context The figures on this page were measured on an Apple M3 Ultra running Node 24.12.0 on arm64-darwin, at roughly 3.86 GHz. Treat them as useful comparisons, not promises: your runtime, CPU, payload shape, and task itself will change the result. --- ## How payloads move Most performance differences come down to how much data has to cross the worker boundary and whether that data is copied. - **Small values.** Numbers, booleans, short strings, and similar values fit in the call header, so encoding and decoding are very cheap. - **Small binary payloads.** Typed arrays and other small values use the transport's preallocated space and usually stay fast. - **Larger payloads.** The transport may need to allocate more space and copy the data. This is still efficient, but the cost becomes visible in a hot loop. - **Shared memory.** `SharedArrayBuffer` and `ProcessSharedBuffer` avoid copying the bytes. Knitting passes a handle instead, so the cost of sending a 1 KiB buffer and a 64 MiB buffer is roughly the same. `BufferReference` (`knitting/unsafe`) is the thread-only move variant: it detaches the source and hands the bytes to the worker. --- ## Typical cost by value type These ratings describe one value passed in a single call. They are useful for choosing a representation, but measure your real workload before tuning around them. | Value | Typical cost | | --- | --- | | Primitives: `boolean`, `undefined`, `null` | Best | | Numbers: `number` | Best | | Time/IDs: `Date` | Best | | Strings: small `string` | Best | | Symbols: `Symbol.for` | Fast | | BigInt: small `bigint` | Best | | BigInt: large `bigint` | Fast | | Binary: typed arrays | Best | | Views: `DataView` | Good | | Structured: JSON object | Good | | Structured: JSON array | Good | | Errors: `Error` | Slow | --- ## Tuning the pool ### Thread count Thread count depends on what the host thread must still do. For a mixed HTTP service, start with one worker: the host still accepts connections, routes, encodes requests, and sends responses. Add workers only when CPU-heavy calls queue and lower tail latency is worth the extra coordination. For independent batch compute, `os.availableParallelism() - 1` is a reasonable first trial. Every worker also consumes memory for payload buffers, shared locks, and cancellation state. Adding threads beyond the available cores usually stops helping. See [Multi-threading](/guides/multi-threading/) for a server-focused selection process and timer configuration. ### Native work stealing Compatible multi-worker pools use native work stealing by default. The host publishes calls to one shared submit region, and workers claim available tasks as they become free. Responses still use private return lanes, so the task API and promise behavior do not change. This helps most when many CPU-bound calls compete for workers and task durations are uneven. It is unlikely to be the main lever for tasks that mostly wait on databases, networks, or other external I/O. Set `host.steal: false` for a private-lane baseline, or tune `host.stealRegionLanes` for the workload: - Wider regions reduce arbitration overhead for many cheap, similarly sized calls. - Narrower regions expose more independent work for expensive or uneven calls. `1` is a useful starting point when one task can take much longer than another. The host `doorbell` is a separate completion optimization. Node.js and Bun thread pools can use `Atomics.waitAsync` instead of repeatedly polling for responses; Deno, process, compiled, and browser pools use the polling fallback. Compare like with like when benchmarking: keep the worker topology fixed and change one of `host.steal` or `host.doorbell` at a time. See [Work stealing](/guides/work-stealing/) for the full option and compatibility details. ### Inliner The inliner runs eligible calls without sending them through a worker, skipping transport encoding and decoding. It is most useful for very small tasks, such as simple arithmetic. For example, `inliner: { position: "last", batchSize: 64 }` can improve throughput. See the [Inliner guide](/guides/inliner). ### Permissions Strict permissions add a small amount of startup work for each worker, such as generating flags and resolving lock files. They do not add measurable overhead to individual calls once the workers are running. ### Payload sizing `payload.payloadInitialBytes`, `payload.payloadMaxByteLength`, and `payload.maxPayloadBytes` control the transport buffer used by each worker. Increasing the initial size avoids growth later, but uses more memory up front. For consistently small payloads such as primitives and short strings, the defaults—`4 MiB` initially, `64 MiB` maximum length, and an `8 MiB` dynamic payload cap—are usually enough. --- ## Choosing threads or processes Threads and processes have different memory boundaries, so the best zero-copy option depends on which one you use: | Runtime | Isolation | Zero-copy tools | | --- | --- | --- | | `thread` (default) | shares the host address space | `SharedArrayBuffer`, `BufferReference` (move) | | `process` | separate memory and permissions | `ProcessSharedBuffer` (OS shared memory) | A `SharedArrayBuffer` or `BufferReference` cannot cross a process boundary. Use `ProcessSharedBuffer` when the worker runs in a separate process. See [Shared memory](/guides/shared-memory/) and [Buffer reference](/guides/buffer-reference/). Because a process worker is a child process, you can start it through another tool using `worker.processCommandPrefix`. This is useful for a sandbox such as `bwrap` or a container such as Docker: ```ts worker: { runtime: "process", processCommandPrefix: ["bwrap", "--unshare-all", "--ro-bind", "/", "/"], } ``` See [Process workers](/guides/process-workers/) for the full wrapper recipes. > Caution: For security-sensitive work, define the task with > [`importTask`](/guides/defining-tasks/#importing-worker-side-code-with-importtask). > The host holds only a typed wrapper and **never imports or evaluates the > module**. The task runs only inside the worker, under the worker's permissions, > and never at host scope. --- ## Keep request bodies off the main thread `call.*()` accepts promises for supported inputs. In an HTTP handler, you can pass the request body's promise straight to a task instead of awaiting it on the request thread: ```ts app.post("/jwt", async (c) => { const responseJson = await handlers.call.issueJwt(c.req.arrayBuffer()); return c.body(responseJson ?? "Bad request", responseJson ? 200 : 400, { "content-type": "application/json; charset=utf-8", }); }); ``` The promise itself is not faster. The benefit is that the request thread can hand the body to Knitting immediately: - `c.req.arrayBuffer()` already returns a promise, so forwarding it skips an `await` in the handler. - UTF-8 decoding and JSON parsing happen in the worker, not on the request thread. - `ArrayBuffer` stays on the binary fast path. ### Metadata and body with `Envelope` When a task needs both request metadata and the raw body, put them in an `Envelope`. The header carries the metadata and the payload carries the bytes. Use `.then(...)` to build the envelope from the body promise without awaiting the body on the request thread: ```ts import { Envelope } from "knitting"; app.post("/upload", async (c) => { const result = await handlers.call.storeUpload( c.req.arrayBuffer().then( (body) => new Envelope( { contentType: c.req.header("content-type") ?? "application/octet-stream" }, body, ), ), ); return c.json(result); }); ``` For a large binary body sent to a **thread** worker, use a `BufferReference` instead of an `ArrayBuffer`. The bytes move without a copy; only the body line changes: ```ts import { BufferReference } from "knitting/unsafe"; new Envelope( { contentType: c.req.header("content-type") ?? "application/octet-stream" }, new BufferReference(body), // moves the body bytes without copying ); ``` This works best when the route is mainly forwarding data and the worker does the parsing, such as SSR or JWT issuance. If the main thread needs to inspect or validate the body first, await it there instead. ### Choosing how to return data Returning data has the same costs as sending it, just in the other direction. Choose the return type based on the size of the result and how the data is already represented: - **JSON object / array** — serialized on the worker and parsed again on the host, so it makes two passes over the data. This is fine for small results but expensive for large ones. - **`SharedArrayBuffer` / `ProcessSharedBuffer`** — shared memory is usually the cheapest way to return bytes when the result can use it. `ProcessSharedBuffer` also works with process workers. - **`BufferReference`** — useful for large binary results, especially when a library gives you an ordinary `ArrayBuffer` that you cannot turn into shared memory. Returned references can stay zero-copy on Node with native ownership; the safe default makes one copy on Deno, Bun, and some Node builds. See the [Buffer reference guide](/guides/buffer-reference/) for the explicit borrow option and its lifetime rules. See [Payloads](/guides/payloads/) and [Buffer reference](/guides/buffer-reference/) for the full type list. --- # Inliner URL: https://knittingdocs.vercel.app/guides/inliner/ Run pure compute on the host thread as an extra lane. The `inliner` option adds the host (main thread) as one extra lane in the pool. Instead of sending everything to workers, the host participates in execution too -- useful when every task is pure computation and you want to squeeze one more core out of the machine. ```ts const pool = createPool({ threads: 4, inliner: { position: "last", batchSize: 16 }, })({ add }); ``` > Note: The inliner is a lane, not a replacement for worker threads. > Use it to complement workers, not to avoid them. ## When to use it The inliner shines for **math and pure-compute workloads** that run on the host without touching the network or filesystem: - Number crunching, scoring, hashing, matrix ops. - Batch transforms over arrays of primitives. - Short, synchronous functions where IPC overhead matters more than the work itself. - Bursty queues where one extra lane helps drain work faster. Because inline tasks skip worker IPC entirely (no encode/decode round-trip), they can be significantly faster for tiny payloads. ## When NOT to use it - **HTTP / networking** -- inline tasks run on the main thread, so any I/O blocks the event loop and defeats the isolation that workers provide. If you need request handling, keep it in workers. - **File system or database calls** -- same problem. Anything that awaits external I/O will stall timers, sockets, and other pools sharing the host. - **Long-running async work** -- the inliner is designed around fast, ideally synchronous functions. Async tasks that take tens of milliseconds or more will block the batch loop and starve other inline slots. - **Isolation-sensitive code** -- if a task can throw or corrupt shared state, run it in a worker where a crash stays contained. > Warning: Inline tasks run on the main thread. Anything that blocks -- network calls, disk > reads, heavy async chains -- will freeze the event loop. Stick to pure math > and transforms. ## How it runs Inline execution is not immediate in the `call.*()` path. Calls are queued, then processed when the macro-queue turn runs. This delay is intentional: by that point, the dispatcher has had a chance to send/receive worker tasks first, then the host drains inline work. `batchSize` controls how many inline tasks run per macro-queue turn. For compute workloads, higher values let the host churn through more work per tick without yielding back to the event loop unnecessarily. ## Options ```ts createPool({ threads: number, inliner: { position?: "first" | "last", batchSize?: number, dispatchThreshold?: number, }, balancer?: "roundRobin" | "robinRound" | "firstIdle" | "randomLane" | "firstIdleOrRandom", }) ``` ### position Controls where the inline lane sits relative to worker lanes. - `"first"` -- host lane is considered before workers. Good when inline work is cheaper than IPC and you want the host to grab tasks first. - `"last"` -- host lane is considered after workers. Workers get priority; the host only picks up overflow. For most compute pools, `"last"` is the safe default. ### batchSize How many inline tasks are processed per macro-queue turn. - Higher values = better throughput for pure math (fewer yields to the event loop). - Lower values = more responsive host (other timers and callbacks get a chance to run between batches). Defaults to `1` when the inliner is enabled. For compute-heavy pools you typically want a much higher value (16, 64, 128+). ### dispatchThreshold Minimum in-flight calls before the inline lane becomes eligible for scheduling. - `1` (default) -- inline lane is immediately eligible. - Higher values -- host lane stays excluded until concurrency rises past the threshold, then joins to help drain the burst. This is a pressure-relief valve: at low concurrency, workers handle everything; once a burst builds up, the host pitches in. ## Exact behavior - The scheduler tracks `inFlight` calls per task invoker. - On each call, `inFlight` increments before lane selection. - If `inFlight < dispatchThreshold`, scheduling uses worker-only lanes (inline lane excluded). - If `inFlight >= dispatchThreshold`, scheduling uses all lanes (workers + inline lane). - `inFlight` decrements on resolve, reject, or synchronous throw. - The configured balancer strategy applies to whichever lane set is currently active. ## Internals The inline executor uses typed arrays (`Int32Array`, `Int8Array`) to manage execution slots and a `RingQueue` for pending work. It coordinates with the event loop through a `MessageChannel` (macro-task boundary) and `queueMicrotask` (micro-task fast path). The first dispatch in a burst resolves in microtasks; overflow beyond `batchSize` defers to the next macro-task turn. Promise arguments are awaited before execution (unlike thenables, which are passed through as-is). Timeout specs from `task()` are applied via a `Promise.race` wrapper only when the task returns a `Promise`. Abort signals on inline tasks use a static toolkit where `hasAborted()` always returns `false` -- inline tasks cannot be individually aborted since they share the host thread. The toolkit still carries `now()`, which behaves normally and reads the host's clock. ## Balancer guidance - `roundRobin` -- simple rotation across all lanes. Works well when tasks are uniform. - `robinRound` -- legacy alias of `roundRobin`. - `firstIdle` -- picks the first idle lane. Prioritizes workers when `position: "last"`. - `firstIdleOrRandom` or `randomLane` -- useful for pools with many registered tasks or uneven load. ## Examples ### Math pipeline High batch size, host joins after workers: ```ts using pool = createPool({ threads: 4, inliner: { position: "last", batchSize: 64 }, balancer: "firstIdleOrRandom", })({ scoreChunk }); const results = await Promise.all( chunks.map((chunk) => pool.call.scoreChunk(chunk)), ); ``` ### Burst drain with threshold Host stays out until concurrency spikes: ```ts using pool = createPool({ threads: 2, inliner: { position: "last", batchSize: 32, dispatchThreshold: 16, }, balancer: "roundRobin", })({ hash }); ``` ### Single-thread + inliner Useful when you want one worker for isolation but the host can handle the easy math too: ```ts using pool = createPool({ threads: 1, inliner: { position: "first", batchSize: 8 }, balancer: "roundRobin", })({ add }); ``` --- # Multi-threading URL: https://knittingdocs.vercel.app/guides/multi-threading/ Choose a worker count for HTTP services and CPU-bound batch work without starving the host thread. Moving work into a worker does not remove it from the host. Only the task body leaves the request thread: the host still accepts the request, routes it, encodes the arguments, submits the call, receives the result, and writes the response. All of that costs CPU, so pick a worker count that leaves room for it. ## Start with the workload | Workload | Starting point | Why | | --- | --- | --- | | HTTP service with a few CPU-heavy routes | `threads: 1` | Keeps expensive work off the event loop without taking many cores from the host. | | HTTP service with a growing queue of heavy calls | `threads: 2` to `4` | More workers can improve tail latency while requests wait for CPU. | | Independent CPU-bound batch jobs | `available cores - 1` | The worker bodies can execute in parallel; leave a core for the host and OS. | | Mostly network, database, or filesystem work | Keep it async; do not add workers just for I/O | A worker does not make an external dependency faster. | Treat these as places to start. Task duration, payload size, request rate, and the runtime itself all move the answer. ## A server starting point Move the expensive route work into tasks and leave the rest of the handler on the request thread. One worker is usually enough to clear the head-of-line blocking, and it keeps most of the machine available to the host. ```ts import { createPool } from "knitting"; const threads = 1; const handlers = createPool({ threads, })({ renderSsrPage, issueJwt }); ``` The Hono example is built this way: SSR and JWT run in the worker, `/ping` stays on the request thread. Under a saturating mixed load, `/ping` served 201% more requests than the single-threaded version, because it was no longer stuck behind a render. See the [Hono server example](/examples/data_transforms/rendering_output/hono_server/) and its [16-core measurements](/documents/hono-16core-benchmark.md). ## When a second worker helps A second worker earns its place when heavy calls are queueing behind the first one and the host still has CPU to spare. Each worker you add competes with the host for cores and brings its own scheduling, memory, and result draining. In the Hono workload, one worker gave the highest throughput under saturation. Larger pools did cut the p99 of the heavy routes at a fixed offered rate, but they could not make the host a faster producer, and that is what a saturated run measures. This is the usual outcome when a single event loop feeds the pool. Past one compatible worker, [native work stealing](/guides/work-stealing/) is automatic. Workers pull from shared work instead of each sitting on a private backlog, so whichever one is free takes the next task. That evens out uneven task durations. The host-side cost stays where it is. ## What a worker does when it runs out of work Idle CPU is a resource the host can use, so a worker that has caught up spends as little of it as it can, in this order: 1. Finished results go out first. In a stealing pool, a worker sends its response before it tries to claim more work. 2. Safe return-side releases are drained. A `BufferReference` is released only once its result has been consumed or released explicitly. 3. The worker looks for available shared work. If there is none, it makes a best-effort `maybeGc()` call where the runtime offers one, then parks. How it parks depends on the size of the pool: - A single worker spins for 50 µs first, because it sits on the request's critical path. - Multi-worker pools skip the spin and park immediately, leaving the CPU to the host or to an awake peer. Both are defaults. `worker.timers` is there if you need to override them. Each bar is total server CPU divided by completed requests under saturating mixed Hono load, so a shorter bar leaves more CPU for the host and for everything else on the machine. One worker is the efficient point for this workload: 100 CPU-µs per request at 18,014 RPS. A larger pool can still be the right call if what you are buying with that CPU is tail latency. Read those bars knowing that nothing above one worker is saturated here. The host is the only producer, so the extra workers spend most of their time waiting for a call that may never arrive. A pool that spun through that wait would look busy without being useful: at fifteen workers, the old 750 µs spin budget burned 6.4 idle cores and 954 CPU-µs per request to serve fewer requests than the current policy serves at 236. The bars stay short because a worker with nothing to do parks instead of billing the machine for waiting. ## Measure the choice Two runs per candidate thread count are enough: 1. **Saturating mixed load** shows the ceiling, and whether cheap endpoints are still delayed by expensive ones. 2. **Fixed offered rate** makes p50 and p99 comparable, since every candidate receives the same work. Record server CPU, idle CPU, throughput, and p50/p99 for each route. Total RPS on its own will mislead you: a configuration that cuts a heavy route's tail latency while giving up some throughput is often the one you want. For the lower-level options, see [Creating pools](/guides/creating-pools/), [Performance](/guides/performance/), and [Work stealing](/guides/work-stealing/). --- # Compiled workers URL: https://knittingdocs.vercel.app/guides/compiled-workers/ Compile supported tasks into native Porffor workers. If you want to experiment with compiling a task into native code, compiled workers let you do that with [Porffor](https://github.com/CanadaHonk/porffor). Porffor compiles the task module and Knitting runs the result as a child process. Your task code and the `pool.call.*()` API stay the same. The only change is the worker configuration: > Caution: Porffor is still early-stage, so expect its compiler and artifact format to > change. The generated program is native code and is not sandboxed. This is a > feature to try with code you trust, not a security boundary for untrusted > tasks. ## Enable Porffor Let's say your task module is `tasks.ts`. Knitting will look beside it for two files: the compiled program, `tasks.knt`, and its manifest, `tasks.knt.json`. For the usual setup, use both settings below. They tell Knitting to use Porffor and to reuse the compiled file whenever it is still valid: ```ts import { createPool, isMain } from "knitting"; export const hello = (name: string) => "Hello " + name; if (isMain) { using pool = createPool({ worker: { runtime: "compiled", processRuntime: "porffor" }, })({ hello }); console.log(await pool.call.hello("World!")); // Hello World! } ``` Each setting has a separate job: - `runtime: "compiled"` reuses a compatible `.knt` file and rebuilds it if it is missing, stale, or no longer matches the current setup. - `processRuntime: "porffor"` selects Porffor. Used by itself, it rebuilds once for each pool. If the pool has several workers, Knitting still compiles the module only once; all of the native workers start from the same artifact. ## Keep the task module self-contained Porffor bundles the task module and everything it imports. That works well for small, self-contained computations, but it means the task module has to stay simple. There are two rules worth knowing: - From `knitting`, import only `task`, `isMain`, and `createPool`. APIs such as `Envelope`, `importTask`, and `checkCompiledWorker` do not have compiled-worker equivalents and will make the build fail. Keep host-only code in a separate module when you need it. - Do not import Node.js built-ins such as `node:fs` or `node:crypto`. Porffor does not provide them, and some unsupported imports can hang the compiler instead of producing a useful error. Plain local modules are fine. The short version: compiled workers are a good match for focused, synchronous computation. If a task needs I/O, timers, or the wider Node.js runtime, a regular thread or process worker will be a better fit. ## Build artifacts ahead of time When Knitting needs to build automatically, it looks for Porffor in this order: 1. `worker.compiled.compiler` 2. `PORFFOR_MAIN` or `PORF` 3. `porf` on `PATH` If it cannot find one, Knitting downloads a pinned compiler to `$XDG_CACHE_HOME/knitting` or, when that variable is not set, `~/.cache/knitting`. For deployments, you may prefer to build the artifact in CI or during your release step. That keeps production from invoking a compiler at startup: ```bash bun run build:compiled --module tasks.ts --out tasks.knt --tasks addOne ``` ```ts const pool = createPool({ worker: { runtime: "compiled", compiled: { artifact: "./build/tasks-linux-x64.knt", build: false, }, }, })({ addOne }); ``` The `worker.compiled.build` option controls when Knitting is allowed to build: - `true` — build only when the artifact cannot be reused; - `false` — never build; fail if the artifact is unavailable; - `"always"` — rebuild every time. Use `worker.compiled.manifest` when the sidecar manifest is somewhere else. Before starting a worker, Knitting checks that the manifest still matches the protocol version, platform, architecture, source module, source timestamp, and requested task names. If you want to check this yourself, `checkCompiledWorker(task, options)` reports the compatibility state without building the artifact or running the task. ## Supported values and limits Compiled workers support fewer payload types than regular thread and process workers: | Value | Behavior | | --- | --- | | JSON primitives, arrays, and plain objects | Copied; limited to 1 MiB per call | | `ArrayBuffer`, `DataView`, and typed arrays | Copied | | `ProcessSharedBuffer` | Mapped by the worker instead of copied | | `Promise` | Resolved on the host before dispatch | | `Envelope`, `BufferReference`, and BigInt typed arrays | Rejected | There are two other limits to keep in mind: task functions must be synchronous, so returning a promise is not supported, and strings may contain BMP characters but not supplementary Unicode code points yet. For large binary data, use `ProcessSharedBuffer`. It is the one payload type that a compiled worker maps directly instead of copying through the call frame. See [Shared memory](/guides/shared-memory/) to get started with it. ## Abort signals You can also use cooperative cancellation on POSIX systems. Knitting publishes the abort state in named shared memory, and the compiled worker reads it directly instead of asking the host on every check. ```ts import { task } from "knitting"; export const search = task({ abortSignal: true, f: (limit: number, signal) => { for (let i = 0; i < limit; i++) { if (signal.hasAborted()) return i; } return limit; }, }); ``` `signal.now()` returns a monotonic millisecond clock for measuring elapsed time inside the task. Cancellation is cooperative: if the task never checks the signal, it continues until it finishes. Windows support is not available yet. ## What is not supported yet Porffor is intentionally a smaller backend for now. The options below fail during pool creation or invocation instead of quietly switching to another worker type: | Option | Why it is unsupported | | --- | --- | | `inliner`, `host` | There is no compiled equivalent of a host-side lane. | | `permission` | Native artifacts do not use runtime permission flags. | | `worker.bootstrap` | The worker does not import a host bootstrap module. | | `worker.timers`, task `timeout` | The compiled worker has no timer scheduler. | | `importTask` | The compiler needs the task body at build time. | | `payload`, `unsafe`, `source`, `workerExecArgv` | These options belong to the regular frame transport. | | `worker.processCommandPrefix`, `worker.processSharedMemory`, `worker.resolveAfterFinishingAll` | These are process-worker options with no compiled equivalent. | `worker.hardTimeoutMs` is the exception: it works because the host enforces it from outside the compiled worker. If you need one of these features, use a regular process worker instead. See [Process workers](/guides/process-workers/). ## Is this a good fit? Porffor is worth trying when your task is small, synchronous, self-contained, and mostly CPU work. It is less suitable when the task depends on libraries, filesystem or network access, timers, rich payloads, or runtime-specific APIs. If you are unsure, start with a regular thread or process worker first. Once the task works there, compiling it with Porffor is a straightforward experiment. --- # Browser URL: https://knittingdocs.vercel.app/guides/browser/ Run Knitting in a browser — download the build, see what works over web workers, and learn what a page cannot do. `knitting/browser` runs the same pool API on web workers and `SharedArrayBuffer`. Tasks, typed-array payloads, parallel calls, abort signals, and shutdown all work the way they do on Node, Deno, and Bun. What follows is the part that does not carry over: what a page cannot do, what you see when you try, and why. [Download Latest Stable Browser Build](/knitting.js) [Open Quick Start](/start/quick-start/) > Danger: Do not use this website as a resource. `/knitting.js` is served open > (`Access-Control-Allow-Origin: *`) so the smoke test can run, not as a CDN. It is > unversioned, it changes whenever these docs are rebuilt, and it carries no uptime, > caching, or integrity guarantees. Download the file or install `knitting` from npm, > then serve it from your own origin. The [browser smoke test](/guides/browser-smoke-test/) runs the hosted bundle in your own browser. It starts real one-thread pools through both `task()` and `importTask()`, so you can watch the build work before you install anything. ## Two things your page must do ### Serve the page cross-origin isolated A browser only hands out `SharedArrayBuffer` when the page arrives with both of these headers: ``` Cross-Origin-Opener-Policy: same-origin Cross-Origin-Embedder-Policy: require-corp ``` Without them, `createPool` throws before it starts a worker: > SharedArrayBuffer is unavailable: serve the page cross-origin isolated > (Cross-Origin-Opener-Policy: same-origin, Cross-Origin-Embedder-Policy: > require-corp). That is the browser's rule, not Knitting's, and there is no way around it. No shared memory means no pool. Turning isolation on affects the whole page, not just Knitting. Every cross-origin image, script, or font now needs `Cross-Origin-Resource-Policy` or CORS, or the browser refuses to load it. This site sends both headers in local Astro dev and preview, and on Netlify through `public/_headers`. If you host the bundle somewhere else, set them up there too. When the smoke test reports `crossOriginIsolated: false`, a missing header is almost always the reason. ### Call `setModuleUrl(import.meta.url)` in every task module A worker has to import the module your tasks live in, so Knitting needs that module's URL. On Node, Deno, and Bun it finds the URL by reading the call stack. That does not survive a browser: it depends on `Error.prepareStackTrace`, which only V8 provides, and a bundler rewrites the paths anyway. So in a browser the module says where it lives: ```js import { setModuleUrl, task } from "knitting/browser"; setModuleUrl(import.meta.url); export const square = task({ f: (value) => value * value }); ``` Put the call at the top, above the `task()` and `importTask()` calls it covers. Leave it out and the first of those calls throws, long before you reach a pool: > Unable to determine caller file. This runtime exposes no stack traces (e.g. > Andromeda); call setModuleUrl(import.meta.url) at the top of the module that defines > your tasks before creating a pool. An unbundled page in Chromium works without the call, because there the stack really does name the module. Do not rely on that. It breaks as soon as the page goes through a bundler or opens in Firefox or Safari. ## Basic URL import example You do not need a package import in a browser. Load the bundle from a URL, then point `importTask()` at a second URL that holds your tasks. ```ts const knittingUrl = new URL("/knitting.js", window.location.origin).href; const taskModuleUrl = new URL("/example-task.mjs", window.location.origin).href; const { importTask } = await import(knittingUrl); const add = importTask<[number, number], number>({ href: taskModuleUrl, name: "add", }); const pool = add.createPool({ threads: 1 }); try { console.log(await pool.call([2, 3])); // 5 } finally { await pool.shutdown(); } ``` The module on the other end names itself the same way: ```js import { setModuleUrl, task } from "./knitting.js"; setModuleUrl(import.meta.url); export const add = task({ f: ([a, b]) => a + b, }); ``` That is the whole wiring. The page loads `knitting.js` by URL, each task module reports its own URL, and the workers import from there. ## What a page cannot do | Feature | What happens | | --- | --- | | Process workers (`worker.runtime: "process"`, `processRuntime`) | throws `process workers are unavailable in the browser build` | | Compiled / Porffor workers (`runtime: "compiled"`, `.knt` artifacts) | throws `compiled workers are unavailable in the browser build` | | `BufferReference` | throws `BufferReference cannot run in runtime "browser"` | | `ProcessSharedBuffer`, named shared memory | throws `ProcessSharedBuffer is unavailable in the browser build` | | Native addons, FFI, file descriptors | unreachable; nothing in a page can load them | | `checkCompiledWorker` | not exported from `knitting/browser` | | Permissions (`permission: {...}`) | accepted and ignored, see below | All of these need a filesystem, a process to spawn, or FFI. The browser build replaces them with stubs that throw the messages above, so a call fails where you wrote it instead of somewhere deep inside a worker. ### Permissions do nothing here The `permission` option is accepted and then skipped. There is no filesystem to restrict, no process to sandbox, and no runtime flags to pass, so a policy that locks down a Node worker locks down nothing in a page. > Caution: A web worker has the same privileges as the page that started it: same origin, same > `fetch` reach, same storage. Do not run task code in a browser pool that you would not > trust with your origin. Sandboxing that code is the browser's job, through an iframe > on a separate origin, and not something Knitting can do for you. ### Passing a `SharedArrayBuffer` yourself Handing a `SharedArrayBuffer` to a task *as an argument* works on Node, Deno, and Bun. In a browser it throws: ``` KNT_ERROR_3: Unsupported payload type; BufferReference cannot run in runtime "browser" ``` That path shares a buffer by pinning a pointer through FFI, and a page has no FFI. The pool's own transport is unaffected, which is how the workers talk at all. Only buffers you pass yourself are refused. A browser version of this is possible. A cross-origin isolated page can send a `SharedArrayBuffer` straight through `postMessage`, with no pointer involved. It is just not implemented yet. ## Works, but differently - **Workers start from a message, not from `workerData`.** The pool posts the boot payload after it constructs the worker. You cannot see this from the API; it matters if you are reading worker startup code. - **`KNITTING_DEBUG` does nothing.** The env gate reads `Deno.env` or `process.env`, and a page has neither. Pass the `debug` option to `createPool` instead. - **`threads` still defaults to 1.** No runtime picks a thread count for you. In a browser, `navigator.hardwareConcurrency` tells you how many cores you have. - **Every worker parses the whole bundle.** The worker URL is the bundle's own URL, so memory use grows with thread count. ## Not a browser problem These fail the same way on Node, so do not go looking for a browser cause: - `BigInt` payloads: `Do not know how to serialize a BigInt` - `Map` and `Set` payloads: `Unsupported object type` - Functions as payloads: `KNT_ERROR_0: Function is not a valid type` - `Error` values round-trip as `{ name }`, dropping the message ## Browser support Tests run against headless Chromium in two layouts: one bundle holding tasks and library together, and the standalone single-file bundle loaded from a script tag beside a separate task module. Chromium is the only engine covered. Firefox and Safari have the same building blocks, web workers and `SharedArrayBuffer` under cross-origin isolation, so they are expected to work, but nothing has been checked against them. One difference is already known: neither implements `Error.prepareStackTrace`. There, `setModuleUrl(import.meta.url)` is not just good practice, it is the only way a task module can be found. ## How the build is produced `knitting/browser` ships as a single self-contained file. The copy hosted here is the unminified build, about 255 KB on disk and 53 KB over the wire. The published minified build is roughly 103 KB, or 35 KB gzipped. Node-only subsystems are swapped for stubs at bundle time by [`scripts/browser-stubs/plugin.ts`](https://github.com/mimiMonads/knitting/blob/main/scripts/browser-stubs/plugin.ts). Each stub does what the real module would have done in a page anyway: return nothing, answer false, or throw one of the messages above. That is why these errors name the feature you reached for instead of failing generically. --- # Work stealing URL: https://knittingdocs.vercel.app/guides/work-stealing/ How compatible multi-worker pools share one submit region so idle workers can claim waiting tasks. Knitting has two independent host-side scheduling features: - **Native work stealing** changes how requests reach workers. Compatible multi-worker pools publish work to one shared submit region, so workers can claim tasks as they become available. - **The host doorbell** changes how the host waits for responses. Supported runtimes can arm an asynchronous waiter instead of repeatedly polling the response mailbox. Both exist for the same reason. Threading in JavaScript usually buys its throughput with CPU, and Knitting is built to keep that bill small: a worker with nothing to claim parks instead of spinning, and a host with nothing to drain waits on the doorbell instead of polling. An idle thread should cost close to nothing. [Multi-threading](/guides/multi-threading/) shows what that looks like in measurements. Neither feature changes the task API. Tasks are still exported functions or `task()` definitions, and calls still look like `await pool.call.name(input)`. ## Quick start ```ts import { createPool, isMain, task } from "knitting"; export const transform = task({ f: (value) => value.toUpperCase(), }); if (isMain) { using pool = createPool({ threads: 4, host: { steal: true, doorbell: true, }, })({ transform }); console.log(await pool.call.transform("hello")); } ``` The explicit options are useful when documenting or benchmarking a topology. For compatible multi-worker pools, native stealing is selected automatically; the task code does not need to opt in from the worker side. ## Work stealing With private request lanes, the host chooses a worker and publishes the call to that worker's mailbox. A busy worker can therefore hold queued work while a different worker is idle. With native stealing, the pool uses a shared-submit/private-return layout: 1. The host publishes a request to one shared submit region. 2. A worker claims a region when it can make progress on the work there. 3. Each worker keeps a private return region for its responses. 4. The worker that claims a task owns its response; the host's pending registry settles the corresponding promise. This lets a worker that finishes early claim more available work without the host predicting which worker will be free next. Completion order remains task completion order, not submission order. ### Automatic selection Native stealing is selected automatically when the pool has multiple workers and its configuration is compatible with the shared-submit topology. A one-worker pool keeps its ordinary private lane because there is nothing to steal from. An explicit balancer, private-lane dispatcher, inliner, compiled worker, or an unsupported worker count can change that compatibility decision. Use `host.steal: false` to force private request lanes. Use `host.steal: true` only when the resulting topology is supported and is what you intend to measure. ### `stealRegionLanes` The submit region has 32 slots. Stealing divides those slots into regions, and one claiming handshake takes one whole region. `stealRegionLanes` is the region width: ```text number of regions = 32 / stealRegionLanes ``` Use a positive power of two. Knitting chooses the widest valid region by default, leaving a spare region alongside the worker claimants. Wider regions amortise arbitration and suit many cheap, similarly sized calls. Narrower regions expose more independent work and suit expensive or uneven calls. This is a throughput/load-balancing knob, not a correctness knob. ## The host doorbell The doorbell is a host completion mechanism. It is separate from the worker loop that waits for new requests. When the host has drained all visible responses, a polling dispatcher schedules another notification and eventually backs off with a timer. With the doorbell, the host arms an asynchronous wait on the response mailbox's shared signal. When a worker publishes a response, it rings that signal and the host schedules another drain. If the runtime cannot arm the wait, Knitting falls back to polling. Polling spends host CPU whether or not a response is waiting, and it spends it on the same thread that produces the work. The doorbell removes the empty checks: nothing is scheduled until a worker has something to hand back. The less loaded the pool, the larger the share of checks that were empty, which is why the doorbell matters most on a pool that is not saturated. The doorbell is enabled by default only when all of these are true: - the runtime is Node.js or Bun; - `Atomics.waitAsync` is available; - `host.doorbell` is not `false`; and - workers are not in another process. | Runtime or topology | Completion behavior | | --- | --- | | Node.js or Bun thread workers | Doorbell when `Atomics.waitAsync` is available | | Deno | Polling | | Process workers | Polling; another process cannot ring the host isolate's waiter | | Compiled workers | Host options are unsupported by the compiled-worker path | | Browser web workers | Polling; the browser doorbell is intentionally disabled | Setting `host.doorbell: true` cannot override a runtime limitation. Set it to `false` for an apples-to-apples polling benchmark or when the host is competing with workers for every CPU core. ## Tuning and benchmarking Start with the defaults. Tune `stealRegionLanes` when task durations are highly uneven, and compare `host.doorbell: false` with the supported default when host CPU or completion latency matters. Keep the topology constant while benchmarking. Do not compare a polling private-lane pool with a doorbell stealing pool and attribute the entire difference to the doorbell. These settings matter most when many calls compete for local CPU; they are unlikely to dominate a workload that mostly waits on external I/O. --- # Examples URL: https://knittingdocs.vercel.app/examples/intro_examples/ Worked examples you can copy into a real project. These examples are real patterns you can copy into your project -- not toy snippets. Most pages also point to an optional host-vs-worker benchmark script, but the example code is the main thing to copy. The Hono server page is the one that keeps the fuller performance story. ## Pick the right example for your use case **"I have a web server and want to keep it responsive under load"** Start with [Hono server routes](/examples/data_transforms/rendering_output/hono_server). It's the most realistic example -- a real HTTP server with SSR, JWT, and health check routes. **"I do SSR and want to offload rendering"** [React SSR](/examples/data_transforms/rendering_output/react_ssr) shows the basic pattern. [React SSR compression](/examples/data_transforms/rendering_output/react_ssr_compress) adds Brotli and tests where compression should live (worker vs host). **"I need to validate or parse lots of data"** [Schema validation](/examples/data_transforms/validation/schema_validate) (Zod), [JWT revalidation](/examples/data_transforms/validation/jwt_revalidation) (Web Crypto), or [Salt hashing](/examples/data_transforms/validation/salt_hashing) (PBKDF2) -- pick whichever is closest to your workload. **"I'm building an LLM-powered app"** [Prompt token budgeting](/examples/data_transforms/validation/prompt_token_budgeting) trims prompts to fit a token budget before they hit the API. **"I have a CPU-heavy computation I want to parallelize"** The math examples cover the spectrum: [Monte Carlo pi](/examples/maths/monte_pi) (embarrassingly parallel), [Physics loop](/examples/maths/physics_loop) (variable-work simulation), [Big prime](/examples/maths/big_prime) (long-running search), and [TSP](/examples/maths/tsp_gsa) (NP-hard optimization with parallel restarts). **"I just want to convert some text"** [Markdown to HTML](/examples/data_transforms/rendering_output/markdown_to_html) is the simplest rendering example -- string in, compressed string out. ## All examples ### Math and simulation - [Big prime](/examples/maths/big_prime) -- long-running BigInt search with Miller-Rabin - [Monte Carlo pi](/examples/maths/monte_pi) -- independent sampling and reduction - [Physics loop](/examples/maths/physics_loop) -- branch-heavy simulation with variable work per trial - [TSP (GSA)](/examples/maths/tsp_gsa) -- parallel heuristic restarts for NP-hard optimization ### Data transforms - [Data transforms overview](/examples/data_transforms/intro_data_transforms) -- grouped by validation and rendering - [Schema validation](/examples/data_transforms/validation/schema_validate) -- Zod validation on workers - [JWT revalidation](/examples/data_transforms/validation/jwt_revalidation) -- HMAC verify + renewal with Web Crypto - [Salt hashing](/examples/data_transforms/validation/salt_hashing) -- PBKDF2 password hashing (heaviest per-call workload) - [Prompt token budgeting](/examples/data_transforms/validation/prompt_token_budgeting) -- trim LLM prompts to a token budget - [React SSR](/examples/data_transforms/rendering_output/react_ssr) -- render React components on workers - [React SSR compression](/examples/data_transforms/rendering_output/react_ssr_compress) -- SSR + Brotli, worker vs host comparison - [Hono server routes](/examples/data_transforms/rendering_output/hono_server) -- real HTTP server with offloaded routes - [Markdown to HTML](/examples/data_transforms/rendering_output/markdown_to_html) -- simple markdown transform pipeline ## Patterns worth copying 1. **Keep worker tasks focused and deterministic.** One task, one job, predictable output. 2. **Batch calls and await them together** (`Promise.all`) to reduce scheduling overhead. 3. **Return compact summaries from workers** instead of large raw outputs. 4. **Validate on the host.** Recompute critical metrics when needed -- don't trust worker output blindly. 5. **Compare against a host-only baseline** before claiming speedups. 6. **Tune chunk sizes** based on throughput -- there's always a sweet spot between dispatch overhead and load balance. --- # Data transforms URL: https://knittingdocs.vercel.app/examples/data_transforms/intro_data_transforms/ Validation, rendering, and output examples — pick the one closest to your workload. These examples use real payload-carrying workloads on purpose. Primitive-only tasks are often near-instant, which hides coordination and transfer costs. Payload transforms are closer to production traffic, so the host-vs-worker comparisons are more honest. Treat the benchmark blocks in this section as supporting material, not the main event. Most pages are now organized around the runnable example first, with a benchmark command attached if you want to measure the same pattern on your machine. If you want one page that keeps the full performance discussion, start with [Hono server routes](/examples/data_transforms/rendering_output/hono_server). ## Which example should you start with? **If you're validating incoming data** -- API payloads, form submissions, webhook bodies -- start with [Schema validation](/examples/data_transforms/validation/schema_validate). It's the simplest validation pattern and uses Zod, which you're probably already familiar with. **If you're doing auth work** -- token verification, renewal, signature checking -- [JWT revalidation](/examples/data_transforms/validation/jwt_revalidation) shows how to offload HMAC crypto to workers using built-in Web Crypto (no external JWT library). **If you're hashing passwords** -- [Salt hashing](/examples/data_transforms/validation/salt_hashing) is the heaviest per-call example. PBKDF2 is intentionally slow, making it the ideal candidate for offloading. **If you're calling an LLM** -- [Prompt token budgeting](/examples/data_transforms/validation/prompt_token_budgeting) trims prompts to fit a token budget. The budgeting itself is model-agnostic. **If you're rendering HTML** -- [React SSR](/examples/data_transforms/rendering_output/react_ssr) for basic server rendering, [React SSR compression](/examples/data_transforms/rendering_output/react_ssr_compress) to also test Brotli placement, or [Markdown to HTML](/examples/data_transforms/rendering_output/markdown_to_html) for a simpler pipeline without React. **If you want the full picture** -- [Hono server routes](/examples/data_transforms/rendering_output/hono_server) is a real HTTP server that combines SSR, JWT, and health checks. It's the closest to a production setup. Parse incoming strings, validate shape and types, return typed output. [Schema validation](/examples/data_transforms/validation/schema_validate), [JWT revalidation](/examples/data_transforms/validation/jwt_revalidation), [Salt hashing](/examples/data_transforms/validation/salt_hashing), [Prompt token budgeting](/examples/data_transforms/validation/prompt_token_budgeting). Take validated input and produce final output formats. [Hono server routes](/examples/data_transforms/rendering_output/hono_server), [React SSR](/examples/data_transforms/rendering_output/react_ssr), [React SSR compression](/examples/data_transforms/rendering_output/react_ssr_compress), [Markdown to HTML](/examples/data_transforms/rendering_output/markdown_to_html). ## What all these examples have in common - Each includes a host-only baseline so performance comparisons are apples-to-apples. - Worker tasks return compact results (counts, compressed buffers, status flags) -- not large raw payloads. - The same code runs in both host and worker mode, so you can compare behavior without changing logic. --- # Big prime URL: https://knittingdocs.vercel.app/examples/maths/big_prime/ Search for large primes in parallel with a Miller-Rabin test. Scans through large integers (default: 1500-bit) looking for probable primes using Miller-Rabin, distributed across Knitting workers. This is a long-running compute pipeline -- it scans windows of candidates, prints progress, and keeps going until you stop it. ## How it works 1. The host picks a starting odd integer in the chosen bit-width. 2. It scans **windows** of 10,000,000 odd candidates, printing progress after each window. 3. Each window is split across threads -- threads scan disjoint candidate sequences (no overlap, no duplicated work). 4. Each worker runs `candidate -> Miller-Rabin -> next candidate -> ...` in a tight loop. 5. If any worker finds a probable prime, it reports it for that window. 6. The process runs continuously until Ctrl+C. The scanning uses an interleaved stride pattern: thread 0 checks `start + 0`, `start + 2T`, `start + 4T`, ..., thread 1 checks `start + 2`, `start + 2 + 2T`, ..., and so on. This guarantees full coverage with zero overlap. ## Run Expected output: ``` starting at 1500-bit odd integer (4 threads, 8 MR rounds) window 1: scanned 10,000,000 candidates in 4.2s -- no prime found window 2: scanned 10,000,000 candidates in 4.1s -- no prime found window 3: scanned 10,000,000 candidates in 4.3s -- OK probable prime found! digits: 452 hex prefix: 0xA3F7... ``` > Note: At 1500 bits, primes are sparse -- expect to scan through many windows. Lower `--bits` to 65 or 128 for faster results while testing. ## Code `run.ts` ```ts import { createPool, isMain } from "knitting"; import { scanForProbablePrime } from "./prime_scan.ts"; function intArg(name: string, fallback: number) { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const v = Number(process.argv[i + 1]); if (Number.isFinite(v) && v > 0) return Math.floor(v); } return fallback; } const THREADS = intArg("threads", 4); const BITS = intArg("bits", 1500); const WINDOW = intArg("window", 10_000_000); const CHUNK = intArg("chunk", 500_000); const ROUNDS = intArg("rounds", 10); function xorshift32(s: number): number { s |= 0; s ^= s << 13; s ^= s >>> 17; s ^= s << 5; return s | 0; } function makeRandomOdd(bits: number, seed: number): bigint { // Build a BigInt from 3x 32-bit chunks, mask to bits, set top bit, make odd. let s = seed | 0; let x = 0n; for (let k = 0; k < 3; k++) { s = xorshift32(s); x = (x << 32n) | BigInt(s >>> 0); } const mask = (1n << BigInt(bits)) - 1n; x &= mask; x |= 1n << BigInt(bits - 1); x |= 1n; return x; } const seedBase = (Date.now() | 0) ^ 0x9e3779b9; let windowStartOdd = makeRandomOdd(BITS, seedBase); const { call, shutdown } = createPool({ threads: THREADS, balancer: "firstIdle", })({ scanForProbablePrime }); let stopping = false; process.on("SIGINT", () => { if (stopping) return; stopping = true; console.log("\nCtrl+C received. Shutting down..."); shutdown(); process.exit(0); }); function splitCounts(total: number, parts: number): number[] { const base = Math.floor(total / parts); const rem = total % parts; const out = new Array(parts); for (let i = 0; i < parts; i++) out[i] = base + (i < rem ? 1 : 0); return out; } async function scanOneWindow(): Promise< { hit: string | null; tested: number } > { // We interleave odds across threads: thread i tests start+2i, start+2i+2T, ... const stepNum = 2 * THREADS; // Divide WINDOW across threads, and within each thread further divide into CHUNK-sized tasks. const perThread = splitCounts(WINDOW, THREADS); let bestHit: string | null = null; let tested = 0; // For each thread, we run sequential “subtasks” so each thread covers its share of WINDOW. // But all threads run in parallel each wave. const subTasksPerThread = perThread.map((c) => Math.ceil(c / CHUNK)); const maxSubs = Math.max(...subTasksPerThread); for (let sub = 0; sub < maxSubs; sub++) { const jobs: Promise<[number, string, number]>[] = []; for (let t = 0; t < THREADS; t++) { const threadTotal = perThread[t]; const startAt = sub * CHUNK; if (startAt >= threadTotal) continue; const count = Math.min(CHUNK, threadTotal - startAt); // offset in "odd steps": 2*t + 2*THREADS*startAt const offsetNum = 2 * t + stepNum * startAt; jobs.push( call.scanForProbablePrime([ windowStartOdd.toString(), count, stepNum, offsetNum, ROUNDS, ]), ); tested += count; } const results = await Promise.all(jobs); // If any job found a hit, keep the smallest hit (nice for consistency) for (const [found, primeStr] of results) { if (found) { if (bestHit === null) bestHit = primeStr; else { // compare as BigInt safely const a = BigInt(bestHit); const b = BigInt(primeStr); if (b < a) bestHit = primeStr; } } } } return { hit: bestHit, tested }; } async function main() { console.log("Prime hunt (probable primes via Miller–Rabin)"); console.log( "threads:", THREADS, "bits:", BITS, "window:", WINDOW.toLocaleString(), "chunk:", CHUNK.toLocaleString(), "rounds:", ROUNDS, ); console.log("start :", windowStartOdd.toString()); console.log("mode : infinite windows (Ctrl+C to stop)"); let windowsDone = 0; let totalTested = 0n; while (true) { const { hit, tested } = await scanOneWindow(); windowsDone++; totalTested += BigInt(tested); if (hit) { console.log( `[window ${windowsDone}] +${tested.toLocaleString()} tested (total ${totalTested.toString()}) | HIT: ${hit}`, ); } else { console.log( `[window ${windowsDone}] +${tested.toLocaleString()} tested (total ${totalTested.toString()}) | no hit (Ctrl+C to stop)`, ); } // Move start forward by WINDOW odd candidates (i.e., +2*WINDOW) windowStartOdd += 2n * BigInt(WINDOW); } } if (isMain) { main().finally(shutdown); } ``` `prime_scan.ts` ```ts import { task } from "knitting"; /** * Payload-safe: * - args: strings + numbers only * - return: numbers + strings only * This avoids any accidental BigInt/number mixing at the transport boundary. */ // args: [startOddStr, count, stepNum, offsetNum, rounds] export const scanForProbablePrime = task< [string, number, number, number, number], // ret: [found(0/1), primeStrOrEmpty, tested] [number, string, number] >({ f: ([startOddStr, count, stepNum, offsetNum, rounds]) => { // Convert once, keep everything BigInt inside. let x = BigInt(startOddStr) + BigInt(offsetNum); if ((x & 1n) === 0n) x += 1n; const step = BigInt(stepNum); const rds = rounds | 0; for (let i = 0; i < count; i++) { if (isProbablePrime(x, rds)) return [1, x.toString(), i + 1]; x += step; } return [0, "", count]; }, }); function modPow(base: bigint, exp: bigint, mod: bigint): bigint { let r = 1n; let b = base % mod; let e = exp; // must be bigint while (e > 0n) { if ((e & 1n) === 1n) r = (r * b) % mod; e >>= 1n; if (e) b = (b * b) % mod; } return r; } const small = [3n, 5n, 7n, 11n, 13n, 17n, 19n, 23n, 29n, 31n, 37n]; const bases = [2n, 325n, 9375n, 28178n, 450775n, 9780504n, 1795265022n]; function isProbablePrime(n: bigint, rounds: number): boolean { if (n < 2n) return false; if (n === 2n || n === 3n) return true; if ((n & 1n) === 0n) return false; // quick small-prime filter for (const p of small) { if (n === p) return true; if (n % p === 0n) return false; } // n-1 = d * 2^s let d = n - 1n; let s = 0; while ((d & 1n) === 0n) { d >>= 1n; s++; } // good practical bases (still "probable prime" for 65-bit+) const rds = rounds | 0; for (let i = 0; i < rds; i++) { const a = (bases[i % bases.length] % (n - 3n)) + 2n; // [2, n-2] let x = modPow(a, d, n); if (x === 1n || x === n - 1n) continue; let composite = true; for (let r = 1; r < s; r++) { x = (x * x) % n; if (x === n - 1n) { composite = false; break; } } if (composite) return false; } return true; } ``` ## The math behind it **Why primes are findable:** Around a number of size N, prime density is roughly `1/ln(N)`. Primes get rarer as numbers get bigger, but not impossibly rare -- even at 1500 bits, you'll find them. **Probable vs proven primes:** Miller-Rabin is a probabilistic test. With 8 rounds, the false-positive rate is vanishingly small (less than `4^-8`). In practice, this is what cryptographic libraries use for key generation. **Tuning `--rounds`:** More rounds = higher confidence, but more compute per candidate. 8 rounds is a solid default. You can go higher if you need cryptographic-grade confidence. ## CLI knobs - `--bits` -- bit-width of candidates (higher = harder, sparser primes) - `--threads` -- worker count - `--total` -- candidates per window (controls progress granularity) - `--chunk` -- candidates per task (controls scheduling granularity) - `--rounds` -- Miller-Rabin rounds (controls confidence vs speed) ## Things to try 1. Lower `--bits` to 65 and watch primes appear frequently. 2. Increase `--rounds` to 20 and measure the throughput impact. 3. Compare different `--chunk` values -- too small and dispatch overhead dominates, too large and load balance suffers. 4. Modify the task to search for twin primes (p and p+2 both prime) or Sophie Germain primes (p and 2p+1 both prime). --- # React SSR URL: https://knittingdocs.vercel.app/examples/data_transforms/rendering_output/react_ssr/ Render React components to HTML strings on workers, away from the request thread. Renders React components to HTML strings on workers using `renderToString`. If your server does SSR, this is probably the example closest to your real workload -- parse JSON input, normalize data, render a component, return HTML. ## How it works The host generates JSON payload strings. The host path calls `renderUserCardHost` directly (parse + normalize + `renderToString`). The worker path sends the same payloads through `createPool`. Byte totals are compared once to verify the host and worker produce identical output, then `mitata` benchmarks both paths. Three files: - `bench_react_ssr.ts` -- the benchmark itself - `render_user_card.tsx` -- the SSR component and worker task - `utils.ts` -- input payloads and normalization helpers ## Example payload and result Input: ```json { "id": "u42", "name": "Ari Lane", "handle": "@ari", "bio": "Building fast UIs.", "plan": "pro", "location": "Austin, TX", "joinedAt": "2026-01-18", "tags": ["react", "ssr", "workers"], "stats": { "posts": 42, "followers": 1200, "following": 180, "likes": 9800 }, "alerts": { "unread": 3, "lastLogin": "2026-01-18" } } ``` Minimal usage: ```ts const pool = createPool({ threads: 2 })({ renderUserCard }); const html = await pool.call.renderUserCard(payloadJson); ``` Result: ```html
...
``` ## Deno setup (TSX + npm) If you run this with Deno and see `Uncaught SyntaxError: Unexpected token '<'`, set a root `deno.json` so Deno transpiles TSX and resolves npm packages: ```json { "nodeModulesDir": "auto", "compilerOptions": { "jsx": "react-jsx", "jsxImportSource": "react" } } ``` ## Optional benchmark Expected output: ``` byte parity check: host=284,160 worker=284,160 OK match benchmark avg (ns) min ... max (ns) host 18,200 16,800 ... 24,500 knitting 9,400 8,600 ... 14,100 ``` > Note: The byte parity check runs once before benchmarking to confirm host and worker produce identical HTML. If it doesn't match, something is wrong with your setup. ## Code `bench_react_ssr.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { renderUserCard, renderUserCardHost } from "./render_user_card.tsx"; import { buildUserPayloads } from "./utils.ts"; const THREADS = 1; const REQUESTS = 2_000; async function main() { const payloads = buildUserPayloads(REQUESTS); const pool = createPool({ threads: THREADS, inliner: { batchSize: 6, }, })({ renderUserCard }); let sink = 0; try { const hostBytes = runHost(payloads); const knittingBytes = await runWorkers(pool.call.renderUserCard, payloads); if (hostBytes !== knittingBytes) { throw new Error("Host and worker HTML byte totals differ."); } console.log("React SSR benchmark (mitata)"); console.log("workload: parse + normalize + render to HTML"); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("threads:", THREADS, " + inliner"); boxplot(() => { summary(() => { bench(`host (${REQUESTS.toLocaleString()} req)`, () => { sink = runHost(payloads); }); bench( `knitting (${THREADS} thread(s) + main , ${REQUESTS.toLocaleString()} req)`, async () => { sink = await runWorkers(pool.call.renderUserCard, payloads); }, ); }); }); await run(); console.log("last html bytes:", sink.toLocaleString()); } finally { pool.shutdown(); } } function runHost(payloads: string[]): number { let htmlBytes = 0; for (let i = 0; i < payloads.length; i++) { const html = renderUserCardHost(payloads[i]!); htmlBytes += html.length; } return htmlBytes; } async function runWorkers( callRender: (payload: string) => Promise, payloads: string[], ): Promise { const jobs: Promise[] = []; for (let i = 0; i < payloads.length; i++) { jobs.push(callRender(payloads[i]!)); } const results = await Promise.all(jobs); let htmlBytes = 0; for (let i = 0; i < results.length; i++) { htmlBytes += results[i]!.length; } return htmlBytes; } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `render_user_card.tsx` ```tsx import React from "react"; import { renderToString } from "react-dom/server"; import { task } from "knitting"; import { clamp, engagementScore, formatJoinDate, initials, levelForScore, type NormalizedUser, normalizeUser, } from "./utils.ts"; function Stat({ label, value }: { label: string; value: number }) { return (
{value.toLocaleString()} {label}
); } function Badge({ plan }: { plan: "free" | "pro" }) { const text = plan === "pro" ? "PRO" : "FREE"; return {text}; } function UserCard({ user }: { user: NormalizedUser }) { const score = engagementScore(user.stats); const level = levelForScore(score); const joined = formatJoinDate(user.joinedAt); const profileCompleteness = (user.bio ? 30 : 0) + (user.location ? 20 : 0) + (user.tags.length ? 20 : 0) + (user.handle ? 10 : 0) + (user.stats.posts ? 20 : 0); const completeness = clamp(profileCompleteness, 10, 100); const topTags = user.tags.slice(0, 6); const achievementBadges = [ user.stats.followers >= 1000 ? "1k+ followers" : "", user.stats.likes >= 5000 ? "5k+ likes" : "", user.stats.posts >= 50 ? "50+ posts" : "", ].filter(Boolean); return (

{user.name}

{user.handle}
📍 {user.location} • Joined {joined}
Level: {level} Score {score.toLocaleString()}
{user.alerts.unread} unread Last login {user.alerts.lastLogin}

Bio

{user.bio ?

{user.bio}

:

No bio yet.

}

Profile completeness

{completeness}% complete

Highlights

    {achievementBadges.length > 0 ? ( achievementBadges.map((badge) => (
  • {badge}
  • )) ) :
  • Getting started
  • }

Interests

    {topTags.length > 0 ? ( topTags.map((tag) => (
  • {tag}
  • )) ) :
  • No tags
  • }
); } export function renderUserCardHost(payloadJson: string): string { const parsed = JSON.parse(payloadJson) as unknown; const user = normalizeUser(parsed); const html = renderToString(); return html; } export const renderUserCard = task({ f: renderUserCardHost, }); ``` `utils.ts` ```ts export type UserStats = { posts: number; followers: number; following: number; likes: number; }; export type UserAlerts = { unread: number; lastLogin: string; }; export type UserPayload = { id?: string; name?: string; handle?: string; bio?: string; plan?: "free" | "pro"; location?: string; joinedAt?: string; tags?: string[]; stats?: Partial; alerts?: Partial; }; export type NormalizedUser = Required & { stats: UserStats; alerts: UserAlerts; }; const TAGS = [ "react", "ssr", "typescript", "performance", "parallel", "workers", "ui", "web", ]; const LOCATIONS = ["Austin, TX", "Seattle, WA", "Brooklyn, NY", "Denver, CO"]; function pickFrom(arr: T[], index: number): T { return arr[index % arr.length]!; } function toNumber(value: unknown, fallback: number): number { return Number.isFinite(value) ? Number(value) : fallback; } function toStringArray(value: unknown): string[] { if (!Array.isArray(value)) return []; return value.filter((item): item is string => typeof item === "string"); } export function makeUserPayloadJson(i: number): string { const short = i.toString(36); return JSON.stringify({ id: `u${short}`, name: `User ${short.toUpperCase()}`, handle: `@${short}`, bio: `Building fast UIs. Coffee + TypeScript. (${short})`, plan: i % 7 === 0 ? "pro" : "free", location: pickFrom(LOCATIONS, i), joinedAt: `202${(i % 4) + 2}-0${(i % 8) + 1}-1${i % 9}`, tags: [ pickFrom(TAGS, i), pickFrom(TAGS, i + 1), pickFrom(TAGS, i + 2), pickFrom(TAGS, i + 3), ], stats: { posts: (i % 120) + 1, followers: (i * 13) % 50_000, following: (i * 7) % 5_000, likes: (i * 31) % 250_000, }, alerts: { unread: i % 25, lastLogin: `2026-0${(i % 8) + 1}-0${(i % 9) + 1}`, }, }); } export function buildUserPayloads(count: number): string[] { const payloads = new Array(count); for (let i = 0; i < count; i++) payloads[i] = makeUserPayloadJson(i); return payloads; } export function normalizeUser(payload: unknown): NormalizedUser { const obj = (payload ?? {}) as Record; const id = typeof obj.id === "string" ? obj.id : "unknown"; const name = typeof obj.name === "string" ? obj.name : "Anonymous"; const handle = typeof obj.handle === "string" ? obj.handle : `@${id}`; const bio = typeof obj.bio === "string" ? obj.bio : ""; const plan = obj.plan === "pro" ? "pro" : "free"; const location = typeof obj.location === "string" ? obj.location : "Unknown"; const joinedAt = typeof obj.joinedAt === "string" ? obj.joinedAt : "2024-05-01"; const statsRaw = (obj.stats ?? {}) as Record; const alertsRaw = (obj.alerts ?? {}) as Record; const stats: UserStats = { posts: toNumber(statsRaw.posts, 0), followers: toNumber(statsRaw.followers, 0), following: toNumber(statsRaw.following, 0), likes: toNumber(statsRaw.likes, 0), }; const alerts: UserAlerts = { unread: toNumber(alertsRaw.unread, 0), lastLogin: typeof alertsRaw.lastLogin === "string" ? alertsRaw.lastLogin : "2026-01-18", }; const tags = toStringArray(obj.tags); return { id, name, handle, bio, plan, location, joinedAt, tags, stats, alerts, }; } export function formatJoinDate(value: string): string { const date = new Date(value); if (Number.isNaN(date.getTime())) return value; return date.toLocaleDateString("en-US", { month: "short", year: "numeric" }); } export function initials(name: string): string { const parts = name.trim().split(/\s+/); const first = parts[0]?.[0] ?? "U"; const second = parts.length > 1 ? parts[1]?.[0] ?? "" : ""; return (first + second).toUpperCase(); } export function clamp(value: number, min: number, max: number): number { return Math.min(max, Math.max(min, value)); } export function engagementScore(stats: UserStats): number { return Math.round( stats.posts * 2 + stats.likes * 0.05 + stats.followers * 0.4 + stats.following * 0.1, ); } export function levelForScore(score: number): string { if (score >= 5000) return "Legend"; if (score >= 2500) return "Elite"; if (score >= 1000) return "Rising"; if (score >= 300) return "Active"; return "New"; } ``` ## Why SSR is a great fit for workers `renderToString` is synchronous and CPU-bound -- it walks your component tree and builds an HTML string. On a request-per-request basis it's not usually slow, but under load it blocks the event loop and everything else waits. Moving SSR to a worker pool means your main thread stays free to accept requests and handle I/O while rendering happens in parallel. The Hono server example shows what this looks like in a real HTTP server context. --- # Schema validate URL: https://knittingdocs.vercel.app/examples/data_transforms/validation/schema_validate/ Validate JSON payloads against a Zod schema on workers. Parses JSON strings, validates them against a Zod schema, and returns typed results -- either on the main thread or through a Knitting worker pool. If your app already does `JSON.parse` + validation and you want to see what offloading looks like, start here. ## How it works The host generates JSON strings (some valid, some intentionally broken). Each job parses the string, runs it through `UserSchema.safeParse`, and returns `{ ok: true, value }` or `{ ok: false, issues }`. The host aggregates counts and prints sample failures. Three files: - `schema_knitting.ts` -- runs parse+validate in host and Knitting modes - `utils.ts` -- schema logic, payload builders, task exports - `bench_schema_validate.ts` -- host-vs-worker benchmark with `mitata` ## Example payloads ```json // valid { "id": "u_42", "email": "ari@knitting.dev", "displayName": "Ari Lane", "age": 29, "roles": ["admin"], "marketingOptIn": true } // invalid { "id": "u_42", "email": "ari@knitting.dev", "displayName": "x", "age": "unknown", "roles": ["owner"] } ``` The first returns `{ ok: true, value }`. The second returns `{ ok: false, issues }`, with messages like `displayName: String must contain at least 2 character(s)` or `age: Expected number, received string`. ## Run You should see output like: ``` -- host mode -- valid: 800 invalid: 200 sample issue: ["Expected string, received number"] -- knitting mode (2 threads) -- valid: 800 invalid: 200 sample issue: ["Expected string, received number"] ``` ## Optional benchmark The benchmark compares `JSON.parse + safeParse` via direct function imports (`host`) against the same logic dispatched through a worker pool (`knitting`). Batch calls keep per-dispatch overhead predictable. Expected output: ``` benchmark avg (ns) min ... max (ns) host 12,340 11,200 ... 18,400 knitting 6,890 6,100 ... 11,200 ``` > Note: Exact numbers depend on your hardware and batch size. The shape of the result -- workers winning on batched validation -- is what matters. ## Code `schema_knitting.ts` ```ts import { createPool, isMain } from "knitting"; import { buildPayloads, parseAndValidate, parseAndValidateHost, type ParseValidateResult, } from "./utils.ts"; const THREADS = 2; const REQUESTS = 20_000; const INVALID_PERCENT = 15; type Summary = { valid: number; invalid: number; sampleIssues: string[]; }; function summarize(results: ParseValidateResult[]): Summary { let valid = 0; let invalid = 0; const sampleIssues: string[] = []; for (let i = 0; i < results.length; i++) { const result = results[i]!; if (result.ok) { valid++; continue; } invalid++; if (sampleIssues.length < 3 && result.issues.length > 0) { sampleIssues.push(result.issues[0]!); } } return { valid, invalid, sampleIssues }; } function runHost(payloads: string[]): Summary { const results = payloads.map((payload) => parseAndValidateHost(payload)); return summarize(results); } async function runWorkers(payloads: string[]): Promise { const pool = createPool({ threads: THREADS })({ parseAndValidate }); try { const jobs: Promise[] = []; for (let i = 0; i < payloads.length; i++) { jobs.push(pool.call.parseAndValidate(payloads[i]!)); } const results = await Promise.all(jobs); return summarize(results); } finally { pool.shutdown(); } } function printSummary(mode: string, summary: Summary, ms: number): void { const secs = Math.max(1e-9, ms / 1000); const rps = REQUESTS / secs; console.log(mode); console.log("requests :", REQUESTS.toLocaleString()); console.log("invalidRate :", `${INVALID_PERCENT}%`); console.log("valid :", summary.valid.toLocaleString()); console.log("invalid :", summary.invalid.toLocaleString()); console.log("took :", `${ms.toFixed(2)} ms`); console.log("throughput :", `${rps.toFixed(0)} req/s`); if (summary.sampleIssues.length > 0) { console.log("sampleIssues:", summary.sampleIssues.join(" | ")); } } async function main() { const payloads = buildPayloads(REQUESTS, INVALID_PERCENT); const hostStart = performance.now(); const hostSummary = runHost(payloads); const hostMs = performance.now() - hostStart; const workerStart = performance.now(); const workerSummary = await runWorkers(payloads); const workerMs = performance.now() - workerStart; const uplift = (hostMs / Math.max(1e-9, workerMs) - 1) * 100; console.log("JSON parse + schema validation"); console.log(`threads: ${THREADS}`); console.log(""); printSummary("host", hostSummary, hostMs); console.log(""); printSummary("knitting", workerSummary, workerMs); console.log(""); console.log(`uplift: ${uplift.toFixed(1)}%`); } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `bench_schema_validate.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { buildPayloads, makeBatches, mergeValidationSummary, parseAndValidateBatchFast, parseAndValidateBatchFastHost, sameValidationSummary, type ValidationSummary, } from "./utils.ts"; const THREADS = 2; const REQUESTS = 20_000; const INVALID_PERCENT = 15; const BATCH = 64; async function main() { const payloads = buildPayloads(REQUESTS, INVALID_PERCENT); const payloadBatches = makeBatches(payloads, BATCH); const pool = createPool({ threads: THREADS, })({ parseAndValidateBatchFast }); let sink = 0; try { const hostCheck = runHostBatches(payloadBatches); const workerCheck = await runWorkerBatches( pool.call.parseAndValidateBatchFast, payloadBatches, ); if (!sameValidationSummary(hostCheck, workerCheck)) { throw new Error("Host and worker validation counts differ."); } console.log("Schema validation benchmark (mitata)"); console.log("workload: JSON.parse + UserSchema.safeParse"); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("invalid rate:", `${INVALID_PERCENT}%`); console.log("batch size:", BATCH); console.log("threads:", THREADS); boxplot(() => { summary(() => { bench(`host (${REQUESTS.toLocaleString()} req, batch ${BATCH})`, () => { const totals = runHostBatches(payloadBatches); sink = totals.valid; }); bench( `knitting (${THREADS} thread${ THREADS === 1 ? "" : "s" }, ${REQUESTS.toLocaleString()} req, batch ${BATCH})`, async () => { const totals = await runWorkerBatches( pool.call.parseAndValidateBatchFast, payloadBatches, ); sink = totals.valid; }, ); }); }); await run(); console.log("last valid count:", sink.toLocaleString()); } finally { pool.shutdown(); } } function runHostBatches(payloadBatches: string[][]): ValidationSummary { let totals: ValidationSummary = { valid: 0, invalid: 0 }; for (let i = 0; i < payloadBatches.length; i++) { totals = mergeValidationSummary( totals, parseAndValidateBatchFastHost(payloadBatches[i]!), ); } return totals; } async function runWorkerBatches( callBatch: (payloads: string[]) => Promise, payloadBatches: string[][], ): Promise { const jobs: Promise[] = []; for (let i = 0; i < payloadBatches.length; i++) { jobs.push(callBatch(payloadBatches[i]!)); } const results = await Promise.all(jobs); let totals: ValidationSummary = { valid: 0, invalid: 0 }; for (let i = 0; i < results.length; i++) { totals = mergeValidationSummary(totals, results[i]!); } return totals; } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `utils.ts` ```ts import { task } from "knitting"; import { z } from "zod"; const UserSchema = z.object({ id: z.string().min(1), email: z.string().email(), displayName: z.string().min(2).max(80), age: z.number().int().min(13).max(120), roles: z.array(z.enum(["user", "admin", "moderator"])).default(["user"]), marketingOptIn: z.boolean().default(false), }); export type User = z.infer; export type ParseValidateResult = | { ok: true; value: User } | { ok: false; issues: string[] }; export type ValidationSummary = { valid: number; invalid: number; }; export function makeValidPayload(i: number): string { const short = i.toString(36); const role = i % 9 === 0 ? "admin" : "user"; return JSON.stringify({ id: `u_${short}`, email: `${short}@knitting.dev`, displayName: `User ${short.toUpperCase()}`, age: 18 + (i % 60), roles: [role], marketingOptIn: i % 2 === 0, }); } export function makePayload(i: number, invalidPercent: number): string { if (i % 100 >= invalidPercent) return makeValidPayload(i); switch (i % 4) { case 0: return '{"id":"broken"'; case 1: return JSON.stringify({ id: `u_${i}`, displayName: `User ${i}`, age: 33, roles: ["user"], marketingOptIn: true, }); case 2: return JSON.stringify({ id: `u_${i}`, email: `u_${i}@knitting.dev`, displayName: "x", age: "unknown", roles: ["user"], }); default: return JSON.stringify({ id: `u_${i}`, email: `u_${i}@knitting.dev`, displayName: `User ${i}`, age: 31, roles: ["owner"], }); } } export function buildPayloads(count: number, invalidPercent: number): string[] { const cappedInvalid = Math.max(0, Math.min(95, Math.floor(invalidPercent))); const size = Math.max(0, Math.floor(count)); const payloads = new Array(size); for (let i = 0; i < size; i++) payloads[i] = makePayload(i, cappedInvalid); return payloads; } export function makeBatches(values: T[], batchSize: number): T[][] { const size = Math.max(1, Math.floor(batchSize)); const batches: T[][] = []; for (let i = 0; i < values.length; i += size) { batches.push(values.slice(i, i + size)); } return batches; } export function mergeValidationSummary( a: ValidationSummary, b: ValidationSummary, ): ValidationSummary { return { valid: a.valid + b.valid, invalid: a.invalid + b.invalid, }; } export function sameValidationSummary( a: ValidationSummary, b: ValidationSummary, ): boolean { return a.valid === b.valid && a.invalid === b.invalid; } function toIssues(error: z.ZodError): string[] { return error.issues.map((issue) => { const path = issue.path.length > 0 ? issue.path.join(".") : "payload"; return `${path}: ${issue.message}`; }); } export function parseAndValidateHost(rawPayload: string): ParseValidateResult { let parsed: unknown; try { parsed = JSON.parse(rawPayload) as unknown; } catch { return { ok: false, issues: ["payload: invalid JSON string"] }; } const result = UserSchema.safeParse(parsed); if (!result.success) { return { ok: false, issues: toIssues(result.error) }; } return { ok: true, value: result.data }; } export const parseAndValidate = task({ f: parseAndValidateHost, }); export function parseAndValidateFastHost(rawPayload: string): boolean { let parsed: unknown; try { parsed = JSON.parse(rawPayload) as unknown; } catch { return false; } return UserSchema.safeParse(parsed).success; } export function parseAndValidateBatchFastHost( rawPayloads: string[], ): ValidationSummary { let valid = 0; let invalid = 0; for (let i = 0; i < rawPayloads.length; i++) { if (parseAndValidateFastHost(rawPayloads[i]!)) { valid++; } else { invalid++; } } return { valid, invalid }; } export const parseAndValidateBatchFast = task({ f: parseAndValidateBatchFastHost, }); ``` ## When to use this pattern Schema validation is a textbook case for worker offloading: each call is independent, the input/output is small, and Zod's internals are CPU-bound (type checking, error formatting). If you're validating hundreds of payloads per second -- API gateway, webhook ingestion, form processing -- batching them through a pool can free your main thread without changing any validation logic. --- # JWT revalidation URL: https://knittingdocs.vercel.app/examples/data_transforms/validation/jwt_revalidation/ Verify JWTs and reissue them with Web Crypto, without an external JWT library. Verifies JWT signatures with HMAC SHA-256, checks expiry, and optionally reissues tokens -- all using built-in Web Crypto, no external JWT package. This example shows what it looks like to offload auth-related crypto work to a worker pool. ## How it works The host builds a batch of JWTs (mix of valid and expired). Each job verifies the signature, checks the renewal window (`renewWindowSec`, `renewUntil`), and either returns the validated claims or issues a fresh token. The result comes back as a stringified JSON response -- keeping structured-clone overhead low. Three files: - `jwt_knitting.ts` -- runs revalidation in host and Knitting modes - `utils.ts` -- token verification, renewal logic, task exports - `bench_jwt_revalidation.ts` -- host-vs-worker benchmark with `mitata` ## Example request and response Input to the task is a JSON string like this: ```json { "token": "", "nowSec": 1767225600, "ttlSec": 180, "renewWindowSec": 30 } ``` The task returns a JSON string. A successful non-renewed response looks like: ```json { "ok": true, "renewed": false, "token": "", "sub": "user_42", "sid": "session_42", "exp": 1767225750, "canRenew": true } ``` If the token is near expiry but still renewable, the same shape comes back with `renewed: true` and a replacement `token`. Uses built-in Web Crypto APIs -- no extra JWT package required. ## Run Expected output: ``` -- host mode -- verified: 900 renewed: 80 rejected: 20 -- knitting mode (2 threads) -- verified: 900 renewed: 80 rejected: 20 ``` ## Optional benchmark Compares verify + optional renewal + `JSON.stringify` via direct imports (`host`) against the same workload through worker task calls (`knitting`). Batched dispatch keeps the comparison stable. > Note: Both paths use the same logic. The only variable is whether execution happens on the main thread or in a worker. ## Code `jwt_knitting.ts` ```ts import { createPool, isMain } from "knitting"; import { buildDemoRevalidateRequests, type RenewalSummary, revalidateToken, revalidateTokenHost, summarizeJsonResponses, } from "./utils.ts"; const THREADS = 2; const REQUESTS = 25_000; const INVALID_PERCENT = 10; async function runHost(rawRequests: string[]): Promise { const outputs = new Array(rawRequests.length); for (let i = 0; i < rawRequests.length; i++) { outputs[i] = await revalidateTokenHost(rawRequests[i]!); } return summarizeJsonResponses(outputs); } async function runWorkers(rawRequests: string[]): Promise { const pool = createPool({ threads: THREADS })({ revalidateToken }); try { const jobs: Promise[] = []; for (let i = 0; i < rawRequests.length; i++) { jobs.push(pool.call.revalidateToken(rawRequests[i]!)); } const outputs = await Promise.all(jobs); return summarizeJsonResponses(outputs); } finally { pool.shutdown(); } } function printSummary(mode: string, totals: RenewalSummary, ms: number): void { const seconds = Math.max(1e-9, ms / 1000); const rps = REQUESTS / seconds; console.log(mode); console.log("requests :", REQUESTS.toLocaleString()); console.log("invalid rate :", `${INVALID_PERCENT}%`); console.log("accepted :", totals.ok.toLocaleString()); console.log("renewed :", totals.renewed.toLocaleString()); console.log("rejected :", totals.rejected.toLocaleString()); console.log("output bytes :", totals.outputBytes.toLocaleString()); console.log("took :", `${ms.toFixed(2)} ms`); console.log("throughput :", `${rps.toFixed(0)} req/s`); } async function main() { const rawRequests = await buildDemoRevalidateRequests({ count: REQUESTS, invalidPercent: INVALID_PERCENT, }); const hostStart = performance.now(); const hostTotals = await runHost(rawRequests); const hostMs = performance.now() - hostStart; const workerStart = performance.now(); const workerTotals = await runWorkers(rawRequests); const workerMs = performance.now() - workerStart; const uplift = (hostMs / Math.max(1e-9, workerMs) - 1) * 100; console.log("JWT token revalidation"); console.log(`threads: ${THREADS}`); console.log(""); printSummary("host", hostTotals, hostMs); console.log(""); printSummary("knitting", workerTotals, workerMs); console.log(""); console.log(`uplift: ${uplift.toFixed(1)}%`); } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `bench_jwt_revalidation.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { buildDemoRevalidateRequests, makeBatches, mergeRenewalSummary, type RenewalSummary, revalidateTokenBatchFast, revalidateTokenBatchFastHost, sameRenewalSummary, } from "./utils.ts"; const THREADS = 2; const REQUESTS = 25_000; const INVALID_PERCENT = 10; const BATCH = 64; async function runHostBatches(rawBatches: string[][]): Promise { let totals: RenewalSummary = { ok: 0, renewed: 0, rejected: 0, outputBytes: 0, }; for (let i = 0; i < rawBatches.length; i++) { totals = mergeRenewalSummary( totals, await revalidateTokenBatchFastHost(rawBatches[i]!), ); } return totals; } async function runWorkerBatches( callBatch: (rawRequests: string[]) => Promise, rawBatches: string[][], ): Promise { const jobs: Promise[] = []; for (let i = 0; i < rawBatches.length; i++) { jobs.push(callBatch(rawBatches[i]!)); } const results = await Promise.all(jobs); let totals: RenewalSummary = { ok: 0, renewed: 0, rejected: 0, outputBytes: 0, }; for (let i = 0; i < results.length; i++) { totals = mergeRenewalSummary(totals, results[i]!); } return totals; } async function main() { const rawRequests = await buildDemoRevalidateRequests({ count: REQUESTS, invalidPercent: INVALID_PERCENT, }); const rawBatches = makeBatches(rawRequests, BATCH); const pool = createPool({ threads: THREADS })({ revalidateTokenBatchFast }); let sink = 0; try { const hostCheck = await runHostBatches(rawBatches); const workerCheck = await runWorkerBatches( pool.call.revalidateTokenBatchFast, rawBatches, ); if (!sameRenewalSummary(hostCheck, workerCheck)) { throw new Error("Host and worker JWT summaries differ."); } console.log("JWT revalidation benchmark (mitata)"); console.log( "workload: verify token -> renew when allowed -> JSON.stringify", ); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("invalid rate:", `${INVALID_PERCENT}%`); console.log("batch size:", BATCH); console.log("threads:", THREADS); boxplot(() => { summary(() => { bench( `host (${REQUESTS.toLocaleString()} req, batch ${BATCH})`, async () => { const totals = await runHostBatches(rawBatches); sink = totals.outputBytes; }, ); bench( `knitting (${THREADS} thread${ THREADS === 1 ? "" : "s" }, ${REQUESTS.toLocaleString()} req, batch ${BATCH})`, async () => { const totals = await runWorkerBatches( pool.call.revalidateTokenBatchFast, rawBatches, ); sink = totals.outputBytes; }, ); }); }); await run(); console.log("last output bytes:", sink.toLocaleString()); } finally { pool.shutdown(); } } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `utils.ts` ```ts import { task } from "knitting"; export const DEMO_JWT_SECRET = "knitting-docs-demo-hs256-secret"; const DEFAULT_TTL_SEC = 180; const DEFAULT_RENEW_WINDOW_SEC = 30; const MIN_TTL_SEC = 10; const MAX_TTL_SEC = 86_400; const MAX_RENEW_WINDOW_SEC = 3_600; const encoder = new TextEncoder(); const decoder = new TextDecoder(); const keyCache = new Map>(); type JwtHeader = { alg: "HS256"; typ: "JWT"; }; export type JwtClaims = { sub: string; sid: string; scope: string[]; iat: number; exp: number; renewUntil: number; }; export type RevalidateRequest = { token: string; nowSec?: number; ttlSec?: number; renewWindowSec?: number; }; export type RevalidateResponse = | { ok: true; renewed: boolean; token: string; sub: string; sid: string; exp: number; canRenew: boolean; } | { ok: false; renewed: false; reason: string; }; export type RenewalSummary = { ok: number; renewed: number; rejected: number; outputBytes: number; }; export type DemoRequestOptions = { count: number; nowSec?: number; invalidPercent?: number; ttlSec?: number; renewWindowSec?: number; }; export function makeBatches(values: T[], batchSize: number): T[][] { const size = Math.max(1, Math.floor(batchSize)); const batches: T[][] = []; for (let i = 0; i < values.length; i += size) { batches.push(values.slice(i, i + size)); } return batches; } export function mergeRenewalSummary( a: RenewalSummary, b: RenewalSummary, ): RenewalSummary { return { ok: a.ok + b.ok, renewed: a.renewed + b.renewed, rejected: a.rejected + b.rejected, outputBytes: a.outputBytes + b.outputBytes, }; } export function sameRenewalSummary( a: RenewalSummary, b: RenewalSummary, ): boolean { return a.ok === b.ok && a.renewed === b.renewed && a.rejected === b.rejected && a.outputBytes === b.outputBytes; } function clampInt( value: unknown, fallback: number, min: number, max: number, ): number { const numberValue = Number(value); if (!Number.isFinite(numberValue)) return fallback; const integer = Math.floor(numberValue); if (integer < min) return min; if (integer > max) return max; return integer; } function isRecord(value: unknown): value is Record { return typeof value === "object" && value !== null && !Array.isArray(value); } function bytesToBase64Url(bytes: Uint8Array): string { let binary = ""; for (let i = 0; i < bytes.length; i++) { binary += String.fromCharCode(bytes[i]!); } const base64 = btoa(binary); return base64.replace(/\+/g, "-").replace(/\//g, "_").replace(/=+$/g, ""); } function base64UrlToBytes(base64url: string): Uint8Array { const normalized = base64url.replace(/-/g, "+").replace(/_/g, "/"); const pad = (4 - (normalized.length % 4)) % 4; const padded = normalized + "=".repeat(pad); const binary = atob(padded); const out = new Uint8Array(binary.length); for (let i = 0; i < binary.length; i++) { out[i] = binary.charCodeAt(i); } return out; } function safeParseJson(raw: string): unknown | null { try { return JSON.parse(raw) as unknown; } catch { return null; } } function fixedTimeEqual(a: Uint8Array, b: Uint8Array): boolean { if (a.length !== b.length) return false; let diff = 0; for (let i = 0; i < a.length; i++) { diff |= a[i]! ^ b[i]!; } return diff === 0; } function parseClaims(value: unknown): JwtClaims | null { if (!isRecord(value)) return null; const { sub, sid, scope, iat, exp, renewUntil } = value; if (typeof sub !== "string" || sub.length === 0) return null; if (typeof sid !== "string" || sid.length === 0) return null; if (!Array.isArray(scope) || scope.length === 0) return null; if (!scope.every((item) => typeof item === "string" && item.length > 0)) { return null; } if ( !Number.isInteger(iat) || !Number.isInteger(exp) || !Number.isInteger(renewUntil) ) { return null; } if (exp <= iat) return null; if (renewUntil < exp) return null; return { sub, sid, scope: [...scope], iat, exp, renewUntil }; } async function getHmacKey(secret: string): Promise { let keyPromise = keyCache.get(secret); if (!keyPromise) { keyPromise = crypto.subtle.importKey( "raw", encoder.encode(secret), { name: "HMAC", hash: "SHA-256" }, false, ["sign"], ); keyCache.set(secret, keyPromise); } return keyPromise; } async function signInput( signingInput: string, secret: string, ): Promise { const key = await getHmacKey(secret); const signature = await crypto.subtle.sign( "HMAC", key, encoder.encode(signingInput), ); return bytesToBase64Url(new Uint8Array(signature)); } async function verifyToken( token: string, secret: string, ): Promise { try { const parts = token.split("."); if (parts.length !== 3) return null; const [headerPart, payloadPart, signaturePart] = parts; if (!headerPart || !payloadPart || !signaturePart) return null; const headerRaw = safeParseJson( decoder.decode(base64UrlToBytes(headerPart)), ); if (!isRecord(headerRaw)) return null; if (headerRaw.alg !== "HS256") return null; if (headerRaw.typ !== "JWT") return null; const expectedSignature = await signInput( `${headerPart}.${payloadPart}`, secret, ); const expectedBytes = base64UrlToBytes(expectedSignature); const providedBytes = base64UrlToBytes(signaturePart); if (!fixedTimeEqual(expectedBytes, providedBytes)) return null; const payloadRaw = safeParseJson( decoder.decode(base64UrlToBytes(payloadPart)), ); return parseClaims(payloadRaw); } catch { return null; } } function parseRevalidateRequest(rawRequest: string): RevalidateRequest | null { const parsed = safeParseJson(rawRequest); if (!isRecord(parsed)) return null; if (typeof parsed.token !== "string" || parsed.token.length === 0) { return null; } return { token: parsed.token, nowSec: parsed.nowSec as number | undefined, ttlSec: parsed.ttlSec as number | undefined, renewWindowSec: parsed.renewWindowSec as number | undefined, }; } function shouldRenewToken( claims: JwtClaims, nowSec: number, renewWindowSec: number, ): boolean { if (nowSec > claims.renewUntil) return false; return nowSec >= claims.exp - renewWindowSec; } function responseError(reason: string): RevalidateResponse { return { ok: false, renewed: false, reason }; } export async function issueTokenHost( claims: JwtClaims, secret = DEMO_JWT_SECRET, ): Promise { const header: JwtHeader = { alg: "HS256", typ: "JWT" }; const headerPart = bytesToBase64Url(encoder.encode(JSON.stringify(header))); const payloadPart = bytesToBase64Url(encoder.encode(JSON.stringify(claims))); const signingInput = `${headerPart}.${payloadPart}`; const signaturePart = await signInput(signingInput, secret); return `${signingInput}.${signaturePart}`; } export async function revalidateTokenObjectHost( rawRequest: string, secret = DEMO_JWT_SECRET, ): Promise { const request = parseRevalidateRequest(rawRequest); if (!request) { return responseError("payload: expected JSON { token, nowSec? }"); } const nowSec = clampInt( request.nowSec, Math.floor(Date.now() / 1000), 1, 2_147_483_647, ); const ttlSec = clampInt( request.ttlSec, DEFAULT_TTL_SEC, MIN_TTL_SEC, MAX_TTL_SEC, ); const renewWindowSec = clampInt( request.renewWindowSec, DEFAULT_RENEW_WINDOW_SEC, 0, MAX_RENEW_WINDOW_SEC, ); const claims = await verifyToken(request.token, secret); if (!claims) { return responseError("token: invalid signature, claims, or format"); } const renewable = shouldRenewToken(claims, nowSec, renewWindowSec); if (renewable) { const renewedExp = Math.min(nowSec + ttlSec, claims.renewUntil); if (renewedExp > nowSec) { const renewedClaims: JwtClaims = { ...claims, iat: nowSec, exp: renewedExp, }; const renewedToken = await issueTokenHost(renewedClaims, secret); return { ok: true, renewed: true, token: renewedToken, sub: claims.sub, sid: claims.sid, exp: renewedExp, canRenew: nowSec < claims.renewUntil, }; } } if (nowSec > claims.exp) { return responseError("token: expired and outside renewal policy"); } return { ok: true, renewed: false, token: request.token, sub: claims.sub, sid: claims.sid, exp: claims.exp, canRenew: nowSec <= claims.renewUntil, }; } export async function revalidateTokenHost(rawRequest: string): Promise { const response = await revalidateTokenObjectHost(rawRequest); return JSON.stringify(response); } export const revalidateToken = task({ f: revalidateTokenHost, }); function addSummary( totals: RenewalSummary, response: RevalidateResponse, outputBytes: number, ): RenewalSummary { const next: RenewalSummary = { ok: totals.ok, renewed: totals.renewed, rejected: totals.rejected, outputBytes: totals.outputBytes + outputBytes, }; if (!response.ok) { next.rejected += 1; return next; } next.ok += 1; if (response.renewed) next.renewed += 1; return next; } export async function revalidateTokenBatchFastHost( rawRequests: string[], ): Promise { let totals: RenewalSummary = { ok: 0, renewed: 0, rejected: 0, outputBytes: 0, }; for (let i = 0; i < rawRequests.length; i++) { const response = await revalidateTokenObjectHost(rawRequests[i]!); const responseJson = JSON.stringify(response); totals = addSummary(totals, response, responseJson.length); } return totals; } export const revalidateTokenBatchFast = task({ f: revalidateTokenBatchFastHost, }); export function summarizeJsonResponses(rawResponses: string[]): RenewalSummary { const totals: RenewalSummary = { ok: 0, renewed: 0, rejected: 0, outputBytes: 0, }; for (let i = 0; i < rawResponses.length; i++) { const raw = rawResponses[i]!; totals.outputBytes += raw.length; const parsed = safeParseJson(raw); if (!isRecord(parsed) || parsed.ok !== true) { totals.rejected += 1; continue; } totals.ok += 1; if (parsed.renewed === true) totals.renewed += 1; } return totals; } function tamperToken(token: string): string { const chars = token.split(""); const last = chars.length - 1; chars[last] = chars[last] === "a" ? "b" : "a"; return chars.join(""); } export async function buildDemoRevalidateRequests( options: DemoRequestOptions, ): Promise { const count = clampInt(options.count, 1, 1, 5_000_000); const nowSec = clampInt( options.nowSec, Math.floor(Date.now() / 1000), 1, 2_147_483_647, ); const ttlSec = clampInt( options.ttlSec, DEFAULT_TTL_SEC, MIN_TTL_SEC, MAX_TTL_SEC, ); const renewWindowSec = clampInt( options.renewWindowSec, DEFAULT_RENEW_WINDOW_SEC, 0, MAX_RENEW_WINDOW_SEC, ); const invalidPercent = clampInt(options.invalidPercent, 10, 0, 95); const renewable = await issueTokenHost({ sub: "u_demo_renewable", sid: "s_renewable", scope: ["read", "profile"], iat: nowSec - 70, exp: nowSec + Math.max(5, renewWindowSec - 3), renewUntil: nowSec + 900, }); const fresh = await issueTokenHost({ sub: "u_demo_fresh", sid: "s_fresh", scope: ["read"], iat: nowSec - 10, exp: nowSec + 180, renewUntil: nowSec + 900, }); const expiredRenewable = await issueTokenHost({ sub: "u_demo_expired_grace", sid: "s_expired_grace", scope: ["read", "write"], iat: nowSec - 200, exp: nowSec - 4, renewUntil: nowSec + 240, }); const expiredHard = await issueTokenHost({ sub: "u_demo_expired_hard", sid: "s_expired_hard", scope: ["read"], iat: nowSec - 500, exp: nowSec - 30, renewUntil: nowSec - 3, }); const badSignature = tamperToken(fresh); const malformed = "not-a-jwt"; const requests = new Array(count); for (let i = 0; i < count; i++) { const withinInvalid = i % 100 < invalidPercent; let token: string; if (withinInvalid) { token = i % 2 === 0 ? badSignature : malformed; } else { switch (i % 4) { case 0: token = renewable; break; case 1: token = fresh; break; case 2: token = expiredRenewable; break; default: token = expiredHard; } } requests[i] = JSON.stringify({ token, nowSec, ttlSec, renewWindowSec, }); } return requests; } ``` ## Why offload JWT work HMAC verification and key derivation are CPU-bound. On a busy API server handling hundreds of authenticated requests per second, that crypto work competes with route handling on the main thread. Moving it to a pool keeps your event loop responsive -- especially under mixed traffic where some routes are cheap and others hit the auth path hard. --- # Monte Carlo pi URL: https://knittingdocs.vercel.app/examples/maths/monte_pi/ Estimate pi by throwing darts at a circle, split across workers. The classic Monte Carlo dartboard: throw random points into a square, count how many land inside the unit circle, and estimate pi. Every point is independent, making this embarrassingly parallel and a clean demonstration of Knitting's map-reduce pattern. ## How it works Throw lots of random points into $$[-1,1]\times[-1,1]$$. A point is "inside" if $$x^2 + y^2 \le 1$$. Because areas scale nicely: $$ \pi \approx 4 \cdot \frac{\text{inside}}{\text{total}} $$ The host divides the total sample count into chunks, dispatches each chunk to a worker via `piChunk(seed, samples)`, and each worker runs a tight inner loop with a fast deterministic RNG (xorshift32). Each chunk returns a tiny summary: `{ inside, samples }`. The host aggregates and prints the final estimate. ## Run Expected output: ``` threads: 6 samples: 50,000,000 chunk: 1,000,000 dispatching 50 chunks... pi ~ 3.14159842 (error: +0.00000577) elapsed: 0.87s ``` > Note: Monte Carlo error shrinks like `1/sqrtN` -- to halve the error, you need 4x more samples. With 50M samples you'll typically land within ~0.001 of the true value. ## Code `pi.ts` ```ts import { createPool, isMain } from "knitting"; import { piChunk } from "./montecarlo_pi.ts"; function intArg(name: string, fallback: number) { const idx = process.argv.indexOf(`--${name}`); if (idx !== -1 && idx + 1 < process.argv.length) { const v = Number(process.argv[idx + 1]); if (Number.isFinite(v) && v > 0) return Math.floor(v); } return fallback; } // Tunables (pick any numbers you like) const TOTAL_SAMPLES = intArg("samples", 50_000_000_000); const CHUNK_SAMPLES = intArg("chunk", 10_000_000); const THREADS = intArg("threads", 6); const { call, shutdown } = createPool({ threads: THREADS, inliner: { position: "last", batchSize: 8, }, balancer: "firstIdle", })({ piChunk }); async function main() { const jobCount = Math.ceil(TOTAL_SAMPLES / CHUNK_SAMPLES); const jobs = new Array>( jobCount, ); // Seed base: stable-ish, different each run const seedBase = ((Date.now() | 0) ^ 0x9e3779b9) | 0; // Queue one worker task per chunk. for (let i = 0; i < jobCount; i++) { const remaining = TOTAL_SAMPLES - i * CHUNK_SAMPLES; const samples = remaining >= CHUNK_SAMPLES ? CHUNK_SAMPLES : remaining; // Spread seeds so chunks don’t reuse the same random stream const seed = (seedBase + (i * 0x6d2b79f5)) | 0; jobs[i] = call.piChunk([seed, samples]); } const time = performance.now(); const results = await Promise.all(jobs); const finished = performance.now(); let inside = 0; let total = 0; for (const r of results) { inside += r.inside; total += r.samples; } const pi = (4 * inside) / total; // Quick sanity: expected sampling error scales like ~1/sqrt(N) const approxStdErr = 1 / Math.sqrt(total); console.log("Monte Carlo π estimate"); console.log("threads :", THREADS + 1); console.log("total samples:", total.toLocaleString()); console.log("chunk size :", CHUNK_SAMPLES.toLocaleString()); console.log("pi :", pi); console.log("took :", (finished - time).toFixed(3), " ms"); console.log("rough ±err :", `~${(approxStdErr * 4).toExponential(2)}`); } if (isMain) { main().finally(shutdown); } ``` `montecarlo_pi.ts` ```ts import { task } from "knitting"; type ChunkArgs = readonly [seed: number, samples: number]; type ChunkResult = { inside: number; samples: number }; /** * Fast, deterministic RNG: xorshift32. * (Good enough for Monte Carlo demos, and much faster than Math.random in tight loops.) */ function xorshift32(state: number): number { state |= 0; state ^= state << 13; state ^= state >>> 17; state ^= state << 5; return state | 0; } const INV_2_POW_32 = 2.3283064365386963e-10; // 1 / 2^32 export const piChunk = task({ f: ([seed, samples]) => { let s = seed | 0; let inside = 0; for (let i = 0; i < samples; i++) { s = xorshift32(s); const x = ((s >>> 0) * INV_2_POW_32) * 2 - 1; s = xorshift32(s); const y = ((s >>> 0) * INV_2_POW_32) * 2 - 1; const r2 = x * x + y * y; if (r2 <= 1) inside++; } return { inside, samples }; }, }); ``` ## Why this is a good parallel pattern This example captures the core scientific computing pattern: **map** (simulate many independent trials) then **reduce** (combine partial statistics). Each chunk returns a small summary, not raw data -- that's critical for keeping transfer overhead low. The same structure works for numerical integration, uncertainty propagation, risk estimation, statistical physics, and any problem you can phrase as "run many trials and combine results." ## Practical notes **Chunk size matters.** Each task should do enough work to justify dispatch overhead. Start with 100k-5M iterations per chunk depending on your loop cost. Too small and overhead dominates; too big and you lose load balancing. **Seed carefully.** Use one base seed (so runs are reproducible) and derive a different per-chunk seed (so chunks don't share the same random stream). This example does this correctly. **Keep the inner loop tight.** Avoid allocations per iteration. This example uses xorshift32 instead of `Math.random()` for speed. ## Things to try 1. Increase `--samples` to 500M and watch the estimate stabilize. 2. Try different `--chunk` sizes and measure throughput -- there's a sweet spot. 3. Fix the seed and confirm identical output across runs. 4. Modify the worker to also return timing per chunk and plot a histogram. --- # React SSR compression URL: https://knittingdocs.vercel.app/examples/data_transforms/rendering_output/react_ssr_compress/ Render React on a worker and Brotli-compress the HTML, measuring where the compression step belongs. Takes the React SSR example one step further: after rendering HTML, compress it with Brotli. This tests whether it's better to compress on the worker (render + compress in one shot) or on the host (render on worker, compress on main thread). Spoiler: doing both on the worker usually wins because you avoid sending uncompressed HTML back across the thread boundary. ## How it works The host generates JSON payload strings. Both paths render the same user card component, then Brotli-compress the HTML output. The benchmark compares compressed byte totals for parity, then measures throughput. - `bench_react_ssr_compress.ts` -- the benchmark - `render_user_card_compressed.tsx` -- the SSR + compression task - `utils.ts` -- shared payload and compression helpers ## Example input and output Input is the same JSON payload shape as the plain React SSR example. The worker call is: ```ts const compressed = await pool.call.renderUserCardCompressed(payloadJson); ``` Output is a Brotli-compressed `Buffer`, not an HTML string. That is the point of this example: do the expensive render and compression work in one place, then return the smaller payload. Compression uses built-in `node:zlib` -- no extra packages. ## Optional benchmark Expected output: ``` byte parity check: host=42,380 worker=42,380 OK match benchmark avg (ns) min ... max (ns) host 45,600 41,200 ... 58,300 knitting 24,100 21,800 ... 32,400 ``` Brotli is significantly more expensive than `renderToString` alone, so the worker advantage is more pronounced here than in the plain SSR example. ## Code `render_user_card_compressed.tsx` ```tsx import { task } from "knitting"; import { brotliCompressSync } from "node:zlib"; import { renderUserCardHost } from "../react_ssr/render_user_card.tsx"; function compressHtml(html: string) { return brotliCompressSync(html); } export const renderUserCardCompressed = task({ f: (payload: string) => { const html = renderUserCardHost(payload); const compressed = compressHtml(html); return compressed; }, }); ``` `bench_react_ssr_compress.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { renderUserCardHost } from "../react_ssr/render_user_card.tsx"; import { renderUserCardCompressed } from "./render_user_card_compressed.tsx"; import { buildCompressionPayloads, compressHtml, sumCompressedBytes, } from "./utils.ts"; const THREADS = 1; const REQUESTS = 100; async function main() { const payloads = buildCompressionPayloads(REQUESTS); const pool = createPool({ threads: THREADS, inliner: { batchSize: 8, }, })({ renderUserCardCompressed }); let sink = 0; try { runHost(payloads); await runWorkers( pool.call.renderUserCardCompressed, payloads, ); console.log("React SSR + compression benchmark (mitata)"); console.log("workload: parse + normalize + render + brotli"); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("threads:", THREADS, " + main"); boxplot(() => { summary(() => { bench(`host (${REQUESTS.toLocaleString()} req)`, () => { sink = runHost(payloads); }); bench( `knitting (${THREADS} thread(s), ${REQUESTS.toLocaleString()} req)`, async () => { sink = await runWorkers( pool.call.renderUserCardCompressed, payloads, ); }, ); }); }); await run(); console.log("last compressed bytes:", sink.toLocaleString()); } finally { pool.shutdown(); } } function runHost(payloads: string[]): number { let compressedBytes = 0; for (let i = 0; i < payloads.length; i++) { const html = renderUserCardHost(payloads[i]!); compressedBytes += compressHtml(html).byteLength; } return compressedBytes; } async function runWorkers( callRender: (payload: string) => Promise<{ byteLength: number }>, payloads: string[], ): Promise { const jobs: Promise<{ byteLength: number }>[] = []; for (let i = 0; i < payloads.length; i++) { jobs.push(callRender(payloads[i]!)); } const results = await Promise.all(jobs); return sumCompressedBytes(results); } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `utils.ts` ```ts import { brotliCompressSync } from "node:zlib"; import { renderUserCardHost } from "../react_ssr/render_user_card.tsx"; import { buildUserPayloads } from "../react_ssr/utils.ts"; export type CompressionResult = { ms: number; bytes: number; }; export function buildCompressionPayloads(count: number): string[] { return buildUserPayloads(count); } export function compressHtml(html: string) { return brotliCompressSync(html); } export function sumCompressedBytes( chunks: ArrayLike<{ byteLength: number }>, ): number { let total = 0; for (let i = 0; i < chunks.length; i++) { total += chunks[i]!.byteLength; } return total; } export function runHostCompression(payloads: string[]): CompressionResult { const started = performance.now(); let compressedBytes = 0; for (let i = 0; i < payloads.length; i++) { const html = renderUserCardHost(payloads[i]!); compressedBytes += compressHtml(html).byteLength; } return { ms: performance.now() - started, bytes: compressedBytes }; } export function printCompressionMetrics( mode: string, requests: number, ms: number, compressedBytes: number, ): void { const secs = Math.max(1e-9, ms / 1000); const rps = requests / secs; console.log(`${mode} took : ${ms.toFixed(2)} ms`); console.log(`${mode} throughput : ${rps.toFixed(0)} req/s`); console.log(`${mode} compressed : ${compressedBytes.toLocaleString()}`); } ``` ## Compression placement matters Where you compress affects both throughput and data transfer. If you compress on the worker, you send a small compressed buffer back to the host. If you compress on the host, you first transfer the full uncompressed HTML string across the thread boundary, then compress it. For response compression in a real server, doing render + compress on the worker is almost always the right call. --- # Markdown to HTML URL: https://knittingdocs.vercel.app/examples/data_transforms/rendering_output/markdown_to_html/ Convert markdown to HTML and Brotli-compress it in the same worker call. Converts markdown documents to HTML using `marked`, then compresses the output with Brotli. A straightforward transform pipeline -- parse, render, compress -- that shows how well workers handle chained operations. ## How it works The host generates markdown documents. Both the host and worker paths parse the markdown, render HTML, then Brotli-compress the result. Compressed byte totals are compared for parity before the benchmark runs. Three files: - `run_markdown_to_html.ts` -- a small runnable example that compares host and worker output - `bench_markdown_to_html.ts` -- the optional host-vs-worker benchmark - `utils.ts` -- markdown rendering, compression tasks, and shared helpers ## Run Example markdown in: ```md # Knitting markdown example This example renders markdown on the host and in a worker. ## Checklist - Parse markdown - Render HTML - Compare outputs ``` HTML out: ```html

Knitting markdown example

This example renders markdown on the host and in a worker.

Checklist

... ``` ## Optional benchmark Expected output: ``` byte parity check: host=18,720 worker=18,720 OK match benchmark avg (ns) min ... max (ns) host 32,100 29,400 ... 41,200 knitting 17,800 15,600 ... 24,300 ``` ## Code `run_markdown_to_html.ts` ```ts import { createPool, isMain } from "knitting"; import { markdownToHtml, markdownToHtmlHost } from "./utils.ts"; function intArg(name: string, fallback: number): number { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const value = Number(process.argv[i + 1]); if (Number.isFinite(value) && value > 0) return Math.floor(value); } return fallback; } const THREADS = intArg("threads", 2); const SAMPLE_MARKDOWN = [ "# Knitting markdown example", "", "This example renders markdown on the host and in a worker.", "", "## Checklist", "", "- Parse markdown", "- Render HTML", "- Compare outputs", "", "```ts", "const status = 'ready';", "```", ].join("\n"); async function main() { const hostHtml = markdownToHtmlHost(SAMPLE_MARKDOWN); const pool = createPool({ threads: THREADS })({ markdownToHtml }); try { const workerHtml = await pool.call.markdownToHtml(SAMPLE_MARKDOWN); console.log("Markdown -> HTML example"); console.log("threads :", THREADS); console.log("same html :", hostHtml === workerHtml); console.log("html length :", workerHtml.length); console.log("html preview :", workerHtml.slice(0, 120), "..."); } finally { pool.shutdown(); } } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `bench_markdown_to_html.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { buildMarkdownDocs, markdownToHtmlCompressed, markdownToHtmlCompressedHost, sumChunkBytes, } from "./utils.ts"; const THREADS = 2; const DOCS = 2_000; function runHost(markdowns: string[]): number { let compressedBytes = 0; for (let i = 0; i < markdowns.length; i++) { compressedBytes += markdownToHtmlCompressedHost(markdowns[i]!).byteLength; } return compressedBytes; } async function runWorkers( callRender: (markdown: string) => Promise, markdowns: string[], ): Promise { const jobs: Promise[] = []; for (let i = 0; i < markdowns.length; i++) { jobs.push(callRender(markdowns[i]!)); } const chunks = await Promise.all(jobs); return sumChunkBytes(chunks); } async function main() { const markdowns = buildMarkdownDocs(DOCS); const pool = createPool({ threads: THREADS })({ markdownToHtmlCompressed }); let sink = 0; try { await runWorkers( pool.call.markdownToHtmlCompressed, markdowns, ); console.log("Markdown -> HTML benchmark (mitata)"); console.log("workload: parse + render + brotli"); console.log("docs per iteration:", DOCS.toLocaleString()); console.log("threads:", THREADS); boxplot(() => { summary(() => { bench(`host (${DOCS.toLocaleString()} docs)`, () => { sink = runHost(markdowns); }); bench( `knitting (${THREADS} thread(s), ${DOCS.toLocaleString()} docs)`, async () => { sink = await runWorkers( pool.call.markdownToHtmlCompressed, markdowns, ); }, ); }); }); await run(); console.log("last compressed bytes:", sink.toLocaleString()); } finally { pool.shutdown(); } } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `utils.ts` ```ts import { task } from "knitting"; import { marked } from "marked"; import { brotliCompressSync } from "node:zlib"; marked.setOptions({ gfm: true, }); export function markdownToHtmlHost(markdown: string): string { return marked.parse(markdown) as string; } export const markdownToHtml = task({ f: markdownToHtmlHost, }); export function markdownToHtmlCompressedHost(markdown: string) { const html = markdownToHtmlHost(markdown); return brotliCompressSync(html); } export const markdownToHtmlCompressed = task({ f: markdownToHtmlCompressedHost, }); export const TOPICS = [ "workers", "schema", "compression", "batching", "latency", "rendering", "validation", "throughput", ]; function pick(arr: T[], i: number): T { return arr[i % arr.length]!; } export function makeMarkdown(i: number): string { const topicA = pick(TOPICS, i); const topicB = pick(TOPICS, i + 3); const topicC = pick(TOPICS, i + 5); const id = i.toString(36); return [ `# Job ${id.toUpperCase()}`, "", `This page documents a ${topicA} pipeline for ${topicB}.`, "", "## Checklist", "", `- Parse input payload ${id}`, "- Validate required fields and defaults", `- Render output for ${topicC}`, "", "## Sample code", "", "```ts", `const jobId = \"${id}\";`, 'const status = "ready";', "```", "", `Generated at 2026-01-${String((i % 27) + 1).padStart(2, "0")}.`, ].join("\n"); } export function buildMarkdownDocs(count: number): string[] { const docs = new Array(count); for (let i = 0; i < count; i++) docs[i] = makeMarkdown(i); return docs; } export function sumChunkBytes( chunks: ArrayLike<{ byteLength: number }>, ): number { let total = 0; for (let i = 0; i < chunks.length; i++) { total += chunks[i]!.byteLength; } return total; } ``` ## A clean pipeline example This is the simplest rendering example -- no component tree, no JSX, just string in, string out. It's a good reference if you want to understand the worker pattern without the React SSR complexity. The same approach works for any transform chain: parse input, process it, compress or encode the result. --- # Physics loop URL: https://knittingdocs.vercel.app/examples/maths/physics_loop/ A branch-heavy 2D random-walk simulation spread across workers. A physics-style simulation loop: step-by-step state updates with branching and early exits. Unlike the pi example (pure arithmetic), this one simulates a **2D random walk** -- start at the origin, move one unit in a random direction each step, stop when the particle crosses a radius boundary. The work per trial varies (some particles escape early, some don't), which makes chunking important for load balance. ## How it works 1. The host creates a pool and splits the total run count into chunks. 2. Each worker runs many independent particle trials in a tight inner loop: - Initialize `(x, y)` at the origin - Each step: pick a random direction, update position, check escape condition - Track whether the particle escaped and how many steps it took 3. Each chunk returns **compact summary stats**: escaped count, total runs, sum of escape steps. 4. The host aggregates chunk results into final estimates. We measure two things: - **Escape probability:** what fraction of particles reached the boundary within `maxSteps`? - **Mean escape time:** how many steps did it take (for particles that escaped)? ## Run Expected output: ``` threads: 6 runs: 15,000,000 batch: 5,000 steps: 15,000 radius: 100 escape probability: 0.9847 mean escape steps: 7,312 elapsed: 3.41s ``` > Note: With `--radius 100` and `--steps 15000`, most particles escape. Increase the radius or decrease max steps to see the escape probability drop. ## Code `walks_runs.ts` ```ts import { createPool, isMain } from "knitting"; import { walkChunk } from "./walk2d.ts"; function intArg(name: string, fallback: number) { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const v = Number(process.argv[i + 1]); if (Number.isFinite(v) && v > 0) return Math.floor(v); } return fallback; } function numArg(name: string, fallback: number) { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const v = Number(process.argv[i + 1]); if (Number.isFinite(v) && v > 0) return v; } return fallback; } // Tunables (pick any data you like) const THREADS = intArg("threads", 4); const TOTAL_RUNS = intArg("runs", 5_000_000); const RUNS_PER_JOB = intArg("batch", 5_000); const MAX_STEPS = intArg("steps", 15_000); const RADIUS = numArg("radius", 100); const { call, shutdown } = createPool({ threads: THREADS, balancer: "firstIdle", // Optional: inliner helps if each job is too small. // inliner: { position: "last", batchSize: 1 }, })({ walkChunk }); async function main() { const jobsCount = Math.ceil(TOTAL_RUNS / RUNS_PER_JOB); const jobs = new Array< Promise< { escaped: number; totalRuns: number; sumSteps: number; sumSteps2: number; } > >(jobsCount); const seedBase = ((Date.now() | 0) ^ 0x9e3779b9) | 0; for (let j = 0; j < jobsCount; j++) { const remaining = TOTAL_RUNS - j * RUNS_PER_JOB; const runs = remaining >= RUNS_PER_JOB ? RUNS_PER_JOB : remaining; // Spread seeds per job so streams differ const seed = (seedBase + (j * 0x6d2b79f5)) | 0; jobs[j] = call.walkChunk([seed, runs, MAX_STEPS, RADIUS]); } const results = await Promise.all(jobs); let escaped = 0; let total = 0; let sumSteps = 0; let sumSteps2 = 0; for (const r of results) { escaped += r.escaped; total += r.totalRuns; sumSteps += r.sumSteps; sumSteps2 += r.sumSteps2; } const pEscape = escaped / total; let mean = NaN; let stdev = NaN; if (escaped > 0) { mean = sumSteps / escaped; const mean2 = sumSteps2 / escaped; const variance = Math.max(0, mean2 - mean * mean); stdev = Math.sqrt(variance); } console.log("Monte Carlo: 2D random-walk first-exit"); console.log("threads :", THREADS); console.log("total runs :", total.toLocaleString()); console.log("radius :", RADIUS); console.log("max steps :", MAX_STEPS.toLocaleString()); console.log("escape prob :", pEscape); console.log("mean steps :", mean); console.log("stdev steps :", stdev); } if (isMain) { main().finally(shutdown); } ``` `walk2d.ts` ```ts import { task } from "knitting"; type Args = readonly [ seed: number, runs: number, maxSteps: number, radius: number, dirPow2?: number, // optional: directions table size = 2^dirPow2 (default 10 => 1024) ]; type Result = { escaped: number; totalRuns: number; sumSteps: number; // sum of steps taken until escape (only for escaped runs) sumSteps2: number; // sum of steps^2 (only for escaped runs) }; // Fast deterministic RNG (xorshift32) function xorshift32(state: number): number { state |= 0; state ^= state << 13; state ^= state >>> 17; state ^= state << 5; return state | 0; } // Precompute direction tables (module-scope = done once per worker) function makeDirs(pow2: number) { const n = 1 << pow2; const xs = new Float64Array(n); const ys = new Float64Array(n); const twoPi = Math.PI * 2; for (let i = 0; i < n; i++) { const a = (i * twoPi) / n; xs[i] = Math.cos(a); ys[i] = Math.sin(a); } return { xs, ys, mask: n - 1 }; } // Default table: 1024 directions const DEFAULT_DIR_POW2 = 10; let DIRS = makeDirs(DEFAULT_DIR_POW2); export const walkChunk = task({ f: ([seed, runs, maxSteps, radius, dirPow2]) => { if (dirPow2 && dirPow2 !== DEFAULT_DIR_POW2) { // Rare path: allow custom resolution if you want DIRS = makeDirs(dirPow2 | 0); } const r2Limit = radius * radius; let s = seed | 0; let escaped = 0; let sumSteps = 0; let sumSteps2 = 0; const xs = DIRS.xs; const ys = DIRS.ys; const mask = DIRS.mask; for (let run = 0; run < runs; run++) { let x = 0.0; let y = 0.0; for (let step = 1; step <= maxSteps; step++) { s = xorshift32(s); const idx = s & mask; x += xs[idx]; y += ys[idx]; const r2 = x * x + y * y; if (r2 >= r2Limit) { escaped++; sumSteps += step; sumSteps2 += step * step; break; } } } return { escaped, totalRuns: runs, sumSteps, sumSteps2 }; }, }); ``` ## What makes this different from the pi example The pi example does the same amount of work per sample (two multiplies and a compare). This simulation has **variable work per trial** -- particles that escape early are cheap, particles that hit the step limit are expensive. That variability means chunking matters more: well-sized chunks smooth out the variance so no single worker gets stuck with all the hard trials. The inner loop is also branch-heavy (escape checks, direction selection), which is closer to real simulation code than a pure arithmetic kernel. ## The science This is a discrete-time approximation of **Brownian motion** / **diffusion**. The estimates converge by the Law of Large Numbers, and Monte Carlo error shrinks like `1/sqrtN`. $$ \hat{p} = \frac{\text{escaped}}{\text{total runs}} \qquad \hat{\mu} = \frac{\sum \text{steps}}{\text{escaped}} $$ Real applications of this pattern: diffusion/Brownian motion, hitting time problems, Monte Carlo transport (particles through materials), agent-based models, game simulation, and uncertainty propagation. ## Practical notes **Chunk size:** Start with `--batch 5000` to `50000` for heavier loops. Increase if each run is short, decrease if each run is long. **Keep the inner loop tight:** Avoid allocations per step. Precompute direction tables if possible. Use simple numeric types. **Validate invariants:** Check that `0 <= escaped <= totalRuns`, totals add up across chunks, and results are stable under fixed seeds. ## Things to try 1. Increase `--radius` and see how mean escape steps changes. 2. Compare different `--batch` chunk sizes and measure throughput. 3. Add variance reporting and compute a 95% confidence interval. 4. Replace the random walk with a drift term (constant force) and compare escape behavior. --- # Salt hashing URL: https://knittingdocs.vercel.app/examples/data_transforms/validation/salt_hashing/ PBKDF2 password hashing with constant-time verification, kept off the request thread. Derives password hashes with PBKDF2-SHA256 and verifies them with constant-time comparison. This is the heaviest per-call workload in the validation set -- each hash derivation is genuinely CPU-expensive by design (that's the point of PBKDF2). ## How it works 1. Generate a random salt per password. 2. Derive a hash using `PBKDF2-SHA256` via Web Crypto. 3. Store as a compact record: `algorithm$iterations$keyBytes$salt$hash`. 4. Verify login attempts by recomputing and constant-time comparing. The benchmark uses `Uint8Array` payloads to reduce serialization noise in hot loops. Three files: - `salt_knitting.ts` -- runs salting + verification in host and Knitting modes - `utils.ts` -- hashing, verification, and fast-path packet functions - `bench_salt_hashing.ts` -- host-vs-worker benchmark with `mitata` ## Example record shape The stored credential format is intentionally compact: ```txt pbkdf2-sha256$600000$32$$ ``` At a high level the flow is: ```txt password -> derive PBKDF2 hash with random salt -> store compact record login attempt -> derive again with stored salt -> constant-time compare ``` That makes the worker boundary simple: small input, small output, expensive CPU work in the middle. Uses built-in Web Crypto APIs -- no extra crypto package required. ## Run Expected output: ``` -- host mode -- hashed: 100 verified: 100 mismatches: 0 -- knitting mode (2 threads) -- hashed: 100 verified: 100 mismatches: 0 ``` ## Optional benchmark Compares hashing typed-array packets through direct imports (`host`) vs worker task calls (`knitting`). Because PBKDF2 is intentionally slow (high iteration count), this is where workers shine most -- each call does enough real work to easily justify dispatch overhead. ## Code `salt_knitting.ts` ```ts import { createPool, isMain } from "knitting"; import { decodeHashResultPacket, hashPassword, hashPasswordHost, hashPasswordPacketHost, makeHashPacketForIndex, verifyPassword, verifyPasswordHost, } from "./utils.ts"; const THREADS = 2; const REQUESTS = 2_000; const ITERATIONS = 120_000; const MISMATCH_PERCENT = 5; type Summary = { hashed: number; verified: number; mismatched: number; }; function passwordFor(i: number): string { return `user-${i.toString(36)}-password`; } function expectedPassword(i: number): string { if (i % 100 < MISMATCH_PERCENT) return `wrong-${i.toString(36)}-password`; return passwordFor(i); } async function runHost(): Promise { let verified = 0; let mismatched = 0; for (let i = 0; i < REQUESTS; i++) { const password = passwordFor(i); const hashed = await hashPasswordHost({ password, iterations: ITERATIONS }); const checked = await verifyPasswordHost({ password: expectedPassword(i), record: hashed.record, }); if (checked.ok) verified++; else mismatched++; } return { hashed: REQUESTS, verified, mismatched }; } async function runWorkers(): Promise { const pool = createPool({ threads: THREADS })({ hashPassword, verifyPassword, }); try { const hashJobs: Promise<{ record: string }>[] = []; for (let i = 0; i < REQUESTS; i++) { hashJobs.push(pool.call.hashPassword({ password: passwordFor(i), iterations: ITERATIONS, })); } const hashes = await Promise.all(hashJobs); const verifyJobs: Promise<{ ok: boolean }>[] = []; for (let i = 0; i < REQUESTS; i++) { verifyJobs.push(pool.call.verifyPassword({ password: expectedPassword(i), record: hashes[i]!.record, })); } const checks = await Promise.all(verifyJobs); let verified = 0; for (let i = 0; i < checks.length; i++) { if (checks[i]!.ok) verified++; } const mismatched = REQUESTS - verified; return { hashed: REQUESTS, verified, mismatched }; } finally { pool.shutdown(); } } function printSummary(mode: string, summary: Summary, ms: number): void { const seconds = Math.max(1e-9, ms / 1000); const ops = REQUESTS / seconds; console.log(mode); console.log("requests :", REQUESTS.toLocaleString()); console.log("iterations :", ITERATIONS.toLocaleString()); console.log("mismatch rate :", `${MISMATCH_PERCENT}%`); console.log("hashed :", summary.hashed.toLocaleString()); console.log("verified :", summary.verified.toLocaleString()); console.log("mismatched :", summary.mismatched.toLocaleString()); console.log("took :", `${ms.toFixed(2)} ms`); console.log("throughput :", `${ops.toFixed(0)} req/s`); } async function printPacketSample() { const packet = makeHashPacketForIndex(7, ITERATIONS, 32, 16); const result = await hashPasswordPacketHost(packet); const decoded = decodeHashResultPacket(result); console.log("packet sample : iterations", decoded.iterations); console.log("salt(base64) :", decoded.saltBase64); console.log("hash(base64) :", decoded.hashBase64); } async function main() { const hostStart = performance.now(); const host = await runHost(); const hostMs = performance.now() - hostStart; const workerStart = performance.now(); const knitting = await runWorkers(); const workerMs = performance.now() - workerStart; const uplift = (hostMs / Math.max(1e-9, workerMs) - 1) * 100; console.log("Password salting + hashing"); console.log(`threads: ${THREADS}`); console.log(""); printSummary("host", host, hostMs); console.log(""); printSummary("knitting", knitting, workerMs); console.log(""); console.log(`uplift: ${uplift.toFixed(1)}%`); await printPacketSample(); } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `bench_salt_hashing.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { buildDemoHashPackets, type HashBatchSummary, hashPasswordPacketBatchFast, hashPasswordPacketBatchFastHost, } from "./utils.ts"; const THREADS = 2; const REQUESTS = 500; const BATCH = 32; const ITERATIONS = 1_200; const KEY_BYTES = 32; const SALT_BYTES = 16; async function main() { const packets = buildDemoHashPackets({ count: REQUESTS, iterations: ITERATIONS, keyBytes: KEY_BYTES, saltBytes: SALT_BYTES, }); const batches = makeBatches(packets, BATCH); const pool = createPool({ threads: THREADS })({ hashPasswordPacketBatchFast, }); let sink = 0; try { const hostCheck = await runHostBatches(batches); const workerCheck = await runWorkerBatches( pool.call.hashPasswordPacketBatchFast, batches, ); if (!same(hostCheck, workerCheck)) { throw new Error("Host and worker hashing summaries differ."); } console.log("Salt hashing benchmark (mitata)"); console.log("workload: PBKDF2-SHA256 on Uint8Array request packets"); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("iterations:", ITERATIONS.toLocaleString()); console.log("batch size:", BATCH); console.log("threads:", THREADS); boxplot(() => { summary(() => { bench( `host (${REQUESTS.toLocaleString()} req, batch ${BATCH})`, async () => { const totals = await runHostBatches(batches); sink = totals.outputBytes ^ totals.digestXor; }, ); bench( `knitting (${THREADS} thread${ THREADS === 1 ? "" : "s" }, ${REQUESTS.toLocaleString()} req, batch ${BATCH})`, async () => { const totals = await runWorkerBatches( pool.call.hashPasswordPacketBatchFast, batches, ); sink = totals.outputBytes ^ totals.digestXor; }, ); }); }); await run(); console.log("last sink:", sink.toLocaleString()); } finally { pool.shutdown(); } } function makeBatches(packets: Uint8Array[], batchSize: number): Uint8Array[][] { const out: Uint8Array[][] = []; for (let i = 0; i < packets.length; i += batchSize) { out.push(packets.slice(i, i + batchSize)); } return out; } function merge(a: HashBatchSummary, b: HashBatchSummary): HashBatchSummary { return { count: a.count + b.count, outputBytes: a.outputBytes + b.outputBytes, digestXor: a.digestXor ^ b.digestXor, }; } async function runHostBatches( batches: Uint8Array[][], ): Promise { let totals: HashBatchSummary = { count: 0, outputBytes: 0, digestXor: 0 }; for (let i = 0; i < batches.length; i++) { totals = merge(totals, await hashPasswordPacketBatchFastHost(batches[i]!)); } return totals; } async function runWorkerBatches( callBatch: (packets: Uint8Array[]) => Promise, batches: Uint8Array[][], ): Promise { const jobs: Promise[] = []; for (let i = 0; i < batches.length; i++) jobs.push(callBatch(batches[i]!)); const results = await Promise.all(jobs); let totals: HashBatchSummary = { count: 0, outputBytes: 0, digestXor: 0 }; for (let i = 0; i < results.length; i++) totals = merge(totals, results[i]!); return totals; } function same(a: HashBatchSummary, b: HashBatchSummary): boolean { return a.count === b.count && a.outputBytes === b.outputBytes && a.digestXor === b.digestXor; } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `utils.ts` ```ts import { task } from "knitting"; const encoder = new TextEncoder(); const decoder = new TextDecoder(); const DEFAULT_ITERATIONS = 120_000; const DEFAULT_KEY_BYTES = 32; const DEFAULT_SALT_BYTES = 16; const MIN_ITERATIONS = 10_000; const MAX_ITERATIONS = 2_000_000; const MIN_KEY_BYTES = 16; const MAX_KEY_BYTES = 64; const MIN_SALT_BYTES = 8; const MAX_SALT_BYTES = 32; export type HashRequest = { password: string; iterations?: number; keyBytes?: number; saltBase64?: string; }; export type HashResponse = { record: string; algorithm: "pbkdf2-sha256"; iterations: number; keyBytes: number; saltBase64: string; hashBase64: string; }; export type VerifyRequest = { password: string; record: string; }; export type VerifyResponse = { ok: boolean; reason?: string; }; export type HashBatchSummary = { count: number; outputBytes: number; digestXor: number; }; export type DemoPacketOptions = { count: number; iterations?: number; keyBytes?: number; saltBytes?: number; }; type ParsedRecord = { algorithm: "pbkdf2-sha256"; iterations: number; keyBytes: number; salt: Uint8Array; hash: Uint8Array; }; function clampInt( value: unknown, fallback: number, min: number, max: number, ): number { const numeric = Number(value); if (!Number.isFinite(numeric)) return fallback; const integer = Math.floor(numeric); if (integer < min) return min; if (integer > max) return max; return integer; } function bytesToBase64(bytes: Uint8Array): string { let raw = ""; for (let i = 0; i < bytes.length; i++) raw += String.fromCharCode(bytes[i]!); return btoa(raw); } function base64ToBytes(value: string): Uint8Array | null { try { const raw = atob(value); const bytes = new Uint8Array(raw.length); for (let i = 0; i < raw.length; i++) bytes[i] = raw.charCodeAt(i); return bytes; } catch { return null; } } function fixedTimeEqual(a: Uint8Array, b: Uint8Array): boolean { if (a.length !== b.length) return false; let diff = 0; for (let i = 0; i < a.length; i++) diff |= a[i]! ^ b[i]!; return diff === 0; } function assertPassword(password: string): string { if (typeof password !== "string" || password.length < 8) { throw new Error("password must be at least 8 characters"); } return password; } function normalizeSalt( saltBase64: string | undefined, saltBytes: number, ): Uint8Array { if (!saltBase64) return crypto.getRandomValues(new Uint8Array(saltBytes)); const salt = base64ToBytes(saltBase64); if (!salt) throw new Error("saltBase64 is not valid base64"); if (salt.length < MIN_SALT_BYTES || salt.length > MAX_SALT_BYTES) { throw new Error( `salt length must be ${MIN_SALT_BYTES}-${MAX_SALT_BYTES} bytes`, ); } return salt; } async function derivePbkdf2( passwordBytes: Uint8Array, saltBytes: Uint8Array, iterations: number, keyBytes: number, ): Promise { const key = await crypto.subtle.importKey( "raw", passwordBytes, "PBKDF2", false, ["deriveBits"], ); const bits = await crypto.subtle.deriveBits( { name: "PBKDF2", hash: "SHA-256", salt: saltBytes, iterations }, key, keyBytes * 8, ); return new Uint8Array(bits); } function makeRecord( iterations: number, keyBytes: number, salt: Uint8Array, hash: Uint8Array, ): string { return [ "pbkdf2-sha256", String(iterations), String(keyBytes), bytesToBase64(salt), bytesToBase64(hash), ].join("$"); } function parseRecord(record: string): ParsedRecord | null { const parts = record.split("$"); if (parts.length !== 5) return null; if (parts[0] !== "pbkdf2-sha256") return null; const iterations = Number(parts[1]); const keyBytes = Number(parts[2]); if (!Number.isInteger(iterations) || !Number.isInteger(keyBytes)) return null; if (iterations < MIN_ITERATIONS || iterations > MAX_ITERATIONS) return null; if (keyBytes < MIN_KEY_BYTES || keyBytes > MAX_KEY_BYTES) return null; const salt = base64ToBytes(parts[3]!); const hash = base64ToBytes(parts[4]!); if (!salt || !hash) return null; if (salt.length < MIN_SALT_BYTES || salt.length > MAX_SALT_BYTES) return null; if (hash.length !== keyBytes) return null; return { algorithm: "pbkdf2-sha256", iterations, keyBytes, salt, hash, }; } export async function hashPasswordHost( request: HashRequest, ): Promise { const password = assertPassword(request.password); const iterations = clampInt( request.iterations, DEFAULT_ITERATIONS, MIN_ITERATIONS, MAX_ITERATIONS, ); const keyBytes = clampInt( request.keyBytes, DEFAULT_KEY_BYTES, MIN_KEY_BYTES, MAX_KEY_BYTES, ); const saltBytes = clampInt( DEFAULT_SALT_BYTES, DEFAULT_SALT_BYTES, MIN_SALT_BYTES, MAX_SALT_BYTES, ); const salt = normalizeSalt(request.saltBase64, saltBytes); const hash = await derivePbkdf2( encoder.encode(password), salt, iterations, keyBytes, ); const saltBase64 = bytesToBase64(salt); const hashBase64 = bytesToBase64(hash); return { record: makeRecord(iterations, keyBytes, salt, hash), algorithm: "pbkdf2-sha256", iterations, keyBytes, saltBase64, hashBase64, }; } export async function verifyPasswordHost( request: VerifyRequest, ): Promise { const password = assertPassword(request.password); const parsed = parseRecord(request.record); if (!parsed) return { ok: false, reason: "record format is invalid" }; const hash = await derivePbkdf2( encoder.encode(password), parsed.salt, parsed.iterations, parsed.keyBytes, ); return fixedTimeEqual(hash, parsed.hash) ? { ok: true } : { ok: false, reason: "password mismatch" }; } export const hashPassword = task({ f: hashPasswordHost, }); export const verifyPassword = task({ f: verifyPasswordHost, }); function writeU16LE(out: Uint8Array, offset: number, value: number): void { out[offset] = value & 255; out[offset + 1] = (value >>> 8) & 255; } function writeU32LE(out: Uint8Array, offset: number, value: number): void { out[offset] = value & 255; out[offset + 1] = (value >>> 8) & 255; out[offset + 2] = (value >>> 16) & 255; out[offset + 3] = (value >>> 24) & 255; } function readU16LE(input: Uint8Array, offset: number): number { return input[offset]! | (input[offset + 1]! << 8); } function readU32LE(input: Uint8Array, offset: number): number { return ( input[offset]! | (input[offset + 1]! << 8) | (input[offset + 2]! << 16) | (input[offset + 3]! << 24) ) >>> 0; } function toBytes(value: unknown): Uint8Array { if (value instanceof Uint8Array) return value; if (value instanceof ArrayBuffer) return new Uint8Array(value); if (ArrayBuffer.isView(value)) { return new Uint8Array(value.buffer, value.byteOffset, value.byteLength); } if (Array.isArray(value)) { const out = new Uint8Array(value.length); for (let i = 0; i < value.length; i++) out[i] = Number(value[i] ?? 0) & 255; return out; } if (typeof value !== "object" || value === null) { throw new Error("packet is not a byte buffer"); } const candidate = value as { length?: unknown; byteLength?: unknown; data?: unknown; [index: number]: unknown; [key: string]: unknown; }; if (Array.isArray(candidate.data)) { const out = new Uint8Array(candidate.data.length); for (let i = 0; i < candidate.data.length; i++) { out[i] = Number(candidate.data[i] ?? 0) & 255; } return out; } const lengthValue = Number(candidate.length); const byteLengthValue = Number(candidate.byteLength); let size = Number.isFinite(lengthValue) ? Math.max(0, Math.floor(lengthValue)) : Number.isFinite(byteLengthValue) ? Math.max(0, Math.floor(byteLengthValue)) : -1; if (size < 0) { let maxIndex = -1; for (const key of Object.keys(candidate)) { if (/^\d+$/.test(key)) maxIndex = Math.max(maxIndex, Number(key)); } if (maxIndex >= 0) size = maxIndex + 1; } if (size < 0) throw new Error("packet has no length"); const out = new Uint8Array(size); for (let i = 0; i < size; i++) { out[i] = Number(candidate[i] ?? 0) & 255; } return out; } // Compact binary payloads are faster than structured objects for hot loops. // Header: u16 passwordLen, u16 saltLen, u32 iterations, u16 keyBytes. export function encodeHashPacket( passwordBytes: Uint8Array, saltBytes: Uint8Array, iterations: number, keyBytes: number, ): Uint8Array { const headerSize = 10; const out = new Uint8Array( headerSize + passwordBytes.length + saltBytes.length, ); writeU16LE(out, 0, passwordBytes.length); writeU16LE(out, 2, saltBytes.length); writeU32LE(out, 4, iterations); writeU16LE(out, 8, keyBytes); out.set(passwordBytes, headerSize); out.set(saltBytes, headerSize + passwordBytes.length); return out; } function decodeHashPacket(packetLike: unknown): { password: Uint8Array; salt: Uint8Array; iterations: number; keyBytes: number; } { const packet = toBytes(packetLike); if (packet.length < 10) throw new Error("packet too small"); const passwordLen = readU16LE(packet, 0); const saltLen = readU16LE(packet, 2); const iterations = readU32LE(packet, 4); const keyBytes = readU16LE(packet, 8); const expected = 10 + passwordLen + saltLen; if (expected !== packet.length) throw new Error("packet size mismatch"); if (passwordLen < 8) throw new Error("password too short"); if (saltLen < MIN_SALT_BYTES || saltLen > MAX_SALT_BYTES) { throw new Error("salt size invalid"); } if (iterations < MIN_ITERATIONS || iterations > MAX_ITERATIONS) { throw new Error("iterations invalid"); } if (keyBytes < MIN_KEY_BYTES || keyBytes > MAX_KEY_BYTES) { throw new Error("key size invalid"); } const password = packet.slice(10, 10 + passwordLen); const salt = packet.slice(10 + passwordLen, expected); return { password, salt, iterations, keyBytes }; } // Result packet: u16 saltLen, u16 hashLen, u32 iterations, then salt + hash. function encodeHashResultPacket( salt: Uint8Array, hash: Uint8Array, iterations: number, ): Uint8Array { const out = new Uint8Array(8 + salt.length + hash.length); writeU16LE(out, 0, salt.length); writeU16LE(out, 2, hash.length); writeU32LE(out, 4, iterations); out.set(salt, 8); out.set(hash, 8 + salt.length); return out; } export function decodeHashResultPacket(packet: Uint8Array): { iterations: number; saltBase64: string; hashBase64: string; } { if (packet.length < 8) throw new Error("result packet too small"); const saltLen = readU16LE(packet, 0); const hashLen = readU16LE(packet, 2); const iterations = readU32LE(packet, 4); const expected = 8 + saltLen + hashLen; if (expected !== packet.length) { throw new Error("result packet size mismatch"); } const salt = packet.slice(8, 8 + saltLen); const hash = packet.slice(8 + saltLen, expected); return { iterations, saltBase64: bytesToBase64(salt), hashBase64: bytesToBase64(hash), }; } export async function hashPasswordPacketHost( packet: Uint8Array, ): Promise { const decoded = decodeHashPacket(packet); const hash = await derivePbkdf2( decoded.password, decoded.salt, decoded.iterations, decoded.keyBytes, ); return encodeHashResultPacket(decoded.salt, hash, decoded.iterations); } export const hashPasswordPacket = task({ f: hashPasswordPacketHost, }); export async function hashPasswordPacketBatchFastHost( packets: Uint8Array[], ): Promise { let outputBytes = 0; let digestXor = 0; for (let i = 0; i < packets.length; i++) { const hashed = await hashPasswordPacketHost(packets[i]!); outputBytes += hashed.length; digestXor ^= hashed[hashed.length - 1] ?? 0; } return { count: packets.length, outputBytes, digestXor }; } export const hashPasswordPacketBatchFast = task( { f: hashPasswordPacketBatchFastHost, }, ); function fillDeterministicSalt(seed: number, bytes: number): Uint8Array { let x = (seed ^ 0x9e3779b9) >>> 0; const out = new Uint8Array(bytes); for (let i = 0; i < bytes; i++) { x ^= x << 13; x ^= x >>> 17; x ^= x << 5; out[i] = x & 255; } return out; } export function makePasswordBytes(i: number): Uint8Array { return encoder.encode(`password-${i.toString(36)}-knitting`); } export function makeHashPacketForIndex( i: number, iterations: number, keyBytes: number, saltBytes: number, ): Uint8Array { const password = makePasswordBytes(i); const salt = fillDeterministicSalt(i + 1, saltBytes); return encodeHashPacket(password, salt, iterations, keyBytes); } export function buildDemoHashPackets(options: DemoPacketOptions): Uint8Array[] { const count = clampInt(options.count, 1, 1, 2_000_000); const iterations = clampInt( options.iterations, DEFAULT_ITERATIONS, MIN_ITERATIONS, MAX_ITERATIONS, ); const keyBytes = clampInt( options.keyBytes, DEFAULT_KEY_BYTES, MIN_KEY_BYTES, MAX_KEY_BYTES, ); const saltBytes = clampInt( options.saltBytes, DEFAULT_SALT_BYTES, MIN_SALT_BYTES, MAX_SALT_BYTES, ); const packets = new Array(count); for (let i = 0; i < count; i++) { packets[i] = makeHashPacketForIndex(i, iterations, keyBytes, saltBytes); } return packets; } export function hashSummaryFromOutputs( outputs: Uint8Array[], ): HashBatchSummary { let outputBytes = 0; let digestXor = 0; for (let i = 0; i < outputs.length; i++) { const out = outputs[i]!; outputBytes += out.length; digestXor ^= out[out.length - 1] ?? 0; } return { count: outputs.length, outputBytes, digestXor }; } export function utf8(bytes: Uint8Array): string { return decoder.decode(bytes); } ``` ## The ideal offload candidate Password hashing is one of those tasks where workers are almost always worth it. The work is CPU-bound, each call is independent, the input and output are small, and -- critically -- you *want* it to be slow (high iterations = harder to brute force). That's a perfect storm for offloading: expensive per-call work that would otherwise block your event loop during login/signup spikes. --- # Hono server routes URL: https://knittingdocs.vercel.app/examples/data_transforms/rendering_output/hono_server/ Build a Hono server with ping, SSR, and JWT routes in Knitting or host-only mode. ## What is this about This example builds a small Hono API with 3 routes: 1. `GET /ping` for health checks. 2. `POST /ssr` to SSR a user card. 3. `POST /jwt` to issue a valid signed JWT for a user. This example is split into three files: 1. `hono_knitting.ts` (the Hono server, plus a Knitting worker pool). 2. `hono_componets_ssr.tsx` (SSR task: parse + defaults + render). 3. `hono_components_jwt.ts` (JWT task: validate + sign + return JSON string). ## Technologies used (and why) - **Hono**: a small, fast routing layer. It keeps the request path minimal so most overhead is in your actual route work. - **`@hono/node-server`**: a thin adapter that runs a Hono `fetch` handler on Node/Bun. - **React SSR (`react-dom/server`)**: renders a tiny HTML page for `/ssr` so you can simulate CPU-heavy server work. - **`hono/jwt`**: signs a JWT (HS256) for `/jwt` so the example includes common auth-like CPU work. - **Knitting (`knitting`)**: runs selected transforms in a worker pool (threads). This is the core “offload expensive work” technique the example demonstrates. The JWT route uses `hono/jwt` and signs with HS256. Set `JWT_SECRET` in production. ## Deno setup (TSX + npm) This example imports TSX and npm packages. For Deno, keep a root `deno.json` like this: ```json { "nodeModulesDir": "auto", "compilerOptions": { "jsx": "react-jsx", "jsxImportSource": "react" } } ``` Without this, TSX files can fail with: ```txt Uncaught SyntaxError: Unexpected token '<' ``` ## Run ## Route quick checks ```bash # Ping curl -s http://localhost:3000/ping # SSR (returns HTML) curl -s http://localhost:3000/ssr \ -H 'content-type: application/json' \ -d '{"name":"Ari","plan":"pro","bio":"Building on Knitting","projects":17}' # JWT (returns JSON string with token) curl -s http://localhost:3000/jwt \ -H 'content-type: application/json' \ -d '{"user":{"id":"u_42","email":"ari@example.com","role":"admin"},"ttlSec":900}' ``` ## Performance notes Measurements: [Hono 16-core benchmark](/documents/hono-16core-benchmark.md). ### Saturating mixed load — throughput One worker, with `/ping`, `/ssr`, and `/jwt` loaded concurrently. | Route | Hono only | Hono + Knitting | Delta | | --- | ---: | ---: | ---: | | `/ping` | `2,815 rps` | `8,486 rps` | **+201%** | | `/ssr` | `2,815 rps` | `4,816 rps` | **+71%** | | `/jwt` | `2,578 rps` | `4,712 rps` | **+83%** | ### Fixed 6,000 RPS mixed load — latency One worker, with each route offered `2,000 rps`. | Route | Hono only p50 | Knitting p50 | Hono only p99 | Knitting p99 | | --- | ---: | ---: | ---: | ---: | | `/ping` | `0.81ms` | `0.40ms` | `16.93ms` | `2.31ms` | | `/ssr` | `0.84ms` | `2.22ms` | `14.32ms` | `8.71ms` | | `/jwt` | `1.03ms` | `2.58ms` | `25.18ms` | `9.76ms` | ## Why this pattern matters - `ping` stays cheap and synchronous. - You can benchmark workers vs host-only with the same route behavior. - In Knitting mode, the JWT task returns **stringified JSON** to reduce structured-clone overhead. - Heavy route logic stays in one shared file, keeping both server entrypoints small. ## Code `hono_knitting.ts` ```ts import { serve } from "@hono/node-server"; import { createPool } from "knitting"; import { Hono } from "hono"; import { issueJwt } from "./hono_components_jwt.ts"; import { renderSsrPage } from "./hono_componets_ssr.tsx"; const handlers = createPool({ })({ issueJwt, renderSsrPage, }); async function main() { const app = new Hono(); app.get("/ping", (c) => { return c.json({ ok: true, pong: true, runtime: process.release?.name ?? "unknown", ts: new Date().toISOString(), }); }); app.post("/ssr", async (c) => { const html = await handlers.call.renderSsrPage(c.req.arrayBuffer()); return c.html(html); }); app.post("/jwt", async (c) => { const responseJson = await handlers.call.issueJwt(c.req.arrayBuffer()); return c.body(responseJson ?? "Bad request", responseJson ? 200 : 400, { "content-type": "application/json; charset=utf-8", }); }); const server = serve({ fetch: app.fetch, port: 3000 }, (info) => { console.log("GET /ping"); console.log("POST /ssr body: { name?, plan?, bio?, projects? }"); console.log("POST /jwt body: { user: { id, email?, role? }, ttlSec? }"); }); const close = () => { // IMPORTANT TO CLOSE CONNECTION handlers.shutdown(); server.close(); }; process.on("SIGINT", close); process.on("SIGTERM", close); } main().catch((error) => { console.error(error); process.exitCode = 1; }); ``` `hono_componets_ssr.tsx` ```tsx import React from "react"; import { renderToString } from "react-dom/server"; import { task } from "knitting"; import { z } from "zod"; const utf8Decoder = new TextDecoder("utf-8", { fatal: true }); type SsrInput = { name: string; plan: "free" | "pro"; bio: string; projects: number; }; function UserCard({ user }: { user: SsrInput & { updatedAt: string } }) { return ( {`${user.name} - SSR Card`}

{user.name}

{user.bio}

{user.plan.toUpperCase()} plan {user.projects.toLocaleString()} projects Rendered at {user.updatedAt}
); } const ParsedJsonObjectSchema = z.string().transform((raw, ctx) => { try { const parsed = JSON.parse(raw) as unknown; if ( typeof parsed !== "object" || parsed === null || Array.isArray(parsed) ) { ctx.addIssue({ code: z.ZodIssueCode.custom, message: "payload: expected JSON object", }); return z.NEVER; } return parsed as Record; } catch { ctx.addIssue({ code: z.ZodIssueCode.custom, message: "payload: expected JSON object", }); return z.NEVER; } }); const RawSsrInputSchema = z.object({ name: z.preprocess((value) => { if (typeof value !== "string") return undefined; const normalized = value.trim(); return normalized.length > 0 ? normalized : undefined; }, z.string().optional()), plan: z.preprocess( (value) => (value === "free" || value === "pro" ? value : undefined), z.enum(["free", "pro"]).optional(), ), bio: z.preprocess((value) => { if (typeof value !== "string") return undefined; const normalized = value.trim(); return normalized.length > 0 ? normalized : undefined; }, z.string().optional()), projects: z.preprocess((value) => { const numberValue = Number(value); if (!Number.isFinite(numberValue)) return undefined; return Math.max(0, Math.min(100_000, Math.floor(numberValue))); }, z.number().int().optional()), }); const SsrInputSchema = RawSsrInputSchema.transform( (value): SsrInput => ({ name: value.name ?? "Anonymous", plan: value.plan ?? "free", bio: value.bio ?? "No bio yet.", projects: value.projects ?? 0, }), ); export function renderSsrPageHost(rawPayload: ArrayBuffer): string { let decodedPayload = ""; try { decodedPayload = utf8Decoder.decode(rawPayload); } catch { decodedPayload = ""; } const parsed = ParsedJsonObjectSchema.safeParse(decodedPayload); const user: SsrInput = SsrInputSchema.parse( parsed.success ? parsed.data : {}, ); const html = renderToString( , ); return `${html}`; } export const renderSsrPage = task({ f: renderSsrPageHost, }); ``` `hono_components_jwt.ts` ```ts import { sign } from "hono/jwt"; import { task } from "knitting"; import { z } from "zod"; const utf8Decoder = new TextDecoder("utf-8", { fatal: true }); const ParsedJsonObjectSchema = z.string().transform((raw, ctx) => { try { const parsed = JSON.parse(raw) as unknown; if ( typeof parsed !== "object" || parsed === null || Array.isArray(parsed) ) { ctx.addIssue({ code: z.ZodIssueCode.custom, message: "payload: expected JSON object", }); return z.NEVER; } return parsed; } catch { ctx.addIssue({ code: z.ZodIssueCode.custom, message: "payload: expected JSON object", }); return z.NEVER; } }); const JwtUserSchema = z.object({ id: z.string().min(1), email: z.string().email().optional(), role: z.string().min(1).optional(), }); const TtlSecSchema = z.preprocess((value) => { const n = Number(value); if (!Number.isFinite(n)) return 900; return Math.max(30, Math.min(86_400, Math.floor(n))); }, z.number().int()); const JwtPayloadSchema = z.object({ user: JwtUserSchema, ttlSec: TtlSecSchema.optional().default(900), }); export async function issueJwtHost( rawPayload: ArrayBuffer, ): Promise { let decodedPayload: string; try { decodedPayload = utf8Decoder.decode(rawPayload); } catch { return null; } const parsedResult = ParsedJsonObjectSchema.safeParse(decodedPayload); if (!parsedResult.success) { return null; } const payloadResult = JwtPayloadSchema.safeParse(parsedResult.data); if (!payloadResult.success) { return null; } const { user, ttlSec } = payloadResult.data; const now = Math.floor(Date.now() / 1000); const exp = now + ttlSec; const token = await sign( { sub: user.id, email: user.email, role: user.role ?? "member", iat: now, exp, }, process.env.secret ?? "hello", ); return JSON.stringify({ ok: true, token, sub: user.id, exp, }); } export const issueJwt = task({ f: issueJwtHost, }); ``` --- # Prompt token budgeting URL: https://knittingdocs.vercel.app/examples/data_transforms/validation/prompt_token_budgeting/ Trim LLM prompts to fit a token budget before sending them to any model. Shapes LLM prompts to fit within a token budget before they hit the API. Uses `tiktoken` for token counting and drops oldest conversation turns first, then trims the query if still over budget. The budgeting logic itself is model-agnostic -- the example uses `gpt-4o-mini` as the tokenizer target, but the pattern works the same way for any model with a token limit. > Note: This example imports `tiktoken` and `openai`, but it never actually calls the OpenAI API. The `openai` package is optional -- these scripts only prepare prompts and count tokens. The pattern applies to any LLM provider. ## How it works 1. The host generates prompt inputs (same system prefix, different conversation history and queries). 2. Each task builds the full prompt and counts tokens with `tiktoken`. 3. If over budget, it drops the oldest turns first, then trims query tokens. 4. The host aggregates token savings and trim counts. Three files: - `run_prompt_token_budget.ts` -- practical prompt-budgeting example for app logic - `bench_prompt_token_budget.ts` -- dedicated `mitata` benchmark measuring budgeting throughput - `token_budget.ts` -- the budgeting logic itself ## Example budget decision Input: ```ts const input = { model: "gpt-4o-mini", systemPrefix: "You are a docs assistant.", history: [ "Need guidance on schema validation.", "Compare workers and batching.", "Keep the answer short.", ], query: "Give a migration plan and one code example.", maxInputTokens: 900, }; ``` Output shape: ```ts { prompt: "...final prompt string...", rawInputTokens: 1120, inputTokens: 884, trimmedTurns: 1, queryWasTrimmed: false, } ``` The useful part is not just the final `prompt`; you also get the bookkeeping needed to explain why a prompt was trimmed and by how much. `openai` is optional. These examples only prepare prompts and token budgets. `mitata` is only needed for the benchmark script. ## Run Expected output: ``` mode: knitting (2 threads) requests: 2000 budget: 900 tokens (gpt-4o-mini) trimmed: 1,247 / 2,000 requests (62.4%) avg tokens saved: 312 per trimmed request total tokens saved: 389,064 elapsed: 1.82s ``` ## Optional benchmark Compares budgeting throughput on the host vs through workers. Worker tasks return compact totals (token/trim counters), not full prompt strings. > Note: Batch size matters here. Tiny batches mostly measure dispatch overhead; very large batches can increase memory pressure. Start with `32` or `64`, then tune for your hardware. ## Code `run_prompt_token_budget.ts` ```ts import { createPool, isMain } from "knitting"; import { preparePrompt, preparePromptHost, type PromptInput, type PromptPlan, } from "./token_budget.ts"; function intArg(name: string, fallback: number): number { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const value = Number(process.argv[i + 1]); if (Number.isFinite(value)) return Math.floor(value); } return fallback; } function strArg(name: string, fallback: string): string { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { return String(process.argv[i + 1]); } return fallback; } const THREADS = Math.max(1, intArg("threads", 2)); const REQUESTS = Math.max(1, intArg("requests", 20_000)); const MAX_INPUT_TOKENS = Math.max(64, intArg("maxInputTokens", 900)); const MODE = strArg("mode", "knitting"); const MODEL = strArg("model", "gpt-4o-mini"); const SYSTEM_PREFIX = [ "You are a docs assistant.", "Prefer concrete and short answers.", "If data is missing, say it directly.", "Do not invent unsupported behavior.", ].join("\n"); const TOPICS = [ "token budgeting", "prompt caching", "parallel workers", "schema validation", "rendering pipelines", "markdown output", "compression tradeoffs", "latency under load", ]; function pick(arr: T[], i: number): T { return arr[i % arr.length]!; } function makeHistory(i: number): string[] { const turns = 3 + (i % 10); const history = new Array(turns); for (let t = 0; t < turns; t++) { const topic = pick(TOPICS, i + t); history[t] = `Need guidance on ${topic}. Include practical steps and one small code example.`; } return history; } function makeInput(i: number): PromptInput { const topicA = pick(TOPICS, i); const topicB = pick(TOPICS, i + 3); const query = [ `Please compare ${topicA} with ${topicB}.`, "I care about cost per request and response quality.", "Give a short recommendation and a migration path.", ].join(" "); return { model: MODEL, systemPrefix: SYSTEM_PREFIX, history: makeHistory(i), query, maxInputTokens: MAX_INPUT_TOKENS, }; } type Totals = { rawTokens: number; budgetedTokens: number; staticTokens: number; dynamicTokens: number; trimmedRuns: number; queryTrimmedRuns: number; turnsDropped: number; }; function summarize(plans: PromptPlan[]): Totals { let totals: Totals = { rawTokens: 0, budgetedTokens: 0, staticTokens: 0, dynamicTokens: 0, trimmedRuns: 0, queryTrimmedRuns: 0, turnsDropped: 0, }; for (const plan of plans) { totals.rawTokens += plan.rawInputTokens; totals.budgetedTokens += plan.inputTokens; totals.staticTokens += plan.staticTokens; totals.dynamicTokens += plan.dynamicTokens; totals.turnsDropped += plan.trimmedTurns; if (plan.trimmedTurns > 0) totals.trimmedRuns++; if (plan.queryWasTrimmed) totals.queryTrimmedRuns++; } return totals; } function runHost(inputs: PromptInput[]): Totals { const plans = inputs.map((input) => preparePromptHost(input)); return summarize(plans); } async function runWorkers(inputs: PromptInput[]): Promise { const pool = createPool({ threads: THREADS })({ preparePrompt }); try { const jobs: Promise[] = []; for (let i = 0; i < inputs.length; i++) { jobs.push(pool.call.preparePrompt(inputs[i]!)); } const plans = await Promise.all(jobs); return summarize(plans); } finally { pool.shutdown(); } } function percent(saved: number, base: number): string { if (base <= 0) return "0.0%"; return `${((saved / base) * 100).toFixed(1)}%`; } async function main() { const inputs = new Array(REQUESTS); for (let i = 0; i < REQUESTS; i++) inputs[i] = makeInput(i); const started = performance.now(); const totals = MODE === "host" ? runHost(inputs) : await runWorkers(inputs); const finished = performance.now(); const tookMs = finished - started; const secs = Math.max(1e-9, tookMs / 1000); const reqPerSec = REQUESTS / secs; const savedTokens = Math.max(0, totals.rawTokens - totals.budgetedTokens); const cacheableTokensEstimate = totals.staticTokens; console.log("Prompt token budgeting"); console.log("mode :", MODE); console.log("model :", MODEL); console.log("threads :", MODE === "host" ? 0 : THREADS); console.log("requests :", REQUESTS.toLocaleString()); console.log("maxInputTokens :", MAX_INPUT_TOKENS.toLocaleString()); console.log("raw tokens :", totals.rawTokens.toLocaleString()); console.log("budgeted tokens :", totals.budgetedTokens.toLocaleString()); console.log( "saved tokens :", `${savedTokens.toLocaleString()} (${ percent(savedTokens, totals.rawTokens) })`, ); console.log("trimmed runs :", totals.trimmedRuns.toLocaleString()); console.log("query trimmed runs:", totals.queryTrimmedRuns.toLocaleString()); console.log("turns dropped :", totals.turnsDropped.toLocaleString()); console.log("cacheable estimate:", cacheableTokensEstimate.toLocaleString()); console.log("took :", tookMs.toFixed(2), "ms"); console.log("throughput :", reqPerSec.toFixed(0), "req/s"); } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `bench_prompt_token_budget.ts` ```ts import { createPool, isMain } from "knitting"; import { bench, boxplot, run, summary } from "mitata"; import { preparePromptBatchFast, preparePromptBatchFastHost, type PromptBudgetSummary, type PromptInput, } from "./token_budget.ts"; function intArg(name: string, fallback: number): number { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const value = Number(process.argv[i + 1]); if (Number.isFinite(value)) return Math.floor(value); } return fallback; } function strArg(name: string, fallback: string): string { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { return String(process.argv[i + 1]); } return fallback; } const THREADS = Math.max(1, intArg("threads", 2)); const REQUESTS = Math.max(1, intArg("requests", 10)); const MAX_INPUT_TOKENS = Math.max(64, intArg("maxInputTokens", 500)); const BATCH = Math.max(1, intArg("batch", 32)); const MODEL = strArg("model", "gpt-4o-mini"); const SYSTEM_PREFIX = [ "You are a docs assistant.", "Prefer concrete and short answers.", "If data is missing, say it directly.", "Do not invent unsupported behavior.", ].join("\n"); const TOPICS = [ "token budgeting", "prompt caching", "parallel workers", "schema validation", "rendering pipelines", "markdown output", "compression tradeoffs", "latency under load", ]; function pick(arr: T[], i: number): T { return arr[i % arr.length]!; } function makeHistory(i: number): string[] { const turns = 3 + (i % 10); const history = new Array(turns); for (let t = 0; t < turns; t++) { const topic = pick(TOPICS, i + t); history[t] = `Need guidance on ${topic}. Include practical steps and one small code example.`; } return history; } function makeInput(i: number): PromptInput { const topicA = pick(TOPICS, i); const topicB = pick(TOPICS, i + 3); const query = [ `Please compare ${topicA} with ${topicB}.`, "I care about cost per request and response quality.", "Give a short recommendation and a migration path.", ].join(" "); return { model: MODEL, systemPrefix: SYSTEM_PREFIX, history: makeHistory(i), query, maxInputTokens: MAX_INPUT_TOKENS, }; } function makeBatches(values: T[], batchSize: number): T[][] { const batches: T[][] = []; for (let i = 0; i < values.length; i += batchSize) { batches.push(values.slice(i, i + batchSize)); } return batches; } function mergeSummary( a: PromptBudgetSummary, b: PromptBudgetSummary, ): PromptBudgetSummary { return { rawTokens: a.rawTokens + b.rawTokens, budgetedTokens: a.budgetedTokens + b.budgetedTokens, staticTokens: a.staticTokens + b.staticTokens, dynamicTokens: a.dynamicTokens + b.dynamicTokens, trimmedRuns: a.trimmedRuns + b.trimmedRuns, queryTrimmedRuns: a.queryTrimmedRuns + b.queryTrimmedRuns, turnsDropped: a.turnsDropped + b.turnsDropped, }; } function runHostBatches(inputBatches: PromptInput[][]): PromptBudgetSummary { let totals: PromptBudgetSummary = { rawTokens: 0, budgetedTokens: 0, staticTokens: 0, dynamicTokens: 0, trimmedRuns: 0, queryTrimmedRuns: 0, turnsDropped: 0, }; for (let i = 0; i < inputBatches.length; i++) { totals = mergeSummary(totals, preparePromptBatchFastHost(inputBatches[i]!)); } return totals; } async function runWorkerBatches( callBatch: (inputs: PromptInput[]) => Promise, inputBatches: PromptInput[][], ): Promise { const jobs: Promise[] = []; for (let i = 0; i < inputBatches.length; i++) { jobs.push(callBatch(inputBatches[i]!)); } const results = await Promise.all(jobs); let totals: PromptBudgetSummary = { rawTokens: 0, budgetedTokens: 0, staticTokens: 0, dynamicTokens: 0, trimmedRuns: 0, queryTrimmedRuns: 0, turnsDropped: 0, }; for (let i = 0; i < results.length; i++) { totals = mergeSummary(totals, results[i]!); } return totals; } function sameSummary(a: PromptBudgetSummary, b: PromptBudgetSummary): boolean { return a.rawTokens === b.rawTokens && a.budgetedTokens === b.budgetedTokens && a.staticTokens === b.staticTokens && a.dynamicTokens === b.dynamicTokens && a.trimmedRuns === b.trimmedRuns && a.queryTrimmedRuns === b.queryTrimmedRuns && a.turnsDropped === b.turnsDropped; } async function main() { const inputs = new Array(REQUESTS); for (let i = 0; i < REQUESTS; i++) { inputs[i] = makeInput(i); } const inputBatches = makeBatches(inputs, BATCH); const pool = createPool({ threads: THREADS - 1, inliner: { batchSize: 8, }, })({ preparePromptBatchFast }); let sink = 0; try { const hostCheck = runHostBatches(inputBatches); const workerCheck = await runWorkerBatches( pool.call.preparePromptBatchFast, inputBatches, ); if (!sameSummary(hostCheck, workerCheck)) { throw new Error("Host and worker prompt-budget totals differ."); } console.log("Prompt token budgeting benchmark (mitata)"); console.log("workload: build prompt + tokenize + trim to budget"); console.log("model:", MODEL); console.log("requests per iteration:", REQUESTS.toLocaleString()); console.log("max input tokens:", MAX_INPUT_TOKENS.toLocaleString()); console.log("batch size:", BATCH); console.log("threads:", THREADS); boxplot(() => { summary(() => { bench(`host (${REQUESTS.toLocaleString()} req, batch ${BATCH})`, () => { const totals = runHostBatches(inputBatches); sink = totals.budgetedTokens; }); bench( `knitting (${THREADS} thread${ THREADS === 1 ? "" : "s" }, ${REQUESTS.toLocaleString()} req, batch ${BATCH})`, async () => { const totals = await runWorkerBatches( pool.call.preparePromptBatchFast, inputBatches, ); sink = totals.budgetedTokens; }, ); }); }); await run(); console.log("last budgeted tokens:", sink.toLocaleString()); } finally { pool.shutdown(); } } if (isMain) { main().catch((error) => { console.error(error); process.exitCode = 1; }); } ``` `token_budget.ts` ```ts import { task } from "knitting"; import { encoding_for_model } from "tiktoken"; export type PromptInput = { model: string; systemPrefix: string; history: string[]; query: string; maxInputTokens: number; }; export type PromptPlan = { prompt: string; rawInputTokens: number; inputTokens: number; staticTokens: number; dynamicTokens: number; trimmedTurns: number; queryWasTrimmed: boolean; }; export type PromptPlanFast = Omit; export type PromptBudgetSummary = { rawTokens: number; budgetedTokens: number; staticTokens: number; dynamicTokens: number; trimmedRuns: number; queryTrimmedRuns: number; turnsDropped: number; }; const decoder = new TextDecoder(); type Encoder = ReturnType; const MAX_ENCODER_CACHE = 4; const MAX_STATIC_TOKEN_CACHE = 512; const encoderCache = new Map(); const staticTokenCache = new Map(); function normalizeText(value: string): string { return value.replace(/\s+/g, " ").trim(); } function countTokens(enc: Encoder, text: string): number { return enc.encode(text).length; } function touchMapEntry(map: Map, key: string, value: V): void { map.delete(key); map.set(key, value); } function evictOldestEncoderIfNeeded(): void { if (encoderCache.size <= MAX_ENCODER_CACHE) return; const oldest = encoderCache.keys().next().value; if (oldest === undefined) return; const enc = encoderCache.get(oldest); if (enc) enc.free(); encoderCache.delete(oldest); } function evictOldestStaticTokenIfNeeded(): void { if (staticTokenCache.size <= MAX_STATIC_TOKEN_CACHE) return; const oldest = staticTokenCache.keys().next().value; if (oldest !== undefined) staticTokenCache.delete(oldest); } function getEncoder(model: string): Encoder { const cached = encoderCache.get(model); if (cached) { touchMapEntry(encoderCache, model, cached); return cached; } const enc = encoding_for_model(model as never); encoderCache.set(model, enc); evictOldestEncoderIfNeeded(); return enc; } function getStaticTokens( model: string, systemPrefix: string, enc: Encoder, ): number { const key = `${model}\x1f${systemPrefix}`; const cached = staticTokenCache.get(key); if (cached !== undefined) { touchMapEntry(staticTokenCache, key, cached); return cached; } const value = countTokens(enc, systemPrefix); staticTokenCache.set(key, value); evictOldestStaticTokenIfNeeded(); return value; } export function clearPromptBudgetCaches(): void { for (const enc of encoderCache.values()) enc.free(); encoderCache.clear(); staticTokenCache.clear(); } function truncateToTokenBudget( enc: Encoder, text: string, maxTokens: number, ): string { if (maxTokens <= 0) return ""; const tokens = enc.encode(text); if (tokens.length <= maxTokens) return text; const clipped = tokens.slice(0, maxTokens); return decoder.decode(enc.decode(clipped)); } function buildPrompt( systemPrefix: string, history: string[], query: string, ): string { const rows: string[] = []; rows.push(systemPrefix.trim()); rows.push(""); rows.push("Conversation context:"); for (let i = 0; i < history.length; i++) { rows.push(`- Turn ${i + 1}: ${history[i]}`); } rows.push(""); rows.push(`User request: ${query}`); return rows.join("\n"); } export function preparePromptHost(input: PromptInput): PromptPlan { const model = input.model; const maxInputTokens = Math.max(64, input.maxInputTokens); const cleanHistory = input.history.map(normalizeText).filter(Boolean); let history = [...cleanHistory]; let query = normalizeText(input.query); const enc = getEncoder(model); const staticTokens = getStaticTokens(model, input.systemPrefix, enc); let prompt = buildPrompt(input.systemPrefix, history, query); const rawInputTokens = countTokens(enc, prompt); let inputTokens = rawInputTokens; let trimmedTurns = 0; let queryWasTrimmed = false; while (inputTokens > maxInputTokens && history.length > 0) { history.shift(); trimmedTurns++; prompt = buildPrompt(input.systemPrefix, history, query); inputTokens = countTokens(enc, prompt); } if (inputTokens > maxInputTokens) { const promptWithoutQuery = buildPrompt(input.systemPrefix, history, ""); const promptWithoutQueryTokens = countTokens(enc, promptWithoutQuery); const remainingBudget = Math.max( 16, maxInputTokens - promptWithoutQueryTokens, ); const clipped = truncateToTokenBudget(enc, query, remainingBudget); queryWasTrimmed = clipped.length < query.length; query = clipped; prompt = buildPrompt(input.systemPrefix, history, query); inputTokens = countTokens(enc, prompt); } return { prompt, rawInputTokens, inputTokens, staticTokens, dynamicTokens: Math.max(0, inputTokens - staticTokens), trimmedTurns, queryWasTrimmed, }; } export const preparePrompt = task({ f: (input) => preparePromptHost(input), }); export function preparePromptFastHost(input: PromptInput): PromptPlanFast { const plan = preparePromptHost(input); return { rawInputTokens: plan.rawInputTokens, inputTokens: plan.inputTokens, staticTokens: plan.staticTokens, dynamicTokens: plan.dynamicTokens, trimmedTurns: plan.trimmedTurns, queryWasTrimmed: plan.queryWasTrimmed, }; } export function preparePromptBatchFastHost( inputs: PromptInput[], ): PromptBudgetSummary { let totals: PromptBudgetSummary = { rawTokens: 0, budgetedTokens: 0, staticTokens: 0, dynamicTokens: 0, trimmedRuns: 0, queryTrimmedRuns: 0, turnsDropped: 0, }; for (let i = 0; i < inputs.length; i++) { const plan = preparePromptFastHost(inputs[i]!); totals.rawTokens += plan.rawInputTokens; totals.budgetedTokens += plan.inputTokens; totals.staticTokens += plan.staticTokens; totals.dynamicTokens += plan.dynamicTokens; totals.turnsDropped += plan.trimmedTurns; if (plan.trimmedTurns > 0) totals.trimmedRuns++; if (plan.queryWasTrimmed) totals.queryTrimmedRuns++; } return totals; } export const preparePromptBatchFast = task({ f: (inputs) => preparePromptBatchFastHost(inputs), }); ``` ## When this matters Token budgeting is a preflight step that runs on every LLM request. If you're handling high-throughput chat traffic -- multiple users, long conversation histories -- the tokenization and trimming work adds up. Offloading it to workers keeps your main thread focused on routing and I/O while budget calculations happen in parallel. It also gives you predictable input sizes, which helps with cost control and latency. --- # TSP (GSA) URL: https://knittingdocs.vercel.app/examples/maths/tsp_gsa/ Parallel heuristic restarts for a gravity-inspired TSP solver. The **Traveling Salesman Problem**: given N cities, find the shortest tour that visits each one exactly once and returns to the start. This is NP-hard, so we use heuristics -- specifically, a gravity-inspired population search (GSA) followed by 2-opt local refinement, run as many parallel restarts through Knitting. More restarts = better chance of finding a good solution. This is the most computationally intensive example in the set, and it shows what Knitting looks like on a real optimization workload. ## How it works 1. The host generates (or seeds) a set of cities and a distance matrix. 2. The host launches many **independent solver runs** ("restarts") through the worker pool. 3. Each worker runs a gravity-inspired search to produce a candidate tour, then applies 2-opt to sharpen it. 4. Each worker returns `{ bestLen, bestTour }`. 5. The host picks the global best, validates the tour, and recomputes the length for correctness. The key insight: a single heuristic run can get stuck in a local minimum. Running many independent trials in parallel increases the chance that at least one finds a better region of the search space. This is **embarrassingly parallel** -- restarts are independent, results are small, and the host just picks the best. ## Run Expected output: ``` cities: 64 restarts: 64 threads: 7 pop: 10 iters: 10 dispatching 64 restarts... best tour length: 847.32 tour valid: OK (64 cities, no duplicates, all present) recomputed length: 847.32 OK random baseline: 2,341.07 (solver is 2.76x better than random) elapsed: 2.14s ``` ## Code `run_tsp.ts` ```ts import { createPool, isMain } from "knitting"; import { solveTspGsa } from "./tsp_gsa.ts"; function intArg(name: string, fallback: number) { const i = process.argv.indexOf(`--${name}`); if (i !== -1 && i + 1 < process.argv.length) { const v = Number(process.argv[i + 1]); if (Number.isFinite(v) && v > 0) return Math.floor(v); } return fallback; } const THREADS = intArg("threads", 7); const RESTARTS = intArg("restarts", 64); const N = intArg("cities", 64); const POP = intArg("pop", 10); const ITERS = intArg("iters", 10); const worldSeed = intArg("worldSeed", 123456); const seedBase = (Date.now() | 0) ^ 0x9e3779b9; const { call, shutdown } = createPool({ threads: THREADS, balancer: "firstIdle", })({ solveTspGsa }); function validateTour(tour: number[], n: number) { if (tour.length !== n) throw new Error(`tour length ${tour.length} != ${n}`); const seen = new Uint8Array(n); for (const v of tour) { if ((v | 0) !== v) throw new Error(`non-int city id: ${v}`); if (v < 0 || v >= n) throw new Error(`bad city id: ${v}`); if (seen[v]) throw new Error(`duplicate city: ${v}`); seen[v] = 1; } } /* ---------- Host recompute must match worker exactly ---------- */ function xorshift32(s: number): number { s |= 0; s ^= s << 13; s ^= s >>> 17; s ^= s << 5; return s | 0; } const INV_U32 = 2.3283064365386963e-10; // 1 / 2^32 function makeCities(worldSeed: number, n: number): Float64Array { // coords: [x0,y0,x1,y1,...] in [0,1) const coords = new Float64Array(n * 2); let s = worldSeed | 0; for (let i = 0; i < n; i++) { s = xorshift32(s); coords[i * 2 + 0] = (s >>> 0) * INV_U32; s = xorshift32(s); coords[i * 2 + 1] = (s >>> 0) * INV_U32; } return coords; } function makeDistMatrix(coords: Float64Array, n: number): Float32Array { const d = new Float32Array(n * n); for (let i = 0; i < n; i++) { const xi = coords[i * 2 + 0]; const yi = coords[i * 2 + 1]; for (let j = i + 1; j < n; j++) { const dx = xi - coords[j * 2 + 0]; const dy = yi - coords[j * 2 + 1]; const dist = Math.hypot(dx, dy); d[i * n + j] = dist; d[j * n + i] = dist; } } return d; } function recomputeLen(tour: number[], dist: Float32Array, n: number): number { let s = 0; let prev = tour[0]; for (let i = 1; i < n; i++) { const cur = tour[i]; s += dist[prev * n + cur]; prev = cur; } s += dist[prev * n + tour[0]]; return s; } /* ---------- Random pick + random baseline tour ---------- */ function pickRandomIndex(len: number, seed: number): number { // deterministic “random” based on seed const s = xorshift32(seed | 0); return (s >>> 0) % len; } function makeRandomTour(n: number, seed: number): number[] { // Fisher–Yates shuffle const tour = new Array(n); for (let i = 0; i < n; i++) tour[i] = i; let s = seed | 0; for (let i = n - 1; i > 0; i--) { s = xorshift32(s); const j = (s >>> 0) % (i + 1); const tmp = tour[i]; tour[i] = tour[j]; tour[j] = tmp; } return tour; } /* ------------------------------------------------------------ */ async function main() { const jobs: Promise<{ bestLen: number; bestTour: number[] }>[] = []; for (let r = 0; r < RESTARTS; r++) { const runSeed = (seedBase + r * 0x6d2b79f5) | 0; jobs.push(call.solveTspGsa([worldSeed, runSeed, N, POP, ITERS])); } const results = await Promise.all(jobs); if (results.length === 0) throw new Error("no results (unexpected)"); // Find best + worst correctly let best = results[0]; let worst = results[0]; for (const res of results) { if (res.bestLen < best.bestLen) best = res; if (res.bestLen > worst.bestLen) worst = res; } // Build world once on host for verification const coords = makeCities(worldSeed, N); const dist = makeDistMatrix(coords, N); function checkResult( label: string, res: { bestLen: number; bestTour: number[] }, ) { validateTour(res.bestTour, N); const recomputed = recomputeLen(res.bestTour, dist, N); const delta = recomputed - res.bestLen; if (!Number.isFinite(res.bestLen) || res.bestLen < 0) { throw new Error(`${label}: bestLen invalid: ${res.bestLen}`); } if (Math.abs(delta) > 1e-6) { throw new Error( `${label}: length mismatch (delta=${delta}). Host generator != worker generator?`, ); } return { recomputed, delta }; } // Randomly choose ONE run result and verify it too (not just the best) const randIdx = pickRandomIndex(results.length, seedBase ^ 0xA5A5A5A5); const randomRes = results[randIdx]; const bestCheck = checkResult("best", best); const worstCheck = checkResult("worst", worst); const randomCheck = checkResult(`randomRun[#${randIdx}]`, randomRes); // Random baseline tour (not from solver) const randomTour = makeRandomTour(N, seedBase ^ 0xC0FFEE); validateTour(randomTour, N); const randomLen = recomputeLen(randomTour, dist, N); console.log("TSP via gravity (GSA) + 2-opt"); console.log("threads :", THREADS); console.log("restarts :", RESTARTS); console.log("cities :", N); console.log("pop :", POP); console.log("iters :", ITERS); console.log("worldSeed :", worldSeed); console.log("---"); console.log("bestLen :", best.bestLen); console.log("worstLen :", worst.bestLen); console.log(`randomRunLen :`, randomRes.bestLen, `(picked index ${randIdx})`); console.log("randomTourLen:", randomLen); console.log("---"); console.log("best delta :", bestCheck.delta); console.log("worst delta :", worstCheck.delta); console.log("random delta :", randomCheck.delta); console.log( "tour head :", best.bestTour.slice(0, Math.min(16, best.bestTour.length)).join(", "), "...", ); } if (isMain) { main().finally(shutdown); } ``` `tsp_gsa.ts` ```ts import { task } from "knitting"; type Args = readonly [ worldSeed: number, // generates the same city map for all runs runSeed: number, // controls the optimizer randomness nCities: number, popSize: number, iters: number, ]; type Result = { bestLen: number; bestTour: number[]; }; function xorshift32(s: number): number { s |= 0; s ^= s << 13; s ^= s >>> 17; s ^= s << 5; return s | 0; } const INV_U32 = 2.3283064365386963e-10; // 1 / 2^32 function rand01(stateRef: { s: number }): number { stateRef.s = xorshift32(stateRef.s); return (stateRef.s >>> 0) * INV_U32; } function makeCities(worldSeed: number, n: number): Float64Array { // coords: [x0,y0,x1,y1,...] in [0,1) const coords = new Float64Array(n * 2); const st = { s: worldSeed | 0 }; for (let i = 0; i < n; i++) { coords[i * 2 + 0] = rand01(st); coords[i * 2 + 1] = rand01(st); } return coords; } function makeDistMatrix(coords: Float64Array, n: number): Float32Array { const d = new Float32Array(n * n); for (let i = 0; i < n; i++) { const xi = coords[i * 2 + 0]; const yi = coords[i * 2 + 1]; for (let j = i + 1; j < n; j++) { const dx = xi - coords[j * 2 + 0]; const dy = yi - coords[j * 2 + 1]; const dist = Math.hypot(dx, dy); d[i * n + j] = dist; d[j * n + i] = dist; } } return d; } function tourLen(dist: Float32Array, n: number, tour: Int32Array): number { let sum = 0; let prev = tour[0]; for (let i = 1; i < n; i++) { const cur = tour[i]; sum += dist[prev * n + cur]; prev = cur; } sum += dist[prev * n + tour[0]]; return sum; } function decodeKeysToTour( keys: Float64Array, n: number, scratchIdx: number[], outTour: Int32Array, ) { // scratchIdx contains 0..n-1 and is reused scratchIdx.sort((a, b) => keys[a] - keys[b]); for (let i = 0; i < n; i++) outTour[i] = scratchIdx[i]; } const eps = 1e-12; function twoOpt(dist: Float32Array, n: number, tour: Int32Array): number { let best = tourLen(dist, n, tour); while (true) { let improved = false; outer: for (let i = 0; i < n - 1; i++) { for (let k = i + 2; k < n; k++) { const a = tour[i]; const b = tour[(i + 1) % n]; const c = tour[k]; const d = tour[(k + 1) % n]; const before = dist[a * n + b] + dist[c * n + d]; const after = dist[a * n + c] + dist[b * n + d]; if (after + eps < before) { // reverse segment (i+1..k) for (let l = i + 1, r = k; l < r; l++, r--) { const tmp = tour[l]; tour[l] = tour[r]; tour[r] = tmp; } // delta update is valid because we restart scanning immediately best += after - before; improved = true; break outer; } } } if (!improved) break; } // Safety: compute the true length once (guaranteed non-negative if dist is) return tourLen(dist, n, tour); } export const solveTspGsa = task({ f: ([worldSeed, runSeed, nCities, popSize, iters]) => { const n = nCities | 0; const pop = popSize | 0; const T = iters | 0; const coords = makeCities(worldSeed | 0, n); const dist = makeDistMatrix(coords, n); // Agent states const X = new Float64Array(pop * n); const V = new Float64Array(pop * n); const fit = new Float64Array(pop); const mass = new Float64Array(pop); const st = { s: runSeed | 0 }; // Init positions and velocities for (let i = 0; i < pop * n; i++) { X[i] = rand01(st); // [0,1) V[i] = (rand01(st) - 0.5) * 0.1; // small initial velocity } const scratchIdx: number[] = new Array(n); for (let i = 0; i < n; i++) scratchIdx[i] = i; const tmpTour = new Int32Array(n); const bestTour = new Int32Array(n); let bestLen = Infinity; // Helpers const idxPop: number[] = new Array(pop); for (let i = 0; i < pop; i++) idxPop[i] = i; const eps = 1e-9; const G0 = 100.0; const alpha = 20.0; // Main loop for (let t = 0; t < T; t++) { // Evaluate fitness (tour length) for (let i = 0; i < pop; i++) { const base = i * n; decodeKeysToTour(X.subarray(base, base + n), n, scratchIdx, tmpTour); const L = tourLen(dist, n, tmpTour); fit[i] = L; if (L < bestLen) { bestLen = L; bestTour.set(tmpTour); } } // Sort agents by fitness (ascending) idxPop.sort((a, b) => fit[a] - fit[b]); const bestF = fit[idxPop[0]]; const worstF = fit[idxPop[pop - 1]]; const denom = Math.max(eps, worstF - bestF); // Mass for minimization: better fitness => larger mass let sumM = 0; for (let r = 0; r < pop; r++) { const i = idxPop[r]; const m = (worstF - fit[i]) / denom; mass[i] = m; sumM += m; } const invSumM = 1 / Math.max(eps, sumM); for (let i = 0; i < pop; i++) mass[i] *= invSumM; // K-best shrinks over time const K = Math.max(2, (pop * (1 - t / T)) | 0); const G = G0 * Math.exp(-alpha * (t / T)); // Update each agent via gravitational attraction for (let ii = 0; ii < pop; ii++) { const i = idxPop[ii]; const Mi = Math.max(eps, mass[i]); const baseI = i * n; for (let d = 0; d < n; d++) { let Fi = 0; // Pull from top-K agents for (let kk = 0; kk < K; kk++) { const j = idxPop[kk]; if (j === i) continue; // Distance between agent vectors (cheap L2) const baseJ = j * n; let r2 = 0; for (let q = 0; q < n; q++) { const diff = X[baseJ + q] - X[baseI + q]; r2 += diff * diff; } const R = Math.sqrt(r2) + eps; const Mj = mass[j]; const rij = X[baseJ + d] - X[baseI + d]; // random factor to avoid lockstep collapse Fi += rand01(st) * G * (Mi * Mj) * (rij / R); } // a = F / Mi const a = Fi / Mi; // velocity + position update const idx = baseI + d; V[idx] = rand01(st) * V[idx] + a; X[idx] = X[idx] + V[idx]; // keep keys in a reasonable range if (X[idx] < -2) X[idx] = -2; else if (X[idx] > 3) X[idx] = 3; } } } // Local refinement: 2-opt on best tour const refined = bestTour.slice() as Int32Array; const refinedLen = twoOpt(dist, n, refined); if (refinedLen < bestLen) bestLen = refinedLen; // Return as plain JS array for safe payload compatibility const out: number[] = new Array(n); for (let i = 0; i < n; i++) out[i] = refined[i]; return { bestLen, bestTour: out }; }, }); ``` ## How the algorithm works **Representation:** TSP needs a permutation of cities. Each agent stores a real-valued vector -- sorting the keys produces the permutation. This lets "continuous" gravity-like motion work on a discrete problem. **Global search (GSA):** Each agent has a mass proportional to its tour quality. Better tours = higher mass. Agents attract each other like gravity, gradually pulling the population toward better solutions. **Local refinement (2-opt):** After global search, pick two edges, reverse a segment if it shortens the tour, repeat until no improvement. Fast and usually provides large gains -- often the difference between "random" and "competitive." ## Correctness checks This example doesn't just trust the output. It validates: 1. **Tour validity** -- length is N, all integers, no duplicates, all cities present. 2. **Length recomputation** -- recomputes on the host using the same distance matrix, catching bugs like delta-update drift or corrupted permutations. 3. **Random baseline** -- compares against a random tour to confirm the solver is actually finding structure, not returning noise. ## CLI knobs - `--cities` -- increases difficulty sharply (N! possible tours) - `--restarts` -- more restarts = better chance of finding a good tour (cheap parallel wins) - `--iters` -- more iterations per run = deeper exploration per restart - `--pop` -- population size per run = more exploration (but more compute) - `--worldSeed` -- fix the city layout for reproducible comparisons Start small: `cities=32`, `restarts=threads*4`, `iters=200`. Increase quality via restarts first -- that's the cheapest way to improve results. ## Real-world analogs TSP shows up in delivery routing, manufacturing toolpath optimization, PCB drilling, robotics path planning, and scheduling problems. The parallel-restart pattern works for any metaheuristic where independent runs are cheap and you just need the best result. ## Things to try 1. Compare `--restarts` vs `--iters`: which improves quality faster per second of compute? 2. Increase `--cities` to 128 and watch how the search space explodes. 3. Lower `--pop` to 2-3 and see how much worse it gets (and how much faster). 4. Change the distance metric to Manhattan distance and compare behavior. --- # Benchmarks URL: https://knittingdocs.vercel.app/benchmarks/introduction/ Knitting benchmark results across Node.js, Deno, and Bun, with the hardware and methodology behind them. This page gives you the short version of Knitting's benchmark results across Node, Deno, Bun, and Tokio. The tests focus on worker communication, payload sizes, batching, and CPU-heavy work. They are useful for comparing shapes of workloads—not for predicting exactly how fast your application will be. ## The short version - Small calls have low overhead. - Batching improves throughput, especially as the number of calls grows. - Large binary payloads are mainly limited by copying and memory bandwidth. - CPU-heavy work scales across workers, but the result depends on the runtime and the number of threads. When the charts say one option is “faster,” that only applies to the test shown: the payload, batching, runtime, and amount of work all matter. ## Message overhead This chart compares Knitting with worker `postMessage`, WebSocket, and HTTP. In these runs, Knitting is roughly: - `3.5–6×` faster than worker `postMessage`; - `3.5–15×` faster than WebSocket; - `10–57×` faster than HTTP. ## Latency as calls grow This chart shows what happens as each test sends more messages at a time. Knitting is generally around `3–45×` faster than the worker baselines in these runs. The advantage is clearest in small and medium batches, where communication overhead makes up more of the total time. ## Different payload types These charts compare one value at a time with batches of `100` values. Structured and binary values cost more than primitives, and large objects and errors cost more again. Batching improves throughput while keeping Knitting's relative advantage in these tests. ## Moving a 1 MiB payload These are one-way transfer results for a `1 MiB` payload with a batch size of `64`: | Runtime | String (GB/s) | Uint8Array (GB/s) | | --- | ---: | ---: | | Node | `1.44` | `7.50` | | Deno | `3.37` | `5.78` | | Bun | `11.86` | `16.21` | Bun is fastest in this particular 1 MiB test for both strings and binary data. At this size, memory bandwidth and runtime details matter as much as the worker coordination itself. ## CPU-heavy work This test distributes a CPU-intensive prime-number workload across extra threads. In these runs, speedup reaches roughly `3.5–3.8×` with four extra threads, while efficiency stays around `70–77%`. That does not mean every application will scale the same way. I/O-heavy work usually behaves very differently. ## Explore the details This page is for the broad trends. For raw tables and runtime-specific notes, see the dedicated [Node](/benchmarks/node/), [Deno](/benchmarks/deno/), [Bun](/benchmarks/bun/), and [Tokio](/benchmarks/tokio/) pages. ## Run the benchmarks ```bash ./run.sh ``` Results are written to `results/`. To write JSON for plotting or other scripts: ```bash ./run.sh --json ``` JSON output is useful for plotting scripts under `graphs/`. --- # Node.js URL: https://knittingdocs.vercel.app/benchmarks/node/ Knitting on Node.js: message overhead, a head-to-head with a plain worker, latency as calls grow, heavy-task efficiency, and payload type costs. This page summarizes Node.js benchmark runs for Knitting on `node 24.12.0 (arm64-darwin)`. ## IPC (Node) This benchmark compares one round-trip between a main thread and workers using different transports. Knitting has the lowest overhead in this setup: - `1` message: Knitting is about `6x` faster than worker `postMessage`, `15x` faster than websocket, and `57x` faster than HTTP. - `25` messages: Knitting is about `4x` faster than worker `postMessage`. - `50` messages: Knitting is about `3.5x` faster than worker `postMessage`. ### Data `node_ipc.md` ```md clk: ~3.67 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • knitting | avg | min | p75 | p99 | max | | --------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 1 thread → (1) | ` 2.22 µs/iter` | ` 1.12 µs` | ` 2.63 µs` | ` 5.54 µs` | ` 1.27 ms` | | 1 thread → (25) | ` 15.99 µs/iter` | ` 14.33 µs` | ` 16.56 µs` | ` 17.15 µs` | ` 19.87 µs` | | 1 thread → (50) | ` 30.67 µs/iter` | ` 28.07 µs` | ` 31.19 µs` | ` 32.16 µs` | ` 33.21 µs` | clk: ~3.71 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • websocket | avg | min | p75 | p99 | max | | ------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | local → (1) | ` 26.83 µs/iter` | ` 12.83 µs` | ` 26.96 µs` | ` 85.25 µs` | `468.58 µs` | | local → (25) | `183.39 µs/iter` | `130.33 µs` | `195.83 µs` | `272.71 µs` | `976.96 µs` | | local → (50) | `345.79 µs/iter` | `252.42 µs` | `363.46 µs` | `456.25 µs` | ` 1.45 ms` | clk: ~3.70 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • worker | avg | min | p75 | p99 | max | | ------------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | postMessage → (1) | ` 11.49 µs/iter` | ` 7.29 µs` | ` 12.00 µs` | ` 33.21 µs` | `344.50 µs` | | postMessage → (25) | ` 56.73 µs/iter` | ` 55.91 µs` | ` 56.83 µs` | ` 57.24 µs` | ` 57.76 µs` | | postMessage → (50) | `102.10 µs/iter` | ` 68.83 µs` | `105.63 µs` | `221.17 µs` | ` 3.49 ms` | clk: ~3.68 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • http | avg | min | p75 | p99 | max | | ------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | local → (1) | ` 65.42 µs/iter` | ` 38.42 µs` | ` 66.13 µs` | `174.38 µs` | `601.33 µs` | | local → (25) | `966.49 µs/iter` | `859.96 µs` | ` 1.01 ms` | ` 1.22 ms` | ` 2.04 ms` | | local → (50) | ` 1.90 ms/iter` | ` 1.71 ms` | ` 1.99 ms` | ` 2.17 ms` | ` 2.64 ms` | ``` ## Knitting vs Worker (Node) These charts compare the same payload families sent through Knitting and `worker_threads`. ### One message With a single value per call, Knitting is consistently faster: - For small primitives, Knitting is roughly `10-14x` faster than worker `postMessage`. - Worker `postMessage` stays near `~10 µs/iter` even for tiny payloads. - For larger payloads, Knitting still holds around a `4x` advantage (for example, big object: `4.15 µs` vs `15.74 µs`). `node_types.md` ```md clk: ~3.71 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • knitting 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 1.19 µs/iter` | `500.00 ns` | ` 1.54 µs` | ` 2.88 µs` | ` 1.33 ms` | | bigint small -> (1) | ` 1.05 µs/iter` | `551.27 ns` | ` 1.39 µs` | ` 1.65 µs` | ` 1.67 µs` | | bigint large -> (1) | ` 2.36 µs/iter` | ` 1.29 µs` | ` 2.54 µs` | ` 5.83 µs` | ` 1.24 ms` | | boolean true -> (1) | ` 1.10 µs/iter` | `534.90 ns` | ` 1.46 µs` | ` 1.59 µs` | ` 1.61 µs` | | boolean false -> (1) | ` 1.06 µs/iter` | `531.85 ns` | ` 1.49 µs` | ` 1.61 µs` | ` 1.62 µs` | | undefined -> (1) | ` 1.05 µs/iter` | `531.94 ns` | ` 1.44 µs` | ` 1.60 µs` | ` 1.65 µs` | | null -> (1) | ` 1.03 µs/iter` | `534.34 ns` | ` 1.35 µs` | ` 1.67 µs` | ` 1.90 µs` | | string -> (1) | ` 1.57 µs/iter` | `583.00 ns` | ` 2.13 µs` | ` 3.29 µs` | ` 1.26 ms` | | json object -> (1) | ` 3.64 µs/iter` | ` 2.50 µs` | ` 3.75 µs` | ` 8.67 µs` | ` 1.27 ms` | | json array -> (1) | ` 4.73 µs/iter` | ` 4.28 µs` | ` 4.90 µs` | ` 5.21 µs` | ` 5.22 µs` | | Uint8Array -> (1) | ` 3.11 µs/iter` | ` 1.17 µs` | ` 3.83 µs` | ` 8.08 µs` | ` 1.25 ms` | | ArrayBuffer -> (1) | ` 3.24 µs/iter` | ` 1.50 µs` | ` 4.21 µs` | ` 8.50 µs` | ` 1.27 ms` | | Buffer -> (1) | ` 2.87 µs/iter` | ` 1.17 µs` | ` 3.83 µs` | ` 7.25 µs` | ` 1.24 ms` | | string huge -> (1) | ` 3.73 µs/iter` | ` 2.00 µs` | ` 4.38 µs` | ` 8.75 µs` | ` 1.25 ms` | | Int32Array -> (1) | ` 2.84 µs/iter` | ` 1.80 µs` | ` 3.32 µs` | ` 3.78 µs` | ` 3.86 µs` | | Float64Array -> (1) | ` 3.37 µs/iter` | ` 1.46 µs` | ` 4.17 µs` | ` 8.79 µs` | ` 1.27 ms` | | BigInt64Array -> (1) | ` 2.70 µs/iter` | ` 1.77 µs` | ` 3.19 µs` | ` 3.69 µs` | ` 3.87 µs` | | BigUint64Array -> (1) | ` 3.03 µs/iter` | ` 1.21 µs` | ` 3.83 µs` | ` 7.58 µs` | ` 1.26 ms` | | DataView -> (1) | ` 2.89 µs/iter` | ` 1.87 µs` | ` 3.37 µs` | ` 3.77 µs` | ` 3.85 µs` | | Date -> (1) | ` 1.34 µs/iter` | `542.00 ns` | ` 1.63 µs` | ` 2.71 µs` | ` 1.24 ms` | | • knitting 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | ` 27.35 µs/iter` | ` 15.42 µs` | ` 33.79 µs` | ` 60.38 µs` | `222.08 µs` | | bigint small -> (100) | ` 26.77 µs/iter` | ` 24.24 µs` | ` 28.06 µs` | ` 28.61 µs` | ` 31.03 µs` | | bigint large -> (100) | ` 72.11 µs/iter` | ` 53.50 µs` | ` 73.83 µs` | `137.13 µs` | ` 1.35 ms` | | boolean true -> (100) | ` 28.12 µs/iter` | ` 25.30 µs` | ` 29.23 µs` | ` 31.10 µs` | ` 35.62 µs` | | boolean false -> (100) | ` 27.33 µs/iter` | ` 20.98 µs` | ` 29.51 µs` | ` 29.70 µs` | ` 32.56 µs` | | undefined -> (100) | ` 29.03 µs/iter` | ` 22.39 µs` | ` 31.83 µs` | ` 32.06 µs` | ` 32.95 µs` | | null -> (100) | ` 28.28 µs/iter` | ` 25.01 µs` | ` 29.89 µs` | ` 30.11 µs` | ` 30.44 µs` | | string -> (100) | ` 37.10 µs/iter` | ` 33.71 µs` | ` 38.45 µs` | ` 39.78 µs` | ` 40.58 µs` | | json object -> (100) | `125.75 µs/iter` | `106.08 µs` | `128.67 µs` | `216.33 µs` | ` 1.42 ms` | | json array -> (100) | `178.12 µs/iter` | `154.33 µs` | `183.25 µs` | `290.88 µs` | ` 1.55 ms` | | Uint8Array -> (100) | `108.89 µs/iter` | ` 54.00 µs` | `137.63 µs` | `209.29 µs` | ` 2.93 ms` | | ArrayBuffer -> (100) | `119.46 µs/iter` | ` 63.50 µs` | `145.71 µs` | `222.29 µs` | ` 2.59 ms` | | Buffer -> (100) | `103.74 µs/iter` | ` 54.42 µs` | `143.88 µs` | `207.08 µs` | ` 2.88 ms` | | string huge -> (100) | `167.81 µs/iter` | `111.54 µs` | `208.21 µs` | `265.50 µs` | ` 1.44 ms` | | Int32Array -> (100) | `116.92 µs/iter` | ` 55.12 µs` | `143.83 µs` | `214.25 µs` | ` 3.00 ms` | | Float64Array -> (100) | `120.60 µs/iter` | ` 61.96 µs` | `139.67 µs` | `224.54 µs` | ` 2.08 ms` | | BigInt64Array -> (100) | `107.85 µs/iter` | `102.95 µs` | `109.66 µs` | `111.51 µs` | `113.46 µs` | | BigUint64Array -> (100) | `112.46 µs/iter` | ` 56.50 µs` | `139.96 µs` | `207.25 µs` | ` 3.15 ms` | | DataView -> (100) | `109.26 µs/iter` | ` 59.25 µs` | `140.21 µs` | `214.38 µs` | ` 3.14 ms` | | Date -> (100) | ` 33.41 µs/iter` | ` 31.27 µs` | ` 33.97 µs` | ` 35.82 µs` | ` 36.27 µs` | | • worker 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 11.13 µs/iter` | ` 11.09 µs` | ` 11.15 µs` | ` 11.18 µs` | ` 11.19 µs` | | bigint small -> (1) | ` 11.09 µs/iter` | ` 10.99 µs` | ` 11.11 µs` | ` 11.13 µs` | ` 11.15 µs` | | bigint large -> (1) | ` 11.06 µs/iter` | ` 10.99 µs` | ` 11.07 µs` | ` 11.10 µs` | ` 11.12 µs` | | boolean true -> (1) | ` 11.09 µs/iter` | ` 11.02 µs` | ` 11.11 µs` | ` 11.14 µs` | ` 11.18 µs` | | boolean false -> (1) | ` 11.07 µs/iter` | ` 11.02 µs` | ` 11.10 µs` | ` 11.10 µs` | ` 11.11 µs` | | undefined -> (1) | ` 11.05 µs/iter` | ` 11.03 µs` | ` 11.05 µs` | ` 11.07 µs` | ` 11.09 µs` | | null -> (1) | ` 11.09 µs/iter` | ` 11.03 µs` | ` 11.11 µs` | ` 11.13 µs` | ` 11.13 µs` | | string -> (1) | ` 11.09 µs/iter` | ` 11.04 µs` | ` 11.12 µs` | ` 11.13 µs` | ` 11.15 µs` | | json object -> (1) | ` 14.88 µs/iter` | ` 14.63 µs` | ` 14.86 µs` | ` 14.93 µs` | ` 15.82 µs` | | json array -> (1) | ` 18.70 µs/iter` | ` 18.45 µs` | ` 18.65 µs` | ` 18.79 µs` | ` 19.73 µs` | | Uint8Array -> (1) | ` 12.09 µs/iter` | ` 11.91 µs` | ` 12.13 µs` | ` 12.19 µs` | ` 12.35 µs` | | ArrayBuffer -> (1) | ` 11.85 µs/iter` | ` 11.67 µs` | ` 11.90 µs` | ` 12.00 µs` | ` 12.10 µs` | | Buffer -> (1) | ` 14.02 µs/iter` | ` 13.70 µs` | ` 14.18 µs` | ` 14.21 µs` | ` 14.29 µs` | | string huge -> (1) | ` 11.23 µs/iter` | ` 11.17 µs` | ` 11.27 µs` | ` 11.34 µs` | ` 11.34 µs` | | Int32Array -> (1) | ` 12.08 µs/iter` | ` 11.94 µs` | ` 12.11 µs` | ` 12.31 µs` | ` 12.34 µs` | | Float64Array -> (1) | ` 12.12 µs/iter` | ` 11.99 µs` | ` 12.21 µs` | ` 12.25 µs` | ` 12.30 µs` | | BigInt64Array -> (1) | ` 12.11 µs/iter` | ` 11.91 µs` | ` 12.17 µs` | ` 12.33 µs` | ` 12.41 µs` | | BigUint64Array -> (1) | ` 12.10 µs/iter` | ` 11.95 µs` | ` 12.18 µs` | ` 12.23 µs` | ` 12.29 µs` | | DataView -> (1) | ` 12.67 µs/iter` | ` 12.02 µs` | ` 13.01 µs` | ` 13.15 µs` | ` 13.17 µs` | | Date -> (1) | ` 11.41 µs/iter` | ` 5.92 µs` | ` 12.00 µs` | ` 32.67 µs` | `168.46 µs` | | • worker 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | `173.42 µs/iter` | `121.71 µs` | `177.04 µs` | `333.04 µs` | ` 3.54 ms` | | bigint small -> (100) | `178.73 µs/iter` | `122.33 µs` | `183.04 µs` | `333.79 µs` | ` 3.50 ms` | | bigint large -> (100) | `171.77 µs/iter` | `123.63 µs` | `175.42 µs` | `325.08 µs` | ` 3.16 ms` | | boolean true -> (100) | `168.30 µs/iter` | `119.96 µs` | `172.38 µs` | `321.25 µs` | ` 3.08 ms` | | boolean false -> (100) | `168.24 µs/iter` | `119.42 µs` | `171.92 µs` | `327.38 µs` | ` 3.18 ms` | | undefined -> (100) | `174.87 µs/iter` | `122.58 µs` | `180.38 µs` | `324.38 µs` | ` 3.31 ms` | | null -> (100) | `177.45 µs/iter` | `122.75 µs` | `183.29 µs` | `318.17 µs` | ` 3.12 ms` | | string -> (100) | `172.13 µs/iter` | `121.83 µs` | `175.33 µs` | `311.46 µs` | ` 3.44 ms` | | json object -> (100) | `349.71 µs/iter` | `272.13 µs` | `358.25 µs` | `539.25 µs` | ` 3.52 ms` | | json array -> (100) | `518.73 µs/iter` | `409.38 µs` | `542.92 µs` | `754.54 µs` | ` 3.59 ms` | | Uint8Array -> (100) | `242.72 µs/iter` | `171.54 µs` | `253.50 µs` | `506.21 µs` | ` 2.42 ms` | | ArrayBuffer -> (100) | `229.70 µs/iter` | `160.50 µs` | `239.67 µs` | `470.63 µs` | ` 2.41 ms` | | Buffer -> (100) | `354.02 µs/iter` | `222.83 µs` | `376.63 µs` | `974.71 µs` | ` 1.47 ms` | | string huge -> (100) | `197.79 µs/iter` | `141.29 µs` | `197.08 µs` | `371.71 µs` | ` 2.56 ms` | | Int32Array -> (100) | `243.54 µs/iter` | `169.13 µs` | `254.42 µs` | `512.29 µs` | ` 2.56 ms` | | Float64Array -> (100) | `243.76 µs/iter` | `170.29 µs` | `254.00 µs` | `541.63 µs` | ` 2.44 ms` | | BigInt64Array -> (100) | `242.08 µs/iter` | `170.46 µs` | `252.08 µs` | `527.83 µs` | ` 2.48 ms` | | BigUint64Array -> (100) | `241.30 µs/iter` | `168.63 µs` | `253.17 µs` | `481.75 µs` | ` 2.53 ms` | | DataView -> (100) | `240.07 µs/iter` | `169.00 µs` | `250.58 µs` | `529.42 µs` | ` 2.49 ms` | | Date -> (100) | `174.26 µs/iter` | `124.38 µs` | `178.88 µs` | `325.83 µs` | ` 3.36 ms` | ``` ### 100 messages At `100` messages per iteration, the gap remains strong: - Typical primitives stay around `~6-8x` faster with Knitting. - Heavier payloads still keep a clear edge at roughly `~2.4-4.2x` faster. - Batching improves throughput for both, but Knitting remains lower-overhead across payload classes. Same data as `node_types.md` above. ## Call growth throughput (Node) This benchmark increases payload size from `32 B` up to `1,048,576 B` (1 MiB) and reports batched cost with `batch=64`. Using the `1048576 B` row (`avg`) from the `batch=64` run, one-way transfer throughput is: - `string`: `46.54 ms/iter` -> `1.44 GB/s` - `Uint8Array`: `8.95 ms/iter` -> `7.50 GB/s` Interpretation: - In this batched shape, Node's `Uint8Array` path is much stronger than its string path. - For 1 MiB binary payloads, Node lands in the `~7.5 GB/s` range in this run. ### Data `node_call-growth-batch.md` ```md clk: ~3.70 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • call growth batch string (ascii 32..1048576 x4, batch=64) | avg | min | p75 | p99 | max | | --------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 32 B | ` 22.96 µs/iter` | ` 14.58 µs` | ` 28.71 µs` | ` 54.04 µs` | `582.58 µs` | | 128 B | ` 27.04 µs/iter` | ` 24.48 µs` | ` 27.65 µs` | ` 28.63 µs` | ` 28.94 µs` | | 512 B | ` 44.43 µs/iter` | ` 40.30 µs` | ` 46.17 µs` | ` 46.61 µs` | ` 48.79 µs` | | 2048 B | `142.37 µs/iter` | `103.71 µs` | `168.75 µs` | `233.46 µs` | `301.04 µs` | | 8192 B | `399.94 µs/iter` | `339.63 µs` | `418.25 µs` | `513.25 µs` | ` 1.59 ms` | | 32768 B | ` 1.40 ms/iter` | ` 1.28 ms` | ` 1.43 ms` | ` 2.43 ms` | ` 2.56 ms` | | 131072 B | ` 7.07 ms/iter` | ` 5.41 ms` | ` 7.39 ms` | ` 11.23 ms` | ` 11.49 ms` | | 524288 B | ` 23.22 ms/iter` | ` 21.87 ms` | ` 23.57 ms` | ` 24.56 ms` | ` 25.24 ms` | | 1048576 B | ` 46.54 ms/iter` | ` 44.57 ms` | ` 47.59 ms` | ` 47.69 ms` | ` 48.34 ms` | | • call growth batch uint8array (32..1048576 x4, batch=64) | avg | min | p75 | p99 | max | | --------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 32 B | ` 25.21 µs/iter` | ` 15.46 µs` | ` 29.63 µs` | ` 55.04 µs` | `462.75 µs` | | 128 B | ` 26.11 µs/iter` | ` 23.47 µs` | ` 25.99 µs` | ` 30.16 µs` | ` 33.14 µs` | | 512 B | ` 26.26 µs/iter` | ` 24.79 µs` | ` 26.35 µs` | ` 26.92 µs` | ` 32.01 µs` | | 2048 B | ` 81.17 µs/iter` | ` 41.04 µs` | ` 97.46 µs` | `178.63 µs` | ` 4.42 ms` | | 8192 B | `181.59 µs/iter` | ` 91.17 µs` | `190.71 µs` | ` 1.10 ms` | ` 3.30 ms` | | 32768 B | `431.63 µs/iter` | `178.17 µs` | `370.88 µs` | ` 2.35 ms` | ` 3.21 ms` | | 131072 B | ` 1.40 ms/iter` | `777.21 µs` | ` 1.72 ms` | ` 3.20 ms` | ` 3.36 ms` | | 524288 B | ` 4.66 ms/iter` | ` 3.53 ms` | ` 5.08 ms` | ` 6.42 ms` | ` 7.07 ms` | | 1048576 B | ` 8.95 ms/iter` | ` 7.54 ms` | ` 9.32 ms` | ` 11.37 ms` | ` 12.30 ms` | ``` ## Efficiency under heavy tasks (Node) This stress test computes prime numbers over a large range, then serializes and parses large JSON payloads: ```ts const N = 10_000_000; // search range: [1..N] const CHUNK_SIZE = 250_000; ``` Even under this heavier workload, parallel workers scale well: - `main + 1 extra thread`: `~1.7x` faster than main only. - `main + 2 extra threads`: `~2.3x` faster than main only. - `main + 3 extra threads`: `~3.0x` faster than main only. - `main + 4 extra threads`: `~3.5x` faster than main only. ### Data `node_withload.md` ```md clk: ~3.72 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • knitting: primes up to 10,000,000 (chunk=250,000) | avg | min | p75 | p99 | max | | ----------------------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | main | `959.14 ms/iter` | `955.16 ms` | `959.93 ms` | `961.89 ms` | `963.68 ms` | | main + 1 extra threads → full range | `538.28 ms/iter` | `531.37 ms` | `539.87 ms` | `541.65 ms` | `547.02 ms` | | main + 2 extra threads → full range | `401.87 ms/iter` | `395.43 ms` | `404.32 ms` | `406.77 ms` | `408.40 ms` | | main + 3 extra threads → full range | `317.52 ms/iter` | `311.27 ms` | `319.14 ms` | `323.94 ms` | `327.36 ms` | | main + 4 extra threads → full range | `274.21 ms/iter` | `270.63 ms` | `276.88 ms` | `278.30 ms` | `279.69 ms` | ``` ## All types in Knitting This benchmark covers primitive, structured, collection, typed-array, error/date/symbol, promise-arg, and static-vs-dynamic allocator paths. Results are reported for count `1` and count `100` to show both per-call latency and batched throughput. Quick takeaways: - In count `100`, primitive-style payloads are usually in the `~15-30 µs` range, while heavier structured/collection payloads can be `~120-300+ µs`. - The static payload path is usually around `2x-4x` faster than dynamic allocator paths (for example: string `~3.4x`, json `~3.0x`, `Uint8Array` `~3.6x`, symbol `~4.1x` at count `100`). Payload sizes (approximate): | Payload | Size | | --- | ---: | | `jsonObj` | `206 B` | | `jsonArr` | `217 B` | | `mapPayload` | `284 B` | | `Uint8Array` | `1024 B` | | `Int32Array` | `1024 B` | | `Float64Array` | `1024 B` | | `BigInt64Array` | `1024 B` | | `BigUint64Array` | `1024 B` | | `DataView` | `1024 B` | | `smallU8` | `480 B` | | `largeU8` | `481 B` | ### Data `node_types_knitting.md` ```md payload sizes (approx bytes): jsonObj: 206 bytes jsonArr: 217 bytes stringHuge: 1024 bytes Uint8Array: 1024 bytes Int32Array: 1024 bytes Float64Array: 1024 bytes BigInt64Array: 1024 bytes BigUint64Array: 1024 bytes DataView: 1024 bytes clk: ~3.73 GHz cpu: Apple M3 Ultra runtime: node 24.12.0 (arm64-darwin) | • knitting-types 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 1.16 µs/iter` | `458.00 ns` | ` 1.54 µs` | ` 2.75 µs` | ` 1.63 ms` | | bigint small -> (1) | ` 1.05 µs/iter` | `500.00 ns` | ` 1.50 µs` | ` 2.17 µs` | ` 1.26 ms` | | bigint large -> (1) | ` 2.48 µs/iter` | ` 1.25 µs` | ` 2.54 µs` | ` 5.58 µs` | ` 4.13 ms` | | boolean true -> (1) | `901.13 ns/iter` | `526.34 ns` | ` 1.21 µs` | ` 1.58 µs` | ` 1.60 µs` | | boolean false -> (1) | `824.29 ns/iter` | `520.36 ns` | ` 1.03 µs` | ` 1.57 µs` | ` 1.65 µs` | | undefined -> (1) | `984.18 ns/iter` | `526.26 ns` | ` 1.32 µs` | ` 1.58 µs` | ` 1.62 µs` | | null -> (1) | `847.02 ns/iter` | `525.11 ns` | ` 1.12 µs` | ` 1.57 µs` | ` 1.57 µs` | | string -> (1) | ` 1.44 µs/iter` | `583.00 ns` | ` 2.08 µs` | ` 2.75 µs` | ` 1.24 ms` | | json object -> (1) | ` 3.47 µs/iter` | ` 2.46 µs` | ` 3.83 µs` | ` 8.04 µs` | ` 1.27 ms` | | json array -> (1) | ` 4.65 µs/iter` | ` 3.42 µs` | ` 4.96 µs` | ` 9.92 µs` | ` 1.24 ms` | | Uint8Array -> (1) | ` 2.83 µs/iter` | ` 1.13 µs` | ` 3.42 µs` | ` 7.25 µs` | ` 1.28 ms` | | ArrayBuffer -> (1) | ` 2.89 µs/iter` | ` 1.42 µs` | ` 3.71 µs` | ` 7.88 µs` | ` 1.31 ms` | | Buffer -> (1) | ` 2.54 µs/iter` | ` 1.73 µs` | ` 2.95 µs` | ` 3.55 µs` | ` 3.57 µs` | | string huge -> (1) | ` 3.24 µs/iter` | ` 1.92 µs` | ` 4.00 µs` | ` 7.50 µs` | ` 1.24 ms` | | Int32Array -> (1) | ` 2.64 µs/iter` | ` 1.13 µs` | ` 3.33 µs` | ` 6.42 µs` | ` 1.27 ms` | | Float64Array -> (1) | ` 2.76 µs/iter` | ` 1.38 µs` | ` 3.63 µs` | ` 7.42 µs` | ` 1.43 ms` | | BigInt64Array -> (1) | ` 2.53 µs/iter` | ` 1.13 µs` | ` 3.25 µs` | ` 6.25 µs` | ` 1.27 ms` | | BigUint64Array -> (1) | ` 2.58 µs/iter` | ` 1.17 µs` | ` 3.29 µs` | ` 6.29 µs` | ` 1.28 ms` | | DataView -> (1) | ` 2.70 µs/iter` | ` 1.80 µs` | ` 3.08 µs` | ` 3.53 µs` | ` 3.57 µs` | | Date -> (1) | ` 1.10 µs/iter` | `687.13 ns` | ` 1.37 µs` | ` 1.87 µs` | ` 1.96 µs` | | Symbol.for -> (1) | ` 1.62 µs/iter` | ` 1.19 µs` | ` 1.84 µs` | ` 2.15 µs` | ` 2.24 µs` | | • knitting-types 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | ` 25.45 µs/iter` | ` 15.33 µs` | ` 33.38 µs` | ` 62.25 µs` | ` 1.25 ms` | | bigint small -> (100) | ` 27.62 µs/iter` | ` 25.17 µs` | ` 29.19 µs` | ` 30.32 µs` | ` 32.61 µs` | | bigint large -> (100) | ` 70.42 µs/iter` | ` 53.38 µs` | ` 73.75 µs` | `134.50 µs` | ` 1.44 ms` | | boolean true -> (100) | ` 27.65 µs/iter` | ` 23.07 µs` | ` 29.43 µs` | ` 31.76 µs` | ` 32.41 µs` | | boolean false -> (100) | ` 26.79 µs/iter` | ` 24.41 µs` | ` 27.28 µs` | ` 28.78 µs` | ` 29.58 µs` | | undefined -> (100) | ` 26.81 µs/iter` | ` 24.37 µs` | ` 27.73 µs` | ` 29.37 µs` | ` 34.09 µs` | | null -> (100) | ` 24.44 µs/iter` | ` 21.21 µs` | ` 25.08 µs` | ` 26.05 µs` | ` 28.83 µs` | | string -> (100) | ` 35.29 µs/iter` | ` 31.39 µs` | ` 36.22 µs` | ` 38.08 µs` | ` 41.26 µs` | | json object -> (100) | `123.26 µs/iter` | `105.54 µs` | `123.92 µs` | `221.92 µs` | ` 1.39 ms` | | json array -> (100) | `178.94 µs/iter` | `154.96 µs` | `185.13 µs` | `285.50 µs` | `491.88 µs` | | Uint8Array -> (100) | `108.25 µs/iter` | ` 56.88 µs` | `141.25 µs` | `215.38 µs` | ` 3.06 ms` | | ArrayBuffer -> (100) | `123.32 µs/iter` | ` 66.17 µs` | `153.63 µs` | `230.04 µs` | ` 2.51 ms` | | Buffer -> (100) | `102.49 µs/iter` | ` 57.42 µs` | `139.83 µs` | `215.08 µs` | ` 2.97 ms` | | string huge -> (100) | `162.92 µs/iter` | `111.25 µs` | `210.12 µs` | `268.54 µs` | ` 1.44 ms` | | Int32Array -> (100) | `103.37 µs/iter` | ` 58.38 µs` | `141.13 µs` | `212.63 µs` | ` 3.04 ms` | | Float64Array -> (100) | `121.34 µs/iter` | ` 65.71 µs` | `147.25 µs` | `230.96 µs` | ` 1.85 ms` | | BigInt64Array -> (100) | `105.26 µs/iter` | ` 58.75 µs` | `139.79 µs` | `207.92 µs` | ` 3.10 ms` | | BigUint64Array -> (100) | `102.90 µs/iter` | ` 59.21 µs` | `140.08 µs` | `203.58 µs` | ` 3.23 ms` | | DataView -> (100) | `110.45 µs/iter` | ` 61.42 µs` | `145.21 µs` | `217.79 µs` | ` 2.74 ms` | | Date -> (100) | ` 32.51 µs/iter` | ` 29.91 µs` | ` 33.09 µs` | ` 34.62 µs` | ` 38.05 µs` | | Symbol.for -> (100) | ` 40.59 µs/iter` | ` 38.45 µs` | ` 41.20 µs` | ` 41.93 µs` | ` 42.32 µs` | | • knitting-promise-args 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | promise number -> (1) | ` 1.65 µs/iter` | ` 1.20 µs` | ` 1.83 µs` | ` 2.12 µs` | ` 2.15 µs` | | promise object -> (1) | ` 2.69 µs/iter` | ` 2.20 µs` | ` 2.93 µs` | ` 3.15 µs` | ` 3.16 µs` | | • knitting-promise-args 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | promise number -> (100) | ` 33.41 µs/iter` | ` 30.85 µs` | ` 34.19 µs` | ` 37.45 µs` | ` 37.53 µs` | | promise object -> (100) | ` 71.73 µs/iter` | ` 57.46 µs` | ` 77.17 µs` | `139.38 µs` | `253.00 µs` | ``` --- # Deno URL: https://knittingdocs.vercel.app/benchmarks/deno/ Knitting on Deno: message overhead, a head-to-head with a plain worker, latency as calls grow, heavy-task efficiency, and payload type costs. This page summarizes Deno benchmark runs for Knitting on `deno 2.6.6 (aarch64-apple-darwin)`. ## IPC (Deno) This benchmark compares one round-trip between a main thread and workers using different transports. Knitting keeps the lowest overhead in this setup: - `1` message: Knitting is about `3.5x` faster than worker `postMessage`, `3.6x` faster than websocket, and `10x` faster than HTTP. - `25` messages: Knitting is about `9.5x` faster than worker `postMessage`. - `50` messages: Knitting is about `10.7x` faster than worker `postMessage`. ### Data `deno_ipc.md` ```md clk: ~3.60 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • knitting | avg | min | p75 | p99 | max | | --------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 1 thread → (1) | ` 6.24 µs/iter` | ` 1.17 µs` | ` 6.00 µs` | ` 25.21 µs` | `164.71 µs` | | 1 thread → (25) | ` 14.86 µs/iter` | ` 13.43 µs` | ` 15.38 µs` | ` 15.41 µs` | ` 18.16 µs` | | 1 thread → (50) | ` 28.20 µs/iter` | ` 25.98 µs` | ` 28.92 µs` | ` 29.64 µs` | ` 29.98 µs` | clk: ~3.60 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • websocket | avg | min | p75 | p99 | max | | ------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | local → (1) | ` 24.92 µs/iter` | ` 14.00 µs` | ` 27.13 µs` | ` 83.38 µs` | `209.25 µs` | | local → (25) | `142.08 µs/iter` | `109.33 µs` | `145.08 µs` | `252.46 µs` | `331.33 µs` | | local → (50) | `228.86 µs/iter` | `189.08 µs` | `238.92 µs` | `354.13 µs` | `466.83 µs` | clk: ~3.60 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • worker | avg | min | p75 | p99 | max | | ------------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | postMessage → (1) | ` 22.93 µs/iter` | ` 14.63 µs` | ` 23.54 µs` | ` 67.33 µs` | `229.79 µs` | | postMessage → (25) | `181.86 µs/iter` | `137.13 µs` | `191.71 µs` | `311.67 µs` | ` 1.88 ms` | | postMessage → (50) | `353.30 µs/iter` | `280.42 µs` | `373.38 µs` | `544.17 µs` | ` 2.04 ms` | clk: ~3.61 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • http | avg | min | p75 | p99 | max | | ------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | local → (1) | ` 65.63 µs/iter` | ` 40.63 µs` | ` 65.92 µs` | `165.63 µs` | ` 1.08 ms` | | local → (25) | `810.26 µs/iter` | `654.79 µs` | `839.88 µs` | ` 1.77 ms` | ` 1.93 ms` | | local → (50) | ` 1.62 ms/iter` | ` 1.29 ms` | ` 1.73 ms` | ` 2.61 ms` | ` 2.72 ms` | ``` ## Knitting vs Worker (Deno) These charts compare the same payload families sent through Knitting and Deno workers. ### One message With a single value per call, Knitting is consistently faster: - For small primitives, Knitting is roughly `25-30x` faster than workers. - For string/array/object payloads, Knitting is usually around `4-6x` faster. - For larger payloads, Knitting still holds a clear advantage (for example, big object: `6.37 µs` vs `27.45 µs`, about `4.3x` faster). `deno_types.md` ```md clk: ~3.63 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • knitting 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 1.36 µs/iter` | `500.00 ns` | `667.00 ns` | ` 9.21 µs` | `175.46 µs` | | bigint small -> (1) | ` 1.09 µs/iter` | `541.55 ns` | ` 1.34 µs` | ` 2.76 µs` | ` 3.37 µs` | | bigint large -> (1) | ` 6.45 µs/iter` | ` 1.42 µs` | ` 6.29 µs` | ` 24.42 µs` | `137.17 µs` | | boolean true -> (1) | `977.80 ns/iter` | `509.75 ns` | ` 1.25 µs` | ` 2.40 µs` | ` 3.09 µs` | | boolean false -> (1) | `916.76 ns/iter` | `503.31 ns` | ` 1.12 µs` | ` 2.73 µs` | ` 3.44 µs` | | undefined -> (1) | `873.84 ns/iter` | `508.99 ns` | ` 1.09 µs` | ` 2.63 µs` | ` 3.36 µs` | | null -> (1) | `996.82 ns/iter` | `510.53 ns` | ` 1.25 µs` | ` 2.66 µs` | ` 2.80 µs` | | string -> (1) | ` 3.94 µs/iter` | ` 2.98 µs` | ` 4.40 µs` | ` 4.79 µs` | ` 5.37 µs` | | json object -> (1) | ` 6.75 µs/iter` | ` 6.61 µs` | ` 6.78 µs` | ` 6.88 µs` | ` 6.91 µs` | | json array -> (1) | ` 6.85 µs/iter` | ` 6.74 µs` | ` 6.89 µs` | ` 6.97 µs` | ` 6.98 µs` | | Uint8Array -> (1) | ` 6.90 µs/iter` | ` 1.46 µs` | ` 6.67 µs` | ` 24.83 µs` | `116.42 µs` | | ArrayBuffer -> (1) | ` 6.74 µs/iter` | ` 6.58 µs` | ` 6.78 µs` | ` 6.89 µs` | ` 6.93 µs` | | Buffer -> (1) | ` 6.45 µs/iter` | ` 6.26 µs` | ` 6.49 µs` | ` 6.63 µs` | ` 6.68 µs` | | string huge -> (1) | ` 6.38 µs/iter` | ` 6.07 µs` | ` 6.50 µs` | ` 6.59 µs` | ` 6.64 µs` | | Int32Array -> (1) | ` 6.48 µs/iter` | ` 6.33 µs` | ` 6.51 µs` | ` 6.63 µs` | ` 6.66 µs` | | Float64Array -> (1) | ` 6.42 µs/iter` | ` 6.16 µs` | ` 6.46 µs` | ` 6.58 µs` | ` 6.61 µs` | | BigInt64Array -> (1) | ` 6.48 µs/iter` | ` 6.29 µs` | ` 6.54 µs` | ` 6.71 µs` | ` 6.71 µs` | | BigUint64Array -> (1) | ` 6.53 µs/iter` | ` 6.28 µs` | ` 6.60 µs` | ` 6.78 µs` | ` 6.84 µs` | | DataView -> (1) | ` 6.62 µs/iter` | ` 6.45 µs` | ` 6.67 µs` | ` 6.78 µs` | ` 6.78 µs` | | Date -> (1) | ` 3.83 µs/iter` | ` 3.07 µs` | ` 4.20 µs` | ` 4.55 µs` | ` 4.57 µs` | | • knitting 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | ` 26.77 µs/iter` | ` 14.96 µs` | ` 32.96 µs` | ` 62.83 µs` | `248.88 µs` | | bigint small -> (100) | ` 25.42 µs/iter` | ` 21.67 µs` | ` 26.97 µs` | ` 28.80 µs` | ` 30.04 µs` | | bigint large -> (100) | ` 71.73 µs/iter` | ` 52.08 µs` | ` 73.29 µs` | `140.54 µs` | `234.00 µs` | | boolean true -> (100) | ` 27.15 µs/iter` | ` 22.41 µs` | ` 28.91 µs` | ` 29.56 µs` | ` 30.81 µs` | | boolean false -> (100) | ` 28.01 µs/iter` | ` 26.37 µs` | ` 28.83 µs` | ` 29.42 µs` | ` 31.21 µs` | | undefined -> (100) | ` 26.40 µs/iter` | ` 22.54 µs` | ` 27.72 µs` | ` 28.50 µs` | ` 28.67 µs` | | null -> (100) | ` 26.07 µs/iter` | ` 22.90 µs` | ` 28.27 µs` | ` 28.65 µs` | ` 30.03 µs` | | string -> (100) | ` 34.35 µs/iter` | ` 29.21 µs` | ` 36.95 µs` | ` 37.19 µs` | ` 37.72 µs` | | json object -> (100) | `113.58 µs/iter` | ` 97.21 µs` | `114.67 µs` | `205.96 µs` | `290.96 µs` | | json array -> (100) | `143.75 µs/iter` | `119.04 µs` | `149.17 µs` | `250.25 µs` | `323.88 µs` | | Uint8Array -> (100) | `174.82 µs/iter` | ` 94.83 µs` | `194.46 µs` | `275.33 µs` | ` 2.78 ms` | | ArrayBuffer -> (100) | `189.82 µs/iter` | `101.21 µs` | `211.83 µs` | `281.17 µs` | ` 2.17 ms` | | Buffer -> (100) | `177.31 µs/iter` | ` 94.08 µs` | `194.25 µs` | `268.17 µs` | ` 2.75 ms` | | string huge -> (100) | `165.20 µs/iter` | ` 90.88 µs` | `188.13 µs` | `257.17 µs` | `338.71 µs` | | Int32Array -> (100) | `172.75 µs/iter` | ` 85.92 µs` | `194.58 µs` | `279.21 µs` | ` 2.76 ms` | | Float64Array -> (100) | `174.99 µs/iter` | ` 95.50 µs` | `194.75 µs` | `284.25 µs` | ` 2.76 ms` | | BigInt64Array -> (100) | `171.72 µs/iter` | ` 77.33 µs` | `193.50 µs` | `272.21 µs` | ` 2.78 ms` | | BigUint64Array -> (100) | `171.76 µs/iter` | ` 87.29 µs` | `194.71 µs` | `280.58 µs` | ` 2.73 ms` | | DataView -> (100) | `177.36 µs/iter` | ` 98.92 µs` | `198.17 µs` | `278.00 µs` | ` 2.82 ms` | | Date -> (100) | ` 32.52 µs/iter` | ` 30.22 µs` | ` 33.37 µs` | ` 35.28 µs` | ` 35.75 µs` | | • worker 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 22.61 µs/iter` | ` 14.21 µs` | ` 22.96 µs` | ` 66.50 µs` | `204.67 µs` | | bigint small -> (1) | ` 21.14 µs/iter` | ` 20.73 µs` | ` 21.20 µs` | ` 21.43 µs` | ` 22.15 µs` | | bigint large -> (1) | ` 20.98 µs/iter` | ` 20.73 µs` | ` 20.92 µs` | ` 21.17 µs` | ` 22.04 µs` | | boolean true -> (1) | ` 20.89 µs/iter` | ` 20.59 µs` | ` 20.89 µs` | ` 21.04 µs` | ` 21.96 µs` | | boolean false -> (1) | ` 20.85 µs/iter` | ` 20.59 µs` | ` 20.85 µs` | ` 20.99 µs` | ` 21.85 µs` | | undefined -> (1) | ` 20.87 µs/iter` | ` 20.58 µs` | ` 20.91 µs` | ` 21.01 µs` | ` 21.82 µs` | | null -> (1) | ` 20.85 µs/iter` | ` 20.63 µs` | ` 20.78 µs` | ` 21.24 µs` | ` 21.76 µs` | | string -> (1) | ` 20.75 µs/iter` | ` 20.47 µs` | ` 20.77 µs` | ` 20.82 µs` | ` 21.76 µs` | | json object -> (1) | ` 28.59 µs/iter` | ` 28.30 µs` | ` 28.53 µs` | ` 28.87 µs` | ` 30.14 µs` | | json array -> (1) | ` 33.49 µs/iter` | ` 33.12 µs` | ` 33.41 µs` | ` 34.15 µs` | ` 34.69 µs` | | Uint8Array -> (1) | ` 52.26 µs/iter` | ` 35.42 µs` | ` 53.25 µs` | `125.33 µs` | ` 1.77 ms` | | ArrayBuffer -> (1) | ` 22.34 µs/iter` | ` 21.81 µs` | ` 22.31 µs` | ` 22.46 µs` | ` 24.40 µs` | | Buffer -> (1) | ` 52.37 µs/iter` | ` 34.96 µs` | ` 53.63 µs` | `128.33 µs` | ` 1.58 ms` | | string huge -> (1) | ` 21.55 µs/iter` | ` 21.12 µs` | ` 21.67 µs` | ` 22.38 µs` | ` 22.69 µs` | | Int32Array -> (1) | ` 30.78 µs/iter` | ` 30.43 µs` | ` 30.81 µs` | ` 31.37 µs` | ` 31.86 µs` | | Float64Array -> (1) | ` 27.04 µs/iter` | ` 26.64 µs` | ` 27.07 µs` | ` 27.33 µs` | ` 28.19 µs` | | BigInt64Array -> (1) | ` 26.36 µs/iter` | ` 25.82 µs` | ` 26.31 µs` | ` 27.04 µs` | ` 27.58 µs` | | BigUint64Array -> (1) | ` 26.08 µs/iter` | ` 25.71 µs` | ` 26.06 µs` | ` 26.15 µs` | ` 27.60 µs` | | DataView -> (1) | ` 22.45 µs/iter` | ` 22.07 µs` | ` 22.44 µs` | ` 22.73 µs` | ` 23.78 µs` | | Date -> (1) | ` 21.25 µs/iter` | ` 20.98 µs` | ` 21.29 µs` | ` 21.56 µs` | ` 22.26 µs` | | • worker 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | `659.48 µs/iter` | `543.79 µs` | `693.54 µs` | `927.88 µs` | ` 2.21 ms` | | bigint small -> (100) | `657.48 µs/iter` | `549.58 µs` | `698.63 µs` | `876.04 µs` | ` 1.17 ms` | | bigint large -> (100) | `657.62 µs/iter` | `543.92 µs` | `698.63 µs` | `887.42 µs` | ` 2.45 ms` | | boolean true -> (100) | `655.71 µs/iter` | `545.13 µs` | `696.42 µs` | `849.71 µs` | `929.50 µs` | | boolean false -> (100) | `649.65 µs/iter` | `536.96 µs` | `693.75 µs` | `847.71 µs` | `914.79 µs` | | undefined -> (100) | `654.38 µs/iter` | `547.17 µs` | `700.00 µs` | `863.92 µs` | ` 2.38 ms` | | null -> (100) | `654.17 µs/iter` | `538.58 µs` | `697.33 µs` | `850.33 µs` | `990.25 µs` | | string -> (100) | `656.20 µs/iter` | `537.71 µs` | `698.79 µs` | `874.71 µs` | ` 1.24 ms` | | json object -> (100) | `987.76 µs/iter` | `842.42 µs` | ` 1.03 ms` | ` 1.24 ms` | ` 1.77 ms` | | json array -> (100) | ` 1.20 ms/iter` | ` 1.06 ms` | ` 1.26 ms` | ` 1.43 ms` | ` 1.47 ms` | | Uint8Array -> (100) | ` 3.13 ms/iter` | ` 2.82 ms` | ` 3.22 ms` | ` 3.34 ms` | ` 3.36 ms` | | ArrayBuffer -> (100) | `747.43 µs/iter` | `610.54 µs` | `787.71 µs` | ` 1.02 ms` | ` 2.51 ms` | | Buffer -> (100) | ` 3.13 ms/iter` | ` 2.86 ms` | ` 3.21 ms` | ` 3.35 ms` | ` 3.43 ms` | | string huge -> (100) | `690.43 µs/iter` | `567.92 µs` | `730.46 µs` | ` 1.03 ms` | ` 2.32 ms` | | Int32Array -> (100) | ` 1.43 ms/iter` | ` 1.28 ms` | ` 1.47 ms` | ` 1.59 ms` | ` 1.71 ms` | | Float64Array -> (100) | ` 1.11 ms/iter` | `981.88 µs` | ` 1.16 ms` | ` 1.28 ms` | ` 1.37 ms` | | BigInt64Array -> (100) | ` 1.12 ms/iter` | `986.75 µs` | ` 1.16 ms` | ` 1.31 ms` | ` 1.82 ms` | | BigUint64Array -> (100) | ` 1.11 ms/iter` | `968.50 µs` | ` 1.16 ms` | ` 1.29 ms` | ` 1.37 ms` | | DataView -> (100) | `760.87 µs/iter` | `617.21 µs` | `805.96 µs` | ` 1.01 ms` | ` 2.37 ms` | | Date -> (100) | `674.47 µs/iter` | `557.17 µs` | `713.08 µs` | `888.13 µs` | ` 2.40 ms` | ``` ### 100 messages At `100` messages per iteration, the gap remains strong: - Typical primitives stay around `18-35x` faster with Knitting. - Heavier payloads still keep a clear edge at roughly `~7x` faster. - Batching improves throughput for both, but Knitting remains lower-overhead across payload classes. Same data as `deno_types.md` above. ## Call growth throughput (Deno) This benchmark increases payload size from `32 B` up to `1,048,576 B` (1 MiB) and reports batched cost with `batch=64`. Using the `1048576 B` row (`avg`) from the `batch=64` run, one-way transfer throughput is: - `string`: `19.89 ms/iter` -> `3.37 GB/s` - `Uint8Array`: `11.61 ms/iter` -> `5.78 GB/s` Interpretation: - Deno stays in the middle of the pack under batching, with binary ahead of string in this run. - For 1 MiB `Uint8Array`, throughput lands around `~5.8 GB/s`; string is closer to `~3.4 GB/s`. ### Data `deno_call-growth-batch.md` ```md clk: ~3.61 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • call growth batch string (ascii 32..1048576 x4, batch=64) | avg | min | p75 | p99 | max | | --------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 32 B | ` 30.31 µs/iter` | ` 16.42 µs` | ` 39.71 µs` | ` 77.08 µs` | `302.17 µs` | | 128 B | ` 30.07 µs/iter` | ` 26.71 µs` | ` 30.92 µs` | ` 33.21 µs` | ` 33.46 µs` | | 512 B | ` 31.72 µs/iter` | ` 28.87 µs` | ` 32.66 µs` | ` 32.95 µs` | ` 32.95 µs` | | 2048 B | `110.51 µs/iter` | ` 57.67 µs` | `120.29 µs` | `187.04 µs` | `960.67 µs` | | 8192 B | `154.78 µs/iter` | ` 80.58 µs` | `173.04 µs` | `254.96 µs` | `316.83 µs` | | 32768 B | `292.53 µs/iter` | `227.79 µs` | `312.42 µs` | `423.71 µs` | `500.46 µs` | | 131072 B | ` 4.18 ms/iter` | ` 2.47 ms` | ` 5.65 ms` | ` 8.42 ms` | ` 9.31 ms` | | 524288 B | ` 11.47 ms/iter` | ` 9.01 ms` | ` 11.40 ms` | ` 14.34 ms` | ` 15.34 ms` | | 1048576 B | ` 19.89 ms/iter` | ` 17.47 ms` | ` 20.50 ms` | ` 23.38 ms` | ` 25.33 ms` | | • call growth batch uint8array (32..1048576 x4, batch=64) | avg | min | p75 | p99 | max | | --------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 32 B | ` 27.44 µs/iter` | ` 17.92 µs` | ` 31.54 µs` | ` 63.96 µs` | `224.21 µs` | | 128 B | ` 27.83 µs/iter` | ` 25.18 µs` | ` 28.79 µs` | ` 29.28 µs` | ` 29.52 µs` | | 512 B | ` 31.88 µs/iter` | ` 30.56 µs` | ` 32.45 µs` | ` 33.35 µs` | ` 35.16 µs` | | 2048 B | `121.18 µs/iter` | ` 51.75 µs` | `132.00 µs` | `213.21 µs` | ` 2.91 ms` | | 8192 B | `230.85 µs/iter` | `112.54 µs` | `219.79 µs` | ` 1.66 ms` | ` 3.09 ms` | | 32768 B | `527.39 µs/iter` | `241.25 µs` | `457.08 µs` | ` 2.48 ms` | ` 3.33 ms` | | 131072 B | ` 1.79 ms/iter` | ` 1.04 ms` | ` 1.91 ms` | ` 4.85 ms` | ` 5.29 ms` | | 524288 B | ` 6.32 ms/iter` | ` 4.51 ms` | ` 7.14 ms` | ` 8.85 ms` | ` 10.38 ms` | | 1048576 B | ` 11.61 ms/iter` | ` 9.68 ms` | ` 12.61 ms` | ` 14.15 ms` | ` 14.24 ms` | ``` ## Efficiency under heavy tasks (Deno) This stress test computes prime numbers over a large range, then serializes and parses large JSON payloads: ```ts const N = 10_000_000; // search range: [1..N] const CHUNK_SIZE = 250_000; ``` Even under this heavier workload, parallel workers scale well: - `main + 1 extra thread`: `~1.8x` faster than main only. - `main + 2 extra threads`: `~2.5x` faster than main only. - `main + 3 extra threads`: `~3.2x` faster than main only. - `main + 4 extra threads`: `~3.7x` faster than main only. ### Data `deno_withload.md` ```md clk: ~3.61 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • knitting: primes up to 10,000,000 (chunk=250,000) | avg | min | p75 | p99 | max | | ----------------------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | main | `956.62 ms/iter` | `950.73 ms` | `957.55 ms` | `959.96 ms` | `963.62 ms` | | main + 1 extra threads → full range | `496.69 ms/iter` | `492.92 ms` | `497.48 ms` | `501.54 ms` | `501.70 ms` | | main + 2 extra threads → full range | `360.28 ms/iter` | `355.94 ms` | `361.84 ms` | `364.48 ms` | `366.29 ms` | | main + 3 extra threads → full range | `292.60 ms/iter` | `288.34 ms` | `293.62 ms` | `296.29 ms` | `300.23 ms` | | main + 4 extra threads → full range | `254.74 ms/iter` | `251.74 ms` | `256.04 ms` | `256.66 ms` | `258.82 ms` | ``` ## All types in Knitting This benchmark covers primitive, structured, collection, typed-array, error/date/symbol, promise-arg, and static-vs-dynamic allocator paths. Results are reported for count `1` and count `100` to show both per-call latency and batched throughput. Quick takeaways: - In count `100`, primitive-style payloads are usually in the `~20-50 µs` range, while heavier structured/collection payloads can be `~160 µs` to multi-millisecond outliers. - The static payload path is usually around `~1.8x-2.5x` faster than dynamic allocator paths (for example: string `~2.2x`, json `~2.5x`, `Uint8Array` `~1.8x`, symbol `~2.2x` at count `100`). Payload sizes (approximate): | Payload | Size | | --- | ---: | | `jsonObj` | `206 B` | | `jsonArr` | `217 B` | | `mapPayload` | `284 B` | | `Uint8Array` | `1024 B` | | `Int32Array` | `1024 B` | | `Float64Array` | `1024 B` | | `BigInt64Array` | `1024 B` | | `BigUint64Array` | `1024 B` | | `DataView` | `1024 B` | | `smallU8` | `480 B` | | `largeU8` | `481 B` | ### Data `deno_types_knitting.md` ```md payload sizes (approx bytes): jsonObj: 206 bytes jsonArr: 217 bytes stringHuge: 1024 bytes Uint8Array: 1024 bytes Int32Array: 1024 bytes Float64Array: 1024 bytes BigInt64Array: 1024 bytes BigUint64Array: 1024 bytes DataView: 1024 bytes clk: ~3.61 GHz cpu: Apple M3 Ultra runtime: deno 2.6.6 (aarch64-apple-darwin) | • knitting-types 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 1.25 µs/iter` | `500.00 ns` | `667.00 ns` | ` 8.33 µs` | `159.96 µs` | | bigint small -> (1) | ` 1.35 µs/iter` | `500.00 ns` | `667.00 ns` | ` 8.42 µs` | `151.29 µs` | | bigint large -> (1) | ` 6.53 µs/iter` | ` 1.42 µs` | ` 6.25 µs` | ` 25.58 µs` | `154.50 µs` | | boolean true -> (1) | `915.37 ns/iter` | `513.63 ns` | ` 1.08 µs` | ` 2.64 µs` | ` 2.92 µs` | | boolean false -> (1) | `946.00 ns/iter` | `505.40 ns` | ` 1.21 µs` | ` 2.69 µs` | ` 3.18 µs` | | undefined -> (1) | `978.82 ns/iter` | `517.35 ns` | ` 1.28 µs` | ` 2.72 µs` | ` 3.46 µs` | | null -> (1) | `766.20 ns/iter` | `513.38 ns` | `700.95 ns` | ` 2.26 µs` | ` 2.50 µs` | | string -> (1) | ` 4.39 µs/iter` | ` 3.27 µs` | ` 4.73 µs` | ` 5.35 µs` | ` 5.58 µs` | | json object -> (1) | ` 6.90 µs/iter` | ` 6.81 µs` | ` 6.95 µs` | ` 6.99 µs` | ` 7.08 µs` | | json array -> (1) | ` 6.91 µs/iter` | ` 6.80 µs` | ` 6.94 µs` | ` 6.98 µs` | ` 7.04 µs` | | Uint8Array -> (1) | ` 7.04 µs/iter` | ` 1.42 µs` | ` 6.83 µs` | ` 25.00 µs` | `144.58 µs` | | ArrayBuffer -> (1) | ` 6.79 µs/iter` | ` 6.58 µs` | ` 6.87 µs` | ` 7.16 µs` | ` 7.58 µs` | | Buffer -> (1) | ` 6.50 µs/iter` | ` 6.39 µs` | ` 6.55 µs` | ` 6.66 µs` | ` 6.72 µs` | | string huge -> (1) | ` 6.40 µs/iter` | ` 6.24 µs` | ` 6.47 µs` | ` 6.53 µs` | ` 6.54 µs` | | Int32Array -> (1) | ` 6.54 µs/iter` | ` 6.38 µs` | ` 6.58 µs` | ` 6.84 µs` | ` 6.86 µs` | | Float64Array -> (1) | ` 6.47 µs/iter` | ` 6.29 µs` | ` 6.54 µs` | ` 6.65 µs` | ` 6.69 µs` | | BigInt64Array -> (1) | ` 6.60 µs/iter` | ` 6.41 µs` | ` 6.66 µs` | ` 6.80 µs` | ` 6.86 µs` | | BigUint64Array -> (1) | ` 6.59 µs/iter` | ` 6.38 µs` | ` 6.68 µs` | ` 6.83 µs` | ` 6.89 µs` | | DataView -> (1) | ` 6.61 µs/iter` | ` 6.47 µs` | ` 6.69 µs` | ` 6.78 µs` | ` 6.82 µs` | | Date -> (1) | ` 4.17 µs/iter` | ` 3.22 µs` | ` 4.48 µs` | ` 4.77 µs` | ` 4.95 µs` | | Symbol.for -> (1) | ` 4.45 µs/iter` | ` 3.74 µs` | ` 4.75 µs` | ` 5.05 µs` | ` 5.09 µs` | | • knitting-types 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | ` 24.18 µs/iter` | ` 14.88 µs` | ` 32.96 µs` | ` 60.75 µs` | `215.58 µs` | | bigint small -> (100) | ` 24.50 µs/iter` | ` 22.23 µs` | ` 25.18 µs` | ` 26.85 µs` | ` 28.16 µs` | | bigint large -> (100) | ` 73.66 µs/iter` | ` 53.08 µs` | ` 75.79 µs` | `145.08 µs` | `236.29 µs` | | boolean true -> (100) | ` 25.91 µs/iter` | ` 22.28 µs` | ` 27.88 µs` | ` 29.93 µs` | ` 30.14 µs` | | boolean false -> (100) | ` 25.54 µs/iter` | ` 22.95 µs` | ` 26.74 µs` | ` 27.98 µs` | ` 28.97 µs` | | undefined -> (100) | ` 25.79 µs/iter` | ` 22.26 µs` | ` 27.06 µs` | ` 28.38 µs` | ` 32.01 µs` | | null -> (100) | ` 25.27 µs/iter` | ` 22.40 µs` | ` 26.41 µs` | ` 28.21 µs` | ` 29.08 µs` | | string -> (100) | ` 33.66 µs/iter` | ` 31.45 µs` | ` 34.56 µs` | ` 35.26 µs` | ` 36.58 µs` | | json object -> (100) | `113.57 µs/iter` | ` 98.63 µs` | `114.63 µs` | `207.79 µs` | `311.75 µs` | | json array -> (100) | `146.72 µs/iter` | `119.58 µs` | `154.08 µs` | `257.83 µs` | `472.75 µs` | | Uint8Array -> (100) | `225.34 µs/iter` | `114.58 µs` | `238.04 µs` | `343.42 µs` | ` 2.87 ms` | | ArrayBuffer -> (100) | `235.43 µs/iter` | `122.13 µs` | `252.79 µs` | `343.42 µs` | ` 2.11 ms` | | Buffer -> (100) | `223.72 µs/iter` | `116.79 µs` | `236.96 µs` | `333.08 µs` | ` 2.81 ms` | | string huge -> (100) | `208.01 µs/iter` | `101.08 µs` | `221.21 µs` | `312.46 µs` | `412.83 µs` | | Int32Array -> (100) | `231.01 µs/iter` | ` 99.58 µs` | `240.63 µs` | `360.17 µs` | ` 2.84 ms` | | Float64Array -> (100) | `228.32 µs/iter` | `113.75 µs` | `240.58 µs` | `356.08 µs` | ` 2.99 ms` | | BigInt64Array -> (100) | `236.59 µs/iter` | `116.83 µs` | `244.21 µs` | `377.42 µs` | ` 2.82 ms` | | BigUint64Array -> (100) | `228.29 µs/iter` | `121.71 µs` | `240.17 µs` | `350.04 µs` | ` 2.78 ms` | | DataView -> (100) | `229.76 µs/iter` | `120.04 µs` | `240.04 µs` | `330.33 µs` | ` 2.77 ms` | | Date -> (100) | ` 33.88 µs/iter` | ` 31.90 µs` | ` 35.10 µs` | ` 36.25 µs` | ` 37.30 µs` | | Symbol.for -> (100) | ` 37.24 µs/iter` | ` 35.52 µs` | ` 37.89 µs` | ` 38.57 µs` | ` 39.37 µs` | | • knitting-promise-args 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | promise number -> (1) | ` 5.82 µs/iter` | ` 5.61 µs` | ` 5.88 µs` | ` 6.22 µs` | ` 6.56 µs` | | promise object -> (1) | ` 6.83 µs/iter` | ` 6.60 µs` | ` 6.93 µs` | ` 7.11 µs` | ` 7.31 µs` | | • knitting-promise-args 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | promise number -> (100) | ` 47.74 µs/iter` | ` 45.95 µs` | ` 48.23 µs` | ` 48.55 µs` | ` 49.80 µs` | | promise object -> (100) | ` 89.87 µs/iter` | ` 64.33 µs` | ` 96.21 µs` | `194.17 µs` | `345.00 µs` | ``` --- # Bun URL: https://knittingdocs.vercel.app/benchmarks/bun/ Knitting on Bun: message overhead, a head-to-head with a plain worker, latency as calls grow, heavy-task efficiency, and payload type costs. This page summarizes Bun benchmark runs for Knitting on `bun 1.3.6 (arm64-darwin)`. ## IPC (Bun) This benchmark compares one round-trip between a main thread and workers using different transports. Knitting keeps the lowest overhead in this setup: - `1` message: Knitting is about `5.7x` faster than worker `postMessage`, `9.5x` faster than websocket, and `21x` faster than HTTP. - `25` messages: Knitting is about `7.4x` faster than worker `postMessage`. - `50` messages: Knitting is about `7x` faster than worker `postMessage`. ### Data `bun_ipc.md` ```md clk: ~3.70 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • knitting | avg | min | p75 | p99 | max | | --------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 1 thread → (1) | ` 2.68 µs/iter` | ` 1.97 µs` | ` 2.91 µs` | ` 3.42 µs` | ` 3.55 µs` | | 1 thread → (25) | ` 14.62 µs/iter` | ` 12.43 µs` | ` 15.19 µs` | ` 15.97 µs` | ` 16.67 µs` | | 1 thread → (50) | ` 26.43 µs/iter` | ` 23.69 µs` | ` 27.47 µs` | ` 28.26 µs` | ` 28.71 µs` | clk: ~3.70 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • websocket | avg | min | p75 | p99 | max | | ------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | local → (1) | ` 21.65 µs/iter` | ` 8.04 µs` | ` 21.33 µs` | ` 69.00 µs` | `312.63 µs` | | local → (25) | ` 73.20 µs/iter` | ` 52.08 µs` | ` 75.04 µs` | `183.00 µs` | `459.58 µs` | | local → (50) | `128.42 µs/iter` | ` 91.71 µs` | `132.17 µs` | `275.25 µs` | `543.67 µs` | clk: ~3.70 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • worker | avg | min | p75 | p99 | max | | ------------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | postMessage → (1) | ` 11.68 µs/iter` | ` 7.38 µs` | ` 11.75 µs` | ` 37.75 µs` | `433.88 µs` | | postMessage → (25) | ` 79.84 µs/iter` | ` 47.92 µs` | ` 79.54 µs` | `342.04 µs` | ` 1.20 ms` | | postMessage → (50) | `139.38 µs/iter` | ` 91.58 µs` | `147.17 µs` | `516.08 µs` | ` 1.16 ms` | clk: ~3.66 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • http | avg | min | p75 | p99 | max | | ------------ | ---------------- | ----------- | ----------- | ----------- | ----------- | | local → (1) | ` 44.32 µs/iter` | ` 29.67 µs` | ` 43.79 µs` | `121.21 µs` | `355.92 µs` | | local → (25) | `343.94 µs/iter` | `243.54 µs` | `393.46 µs` | `580.08 µs` | `717.88 µs` | | local → (50) | `724.11 µs/iter` | `492.92 µs` | `792.42 µs` | ` 1.03 ms` | ` 1.13 ms` | ``` ## Knitting vs Worker (Bun) These charts compare the same payload families sent through Knitting and Bun workers. ### One message With a single value per call, Knitting is consistently faster: - For small primitives, Knitting is roughly `~6-17x` faster than workers. - For string/array/object payloads, Knitting is usually around `~4-6x` faster. - For larger payloads, Knitting still holds a clear advantage (for example, big object: `3.97 µs` vs `16.38 µs`, about `4.1x` faster). `bun_types.md` ```md clk: ~3.67 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • knitting 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 2.13 µs/iter` | `500.00 ns` | ` 2.96 µs` | ` 8.71 µs` | `362.42 µs` | | bigint small -> (1) | ` 2.14 µs/iter` | ` 1.35 µs` | ` 2.49 µs` | ` 3.15 µs` | ` 3.18 µs` | | bigint large -> (1) | ` 4.10 µs/iter` | ` 2.08 µs` | ` 4.13 µs` | ` 11.83 µs` | ` 1.64 ms` | | boolean true -> (1) | ` 1.26 µs/iter` | `609.65 ns` | ` 1.78 µs` | ` 2.82 µs` | ` 2.86 µs` | | boolean false -> (1) | ` 1.35 µs/iter` | `628.39 ns` | ` 1.85 µs` | ` 2.93 µs` | ` 2.95 µs` | | undefined -> (1) | ` 1.40 µs/iter` | `621.26 ns` | ` 1.93 µs` | ` 2.69 µs` | ` 2.77 µs` | | null -> (1) | ` 1.44 µs/iter` | `634.12 ns` | ` 1.97 µs` | ` 2.71 µs` | ` 2.97 µs` | | string -> (1) | ` 2.75 µs/iter` | `625.00 ns` | ` 3.21 µs` | ` 8.79 µs` | `384.17 µs` | | json object -> (1) | ` 4.50 µs/iter` | ` 2.33 µs` | ` 5.00 µs` | ` 16.29 µs` | ` 1.48 ms` | | json array -> (1) | ` 4.79 µs/iter` | ` 4.23 µs` | ` 4.92 µs` | ` 5.31 µs` | ` 5.31 µs` | | Uint8Array -> (1) | ` 4.17 µs/iter` | ` 1.58 µs` | ` 4.38 µs` | ` 12.25 µs` | ` 3.37 ms` | | ArrayBuffer -> (1) | ` 4.37 µs/iter` | ` 1.38 µs` | ` 4.79 µs` | ` 12.42 µs` | ` 3.04 ms` | | Buffer -> (1) | ` 3.57 µs/iter` | ` 2.61 µs` | ` 3.85 µs` | ` 4.43 µs` | ` 4.46 µs` | | string huge -> (1) | ` 3.52 µs/iter` | ` 2.57 µs` | ` 3.86 µs` | ` 4.07 µs` | ` 4.25 µs` | | Int32Array -> (1) | ` 3.81 µs/iter` | ` 2.85 µs` | ` 4.16 µs` | ` 4.82 µs` | ` 4.94 µs` | | Float64Array -> (1) | ` 3.50 µs/iter` | ` 1.13 µs` | ` 3.75 µs` | ` 10.75 µs` | ` 1.51 ms` | | BigInt64Array -> (1) | ` 4.03 µs/iter` | ` 2.80 µs` | ` 4.52 µs` | ` 4.94 µs` | ` 5.04 µs` | | BigUint64Array -> (1) | ` 4.04 µs/iter` | ` 2.79 µs` | ` 4.58 µs` | ` 5.14 µs` | ` 5.19 µs` | | DataView -> (1) | ` 4.02 µs/iter` | ` 2.93 µs` | ` 4.40 µs` | ` 5.01 µs` | ` 5.04 µs` | | Date -> (1) | ` 2.29 µs/iter` | ` 1.25 µs` | ` 2.71 µs` | ` 3.08 µs` | ` 3.11 µs` | | • knitting 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | ` 31.40 µs/iter` | ` 14.96 µs` | ` 41.63 µs` | ` 86.79 µs` | ` 4.11 ms` | | bigint small -> (100) | ` 32.56 µs/iter` | ` 29.93 µs` | ` 34.02 µs` | ` 34.18 µs` | ` 34.25 µs` | | bigint large -> (100) | `116.20 µs/iter` | ` 82.33 µs` | `114.04 µs` | `366.38 µs` | ` 1.67 ms` | | boolean true -> (100) | ` 31.78 µs/iter` | ` 24.84 µs` | ` 33.50 µs` | ` 34.93 µs` | ` 35.44 µs` | | boolean false -> (100) | ` 31.52 µs/iter` | ` 27.25 µs` | ` 32.22 µs` | ` 35.46 µs` | ` 36.06 µs` | | undefined -> (100) | ` 32.86 µs/iter` | ` 27.91 µs` | ` 34.44 µs` | ` 35.93 µs` | ` 38.27 µs` | | null -> (100) | ` 32.75 µs/iter` | ` 27.48 µs` | ` 35.88 µs` | ` 37.04 µs` | ` 37.32 µs` | | string -> (100) | ` 34.33 µs/iter` | ` 29.95 µs` | ` 36.04 µs` | ` 37.41 µs` | ` 39.64 µs` | | json object -> (100) | `121.78 µs/iter` | ` 89.04 µs` | `125.21 µs` | `221.00 µs` | `544.79 µs` | | json array -> (100) | `146.27 µs/iter` | `114.50 µs` | `151.71 µs` | `263.79 µs` | ` 1.62 ms` | | Uint8Array -> (100) | `113.40 µs/iter` | ` 53.58 µs` | `113.33 µs` | `947.38 µs` | ` 3.12 ms` | | ArrayBuffer -> (100) | `108.31 µs/iter` | ` 47.63 µs` | `108.71 µs` | `924.00 µs` | ` 3.18 ms` | | Buffer -> (100) | `106.11 µs/iter` | ` 45.54 µs` | `114.92 µs` | `879.42 µs` | ` 1.79 ms` | | string huge -> (100) | ` 79.03 µs/iter` | ` 38.88 µs` | ` 92.92 µs` | `191.17 µs` | ` 1.68 ms` | | Int32Array -> (100) | `103.97 µs/iter` | ` 49.96 µs` | ` 98.42 µs` | `949.13 µs` | ` 3.45 ms` | | Float64Array -> (100) | ` 70.94 µs/iter` | ` 37.46 µs` | ` 81.75 µs` | `158.42 µs` | ` 1.69 ms` | | BigInt64Array -> (100) | `106.06 µs/iter` | ` 49.83 µs` | ` 99.50 µs` | `981.92 µs` | ` 3.84 ms` | | BigUint64Array -> (100) | `104.32 µs/iter` | ` 50.96 µs` | ` 99.50 µs` | `946.75 µs` | ` 4.07 ms` | | DataView -> (100) | `100.39 µs/iter` | ` 51.21 µs` | ` 98.75 µs` | `933.13 µs` | ` 3.82 ms` | | Date -> (100) | ` 31.76 µs/iter` | ` 26.90 µs` | ` 33.15 µs` | ` 33.81 µs` | ` 37.27 µs` | | • worker 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 11.23 µs/iter` | ` 6.08 µs` | ` 11.42 µs` | ` 37.38 µs` | `397.63 µs` | | bigint small -> (1) | ` 11.48 µs/iter` | ` 11.35 µs` | ` 11.52 µs` | ` 11.57 µs` | ` 11.61 µs` | | bigint large -> (1) | ` 11.42 µs/iter` | ` 11.29 µs` | ` 11.47 µs` | ` 11.47 µs` | ` 11.48 µs` | | boolean true -> (1) | ` 11.58 µs/iter` | ` 11.53 µs` | ` 11.60 µs` | ` 11.61 µs` | ` 11.67 µs` | | boolean false -> (1) | ` 11.53 µs/iter` | ` 11.49 µs` | ` 11.54 µs` | ` 11.57 µs` | ` 11.58 µs` | | undefined -> (1) | ` 11.52 µs/iter` | ` 11.47 µs` | ` 11.54 µs` | ` 11.56 µs` | ` 11.60 µs` | | null -> (1) | ` 11.57 µs/iter` | ` 11.47 µs` | ` 11.60 µs` | ` 11.63 µs` | ` 11.64 µs` | | string -> (1) | ` 11.62 µs/iter` | ` 11.53 µs` | ` 11.65 µs` | ` 11.68 µs` | ` 11.69 µs` | | json object -> (1) | ` 15.02 µs/iter` | ` 14.88 µs` | ` 15.07 µs` | ` 15.19 µs` | ` 15.27 µs` | | json array -> (1) | ` 17.80 µs/iter` | ` 17.65 µs` | ` 17.92 µs` | ` 17.95 µs` | ` 18.00 µs` | | Uint8Array -> (1) | ` 12.96 µs/iter` | ` 12.84 µs` | ` 13.00 µs` | ` 13.04 µs` | ` 13.06 µs` | | ArrayBuffer -> (1) | ` 12.41 µs/iter` | ` 12.31 µs` | ` 12.45 µs` | ` 12.51 µs` | ` 12.63 µs` | | Buffer -> (1) | ` 12.94 µs/iter` | ` 12.86 µs` | ` 12.99 µs` | ` 13.01 µs` | ` 13.07 µs` | | string huge -> (1) | ` 11.68 µs/iter` | ` 11.57 µs` | ` 11.69 µs` | ` 11.73 µs` | ` 11.76 µs` | | Int32Array -> (1) | ` 12.99 µs/iter` | ` 12.85 µs` | ` 13.04 µs` | ` 13.07 µs` | ` 13.08 µs` | | Float64Array -> (1) | ` 12.97 µs/iter` | ` 12.81 µs` | ` 12.98 µs` | ` 13.03 µs` | ` 13.14 µs` | | BigInt64Array -> (1) | ` 12.99 µs/iter` | ` 12.87 µs` | ` 13.00 µs` | ` 13.03 µs` | ` 13.14 µs` | | BigUint64Array -> (1) | ` 12.92 µs/iter` | ` 12.83 µs` | ` 12.93 µs` | ` 12.97 µs` | ` 13.00 µs` | | DataView -> (1) | ` 12.92 µs/iter` | ` 12.82 µs` | ` 12.95 µs` | ` 12.97 µs` | ` 12.97 µs` | | Date -> (1) | ` 11.36 µs/iter` | ` 11.26 µs` | ` 11.40 µs` | ` 11.41 µs` | ` 11.42 µs` | | • worker 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | `186.16 µs/iter` | `120.42 µs` | `198.42 µs` | `677.46 µs` | ` 1.44 ms` | | bigint small -> (100) | `225.28 µs/iter` | `150.83 µs` | `244.08 µs` | `766.75 µs` | ` 1.33 ms` | | bigint large -> (100) | `228.58 µs/iter` | `151.13 µs` | `247.13 µs` | `766.54 µs` | ` 1.57 ms` | | boolean true -> (100) | `186.21 µs/iter` | `119.21 µs` | `198.42 µs` | `688.25 µs` | ` 1.61 ms` | | boolean false -> (100) | `188.23 µs/iter` | `120.71 µs` | `200.21 µs` | `710.79 µs` | ` 1.69 ms` | | undefined -> (100) | `185.35 µs/iter` | `121.92 µs` | `197.79 µs` | `672.88 µs` | ` 1.51 ms` | | null -> (100) | `187.17 µs/iter` | `120.04 µs` | `198.88 µs` | `704.04 µs` | ` 1.55 ms` | | string -> (100) | `192.56 µs/iter` | `121.58 µs` | `205.83 µs` | `747.04 µs` | ` 1.57 ms` | | json object -> (100) | `409.65 µs/iter` | `286.46 µs` | `431.88 µs` | `971.58 µs` | ` 1.39 ms` | | json array -> (100) | `538.27 µs/iter` | `394.75 µs` | `554.08 µs` | ` 1.18 ms` | ` 1.39 ms` | | Uint8Array -> (100) | `329.11 µs/iter` | `207.67 µs` | `325.63 µs` | ` 1.41 ms` | ` 5.20 ms` | | ArrayBuffer -> (100) | `291.22 µs/iter` | `187.96 µs` | `296.54 µs` | ` 1.12 ms` | ` 3.67 ms` | | Buffer -> (100) | `323.56 µs/iter` | `203.96 µs` | `323.38 µs` | ` 1.49 ms` | ` 2.07 ms` | | string huge -> (100) | `205.38 µs/iter` | `124.54 µs` | `213.13 µs` | `779.42 µs` | ` 1.69 ms` | | Int32Array -> (100) | `323.10 µs/iter` | `203.96 µs` | `322.33 µs` | ` 1.42 ms` | ` 5.23 ms` | | Float64Array -> (100) | `324.65 µs/iter` | `203.08 µs` | `323.54 µs` | ` 1.36 ms` | ` 2.38 ms` | | BigInt64Array -> (100) | `319.81 µs/iter` | `205.42 µs` | `321.04 µs` | ` 1.56 ms` | ` 2.17 ms` | | BigUint64Array -> (100) | `323.90 µs/iter` | `203.54 µs` | `322.75 µs` | ` 1.49 ms` | ` 2.21 ms` | | DataView -> (100) | `317.78 µs/iter` | `199.58 µs` | `317.50 µs` | ` 1.53 ms` | ` 2.27 ms` | | Date -> (100) | `227.29 µs/iter` | `149.71 µs` | `242.04 µs` | `742.83 µs` | ` 1.72 ms` | ``` ### 100 messages At `100` messages per iteration, the gap remains strong: - Typical primitives stay around `~7-9x` faster with Knitting. - Heavier payloads still keep a clear edge at roughly `~3.7-3.9x` faster. - Batching improves throughput for both, but Knitting remains lower-overhead across payload classes. Same data as `bun_types.md` above. ## Call growth throughput (Bun) This benchmark increases payload size from `32 B` up to `1,048,576 B` (1 MiB) and reports batched cost with `batch=64`. Using the `1048576 B` row (`avg`) from the `batch=64` run, one-way transfer throughput is: - `string`: `5.66 ms/iter` -> `11.86 GB/s` - `Uint8Array`: `4.14 ms/iter` -> `16.21 GB/s` Interpretation: - Bun remains the fastest runtime in this batched 1 MiB shape for both string and binary payloads. - The `Uint8Array` path clears `16 GB/s` one-way equivalent in this run, with string close to `12 GB/s`. ### Data `bun_call-growth-batch.md` ```md clk: ~3.68 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • call growth batch string (ascii 32..1048576 x4, batch=64) | avg | min | p75 | p99 | max | | --------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 32 B | ` 25.85 µs/iter` | ` 11.46 µs` | ` 29.63 µs` | ` 80.08 µs` | ` 1.94 ms` | | 128 B | ` 22.32 µs/iter` | ` 20.54 µs` | ` 22.99 µs` | ` 23.68 µs` | ` 25.10 µs` | | 512 B | ` 28.31 µs/iter` | ` 25.67 µs` | ` 29.32 µs` | ` 29.67 µs` | ` 30.50 µs` | | 2048 B | ` 67.00 µs/iter` | ` 28.96 µs` | ` 70.29 µs` | `247.21 µs` | ` 2.73 ms` | | 8192 B | `130.66 µs/iter` | ` 71.17 µs` | `142.79 µs` | `323.83 µs` | ` 1.84 ms` | | 32768 B | `208.32 µs/iter` | `123.29 µs` | `214.96 µs` | `652.54 µs` | ` 1.84 ms` | | 131072 B | `592.88 µs/iter` | `418.63 µs` | `592.92 µs` | ` 1.88 ms` | ` 2.25 ms` | | 524288 B | ` 2.82 ms/iter` | ` 1.77 ms` | ` 3.36 ms` | ` 6.20 ms` | ` 6.64 ms` | | 1048576 B | ` 5.66 ms/iter` | ` 3.85 ms` | ` 5.85 ms` | ` 10.93 ms` | ` 10.98 ms` | | • call growth batch uint8array (32..1048576 x4, batch=64) | avg | min | p75 | p99 | max | | --------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | 32 B | ` 25.21 µs/iter` | ` 13.38 µs` | ` 30.13 µs` | ` 56.04 µs` | ` 1.53 ms` | | 128 B | ` 23.71 µs/iter` | ` 20.35 µs` | ` 25.02 µs` | ` 25.26 µs` | ` 26.54 µs` | | 512 B | ` 22.87 µs/iter` | ` 20.79 µs` | ` 23.68 µs` | ` 23.82 µs` | ` 24.01 µs` | | 2048 B | ` 67.13 µs/iter` | ` 33.50 µs` | ` 70.92 µs` | `132.92 µs` | ` 7.98 ms` | | 8192 B | `116.93 µs/iter` | ` 62.04 µs` | `114.25 µs` | `221.04 µs` | ` 3.13 ms` | | 32768 B | `193.72 µs/iter` | `131.08 µs` | `173.92 µs` | ` 1.01 ms` | ` 2.90 ms` | | 131072 B | `571.40 µs/iter` | `348.29 µs` | `523.71 µs` | ` 2.01 ms` | ` 2.32 ms` | | 524288 B | ` 2.37 ms/iter` | ` 1.59 ms` | ` 3.00 ms` | ` 5.19 ms` | ` 6.12 ms` | | 1048576 B | ` 4.14 ms/iter` | ` 3.08 ms` | ` 4.64 ms` | ` 6.30 ms` | ` 9.65 ms` | ``` ## Efficiency under heavy tasks (Bun) This stress test computes prime numbers over a large range, then serializes and parses large JSON payloads: ```ts const N = 10_000_000; // search range: [1..N] const CHUNK_SIZE = 250_000; ``` Even under this heavier workload, parallel workers scale well: - `main + 1 extra thread`: `~1.8x` faster than main only. - `main + 2 extra threads`: `~2.5x` faster than main only. - `main + 3 extra threads`: `~3.3x` faster than main only. - `main + 4 extra threads`: `~3.8x` faster than main only. ### Data `bun_withload.md` ```md clk: ~3.69 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • knitting: primes up to 10,000,000 (chunk=250,000) | avg | min | p75 | p99 | max | | ----------------------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | main | `572.57 ms/iter` | `569.05 ms` | `573.73 ms` | `576.83 ms` | `577.70 ms` | | main + 1 extra threads → full range | `304.05 ms/iter` | `299.51 ms` | `305.94 ms` | `310.39 ms` | `310.63 ms` | | main + 2 extra threads → full range | `224.68 ms/iter` | `220.21 ms` | `228.08 ms` | `230.10 ms` | `230.84 ms` | | main + 3 extra threads → full range | `179.33 ms/iter` | `174.37 ms` | `181.15 ms` | `181.78 ms` | `182.35 ms` | | main + 4 extra threads → full range | `157.48 ms/iter` | `153.46 ms` | `159.02 ms` | `159.85 ms` | `160.47 ms` | ``` ## All types in Knitting This benchmark covers primitive, structured, collection, typed-array, error/date/symbol, promise-arg, and static-vs-dynamic allocator paths. Results are reported for count `1` and count `100` to show both per-call latency and batched throughput. Quick takeaways: - In count `100`, primitive-style payloads are usually in the `~16-25 µs` range, while heavier structured/collection payloads are often `~60-270 µs` with occasional larger spikes. - The static payload path is usually around `~1.3x-3.1x` faster than dynamic allocator paths (for example: string `~1.8x`, json `~2.1x`, `Uint8Array` `~1.3x`, symbol `~3.1x` at count `100`). Payload sizes (approximate): | Payload | Size | | --- | ---: | | `jsonObj` | `206 B` | | `jsonArr` | `217 B` | | `mapPayload` | `284 B` | | `Uint8Array` | `1024 B` | | `Int32Array` | `1024 B` | | `Float64Array` | `1024 B` | | `BigInt64Array` | `1024 B` | | `BigUint64Array` | `1024 B` | | `DataView` | `1024 B` | | `smallU8` | `480 B` | | `largeU8` | `481 B` | ### Data `bun_types_knitting.md` ```md payload sizes (approx bytes): jsonObj: 206 bytes jsonArr: 217 bytes stringHuge: 1024 bytes Uint8Array: 1024 bytes Int32Array: 1024 bytes Float64Array: 1024 bytes BigInt64Array: 1024 bytes BigUint64Array: 1024 bytes DataView: 1024 bytes clk: ~3.69 GHz cpu: Apple M3 Ultra runtime: bun 1.3.6 (arm64-darwin) | • knitting-types 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (1) | ` 1.88 µs/iter` | `459.00 ns` | ` 2.83 µs` | ` 8.17 µs` | `433.21 µs` | | bigint small -> (1) | ` 2.19 µs/iter` | ` 1.50 µs` | ` 2.48 µs` | ` 3.01 µs` | ` 3.07 µs` | | bigint large -> (1) | ` 4.01 µs/iter` | ` 2.04 µs` | ` 4.04 µs` | ` 12.50 µs` | ` 1.73 ms` | | boolean true -> (1) | ` 1.33 µs/iter` | `707.16 ns` | ` 1.68 µs` | ` 2.48 µs` | ` 2.74 µs` | | boolean false -> (1) | ` 1.09 µs/iter` | `710.70 ns` | ` 1.35 µs` | ` 2.14 µs` | ` 2.47 µs` | | undefined -> (1) | ` 1.16 µs/iter` | `648.72 ns` | ` 1.51 µs` | ` 2.57 µs` | ` 2.89 µs` | | null -> (1) | ` 1.03 µs/iter` | `649.21 ns` | ` 1.30 µs` | ` 2.70 µs` | ` 2.92 µs` | | string -> (1) | ` 2.39 µs/iter` | `625.00 ns` | ` 2.92 µs` | ` 7.88 µs` | `448.88 µs` | | json object -> (1) | ` 3.59 µs/iter` | ` 3.00 µs` | ` 3.82 µs` | ` 4.37 µs` | ` 4.37 µs` | | json array -> (1) | ` 4.61 µs/iter` | ` 4.15 µs` | ` 4.76 µs` | ` 5.04 µs` | ` 5.09 µs` | | Uint8Array -> (1) | ` 4.34 µs/iter` | ` 1.54 µs` | ` 4.67 µs` | ` 14.04 µs` | ` 3.75 ms` | | ArrayBuffer -> (1) | ` 3.98 µs/iter` | ` 1.33 µs` | ` 4.42 µs` | ` 11.83 µs` | ` 3.49 ms` | | Buffer -> (1) | ` 3.15 µs/iter` | ` 2.44 µs` | ` 3.46 µs` | ` 3.79 µs` | ` 3.84 µs` | | string huge -> (1) | ` 3.51 µs/iter` | ` 1.29 µs` | ` 3.79 µs` | ` 11.96 µs` | ` 1.59 ms` | | Int32Array -> (1) | ` 3.55 µs/iter` | ` 2.61 µs` | ` 3.92 µs` | ` 4.21 µs` | ` 4.25 µs` | | Float64Array -> (1) | ` 3.00 µs/iter` | ` 1.08 µs` | ` 3.38 µs` | ` 8.79 µs` | ` 1.68 ms` | | BigInt64Array -> (1) | ` 3.48 µs/iter` | ` 2.54 µs` | ` 3.86 µs` | ` 4.45 µs` | ` 4.65 µs` | | BigUint64Array -> (1) | ` 3.57 µs/iter` | ` 2.58 µs` | ` 4.05 µs` | ` 4.84 µs` | ` 4.90 µs` | | DataView -> (1) | ` 3.24 µs/iter` | ` 2.50 µs` | ` 3.50 µs` | ` 4.65 µs` | ` 5.04 µs` | | Date -> (1) | ` 1.68 µs/iter` | `541.00 ns` | ` 1.96 µs` | ` 5.21 µs` | `466.54 µs` | | Symbol.for -> (1) | ` 2.41 µs/iter` | ` 1.76 µs` | ` 2.64 µs` | ` 3.04 µs` | ` 3.04 µs` | | • knitting-types 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | number -> (100) | ` 32.00 µs/iter` | ` 13.96 µs` | ` 27.54 µs` | `215.88 µs` | ` 3.23 ms` | | bigint small -> (100) | ` 30.26 µs/iter` | ` 27.43 µs` | ` 31.73 µs` | ` 32.18 µs` | ` 35.93 µs` | | bigint large -> (100) | `122.70 µs/iter` | ` 84.33 µs` | `120.96 µs` | `425.79 µs` | ` 1.91 ms` | | boolean true -> (100) | ` 27.23 µs/iter` | ` 24.14 µs` | ` 27.49 µs` | ` 29.92 µs` | ` 34.21 µs` | | boolean false -> (100) | ` 28.04 µs/iter` | ` 23.23 µs` | ` 29.87 µs` | ` 32.86 µs` | ` 36.96 µs` | | undefined -> (100) | ` 28.24 µs/iter` | ` 22.49 µs` | ` 29.58 µs` | ` 30.68 µs` | ` 32.45 µs` | | null -> (100) | ` 25.01 µs/iter` | ` 22.17 µs` | ` 26.19 µs` | ` 27.17 µs` | ` 27.45 µs` | | string -> (100) | ` 28.18 µs/iter` | ` 23.32 µs` | ` 30.81 µs` | ` 33.08 µs` | ` 41.07 µs` | | json object -> (100) | `122.98 µs/iter` | ` 93.67 µs` | `125.88 µs` | `223.46 µs` | ` 1.62 ms` | | json array -> (100) | `149.52 µs/iter` | `113.92 µs` | `154.54 µs` | `279.08 µs` | ` 1.66 ms` | | Uint8Array -> (100) | `112.02 µs/iter` | ` 54.00 µs` | `112.13 µs` | `928.79 µs` | ` 3.12 ms` | | ArrayBuffer -> (100) | `115.93 µs/iter` | ` 61.33 µs` | `113.54 µs` | `922.46 µs` | ` 3.59 ms` | | Buffer -> (100) | `104.68 µs/iter` | ` 44.46 µs` | `115.63 µs` | `747.96 µs` | ` 2.11 ms` | | string huge -> (100) | ` 93.47 µs/iter` | ` 46.04 µs` | `110.38 µs` | `207.17 µs` | ` 1.73 ms` | | Int32Array -> (100) | `101.19 µs/iter` | ` 49.92 µs` | ` 98.04 µs` | `900.71 µs` | ` 3.78 ms` | | Float64Array -> (100) | ` 71.52 µs/iter` | ` 37.50 µs` | ` 82.33 µs` | `167.50 µs` | ` 1.58 ms` | | BigInt64Array -> (100) | `103.28 µs/iter` | ` 50.00 µs` | ` 98.63 µs` | `920.00 µs` | ` 3.84 ms` | | BigUint64Array -> (100) | `101.80 µs/iter` | ` 50.13 µs` | ` 98.13 µs` | `908.42 µs` | ` 4.21 ms` | | DataView -> (100) | `102.14 µs/iter` | ` 50.29 µs` | ` 98.92 µs` | `895.21 µs` | ` 3.87 ms` | | Date -> (100) | ` 31.00 µs/iter` | ` 27.10 µs` | ` 32.35 µs` | ` 34.37 µs` | ` 38.53 µs` | | Symbol.for -> (100) | ` 37.43 µs/iter` | ` 21.46 µs` | ` 47.08 µs` | ` 83.50 µs` | ` 1.56 ms` | | • knitting-promise-args 1 | avg | min | p75 | p99 | max | | --------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | promise number -> (1) | ` 2.44 µs/iter` | ` 1.76 µs` | ` 2.78 µs` | ` 3.16 µs` | ` 3.20 µs` | | promise object -> (1) | ` 3.13 µs/iter` | ` 2.35 µs` | ` 3.34 µs` | ` 3.74 µs` | ` 3.79 µs` | | • knitting-promise-args 100 | avg | min | p75 | p99 | max | | ----------------------- | ---------------- | ----------- | ----------- | ----------- | ----------- | | promise number -> (100) | ` 43.54 µs/iter` | ` 39.69 µs` | ` 45.25 µs` | ` 47.23 µs` | ` 48.48 µs` | | promise object -> (100) | ` 61.77 µs/iter` | ` 38.13 µs` | ` 68.83 µs` | `122.88 µs` | ` 1.49 ms` | ``` --- # Tokio URL: https://knittingdocs.vercel.app/benchmarks/tokio/ How Knitting compares with Tokio on small values, large payloads, and shared byte buffers. This page compares Tokio and Knitting on Bun, Node.js, and Deno using the same batch-oriented echo benchmark. The short version: JavaScript runtimes are quicker for some tiny values, while Tokio becomes much more competitive as the payload gets larger. The charts below show where that change happens. Benchmark source: [`mimiMonads/knitting-vs-tokio-bench`](https://github.com/mimiMonads/knitting-vs-tokio-bench). ## What this benchmark measures Each test sends a value to a worker, echoes it back, and measures the full batch round trip. The payloads are: - `f64` values - `1 MiB` strings - `1 MiB` `Uint8Array` values - `Uint8Array` values from `8 B` to `1 MiB` - a separate `Arc>` sweep for small shared-byte payloads The results are from one run on Ubuntu 23.10, `x86_64`, with an AMD Ryzen 7 4700U. Every runtime uses the same batch sizes (`1`, `10`, and `100`), warmup, measured iterations, and percentile reporting. These numbers describe this benchmark and this machine, not a universal ranking of runtimes. ## Small payloads: `f64` For small scalar values, the JavaScript runtimes are faster on average in this run: - At `n=1`, Node.js is lowest at `6.63 µs`, followed by Bun, Tokio, and Deno. - At `n=10`, Deno is lowest at `11.97 µs`; Bun and Node.js follow, with Tokio at `27.50 µs`. - At `n=100`, Node.js and Deno remain ahead of Tokio, while Bun sits between them. Tokio does have the best `p99` at `n=1`. Bun has the best `p99` at `n=10` and `n=100`. The chart uses a logarithmic y-axis. Lower is better; the exact values are in the raw report at the end of this page. ## Large payloads: `1 MiB` The picture changes once each message is `1 MiB`: - Tokio is fastest for the large-string test at `n=1` and `n=100`. Bun is just ahead at `n=10`. - Tokio leads the `Uint8Array` test at every batch size: `272.81 µs`, `4.64 ms`, and `37.83 ms`. - Node.js and Deno fall further behind as payload materialization becomes a larger part of the round trip. These timings include the trip to the worker and the trip back, including the work needed to materialize the payload on both sides. ## How the payload size changes things This sweep keeps the batch at `100` and grows a binary payload from `8 B` to `1 MiB`. Bun is fastest from `8 B` through `512 B`. Tokio takes the lead from `1 KiB` through `16 KiB`, then Bun and Node.js move slightly ahead through `512 KiB`. At `1 MiB`, Tokio is fastest again at `29.48 ms`. Both axes are logarithmic, so the chart is showing changes across several orders of magnitude. The important shape is the transition from fixed per-message overhead to payload-copying cost. ## A separate shared-ownership reference The next chart is intentionally different. Tokio uses `Arc>`, so sending a value mostly increments a reference count instead of copying the bytes. This is useful as an upper-bound reference for shared ownership, but it is not the default apples-to-apples byte comparison. At `8 B` through `256 B`, Bun is still faster than the Tokio `Arc` path. At `512 B`, the two are close: `74.78 µs` for Bun versus `79.51 µs` for Tokio. Read this chart as “how close can normal transport get to shared ownership?” For the fair default comparison, use the regular `Uint8Array` sweep above. ## How fair is the comparison? The benchmark keeps the important parts aligned: - Both sides create the work up front and wait for the whole batch to finish. - Both use one worker thread, so the sender cannot quietly spread across a larger worker pool. - The default string and byte paths pay for a full round trip. Tokio clones on send and clones again for the worker reply instead of taking a cheaper move-only shortcut. The deliberate difference is memory management. Tokio's `String` and `Vec` paths allocate and copy with `clone()`. Knitting copies into preallocated shared memory and manages those regions itself. That cost is part of what this benchmark is trying to measure, so it is left in the numbers rather than normalized away. The `Arc>` test is kept separate because shared ownership does less work than the normal byte path. ## Why the results change with payload size Knitting keeps transport data in reused typed-array-backed buffers. Small primitive values fit in the per-call header, while larger values spill into a preallocated shared payload buffer. Fixed worker lanes and a small allocator-like bookkeeping layer reduce general-purpose allocation in the hot path. That design has a tradeoff: it adds implementation complexity and memory bookkeeping. The payoff shows up most clearly when copying and allocating a large payload would otherwise dominate the round trip. ## Full raw report The generated report contains the exact average and `p99` tables, ratios, machine details, and methodology notes used for the charts. `summary.md` ````md # Benchmark Summary ## Sources - tokio: `results/tokio-1773827721825.csv` - bun: `results/knitting-bun-1773827812276.csv` - node: `results/knitting-node-1773828074699.csv` - deno: `results/knitting-deno-1773827925019.csv` ## Machine Specs - OS: Ubuntu 23.10 - Kernel: 6.5.0-44-generic - Architecture: x86_64 - CPU: AMD Ryzen 7 4700U with Radeon Graphics - Topology: 8 logical CPUs, 1 socket(s), 8 core(s)/socket, 1 thread(s)/core - Memory: 15.1 GiB - Swap: 4.0 GiB ## Methodology Notes - The main string and byte benchmarks are intended to compare the same logical round trip on both sides: send payload, receive it in the worker, echo it back, receive it again on the caller, then wait for the whole batch. - In `src/main.ts`, the `string` and `Uint8Array` paths go through knitting transport in both directions. That transport materializes a fresh payload on receive, so the round trip includes payload work on both the request side and the reply side. - To keep the Tokio baseline fair, `src/main.rs` clones `String` and `Vec` on send and also clones again on the worker reply. The reply clone is intentional. Without it, Tokio would be measuring a cheaper return-path move while the JS runtimes were still paying for fresh payload materialization on the way back. - The `Arc>` sweep is intentionally separate and is not the default apples-to-apples byte benchmark. It exists as an upper-bound shared-bytes reference for small payloads. `Arc::clone` only bumps a refcount, so it is expected to be cheaper than copying bytes. - This means the default `string` and `Uint8Array` tables should be read as the fairer comparison, while the Arc section should be read as "how close does the normal transport get to shared ownership for small values?" ## Batch Avg Latency (less is better) ```text benchmark | batch | tokio | bun | node | deno ---------------------+-------+-----------+----------+----------+--------- number f64 (8 bytes) | n=1 | 13.01 us | 7.35 us | 6.63 us | 21.54 us number f64 (8 bytes) | n=10 | 27.50 us | 13.41 us | 17.28 us | 11.97 us number f64 (8 bytes) | n=100 | 89.55 us | 80.61 us | 62.33 us | 63.28 us large string 1 MiB | n=1 | 221.35 us | 1.19 ms | 2.85 ms | 1.38 ms large string 1 MiB | n=10 | 6.20 ms | 6.01 ms | 10.79 ms | 10.38 ms large string 1 MiB | n=100 | 37.93 ms | 50.16 ms | 84.66 ms | 83.90 ms Uint8Array 1 MiB | n=1 | 272.81 us | 1.30 ms | 2.35 ms | 1.14 ms Uint8Array 1 MiB | n=10 | 4.64 ms | 5.27 ms | 5.22 ms | 6.36 ms Uint8Array 1 MiB | n=100 | 37.83 ms | 47.76 ms | 54.95 ms | 60.04 ms ``` ## Batch P99 Latency (less is better) ```text benchmark | batch | tokio | bun | node | deno ---------------------+-------+-----------+----------+-----------+---------- number f64 (8 bytes) | n=1 | 16.85 us | 18.70 us | 26.25 us | 160.57 us number f64 (8 bytes) | n=10 | 40.83 us | 36.58 us | 81.26 us | 111.89 us number f64 (8 bytes) | n=100 | 203.57 us | 92.37 us | 314.11 us | 263.05 us large string 1 MiB | n=1 | 371.10 us | 3.74 ms | 3.73 ms | 2.97 ms large string 1 MiB | n=10 | 8.44 ms | 8.41 ms | 15.75 ms | 14.45 ms large string 1 MiB | n=100 | 40.81 ms | 61.62 ms | 105.51 ms | 106.67 ms Uint8Array 1 MiB | n=1 | 400.45 us | 3.03 ms | 5.58 ms | 5.81 ms Uint8Array 1 MiB | n=10 | 8.41 ms | 7.80 ms | 9.52 ms | 14.37 ms Uint8Array 1 MiB | n=100 | 43.71 ms | 59.77 ms | 72.39 ms | 81.71 ms ``` ## Avg Ratio Vs Tokio ```text benchmark | batch | bun/tokio | node/tokio | deno/tokio ---------------------+-------+-----------+------------+----------- number f64 (8 bytes) | n=1 | 0.56x | 0.51x | 1.66x number f64 (8 bytes) | n=10 | 0.49x | 0.63x | 0.44x number f64 (8 bytes) | n=100 | 0.90x | 0.70x | 0.71x large string 1 MiB | n=1 | 5.36x | 12.86x | 6.22x large string 1 MiB | n=10 | 0.97x | 1.74x | 1.68x large string 1 MiB | n=100 | 1.32x | 2.23x | 2.21x Uint8Array 1 MiB | n=1 | 4.75x | 8.62x | 4.18x Uint8Array 1 MiB | n=10 | 1.13x | 1.12x | 1.37x Uint8Array 1 MiB | n=100 | 1.26x | 1.45x | 1.59x ``` ## Uint8Array Size Sweep Avg Latency (less is better) ```text size | tokio | bun | node | deno --------+-----------+-----------+-----------+---------- 8 B | 82.99 us | 62.80 us | 88.20 us | 107.01 us 16 B | 81.91 us | 56.24 us | 65.37 us | 95.78 us 32 B | 85.70 us | 49.48 us | 65.76 us | 85.05 us 64 B | 76.98 us | 42.68 us | 66.88 us | 78.27 us 128 B | 92.53 us | 53.53 us | 79.28 us | 84.39 us 256 B | 99.70 us | 63.42 us | 83.89 us | 100.44 us 512 B | 86.67 us | 68.55 us | 97.07 us | 118.03 us 1 KiB | 101.42 us | 171.09 us | 157.61 us | 169.50 us 2 KiB | 191.25 us | 194.62 us | 220.68 us | 233.39 us 4 KiB | 195.56 us | 260.16 us | 324.39 us | 391.43 us 8 KiB | 208.84 us | 397.05 us | 465.89 us | 539.98 us 16 KiB | 279.25 us | 649.18 us | 741.47 us | 899.81 us 32 KiB | 1.48 ms | 1.14 ms | 1.27 ms | 1.41 ms 64 KiB | 2.71 ms | 2.38 ms | 2.49 ms | 2.89 ms 128 KiB | 5.66 ms | 5.02 ms | 5.11 ms | 6.08 ms 256 KiB | 11.92 ms | 10.56 ms | 10.11 ms | 11.96 ms 512 KiB | 23.06 ms | 22.33 ms | 22.66 ms | 25.26 ms 1 MiB | 29.48 ms | 46.97 ms | 52.53 ms | 55.77 ms ``` ## Arc Comparison Size Sweep Avg Latency (less is better) Tokio uses `Arc>` here as a separate shared-bytes reference point, not the default apples-to-apples byte path. ```text size | tokio | bun | node | deno ------+----------+----------+----------+---------- 8 B | 80.76 us | 70.31 us | 86.19 us | 97.79 us 16 B | 79.35 us | 60.94 us | 73.73 us | 77.46 us 32 B | 81.48 us | 57.04 us | 70.26 us | 77.03 us 64 B | 80.14 us | 54.44 us | 75.94 us | 78.81 us 128 B | 79.89 us | 68.50 us | 82.51 us | 85.95 us 256 B | 79.48 us | 50.59 us | 85.94 us | 100.10 us 512 B | 79.51 us | 74.78 us | 97.23 us | 123.11 us ``` ## Arc Comparison Avg Ratio Vs Tokio ```text size | bun/tokio | node/tokio | deno/tokio ------+-----------+------------+----------- 8 B | 0.87x | 1.07x | 1.21x 16 B | 0.77x | 0.93x | 0.98x 32 B | 0.70x | 0.86x | 0.95x 64 B | 0.68x | 0.95x | 0.98x 128 B | 0.86x | 1.03x | 1.08x 256 B | 0.64x | 1.08x | 1.26x 512 B | 0.94x | 1.22x | 1.55x ``` ```` --- # Why Knitting URL: https://knittingdocs.vercel.app/extras/why/ Why Knitting exists: move hot JavaScript functions off the main thread without turning them into services. The first time one function gets expensive, the usual advice is to pull it out into a service. That is a lot of machinery for one function. A parser grows. A validator picks up rules. A hash, render, or compression step starts taking more of the event loop than it should. Nothing else about the app has changed: one function got heavy, and the fix on offer is new infrastructure. Knitting is for that gap. Keep the code where it is, move the work off the main thread. ## The hot-function problem CPU pressure in a JavaScript app rarely spreads evenly. Real codebases often have thousands of functions, but only a few of them create most of the cost. That changes the size of the fix. When the cost is concentrated in a handful of functions, the response can be concentrated too — moving the expensive part does not have to mean moving the application around it. ## The usual choices When the main thread starts paying for those functions, JavaScript developers usually reach for one of three options. **Do nothing.** It is simple, and for a while it works. But every millisecond a hot function spends on the main thread is a millisecond the server cannot spend on everything else. Cheap routes queue behind expensive ones. Scaling out dilutes the pain, but it does not remove the hot path from the event loop. **Workers.** Closer to the mark: the work leaves the main thread and stays on the machine. What makes it painful is everything around it. The runtime hands you `postMessage` and leaves the rest to you — request IDs, routing, errors, promises, lifecycle, payload choices. That plumbing is the original reason Knitting exists. **Make it a service.** Sometimes that is the right call. Separate ownership, deploy cadence, and failure domains are all real reasons to cross the network. But for one hot function, owned by the same team, in the same repo, the bill is strange: a network hop, serialization, deployment, monitoring, and one more thing to operate. The function did not ask for a hostname. It asked for somewhere else to run. ## The fourth shape Knitting adds a smaller boundary: the code stays where it is, and only the execution moves. The function stays in your repo, in your language, in your types — one `import` away from its caller. What changes is *where it runs* and *what it can touch*. You export it, hand it to a pool, and call it like the async function it already was. Underneath, it runs on a real thread or an isolated process. ```ts import { createPool, isMain } from "knitting"; export const hello = (name: string) => "Hello " + name; if (isMain) { using pool = createPool({})({ hello }); console.log(await pool.call.hello("World!")); } ``` That is the whole boundary: one export, one pool, one call. `hello` still reads like a plain function, but it no longer runs on the main thread. Call that a **function-level execution boundary**: a way to move the expensive part without moving the whole app. Compare what each option costs for the same few hot functions: | | In-process | Hand-rolled workers | Microservice | Knitting | |---|---|---|---|---| | Main thread protected | no | yes | yes | yes | | Code stays in the app | yes | mostly | no | yes | | Call-site ergonomics | function call | protocol you wrote | HTTP client | function call | | Transport cost | none | `postMessage` per call | network + JSON | shared-memory mailboxes | | Isolation available | none | thread only | full, always-on | a dial, per pool | | New deploy unit | no | no | yes | no | Two things make that boundary worth having. **Transport cost matters.** Workers can feel disappointing when reaching the worker costs more than the work. Knitting uses shared-memory mailboxes instead of the runtime's message queue, so small and medium calls stay practical. The [Architecture](/extras/architecture/) page explains the mechanism, and the [benchmarks](/benchmarks/introduction/) show where the shape wins and where it does not. **Isolation should match the task.** A microservice gives you process isolation whether or not you need it. Knitting lets each pool pick its own boundary: in-process guards, a bootstrap hook, runtime-native permissions, or a real OS-sandboxed process. Trusted math can stay on cheap threads. An untrusted plugin can run behind `bwrap`, and [`importTask`](/guides/defining-tasks/#module-loading) keeps its code off the host entirely. Knitting does not replace distributed systems. If you need cross-machine scale, separate failure domains, or team-level ownership boundaries, you still need the network. Knitting is for the moment before that, when the code belongs in the app but the work no longer belongs on the main thread. ## Where it's going > Note: Everything below is roadmap as of June 2026. None of it is shipped API, and the shape may change or be cut as the design matures. Build against what the [guides](/guides/defining-tasks/) document today. Knitting today is task-call oriented: one request in, one response out. The transport underneath it is more general than that. Mailboxes, payload buffers, and named mappings leave room for other shapes. 01 Nearer term

Channels

Some same-host work is not a task call: progress, long-lived coordination, producer/consumer flows. Today the honest answer is MessagePort. A future channel API would keep that shape on the same shared-memory transport, without pretending every conversation is request/response. 02 Longer term

Cross-language, same-host

The mailbox protocol is not tied to JavaScript. Process workers already open named mappings, which makes the boundary more about memory than a specific runtime. A later version could let JavaScript hand large payloads to another runtime on the same machine without copying through JSON or re-marshalling at every hop. That is why the transport is built with more care than a worker pool strictly needs. Until those APIs ship, though, judge Knitting on what it does now. ## Read next - [Quick Start](/start/quick-start/) -- the ten-line version of everything argued above. - [Architecture](/extras/architecture/) -- the mailbox transport, lane model, and safety layers in detail. - [Process workers](/guides/process-workers/) -- the hard end of the isolation dial: sandboxes and containers. - [Examples](/examples/intro_examples/) -- the hot functions behind this argument, as copyable code. --- # Architecture URL: https://knittingdocs.vercel.app/extras/architecture/ How Knitting's shared-memory transport, mailbox protocol, and lane model fit together. This page follows one `call.myTask(args)` from your code down to shared memory and back. You don't need any of it to *use* Knitting -- the [Quick start](/start/quick-start/) and [Creating pools](/guides/creating-pools/) guides have you covered -- but it helps when you are curious why things are fast, chasing a strange bug, or reaching for the advanced knobs. The real implementation has many more moving parts -- fast paths, edge cases, platform quirks -- left out here so the big picture stays readable. Take the numbers and steps below as the shape of things rather than a contract. ## High-level picture Knitting has three layers: 1. **API layer** -- `task()`, `createPool()`, `call.*()`, `shutdown()`. This is what your code touches. 2. **Dispatch layer** -- the host handler: lane routing, the optional balancer strategy, native work stealing, and the inliner. This decides *where* a call runs. 3. **Transport layer** -- shared-memory mailboxes, payload buffers, wakeups. This moves data between threads without going through the runtime's message queue. Knitting runs **thread** workers by default. It can also run each worker as a **separate process** (for sandboxing or containers): the transport is the same shared-memory idea; a process worker just reaches the memory through a *named* mapping instead of an inherited handle. See [Process workers](/guides/process-workers/) and [Shared memory](/guides/shared-memory/). ## Transport: shared-memory mailboxes This is the core of Knitting's speed advantage. Instead of using `postMessage` (which serializes data, queues it, and deserializes on the other side), Knitting writes directly to `SharedArrayBuffer` regions visible to both threads. ### The mailbox model In the private-lane topology, each worker has two independent mailboxes: - **Request mailbox** (host -> worker): the host writes a call header here, the worker reads it. - **Response mailbox** (worker -> host): the worker writes the result header here, the host reads it. Compatible multi-worker pools can use native work stealing instead. The host publishes requests into one shared submit region, workers claim available work from that region, and each worker retains a private response mailbox. This shared-submit/private-return layout prevents a busy worker's private request lane from holding work while another worker is idle. The payload format and response registry are unchanged. Each mailbox has **32 slots** -- the slot index is a 5-bit field, so 2⁵ = 32 slots per direction. A slot is a small fixed-size region that holds: - Task function ID (`Uint16`, supports up to 65,536 tasks per pool) - Payload type tag - Small inline values (numbers, booleans, short strings fit directly in the header) ### Slot ownership via a two-word lock Slot state lives in two 32-bit atomic words: `hostBits` and `workerBits`. They sit on **separate 64-byte cache lines**, so the host and worker never fight over the same line -- the *false sharing* that would otherwise bounce a cache line between cores and collapse throughput. A slot is **busy** when the two words disagree on its bit and **free** when they agree: $$ \text{free} = \sim(H \oplus W)\ \&\ \text{MASK} $$ In the private-lane topology, publishing work is a single atomic toggle of one bit -- not a message copy. Because each direction has exactly one writer (the host writes requests, the worker writes responses), this is a single-producer/single-consumer queue per direction: no mutex, no critical section, just write the payload, then publish the bit. The reader acquires the bit before it trusts the bytes. With native stealing, the shared submit region adds a claim handshake so multiple workers can compete for available regions. Once a worker claims one, the payload publication and private response path use the same shared-memory protocol. For private request lanes, the lifecycle of a slot is: 1. Host atomically claims a free slot in the request mailbox. 2. Host writes the call header (task ID, payload tag, inline data or payload offset). 3. Host notifies the worker. 4. Worker reads the slot, executes the task. 5. Worker claims a slot in the response mailbox, writes the result. 6. Worker notifies the host. 7. Host reads the result, releases the response slot, and resolves the promise. ### Wakeups When the host writes a request, it needs to wake a potentially parked worker. When a worker writes a response, it needs to wake the host handler. Knitting uses `Atomics.notify` (futex-style wakeups) for this. Workers finish available work before waiting. The idle cycle is a bounded spin when one worker is on the request's critical path, then a park: _Diagram: Worker idle cycle: spin, then park, waking to run work and looping back_ 1. **Spin** -- a one-worker pool checks the mailbox for 50 µs. Multi-worker pools skip the spin, leaving the CPU for the host or an awake peer. `Atomics.pause` reduces power draw when spinning. 2. **Park** -- call `Atomics.wait` with a timeout of `parkMs`. The thread sleeps until notified or the timeout expires. 3. **Wake** -- `Atomics.notify` from the host breaks the park immediately. Before this cycle, workers flush completed results, release return-side payload resources when safe, and make a best-effort GC call when the runtime exposes one. In a stealing pool, a completed response is flushed before another claim attempt. This keeps idle workers cheap without delaying work already finished. ## Data path: payload buffers Values that don't fit in the mailbox header slot need a separate path. Each worker allocates two `SharedArrayBuffer` regions: - **Request payload buffer** -- for arguments sent host -> worker - **Return payload buffer** -- for results sent worker -> host Sizes are controlled by `payloadInitialBytes` (default 4 MiB) and `payloadMaxByteLength` (default 64 MiB). When growable `SharedArrayBuffer` is available in the runtime, buffers start small and grow on demand. Otherwise, they're allocated at max size upfront. ### Three encoding paths (fast -> slow) | Path | When it's used | Cost | | --- | --- | --- | | **Header-only** | Primitives (`boolean`, `null`, `undefined`, `number`, small `string`, small `bigint`, `Date`) | Near zero. Value fits in the slot header. | | **Static payload** | Small typed arrays, short strings that overflow the header | Low. Copies into a fixed region of the payload buffer. | | **Dynamic payload** | Objects, arrays, large strings, `Error` | Higher. Requires allocation, encoding (similar to JSON serialization), and copying. | Shared-memory types (`SharedArrayBuffer`, `ProcessSharedBuffer`) and the ownership-move `BufferReference` skip all three: the bytes aren't copied, only a small descriptor crosses. The [Performance guide](/guides/performance) has exact tier ratings per type. ## Dispatch: lanes and routing A **lane** is an execution target -- either a worker thread or the inline lane. The total lane count is `threads + (inliner ? 1 : 0)`. When you call `call.myTask(args)`, the host **handler**: 1. Resolves any promise arguments on the host. 2. Picks a private lane using the balancer strategy, or publishes to the shared submit region when native work stealing is active. 3. Encodes the arguments and writes them to the selected mailbox or shared region (or queues them for the inliner). 4. Returns a promise that resolves when the response arrives. The handler runs on *every* call; the **balancer** only decides *which* lane it lands on. With a single worker and no inliner there is nothing to balance, so the balancer is bypassed entirely. ### Native work stealing With stealing enabled, the host does not predict which worker will execute a call. A worker claims a submit region when it can make progress, executes the task, and writes the response to its own return mailbox. The host uses the pool-global pending registry to settle the promise for the task's caller. Stealing changes request ownership, not task semantics. Completion order still follows task completion rather than submission order, so use `Promise.all()` or another explicit ordering strategy when order matters. ### Balancer strategies These are the values of the `balancer` option on `createPool`: | Strategy | Behavior | | --- | --- | | `roundRobin` (default) | Rotates through lanes in order. Simple, fair, predictable. | | `firstIdle` | Picks the first lane with no in-flight work, falls back to round-robin. | | `randomLane` | Picks a random lane. Good for uneven task durations. | | `firstIdleOrRandom` | First idle lane, else random. Balances fairness with load distribution. | ### Handler backoff The host handler has its own stall-avoidance logic: - `stallFreeLoops` (default 128): how many immediate notify-check loops run before backoff starts. - `maxBackoffMs` (default 10): ceiling for exponential backoff delay. Under sustained high load, the handler stays in tight loops. Under intermittent load, it backs off to avoid burning CPU while idle. ## Inliner: the host as a lane The optional **inliner** adds the host thread itself as an execution lane. Inline tasks skip the entire transport layer -- no encode, no mailbox write, no decode. The task function runs directly on the main thread. Inline execution is deferred to a macro-task boundary (via `MessageChannel`) so the handler can service worker sends/receives first, then the host drains inline work. Key details: - `position: "first" | "last"` -- where the inline lane sits in the balancer's lane order. - `batchSize` -- how many inline tasks run per event-loop tick. - `dispatchThreshold` -- minimum in-flight calls before the inline lane is eligible. - Abort signals on inline tasks use a static toolkit where `hasAborted()` always returns `false` (inline tasks can't be individually cancelled since they share the host thread), while `now()` behaves normally and reads the host's clock. See [Inliner guide](/guides/inliner) for when to use it and when to avoid it. ## Task lifecycle (end to end) Once the handler has picked a lane, that lane's **tx-queue** drives the round trip. Here's a single `call.add([1, 2])`: _Diagram: Lifecycle of one pool.call invocation across user code, host tx-queue, shared memory, and the worker loop_ 1. **User code (host):** `pool.call.add([1, 2])` returns a `Promise` immediately -- the input isn't a promise, so nothing is awaited first. 2. **Host tx-queue:** encodes the call header + payload into a free request slot. (`[1, 2]` is a small tuple, so it takes the static-payload path at a known offset.) 3. **Host tx-queue:** toggles `hostBits` to publish the slot. 4. **Worker loop:** had been spinning, then parked on the signal word. The signal word changed, so it wakes. 5. **Worker loop:** decodes the args and runs the task fn -- `([a, b]) => a + b` returns `3`. 6. **Worker loop:** writes the result into a response slot and toggles `workerBits` to publish. 7. **Host:** drains the result from shared memory and toggles its bit to release the slot. 8. **Host:** resolves (or rejects) the `Promise` returned in step 1. If the task throws, step 6 writes an error result instead, and step 8 rejects the promise. ## Pool lifecycle ### Startup (`createPool`) 1. Allocates shared memory regions (mailboxes + payload buffers) for each worker. 2. Spawns `threads` workers -- threads by default, or separate processes when configured. Each worker imports the task module to discover exported `task()` values. 3. Workers enter their idle spin/park loop, waiting for work. 4. If `permission` is set, generates runtime-specific CLI flags and passes them via `workerExecArgv`. 5. Returns the pool -- a typed `{ call, shutdown }` object that is also disposable, so a `using` declaration can close it for you. ### Running - Calls flow through the dispatch layer continuously. - Workers spin briefly after completing work (in case more arrives), then park. - The host handler manages backoff independently. ### Shutdown A `using` pool shuts down on its own when the scope ends. Calling `shutdown()` yourself runs the same teardown -- just earlier, the moment you ask for it. Either way: 1. `shutdown()` signals all workers to stop. 2. If `resolveAfterFinishingAll` is `true`, workers finish all pending promises before exiting. 3. All in-flight `call.*()` promises for abort-aware tasks reject with `"Thread closed"`. 4. Worker threads terminate. ## Abort signal pool Tasks defined with `abortSignal: true` or `abortSignal: { hasAborted: true }` use a shared-memory bitset to track cancellation state. The pool has a fixed capacity (default 258, tunable via `abortSignalCapacity`). When the host calls `.reject()` on an abort-aware promise, it flips a bit in the shared bitset. The worker can poll `signal.hasAborted()` to check that bit and bail out early. This is cooperative, not preemptive -- the worker must check. If it doesn't, the host promise still rejects immediately, but the worker task runs to completion in the background. ## Memory layout (per worker) | Region | Default size | Purpose | | --- | --- | --- | | Request mailbox | Fixed (32 slots) | Call headers, host -> worker | | Response mailbox | Fixed (32 slots) | Result headers, worker -> host | | Request payload buffer | 4 MiB initial, 64 MiB max | Argument data | | Return payload buffer | 4 MiB initial, 64 MiB max | Result data | | Abort signal bitset | Scales with `abortSignalCapacity` | Cancellation flags | With default settings and 4 workers, the shared memory footprint is roughly: `4 workers x (2 mailboxes + 2 x 4 MiB payload buffers) ~ 32 MiB initial` Payload buffers grow on demand up to `payloadMaxByteLength` if the runtime supports growable `SharedArrayBuffer`. ## Safety layers Workers run code the host may or may not trust, so isolation is a dial, not a switch. Knitting stacks four independent layers -- cheapest and softest first, costliest and hardest last. They compose: trusted local tasks can stop at Layer 1, while an untrusted plugin can be pushed all the way to a sandboxed process. This is defence in depth; no single layer is assumed sufficient on its own. _Diagram: The four safety layers, from in-process guards to a full OS sandbox_ - **Layer 1 -- in-process guards (always on).** Before any task module loads, the worker neutralizes the most dangerous calls: `process.exit`, `process.kill`, `process.abort`, and `Deno.exit` are redefined to throw, and the raw shared-memory handles are scrubbed from the data object visible to task code. Cheap, but a guardrail against accidental misuse -- not a wall against a hostile co-resident. - **Layer 2 -- bootstrap hook.** `worker.bootstrap` runs a privileged module once per worker, *before* task imports -- the right place to strip env vars, install your own guards, or freeze globals. - **Layer 3 -- runtime permissions (strict by default).** The policy is translated into each runtime's native enforcement (Deno's permission flags, Node's permission model, Bun's equivalents), so the boundary is the runtime's, not a library check task code could bypass. - **Layer 4 -- process + real sandbox.** The only OS-enforced boundary: `worker.runtime: "process"`, optionally launched through `bwrap`, Docker, or `systemd-run` via `processCommandPrefix`. Pair it with `importTask` so the isolated code never loads in the host at all. See [Permissions](/guides/permissions/) and [Process workers](/guides/process-workers/) for the full configuration. ## What Knitting does NOT do Knowing where Knitting stops is as useful as knowing what it does: - **No message passing protocol.** Knitting is task-call oriented. If you need pub/sub or event-style messaging, use `postMessage` / `MessagePort`. - **No preemption.** A long-running task blocks its lane until it finishes. Use `abortSignal` with `hasAborted()` polling for cooperative cancellation. - **Same host only.** Workers can be threads *or* separate processes on one machine, but the shared memory never crosses the network -- there is no cross-machine transport. Pair Knitting with one if you need to scale out. - **No automatic scaling.** Worker count is fixed at pool creation. You choose the parallelism level upfront. - **Browser support is limited.** There is a [browser build](/guides/browser/), but shared memory there needs cross-origin isolation (`COOP` / `COEP` headers) because of Spectre-era constraints. Most Knitting work still happens server-side.