Internals: Durable workers
How runtime:workers behaves underneath: what "durable" is actually promising, where the state lives, who is allowed to open it, and what each decision costs.
For signatures see the API reference. The reasoning and the rejected alternatives are DECISIONS D80.
The gate is the design
A durable worker's method can write and return. The write is committed to SQLite asynchronously; the return value is not handed to the caller until that commit has happened.
async add(item) { const items = this.state.get("items") ?? []; items.push(item); this.state.set("items", items); // not awaited return items.length; // the caller waits for the commit anyway }
That ordering is the whole promise. A process that dies mid-call is a call that never returned — not one that returned something the disk never heard about. It is also what makes coalescing safe: several sets in one call become one transaction, because nothing has left the process yet to be contradicted.
What the gate does not cover is a side effect issued in the middle of a call — a fetch, a message on a port, a row written to another database. Those leave before the method returns, so the gate has not run yet.
await this.state.sync(); // …then the effect await fetch(webhook, { method: "POST", body });
This is the same shape Cloudflare's output gate has, and the same caveat.
Measured. A test kills a real esrun with SIGKILL after five acknowledged appends and finds five on restart. It is in crates/runtime-cli/tests/durable_workers.rs, and it fails if the gate is removed.
Why the state is resident, and therefore capped
state.get(k) is synchronous — a Map lookup, not an await. That is possible because the whole key/value page is read when the worker is opened and kept in its heap.
Resident state has a memory cost that somebody has to bound, so it is bounded out loud: 1 MiB per worker, 128 KiB per value, refused at the set with ERR_DURABLE_STATE_TOO_LARGE rather than nudged with a warning. A soft cap on a resident cache is an unbounded cache with a comment.
The consequence is a boundary you can state: durable-worker state is the small, hot thing a request needs immediately — a cart, a session, a counter, a cursor. Anything that accumulates belongs in runtime:db, which is in the same runtime and needs the same grants.
The value format
Values are stored as structured-clone bytes — the same serialization postMessage and structuredClone already use — so a Date comes back a Date, a Map a Map, a BigInt a BigInt, and a cyclic object cyclic.
Date | Map | Set | BigInt | cycles | |
|---|---|---|---|---|---|
| JSON | string | {} | [] | throws | throws |
MessagePack (runtime:serialization) | string | object | array | throws | — |
| Structured clone | Date | Map | Set | BigInt | kept |
Those MessagePack rows are measured, not assumed, which is why the choice was made once rather than deferred: changing the format later means rewriting every worker's file, and a data migration is a far worse thing to owe than a decision.
A codec tag is stored beside every value. State written by a newer runtime, in a format this build does not know, is refused by name (ERR_DURABLE_STATE_FORMAT) rather than handed to a deserializer that will misread it.
A set that stores what is already stored writes nothing: the encoded bytes are compared with what is on disk first. That comparison is worth its cost because the storing is the expensive half — a commit that changes a page costs milliseconds against one that changes none — and "read it, put it back" is what a handler written over resident state does all day.
One process owns a directory
Each worker's state is its own SQLite database, under <dir>/<class>/<xx>/<hash>.db, with a _registry.db beside them holding the catalog of what exists.
The file name is a hash of the id, never the id. An id is a string a program chose: it may hold slashes, it may be 400 characters, and on macOS and Windows Cart and cart are the same file. The id itself is stored inside the file and checked when it is opened, so a hash collision is an error rather than two workers quietly sharing a state.
A directory belongs to one process, and nothing in this module arranges that. The embedded engine takes an exclusive lock on a database file for as long as it is open, and the operating system drops it when the process ends. So the guarantee needs no heartbeat, cannot be lost while it is held, and leaves nothing stale behind when a process is killed — the three ways an advisory lock file goes wrong. What the module adds is the sentence: an engine Locking error about a file the caller never named becomes ERR_DURABLE_LOCKED, naming the directory.
The cost is stated rather than hidden: two processes cannot share a durable directory at all, not even to read. A second one is refused until the first exits — which is the correct behaviour for an overlapping redeploy, and a real constraint for anything that wanted two readers.
One connection is one conversation
Each worker's database has exactly one writer — its own flush — but the catalog is shared by every worker in the process, and twelve materializing at once would put twelve statements on one connection. The embedded engine does not refuse that; it panics from its WAL, which panic containment turns into a JavaScript exception a caller cannot act on. So every catalog statement goes through one queue. It is a small thing that took a flaky test to find, and it is an argument for the file-per-worker layout that was not the reason for it.
What waits in that queue together commits together. Opening a worker, changing an alarm and closing a worker all write the catalog, and each statement used to be a commit of its own: on a disk where a commit is a sync, that capped the whole process at the disk's sync rate. Whatever is queued when the connection comes free now runs as one transaction, and each caller is answered when it has committed. The catalog keeps its syncs, even though it is only an index: it is what makes an alarm findable, and a row lost in a crash would be an alarm that never fires.
The mailbox
Calls to one worker queue and run one at a time, in the order they were made. There is no lock to take, because there is nothing to take it against: the state has exactly one reader and one writer, and they are the same call.
"One at a time" is about running code, not about the disk. The next call starts when the previous call's code has finished, while that call's writes are still committing; each result is released once the commit that carries its writes has landed. So a burst of calls on one busy worker (a hot product's stock, say) shares its commits instead of paying one sync per call. A call can read what the previous call wrote before it is on disk, but nothing leaves the worker ahead of the disk: that call's own result waits for a later flush, and flushes are ordered. A side effect issued mid-call is the exception, as it always was, and state.sync() before it covers the previous call's writes too.
The queue is bounded (mailbox, 1024 by default). Past it a call is refused with ERR_DURABLE_BUSY rather than waited on — a queue that grows without limit is a failure that has been hidden rather than reported.
Two workers of the same class are two mailboxes. Work on one is not work on the other, which is the reason to address state by worker rather than by row.
A cycle is refused. A worker calling back into one that is calling it would wait for ever, because a mailbox is strictly one at a time. So every call carries the chain of workers it came through, in runtime:context, which follows awaits and timers and is sent across to a shard. A call to a worker already in the chain throws ERR_DURABLE_CYCLE, naming the chain. A timeout was the alternative, and it would have turned a certain deadlock into an occasional slow failure.
A chain only sees one request. If A is in the middle of calling B for one request while B is in the middle of calling A for another, no chain contains both, and the mailboxes still wait on each other. So the host also keeps a wait-for graph: an edge A → B while a call made from A waits on B. Because a mailbox runs one entry at a time, a cycle in that graph is a deadlock whichever requests built it, and the call that would close one is refused before it waits. The one false alarm: a call a method starts and does not await still counts as its worker waiting.
The same chain gives calls between workers an output gate. A call made from inside a worker is a message leaving it, so it waits until the caller's writes so far are committed, the barrier the reply already waits for. Without that, a worker that wrote an order and then called the inventory could have the sale committed while its own record was still in memory.
Eviction has no timer
A worker that has been idle past evictAfter is closed: its writes are flushed, stop("idle") runs, and its database handle is released. It is opened again, with its state, the next time it is addressed.
That sweep runs when work arrives, never on a timer. This runtime's timers keep the process alive, so a repeating sweep would be a reason a program could never exit — a script that used one durable worker would sit there for the whole idle window with nothing to do. The cost is that a worker alone in a quiet process stays open until something else happens, which costs one file handle.
Collections: a blob the database cannot read, beside columns it can
A collection is a table per name: id, the document as structured-clone bytes, and one real column for each field the class declared, with an index on it.
That is the whole trade, and it is deliberate. The blob is what keeps a Date a Date — the property the keys have and JSON does not — and it is exactly what makes the document opaque to SQL. So what you want to query, you declare, and it is copied out into a column where the database can order it. What you do not declare is still stored, still returned, and simply not indexed.
{ scan: true } is the escape hatch and says what it costs: the rows are read, the documents decoded, and the filter applied here. On a small collection that is honest work; on a large one it is a full read, which is why it cannot happen by accident.
Declaring a field later
The column is added on the first wake after the deploy — and then filled in from the documents already stored. Without that backfill the column would be null for everything written before the field existed, and a query over it would silently return a subset. That read-and-update pass is the cost of declaring a field late, and it is paid once, on the worker's first wake, where it is visible as a slow first call rather than as wrong answers later.
Nothing is ever dropped. A field or a collection removed from the schema keeps its column and its table: a class that stopped asking for something is not the same as data nobody wants, and a schema edit that deletes rows is one nobody can undo.
One connection, two chains
A worker's connection has two users that know nothing about each other: the flush behind a set, and whatever a collection is doing. Both go through one serial chain, for the reason the catalog does — the engine panics rather than refuses when two statements overlap.
A transaction is the exception, and it needs a second chain rather than a bypass. It holds the outer chain for its whole body, so nothing else can reach the connection; statements inside it join an inner chain instead, which keeps them one at a time without waiting for the holder. A set inside a transaction therefore stops scheduling a flush of its own and is written by the transaction, which is what makes state.transaction cover both halves of the storage.
A query is read in one turn rather than streamed while the caller iterates. A cursor would hold the connection across the loop body, and the obvious loop body — await state.set(…) per document — would then be waiting on the connection its own iteration is holding. .limit() bounds the read instead.
Alarms: an index that may be early, never late
A worker's alarm time is stored twice — in its own file, where it is the truth, and in the catalog, where it is what makes the alarm findable without opening every worker to ask. Two files means no transaction spans them, so the order of the two writes is the design:
the catalog, moved earlier only (
MINof what is there and the new time),the worker's own file,
the catalog again, exactly.
A crash anywhere in that sequence therefore leaves the index early, and an early index costs one wake-up that opens a worker, finds nothing due, and puts the index right. A late one would be an alarm that never fires. Opening a worker reconciles the two in the same way, so a directory repairs itself.
The scheduler asks the catalog for the minimum future time among the classes it runs and sleeps until then — no polling loop, no fixed interval. Setting an earlier alarm wakes it immediately; alarmPoll (60s) is only a ceiling, so a clock that jumps cannot leave it asleep for ever.
Why the class list is required
startAlarms({ classes }) will not guess. A class is a JavaScript value in one process's module graph; whether this deployment is the one meant to service a given class is not something the runtime can read off it. If the scheduler had fallen back to "classes something has addressed so far", an alarm would fire on a busy process and not on an idle one — the failure mode that is hardest to see and worst to have. Rows for classes not listed are excluded by the query itself, so they are neither woken nor stepped over.
Firing
An alarm goes through the worker's mailbox, so it cannot interleave with a call, and it reads as cleared inside the handler: a handler that sets the next time repeats, one that sets nothing is done. It is only removed from disk once the handler has finished, though, so a process that dies mid-handler runs it again after a restart. That makes an alarm at-least-once, which is what a scheduler that never drops work has to be; a handler's effects should be idempotent. A failure puts one back — 1s, 2s, 4s, to a five-minute cap — with the attempt count stored beside the alarm, so a restart does not reset it, and the last failure is reported rather than dropped.
The process stays alive while the scheduler runs, and this is the one place this module holds a timer at all. Which is why it is not started for you: eviction deliberately has no timer so that a script exits, and an alarm scheduler that started itself would undo that for anyone who merely set a time.
Sockets that hibernate
A worker that installed its own listeners on a WebSocket would lose them the moment it was evicted, while the socket stayed open: messages would arrive at listeners on an instance that no longer exists. So the only workable choice was to never evict such a worker, which makes a thousand idle chat rooms a thousand resident workers.
Instead, the agent that holds the directory holds the socket. When a worker calls ctx.acceptWebSocket(ws), the listeners are installed here, keyed to the worker, not to its instance. Eviction then frees the instance and leaves the socket connected. When a message arrives for a worker that is not live, it is materialized first (its start() runs), and then webSocketMessage runs as one mailbox entry, serialized with calls and alarms like any of them. Events on one socket are delivered in order.
A socket reaches a worker as an argument because that is how everything reaches a worker. It crosses as a handle rather than a structured clone, since a connection cannot be cloned and the worker needs that very one. On a shard, the handle is a proxy: sends, closes and attachments are messages to the host, which owns the socket as it owns the database. The shard keeps a mirror of its sockets' tags and attachments, so getWebSockets() and deserializeAttachment() stay synchronous.
A send is gated. A message to a client is a message leaving the worker, so ws.send() waits until the worker's writes so far are committed, the same barrier a reply and a call to another worker wait for. A client is never told something the disk has not heard; the test for it ends the process the moment the client hears back, and finds the write.
What survives what. An attachment (at most 16 KiB) and the tags survive hibernation, because they are held with the socket. Neither survives a process restart, and neither does the socket: a TCP connection dies with its process. Cloudflare can keep sockets across a restart because its edge holds them; this runtime has no such tier, so clients reconnect.
Heartbeats do not wake anything. An auto-response is matched on the host against the exact text of a message and answered there, so an application-level ping keeps a connection alive without materializing a worker.
Shards: the code moves, the storage does not
A shard is a Worker that runs the classes' code. Every database handle stays on the agent that owns the directory, and a shard reaches its state by message. Two things follow from that, and both are the reason for it:
One writer per file, still. The single-writer rule above is the engine's file lock, and it holds per process, not per thread. Opening a worker's file from a shard would be a second connection in the same process, which the lock does not stop. Keeping every handle on one agent is what keeps the rule.
A shard needs no filesystem grant. It holds
imports, so it can load the module its classes are in, and nothing else unless the program says so.
What a shard costs is one message each way per call, and one per turn of writes. Next to a commit that is small.
What stays the same
A shard keeps its own resident copy of each worker's keys, so state.get is still a map lookup. Writes are encoded and checked against the ceilings in the shard, sent to the host as encoded bytes, and resolved when the host has committed them. The host never decodes them: a flush, a size and the same-bytes comparison need only the bytes.
The gate is unchanged. The shard waits for its writes to be committed before it answers a call, and the host's gate then finds nothing left to wait for.
A transaction is held open by the host while its body runs in the shard. Everything the body sends in the meantime joins it, and the shard says at the end whether to commit or roll back. A function passed to collection.update cannot cross a thread, so it runs in the shard, between a read and a write. Nothing can come between those two, because the mailbox runs one call at a time.
Placement
A worker's shard is a hash of its class and id. The answer has to be the same every time: a worker is one file, and two shards holding it would be two states. A queue or a least-loaded choice would be fairer and would break that.
Failure
A shard is lost in one of three ways: an error nothing caught, running out of memory, or not returning to its event loop. The first two are the Worker's own error event. The third needs a watchdog, because a synchronous loop never yields to anything that could report it.
The watchdog sends a heartbeat every second to each shard that has work in flight, and ends a shard that misses three. A shard waiting on I/O, or on the host, still answers: its loop is free between awaits. So a miss means the isolate is not getting back to its loop at all, and no amount of waiting will fix that. terminate() can end it. Nothing can end a loop on the agent a server is running on, which is why this protection is only available on a shard.
When a shard is lost:
the call in flight rejects with
ERR_DURABLE_SHARD_LOST;calls queued behind it were never sent, so they run again on a replacement;
a transaction it had open is rolled back.
State is on disk, so its workers come back on their next call. The pool is refilled on its next use rather than straight away, so a module that fails every time it starts produces failing calls rather than a restart loop.
Lifetime
A Worker keeps the process alive until it is terminated. A pool of them would make every sharded script hang after its last call, for the same reason eviction has no timer. So a shard is referenced only while something is waiting on it, and the watchdog runs only while some shard is busy.
A shard starts by saying so
A message sent to a worker before it is listening reaches nobody. So a new shard speaks first: when its half of the protocol is installed it sends "ready", and only then is it told its module and ceilings. A module that never imports runtime:workers, or whose top-level code never finishes, fails the start after 30 seconds instead of leaving a call waiting for ever.
On real storage
Every durable write is a commit, and every commit is a disk sync. So what a durable worker costs is mostly a count of syncs, multiplied by what the storage charges for one. The same release build, on the same machine, on tmpfs and on a spinning disk:
| tmpfs | Spinning disk | |
|---|---|---|
| A call that touches no state | 4.5 µs | 4.4 µs |
| A call that reads resident state | 5.3 µs | 5.1 µs |
| A call that writes, one after another | 82 µs | 11.2 ms |
| The same, 500 in flight on one worker | 8.5 µs | 14.9 µs |
| The first call on a new worker | 2.1 ms | 59 ms |
bench/durable-workers.js, best of 3, release build, Linux, 2026-09-25.The rows that do not write are the same on both, because the state is resident. The row that writes one call at a time is the device's sync rate. The row below it is why a busy worker is not bound by that rate: consecutive calls share their commits (the mailbox section above). Nothing shares a sync across workers, though, because each worker is its own file, so work spread over many workers is bound by the device.
The shop example makes that concrete. It uses a new worker per customer, a handful of writes per journey, and 32 customers at once:
| tmpfs | Spinning disk | |
|---|---|---|
| Journeys a second | 212 | 6 |
| Add to cart, p50 / p99 | 18 ms / 131 ms | 0.9 s / 4.0 s |
| Checkout, p50 / p99 | 66 ms / 154 ms | 2.3 s / 4.4 s |
Acknowledged orders lost to SIGKILL | 0 of 114 | 0 of 101 |
| Checkouts in flight at the kill, finished after restart | 14 of 14 | 15 of 15 |
| Restart | 64 ms | 70 ms |
bench/durable-shop.js (WORKDIR selects the disk), release build, Linux, 2026-09-25.Durability does not depend on the disk; throughput does, by a factor of about thirty here. Durable workers need storage with fast syncs, meaning an SSD or NVMe. A spinning disk keeps every promise and runs at a small fraction of the speed.
What memory costs
A worker's own state is in the JS heap and capped at 1 MiB. The larger cost is native: every open worker keeps its SQLite file open, and the embedded engine gives each open database two buffer arenas of 3 MiB, of which about 2.2 MB ends up resident. That size is fixed inside the engine (turso_core, checked up to 0.8.0-pre.13) and no pragma changes it.
maxLive | Resident memory at 500 / 1000 / 1500 / 2000 workers | JS heap |
|---|---|---|
| 16 | 92 / 107 / 124 / 130 MB | 7 MB |
| 64 | 201 / 217 / 226 / 235 MB | 22 MB |
| 128 | 346 / 362 / 373 / 380 MB | 8 MB |
| 256 | 609 / 627 / 642 / 651 MB | 17 MB |
bench/durable-memory.js, release build, Linux, 2026-09-25: 2,000 new workers, one write each.So memory is set by maxLive, not by how many workers exist: about 2.2 MB per open worker, on top of the process. Workers past maxLive are evicted and cost nothing until they are called again. The slower rise across each row is the catalog's own page cache growing with its rows, and it levels off: a 6,000-worker run stayed at 416–419 MB from the 2,000th worker on. Size maxLive to the memory you have: the default of 128 is about 280 MB of open files.
When shards pay
Shards move the classes' code to other threads. They add a message each way per call (about 1 ms at p50 in the shop) and each costs a V8 isolate's memory. What they buy is isolation: a memory ceiling and a watchdog that end one shard instead of the process, and CPU work that does not block the thread serving requests.
They do not buy throughput for work that waits on storage. The shop, which waits on disk and not on CPU, runs slightly slower with shards (212 journeys a second unsharded, 192 with two, 188 with four). Use shards when a worker's methods do real computation, or when one misbehaving worker must not take the server with it, not to make I/O-bound workers faster.
What this is not
Not an actor model. The non-goal — no process model, scheduler, preemption, mailboxes or supervisors — is about the runtime, and it stands: nothing in the Rust crates gained a scheduler, nothing preempts, and no agent gained a mailbox. A durable worker is a value in guest JavaScript with a queue in front of it. The runtime still does not decide when your code runs.
The module is written in JavaScript over runtime:db, runtime:fs and runtime:hashing, and adds no capability and no Rust. That is the same rule runtime:db's driver kit set for database backends, applied to the layer above them.
Not yet
Durable execution (retries, a step journal, workflows) is deliberately not planned as a second subsystem. With alarms and workers in place it is a class on top of this primitive.