Twenty test servers picked a port with `NNNN + Math.floor(Math.random() * 500)`
and hoped it was free. Several files guessed inside the *same* range —
sdk.test.ts and gates.test.ts both on 19500, three files on 22000, and
webrtc-integration.test.ts used 22000 twice within itself — so with test files
serving in parallel the collision was not unlikely, it was scheduled.
Found from the outside, in Nova, where Shade is vendored: across eight runs of
an unchanged tree, two had failures and six were clean. The visible error was
Failed to start server. Is port 24287 in use? EADDRINUSE
and then a second, louder one: `teardownRig` ran on a rig that setup had never
finished building and died on `rig.alice`, so the TypeError from cleanup is
what you read first and the real cause scrolled past above it.
`Bun.serve({ port: 0 })` lets the OS hand out a free port and `server.port`
reads it back. There is no range to collide inside.
teardownRig now tolerates a rig that was never built. Cleanup must never be the
loudest thing in a failing test.
A suite that fails a quarter of the time is worse than a suite that fails: it
teaches everyone to re-run it, and then a real regression looks like the flake.
Verified: 1198 pass / 0 fail here, and the same fix is going into Nova's
vendored copy so the next sync does not bring the guessing back.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014489bKUtUEY1Zgs9xN9mt7
aliasSession knew that two labels name the same peer, then threw that
knowledge away. The binding lived only in the caller's memory, so a
restart lost it — and the peer could not repair it from its side.
First contact forces the receiver to label a session by the only sender
hint a relay surfaces, an 8-byte signing-key fingerprint (`fp:<hex>`).
Once the peer announces its canonical address, aliasSession moves the
session there. But the peer keeps sending under `fp:<hex>`, because its
transport derives the same label from the same hint every time. After a
restart the session sat under the canonical address, inbound frames
resolved to `fp:<hex>`, and nothing matched. The peer held a valid
session so it never re-ran X3DH: the failure was permanent, and only a
manual re-link cleared it.
Observed in Prism as `No session for address: fp:579c3b335d66e2c0` on
every receive for three days, with a phone whose every RPC timed out.
StorageProvider gains saveSessionAlias / getSessionAlias /
removeSessionAliasesFor, optional so third-party implementations keep
compiling, and implemented across all seven backends. Lookups resolve
through resolveLabel(), which runs BEFORE the peer mutex — locking the
alias while mutating the canonical session would let an aliased and a
canonical caller ratchet the same state concurrently.
A live session under a label always wins over an alias, and prekey
envelopes never resolve: both keep a re-link establishing a fresh
session instead of being redirected into the stale one. Aliases are
dropped in resetSession and acceptIdentityChange, and memoized so the
hot path costs no extra read.
The sdk.test.ts case that asserted a dead fp-label encoded the old
behaviour; it now pins the new contract.
Verified: 1166 tests pass (from 1160). With alias persistence disabled
as a negative control, 5 of the 6 new tests fail, including both
restart cases.
Also drops `baseUrl` from the consumer-strict tsconfig — removed in
TS 6.0, and it was failing the typecheck that gates publishing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two follow-ups to the V4.8.2 duplicate-fan-out fixes Prism filed.
1. `Inbox.acceptBridgeFrame(blob)` + shared 4096-entry msgId LRU.
The relay durably stores blobs and pushes them to every active
delivery channel; without a cross-channel ack the bridge frame
ran first and the next inbox-poll re-dispatched the same blob
~30 s later, tripping on consumed prekeys. Bridge consumers now
plumb pushed frames through `acceptBridgeFrame`, which shares
the dedup gate + ack path with `pollOnce`. Whichever channel
delivers first wins; the other acks-and-skips. Inbox records
the msgId before the ack so a parallel poll can't observe an
in-flight ack window.
2. `Shade.aliasSession(oldLabel, newLabel)`. First-contact forces
the receiver to label the new session by the relay's sender
fingerprint hint (`fp:<senderfp>`); the post-decrypt plaintext
typically announces the peer's real address. Aliasing moves
session, trusted identity, peer-verification, and identity-
version under the canonical label. Holds the per-peer mutex on
both labels (lexicographic order) so concurrent crypto ops can't
observe a half-moved state. Refuses to overwrite an existing
session at the new label.
Wire change: `IncomingMessage.expiresAt?` now surfaces the relay's
expiry so receivers can pass bridge frames straight to
`acceptBridgeFrame` without inventing a TTL.
Tests cover bridge-then-poll, poll-then-bridge, aliasSession happy
path, refuse-to-overwrite, and same-label no-op.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two interlocking robustness fixes for the duplicate-fan-out / first-contact
class of failures Prism reported.
1. `Shade.receive(from, env)` now queues its `manager.decrypt` step
per `from` so concurrent dispatches can't race the SessionManager
ratchet or the StorageProvider (sqlite "database is locked", IDB
transaction conflicts). User message handlers run *outside* the
queue so streams + file-RPC's nested `shade.receive` calls don't
self-deadlock.
2. Bridge WS + SSE handlers now run a per-connection bounded msgId
LRU as defense-in-depth against any flushTo re-entry (event-storm,
future refactor). Pending-flush chains are wrapped in `.catch(() =>
{})` so a transient `ws.send` rejection no longer poisons the
connection's flush loop.
Tests: storming `inbox.blob_stored` 10× per PUT yields exactly one WS/
SSE frame; 8 concurrent `bob.receive('alice', envelope)` calls keep
the ratchet intact and never surface "database is locked".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Expose the local device's 32-byte Ed25519 identity public key on Shade
so apps can hand it to their own backend at enrollment time for
signature verification, key pinning or per-device safety-number
computation. Closes the gap that forced consumers to ship placeholder
random bytes their backend could store but never verify against.
- @shade/sdk Shade.identityPublicKey: Promise<Uint8Array> — getter
mirrors the existing fingerprint accessor. Throws pre-init,
reflects the current key after rotate(), retired key preserved in
retired-identities storage per existing grace-period contract.
Private key remains unreachable.
- Test in shade-sdk/tests/sdk.test.ts: round-trip match against the
underlying storage's signingPublicKey, plus value updates after
rotate().
- Lockstep version bump 4.3.0 → 4.4.0 across all 25 packages.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>