Compare commits

..
Author SHA1 Message Date
molluskandClaude Opus 5 9f06741b99 core/teardown: an unconfirmed stop is not a clean stop
Codex's re-review of the branch returned "approve with follow-ups" — no
blocking findings, five P3s. All five are applied here rather than carried as
debt, since each is a few lines.

The one with user-visible consequences: `stop_host` returned a bare "was
sharing" bool, so the single case where the availability-first policy gives up
(SIGKILL queued, reap never confirmed) still sent `ScreenShareStopped` with
nothing else. The UI would say sharing had ended while pixelpass might still be
alive and fanning out — a claim the user cannot see through. `shutdown` now
returns `StopOutcome`, `stop_host` returns `Option<StopOutcome>`, and an
unconfirmed *user-initiated* stop raises a UI error naming the stray process.
Session and viewer teardown discard the outcome deliberately: nobody is waiting
on an answer there, and the residual risk is already logged.

Also: the three failure diagnoses in `shutdown` (the signal never left, the
child ignored it, the wait itself broke) were collapsed into one log line and
are now distinct — they mean different things to whoever reads the log.

The second cancellation gate is the one worth keeping. The review pointed out
that all cancellation coverage sat in the *graceful* wait, so a mutant that
disarmed the wrapper between the two waits would survive. It was right, with a
wrinkle: the naive mutant does not compile, because the child is borrowed from
`self` for the whole function — the borrow checker is doing real work here. The
restructured form (`self.child.take()` once cooperation has failed) does
compile, and the pre-existing mid-wait test passes it.
`cancelling_shutdown_after_the_kill_leaves_the_fallback_armed` kills it.

Mutation-verified, both new gates: reporting an unconfirmed stop as `Reaped`
fails exactly `a_failed_wait_is_not_treated_as_a_confirmed_reap`; disarming
between the waits fails the new cancellation test (and the failed-wait test,
which also asserts armedness) while leaving the old mid-wait test green — which
is the proof the new test is not redundant. The logging split is diagnostics
only and has no gate; said plainly rather than dressed up as covered.

Docs: the "four mutations" line is now an explicit table naming each target and
its test, with 0c's pair counted under 0c; and the aggregate teardown latency is
recorded as a deferred item with a trigger (a fourth routine child, or a
measured teardown over 5 s) instead of an unwritten known cost.

638 lib tests, clippy clean, fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:43:04 -04:00
molluskandClaude Opus 5 3aa768af52 docs: the 0b DAG row says four mutations, matching §10 round 14
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:11:19 -04:00
molluskandClaude Opus 5 1cfa932fbe docs: record the 0c mechanism probe, and revise 0b's mutation matrix
The 0c mechanism probe was run on this host before any structural work, because
one unverified assumption could have invalidated the whole approach: whether a
hand-created `adapter` node is visible to pipewire-pulse under the name the
capture path depends on. It is. Five gates green, including the two that
mattered — `<node.name>.monitor` is exposed as a Pulse source, and SIGKILL of
the owning connection removes both Pulse-visible names with zero graph residue.
No null-sink module is involved at any point. Every O1 stop condition for 0c is
retired, and the default sink never moved, so the probe is safe on a live
desktop.

The probe also settled the native-sink scope question: it applies to every mode
that owns a capture sink, not only `DesktopExcluding`. That makes the `--repair`
rework load-bearing rather than defensive — repair derives dead PIDs only from
`module-null-sink` entries, so once the sink is native its loopbacks become
undiscoverable orphans.

§10 gains rounds 14 and 15: the 0b matrix drops to four mutations because the
best-effort wake arm is unreachable by construction (the loop owns a sender, and
the biased select would win anyway), and teardown is hoisted to one unconditional
post-loop site instead of being duplicated across one live arm and one dead one.
Round 15 records the two blocking implementation-review findings and the vacuous
gate of my own that the review's test-double critique exposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:11:06 -04:00
molluskandClaude Opus 5 d8b8fd79cf core/teardown: stay armed across the wait, and never call a failed wait a reap
Codex's review of the two commits below returned "changes requested" with two
blocking findings. Both were real.

1. `ReapOnDrop` disarmed itself across the async wait. `shutdown` moved the
   child out of the wrapper with `take()` before the first `.await`, so if that
   future was cancelled or unwound mid-wait, the raw child dropped with nothing
   but `kill_on_drop` (signals, does not reap) while `Drop` found `None` and did
   nothing — the AEC could then unload over a live child. That is precisely the
   hole the type exists to close, left open for the duration of every wait. The
   child now stays owned by `self` across every await and is released only on a
   *confirmed* reap.

2. A failed wait was silently converted into success, and the hard-kill path was
   unbounded. `wait_reaped` discarded `io::Result`, so a wait error made the
   timeout return `Ok` and shutdown returned as though the reap were confirmed;
   meanwhile a process stuck in uninterruptible sleep after SIGKILL could wedge
   the core command loop forever. The trait now preserves the result, both waits
   are bounded, and the conflict case has an explicit written policy: we choose
   availability, leave the child owned so the bounded Drop retry stays armed,
   and log the residual risk rather than hiding it.

Codex also showed the test double was flattering the implementation in four
ways. All four are closed: the fake can now be cancelled mid-wait, can fail its
wait, and can take several polls to die, and the grace is pinned independently.

That last one caught a flaw in my own gate. The elapsed-time assertion compares
against `STOP_GRACE` itself, so setting the constant to zero leaves it vacuously
true — both sides move together. `the_grace_is_a_real_interval` pins the
constant to a band instead, and now kills that mutation directly.

Mutation-verified again, five mutants, each killed by its own gate: disarming
the wrapper (cancellation test), treating a wait error as success (failed-wait
test, exactly one), a zero grace (the new band test), a single poll instead of
the drop loop (delayed-reap test, exactly one), reversed field order (the two
ordering tests).

Also applies the matrix adjudication, which Codex and I reached independently:
teardown moves out of the reliable close arm to ONE unconditional site after the
loop, so every `break` is covered structurally — including any added later —
instead of duplicating teardown across one live arm and one provably dead one.

637 lib tests, clippy clean, fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:09:58 -04:00
molluskandClaude Opus 5 92a64465a4 core/teardown: ask pixelpass to stop before killing it
The peerspeak half of phase 0c (design v3.4 §7.4 item 1). Stop Share and
session teardown both went straight to `Child::kill()`, i.e. SIGKILL, which
skips pixelpass's own cleanup and leaks one null-sink module every time.

`ReapOnDrop::shutdown` now asks first: SIGINT, a bounded 2 s wait, then SIGKILL
only if the child ignored the request. SIGINT specifically, not SIGTERM —
pixelpass installs only a `ctrl_c()` handler, so SIGTERM would take the default
disposition and be indistinguishable from SIGKILL.

Signalling by pid is safe against pid reuse here: we have not reaped the child,
so it is a zombie whose pid the kernel reserves until we wait it, and the pid
cannot name a stranger. (Same reasoning that dismissed pixelpass bug #6.)

The grace is 2 s because it is awaited inline in the core command loop, so it
is also how long a wedged child can delay other commands. A healthy pixelpass
never spends it.

The drop/unwind path deliberately stays a hard kill: `Drop` cannot await, and
there the ordering invariant (§7.1) outranks tidiness. Once the pixelpass half
of 0c lands, the capture sink is connection-owned and that path stops leaking
by construction.

`libc` becomes a direct unix-only dependency, pinned to 0.2.186 — the version
already in the tree via alsa/cpal/tokio — so Cargo.lock gains one line and no
new code enters the build.

Mutation-verified, five mutations, each killing its own gate: no wait (6 fail),
reversed field order (2, reap test green), no reap loop (2, ordering test
green), no SIGINT (6), no SIGKILL fallback (exactly 1 — the wedged-child test).
633 lib tests, clippy clean, fmt clean.

Not yet field-tested: the live Stop Share gate (SIGINT sent, child exits within
the bound, no fallback kill on the normal path) still owes a real run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:26:05 -04:00
molluskandClaude Opus 5 6ba763774d core/teardown: reap the screen-share children before the AEC unloads
Phase 0b, fixes 1-3 of design v3.4 §7.2 (decision D4). The invariant is that
the echo-cancel module must not unload while a pixelpass host is alive and
fanning out; two paths have to honour it and only one is code we get to run.

The explicit path: `ActiveSession::shutdown` now awaits
`ScreenshareTeardown::shutdown_children`, and the reliable command channel's
close arm tears the session down explicitly instead of letting it drop on the
way out of `run_core_loop`.

The drop/unwind path: the ordering-critical fields move out of `ActiveSession`
into `core::teardown::ScreenshareTeardown`, where `echo_cancel` is the LAST
declared field and therefore the last dropped. Previously it was declared
first (`:682`, ahead of `screenshare_host` at `:685`), so an unwind unloaded
the AEC while the host was still live — and unwind is reachable, the core is
full of `unwrap()` and has no `panic=abort` profile.

Killing is not enough. `kill_on_drop(true)` only signals: it hands the child to
the runtime's orphan queue and returns, which an unwinding runtime may never
drain. `ReapOnDrop` blocks on a bounded 250 ms budget until the child is really
gone, because a bounded stall beats unloading the AEC out from under a live
pixelpass.

Everything is generic over a narrow `ChildProcess` trait and over the guard
type, so ordering is unit-testable without spawning processes or loading
PipeWire modules — the seam idiom already used by `replace_viewer_index`.

Mutation-verified, and the plan's demand that mutations 4 and 5 prove
*different* defenses holds: reversing the field order fails only the
AEC-ordering tests and leaves the reap test green; removing the reap loop fails
only the reap tests and leaves the ordering test green. Removing the explicit
wait fails the explicit-path tests. 631 lib tests, clippy clean, fmt clean.

⚠️ Mutations 1 and 2 of the pinned matrix do not both exist: the best-effort
wake arm is unreachable by construction, twice over. Documented at the site;
adjudication owed in the impl plan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:19:07 -04:00
molluskandClaude Opus 5 692ad677d2 docs: record F11-1 closed — boundedness needs a resolved Client
Design v3.7 §6.1.1 gains the round-13 box (the rule, the ordering that is
load-bearing in both directions, and why bridging deliberately still uses the
full union); the phase-5 results file records the close with the measurement
the deferral was waiting for; the impl plan's phase-6 gate note drops F11-1.

pixelpass c78eb2d is the implementation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 04:46:25 -04:00
mollusk bf908adbf0 docs: phase 5 matrix PASSED (13/13) — design round 10, results run 2
The §5.1 dry-run audit gate passes. All 13 rows completed with a non-empty
eligible half in every one, and O5 is re-measured on the fixed graph: worst
recompute 67 µs across every run, 10 µs mean under deliberate churn, busy
fraction 0.0006, readiness 1-2 ms with 18 binds against a 2000 ms budget.

Round 10's finding, and it is the third of exactly the same shape: the
pipewire-pulse PID derivation required a SINGLE repeated pipewire.sec.pid.
WirePlumber repeats one too (two Clients, both sec_pid 1747), so the derivation
returned None permanently on a stock desktop, key 4's suppression never fired,
and every Pulse-emulated node fused into one owner. Row 1's CLEAN control
forwarder and Firefox were both excluded. Fixed in pixelpass 91c4ded by deleting
the heuristic: probe every distinct sec_pid and let /proc/<pid>/comm decide.

Three measured rounds now, all at the observation boundary, none in the
architecture -- and all three were fail-closed and silent, caught only because
§5.1 requires asserting what must remain ELIGIBLE. An exclusion-only checklist
would have passed every one of these builds.

Rows 4-6 are closed through peerspeak's REAL tagging sites (call, mpv, notify,
plus clip) with a hand-launched mpv staying eligible, so the cross-repo contract
is proven end to end on live nodes. Row 10 covers the full sticky lifecycle
including retirement; row 11 is provably non-vacuous (the recycled
node.link-group came back byte-identical and did not inherit taint).

Recorded and NOT claimed as passes: substitutions in rows 8, 9 and 13, and two
reporting-only findings (the audit's sticky flag is uninformative; a bridge's
named key is lost when a leg reappears under a new serial).

Phase 6 remains blocked by F11-1, phases 0b/0c/0d and the Stereo Mix design
call -- this file removes one gate, not all of them.
2026-07-26 02:59:43 -04:00
mollusk b68fca689e Merge phase 1: ownership tagging (SPA-JSON carriers via libspa)
peerspeak-side half of phase 1 of the screenshare audio-exclusion work: every
node peerspeak owns carries two registry-visible ownership carriers, so the
taint engine has a primary root that survives the registry's filtered global
event (design v3.5 section 6.7).

Reviewed by Codex over rounds 10-12; all findings verified and dispositioned.
Round 12's F12-1 (rfind('}') spliced carriers inside a trailing comment, a
fail-open) and F12-2 (depth ceiling taken from the consumer,
pw_properties_update_string, not from the spa-json-dump grammar) are fixed and
live-verified through the real ALSA plugin.

623 lib tests green, fmt clean, clippy clean.
2026-07-26 02:14:25 -04:00
molluskandClaude Opus 5 c82ef07464 audio/ownership: take the depth ceiling from the consumer, not the grammar
Round 12 review, finding 2 — filed as P2, and the interesting part is
that its author retracted it to P3 once we had measurements, while the
remedy it originally proposed would have been a fail-open.

The finding was that our validator rejects nesting `spa-json-dump -s`
accepts, and the suggested fix was a recursive sub-iterator walk to
match the dump tool. Both halves rest on the dump tool being the
reference. It is not. Nothing reads `PIPEWIRE_PROPS` or `PIPEWIRE_ALSA`
with `spa-json-dump`; `pw_properties_update_string` does, in the client
process.

Measured live on this host, against the real ALSA plugin:

    depth 513  dump accept   plugin accept   ours accept
    depth 514  dump accept   plugin accept   ours REJECT
    depth 515  dump accept   plugin REJECT   ours reject
    depth 1000 dump accept   plugin REJECT   ours reject

At 515 the plugin discards the whole object: the node came back as
`alsa_playback.aplay` with no properties at all. So matching the dump
tool would have made us splice carriers into values the consumer throws
away wholesale — losing both, which is the echo this feature exists to
prevent. Over-rejecting costs a routing preference; over-accepting costs
a carrier. Those are not the same price.

What was genuinely wrong is narrower: we sat exactly one level below the
consumer. `pw_properties_update_string` calls `spa_json_container_len`
on a container value, which enters one more sub-iterator before its flat
walk, and that single level is the entire discrepancy. Doing the same
puts the boundaries on the same number.

Codex reached the same three numbers independently by calling
`pw_properties_update_string_checked(NULL, ...)` directly, having
disassembled both call sites; I measured through the live plugin. Two
methods, one table.

The dump differential stays, but it is now labelled a *grammar* oracle
with a warning not to add deep values — it would fail by design. The
acceptance oracle is the new boundary test.

Mutation-verified: removing the container step fails the 514 assertion.
622 -> 623 lib tests, fmt clean, clippy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 01:59:36 -04:00
molluskandClaude Opus 5 9eab6c118d audio/ownership: let libspa say where the object closes
Round 12 review, finding 1 — a measured fail-open, and the third
distinct door into the same failure.

A trailing comment is valid SPA-JSON and *ends the document*
(`case __COMMENT: return 0` in spa/utils/json-core.h), so an object may
close before the last `}` in the string. `merge_pipewire_props` located
the closing brace with `rfind('}')`, which is a byte scan and not a
parse, so for

    { "target.object" = "my-sink" } # trailing }

it selected the comment's brace and spliced both ownership carriers
*into the comment*. The re-validation did not catch it, because the
result parses perfectly well — as `{ target.object = "my-sink" }`, with
neither carrier present. Confirmed against `spa-json-dump -s`.

That is an untagged node, so no taint root, so echo — exactly what
rounds 10 and 11 each closed by a different route. Latent rather than
live: pixelpass's evaluate() is still audit-only, so today it corrupts
an audit classification and becomes a leak when phase 6 consumes
eligibility.

The whole thesis of round 11 was "do not re-implement someone else's
grammar". The scanner went, but this brace hunt stayed behind in the
caller, which is the same defect wearing different clothes.

So spa_object now reports the object's own closer, taken from libspa:
closing a container at depth 0 writes the brace's position back to the
parent iterator, and spa_json_enter made `outer` that parent. Read
before the trailing check, which advances past it.

Also:
- whatever followed the object is preserved, so a user's trailing
  comment survives instead of being silently deleted;
- the output check now asks whether the object closes where we put our
  brace, not merely whether the string parses. A parse-only check is
  what this finding defeated.

Mutation-verified: restoring `rfind` fails the new test, and dropping
the tail fails it on the deleted comment. Honest note in the code —
mutation cannot distinguish the closer comparison or the is-object
test; both are labelled belt-and-braces rather than presented as
tested.

621 -> 622 lib tests, fmt clean, clippy clean, and the ignored
spa-json-dump differential still agrees.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 01:51:14 -04:00
molluskandClaude Opus 5 21ba633825 audio/ownership: validate inherited SPA-JSON with libspa, not a scanner
Round 11 review, findings 2, 3 and 4.

The round-10 fix replaced a brace check with a hand-written scanner. That was
the wrong shape: a second implementation of someone else's grammar drifts in
both directions at once, and measured against `spa-json-dump -s` on this host
it did.

It ACCEPTED `{ "foo" = { garbage } }` (only brackets were balanced, contents
never validated), `{ "a" = "\é" }`, `{ "a" = é }` and `{ "a" = foo\bar }`.
Merging into those put an invalid pair before our carriers, so the daemon
stops at it and drops both -- recreating the exact fail-open the round-10 fix
existed to close. Its own test even pinned `"\é"` as a valid token.

It REJECTED `{ target.object, "my-sink" }`, `{ key == "value" }` and
CR-terminated comments, all valid -- so a user with one of those in their
environment silently lost their routing policy to an overwrite. That half
affects a running Linux user.

Now libspa's own parser validates, and the merge splices into the validated
text instead of re-emitting parsed pairs. Splicing preserves the user's bytes
exactly, which also answers the review's point that re-quoting a bare key can
invent a different one (`foo\bar` -> a string with a \b escape). Three
measured properties make the splice safe -- the last `}` is the object's, a
validated object's brace is never mid-comment, and commas are pure separators
-- and the result is validated again before it is returned.

Mutation testing then deleted the rest: every pairing and recursion check I
had written turned out to be redundant, because spa_json_next already errors
on `{ garbage }` and on nested garbage, and skips containers rather than
descending. ~60 lines of my own grammar logic removed. What remains is gated
by a new differential test against `spa-json-dump -s` over a 27-value corpus
-- the check whose absence caused this round. It found a real disagreement on
its first run (a bare document, which we reject by design, not by accident).

One mutation HUNG rather than failed: dropping the `length < 0` check makes
libspa report the same error without advancing, spinning forever. Kept, now
labelled load-bearing for termination, with a token-count bound beside it.

Finding 4: the ordering test took the first textual match of `fn main`, so a
raw-string decoy above the real function satisfied it while the real one
spawned a thread first. Now requires each of the three anchors to be unique.
Mutation-verified with the review's own decoy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 01:08:17 -04:00
molluskandClaude Opus 5 45b1b97dd8 audio/ownership: pin the no-lost-carrier invariant, and harden the byte scan
Verification round on the round-10 review fixes.

Adds the property the whole of finding 3 is about, stated directly: over
20,000 deterministic inputs built from the exact characters that break
SPA-JSON (braces, brackets, quotes, separators, comment marks, escapes,
newlines, multi-byte characters), the merge always emits both carriers in an
object it can read back. Either outcome — parse and rebuild, or overwrite —
has to end that way, and now nothing can quietly change which.

Also replaces two byte-index steps with character-boundary steps. Both were
correct on the ASCII input they actually see, but `index + 1` after a
reverse find would have split a multi-byte character and panicked the slice.
scan_token gains multi-byte cases for the same reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 00:00:23 -04:00
molluskandClaude Opus 5 ae2e9de523 tests/fixtures: the ownership contract says exact-match, not truthy
Round 10 review, finding 6. The cross-repo contract still documented carrier
1 as "any value other than false/0 is truthy" after R10-4 made pixelpass
match it exactly. A future producer following the fixture could emit "true"
and silently lose the carrier.

Committed byte-identical with pixelpass's copy in the same session, as the
file's own rules require.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 23:57:16 -04:00
molluskandClaude Opus 5 985c63806b audio/ownership: state the playlist policy, and gate main's ordering properly
Round 10 review, findings 2 and 5.

Finding 2 — R10-2's rationale for tagging local playlist audio was factually
wrong. It claimed a local track is "already being broadcast to peers on the
same keypress", but shared listening is opt-in: music_broadcast defaults to
false, play_music_index starts local playback unconditionally, and
broadcast_track returns immediately when can_broadcast_music is false. So a
default-config playlist is not already broadcast.

The tag stays, now as an explicit policy with the real reason: the carriers
reach rodio through PIPEWIRE_ALSA, which is process-wide, and clip_player
and music_player are two ClipPlayer instances in one process — no value of
that variable can tag one and not the other. Exempting the playlist means
giving it a separately taggable stream, which is a large change for a case
with a one-step workaround (play it in any other app). Tagging is not
optional for received clips and peer music, which are the far end's own
audio.

Finding 5 — the ordering test proved only "before run_gui", which a
thread::spawn inserted above the tag still satisfies while making the
set_var a data race. It now requires the tag to be the first executable
statement in main: attributes, `unsafe` and block punctuation are stripped,
and any residue fails. Mutation-verified against a spawn, an unrelated
statement, and the call deleted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 23:51:02 -04:00
molluskandClaude Opus 5 6fc55a286d audio/ownership: parse inherited SPA-JSON instead of trusting its braces
Round 10 review, finding 3. The merge's shape check was the outer braces
only, so an inherited `PIPEWIRE_ALSA='{ garbage }'` was spliced into rather
than overwritten, producing an object the daemon does not accept.

Measured live 2026-07-25, and the failure is worse than a rejection: with
PIPEWIRE_ALSA set to the old merge's output, a real aplay node came up as
node.name=alsa_playback.aplay, no peerspeak.owned, and a junk property
`garbage = "peerspeak.owned"` — the lenient parser ate our key as their
value and stopped. Both ownership carriers lost on a live
Stream/Output/Audio node, which is an echo.

So: parse the inherited object and REBUILD it with our pairs last, rather
than splicing before the closing brace. Rebuilding is what makes the result
independent of the input's formatting — a value ending in a `#` comment
would otherwise swallow everything appended after it.

The three values the new merge emits were verified against the live daemon
(user props preserved, both carriers present) and are pinned byte-for-byte.
scan_token is gated on its own postcondition: at the object level an
unterminated string is also caught by "the object never closed", so the two
implementations only disagree at the seam.

Also parameterizes the malformed-value warning, which always named
PIPEWIRE_PROPS even when PIPEWIRE_ALSA was the malformed one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 23:47:32 -04:00
molluskandClaude Opus 5 d63db68318 audio/ownership: apply the merge rule to the ALSA carrier too
Verification round on round 10's own fixes, not on the next layer.

R10-5 preserved a user's PULSE_PROP and PIPEWIRE_PROPS but
tag_this_process_alsa_audio still clobbered their PIPEWIRE_ALSA, which is
the same kind of routing policy and deserves the same treatment. Both it
and tag_child now merge.

MEASURED, rather than assumed, because "our pairs go last so they win"
was load-bearing for the whole merge design and was never checked:
  PIPEWIRE_PROPS='{ "node.name"="theirs_first", "media.role"="music",
                    "node.name"="ours_last" }' on pw-play
    -> node.name=ours_last, media.role preserved.
  The PULSE_PROP equivalent on paplay -> the same.
So last-wins holds on both grammars: a user who already sets node.name
cannot silently untag us, and their other keys survive.

That also makes tag_child's ALSA carrier merge from the inherited value
safely: in production main has already put this process's `clip` tag
there, and the child's own role now overrides it by coming last. The
existing row could not see this — the test binary never runs main, so it
only ever exercised the merge-into-nothing case. Added a row that drives
the real shape directly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 21:07:33 -04:00
molluskandClaude Opus 5 e7923a1b5c audio/ownership: merge inherited player env vars instead of clobbering
PULSE_PROP and PIPEWIRE_PROPS can legitimately carry a user's own routing
policy — media.role, a target sink — and replacing them changes where the
user's audio goes as a side effect of a tagging mechanism that is
supposed to be behaviourally invisible.

PULSE_PROP is space-separated key=value, so merging is appending;
PIPEWIRE_PROPS is a SPA-JSON object, so it is an insert before the
closing brace. Our pairs go last in both, so they win a duplicate key —
without that, a user with node.name already set would silently untag us.
A value that does not match the expected shape is logged and overwritten:
a half-merged string that fails to parse would drop the tag silently,
which is worse than losing a routing preference. No full SPA-JSON parser,
which would be over-engineering for a case with no live consumer
(measured: neither variable is set anywhere in this user's env or config).

Also sets PIPEWIRE_ALSA on the child, with the child's own role. A player
configured for ALSA output is reached by neither of the other two
variables, so this closes a real gap rather than only a cosmetic one —
and without it such a child would inherit this process's `clip` tag from
tag_this_process_alsa_audio and report the wrong role in the audit.

Corrects a stale doc comment on OWNED_PROP_VALUE that still claimed
pixelpass accepts any truthy value; R10-4 made the match exact. Codex's
F5 was reasoned partly from a stale comment of mine, so these are worth
fixing on sight.

Codex phase-1 review F4. Round 10, R10-5.
8 new rows; 5 mutations verified (clobber PULSE_PROP, our pairs first,
naive object concat, doubled trailing comma, drop the ALSA carrier).
All 4 live ownership gates re-run green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 21:03:32 -04:00
molluskandClaude Opus 5 b5569fe2c6 audio: tag the fourth playback path, rodio's ClipPlayer
ClipPlayer opens a rodio default sink, which on Linux reaches the graph
through PipeWire's ALSA plugin. It was untagged through all of phase 1,
and it is a real echo path: B broadcasts music, A tunes in, A shares
their desktop, B hears their own track played back at them. Confirmed
live as `alsa_playback.peerspeak-...` with no ownership properties.

rodio exposes no way to set PipeWire node properties, so the carrier is
PIPEWIRE_ALSA, set once at the top of main while still single-threaded.

Measured, with PIPEWIRE_PROPS and PULSE_PROP unset, to establish that
setting it process-wide is safe:
  - aplay (ALSA plugin)   -> both carriers land. Confirms the mechanism.
  - pw-play (native)      -> untouched. Our own call-playback and capture
                             streams are native, so they keep their own
                             explicit tagging and are unaffected.
  - arecord (ALSA capture)-> IS tagged, on a Stream/Input/Audio. Not
                             surgical in the role dimension; harmless only
                             because R10-1 honours the carriers on
                             producers alone. This is why R10-1 lands first.

Local playlist tracks are tagged too, not just inbound peer audio. A
local track is already broadcast to peers over the call on the same
keypress, so sharing it again through the screen share would send the far
end two copies at differing latency. That is a defect, not a feature.

Codex phase-1 review F1. Round 10, R10-2.

New live exit-gate row drives the real ClipPlayer; mutation-verified
(drop the tag -> no node within 5s). The wiring guard is mutation-
verified too, and its first version was WRONG: it searched raw source and
passed against a main with the call deleted, because the comment above it
named the function. It strips comments now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 20:52:22 -04:00
molluskandClaude Opus 5 503f78153b audio/ownership: refuse an ambiguous contract fixture
Producer half of the same fix (Codex phase-1 review, finding 3, P2).
This side collected fixture lines into a map, so a duplicated key
silently took the last value while pixelpass took the first — both
repos green on different contracts.

Mutation-verified in both repos with a duplicated `prop_value`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 20:25:27 -04:00
molluskandClaude Opus 5 d40385f85c notify: correct a measured claim about the aplay fallback
The comment said aplay ignores PULSE_PROP/PIPEWIRE_PROPS. Measured:
it reaches the graph through PipeWire's ALSA plugin and carries both
carriers exactly like pw-play and paplay. Comment only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 20:07:08 -04:00
molluskandClaude Opus 5 bcf1343a55 phase 1: tag every audio node peerspeak owns, on both carriers
Zero behaviour change. This is what makes the screenshare exclusion
engine able to see us at all (plan §5.1, impl plan §3): pixelpass must
refuse to fan out our own playback, and until now it had no way to
recognise it.

Two carriers, matched by pixelpass as a union — `peerspeak.owned=1`
and a `node.name` prefix `peerspeak_owned_<role>_<pid>`. Round 8 added
the second after the phase-5 audit found a node property is invisible
to the PipeWire registry `global` event and recoverable only by
binding the node; the prefix is announced directly. A union is also
the fail-closed direction: a missed tag leaks call audio into a share,
a spurious one only over-excludes.

Three tagging sites, all three verified live on this host:
  - native call playback  → props on the stream dict
  - screenshare mpv/VLC   → PULSE_PROP + PIPEWIRE_PROPS on the child
  - notification chimes   → same, on pw-play/paplay

The literals are a cross-repo wire contract, so they appear once here
as named constants and are pinned in a fixture committed byte-identical
in both repos (tests/fixtures/ownership-tag-contract.txt). The contract
test is black-box: it builds a real child `Command` and reads back the
environment it would carry, rather than testing our own formatter.

Three live `#[ignore]`d exit-gate tests drive the real call sites and
poll `pw-dump` for the resulting node — the plan requires the tag be
shown landing on a live node, not just in the env. All three
mutation-verified (drop either carrier, or the role, and the matching
gate fails).

Measured while verifying: mpv, VLC, pw-play and paplay all honour
`node.name` from those env vars. The native stream set neither
`application.name` nor a description, so a mixer fell back to
`node.name` — which the tag turns into an internal identifier. Added
an explicit `node.description = "PeerSpeak"` there, which keeps the
plan's rule (the prefix must not reach `node.description`) while
preserving its intent: mixers stay readable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 19:19:25 -04:00
mollusk 6773a3882b docs: round 9 — uncertainty is not history (design v3.6)
Phase 3r landed §6.7 and the phase-5 audit was re-run against it immediately.
It found a second measured defect within minutes: a real hardware sink
carrying `unresolved-ancestry` permanently, from one link observed while its
output node was still unbound during enumeration. Round 8 made that
systematic rather than rare, because every node is now withheld until its
bind resolves.

New §6.8: sticky taint is a claim about history, and uncertainty is not
history. Retiring by reason code would not be enough — an unresolved node
propagates `TaintedUpstream`, which is indistinguishable from real
contamination once recorded — so the split is by provenance: the engine runs
its fixpoint twice, and only the evidence-only pass may feed sticky state.
Decisions are unchanged and still fail closed.

Also recorded in §6.8, both from Codex's round-9 review and both pre-existing:
hardware playback-to-capture paths ("Stereo Mix") defeat the `session_device`
classifier in a way the driver denylist cannot detect — a real echo path
needing a design call — and the 2 s readiness budget has no calibration
argument beyond one measurement on one idle desktop.

Impl plan: phase 3r marked built and merged with its gate results, including
the extra Device-side live gate and why row 1 alone could not cover it.
2026-07-25 18:51:21 -04:00
molluskandClaude Opus 5 1cd19b355f docs: design round 8 — the observation boundary (v3.5)
The phase-5 dry-run gate failed on its first live run: the engine built to
v3.4 could not see its own primary taint root (echo, AEC off) while excluding
every stream on the machine (silence). One cause — the PipeWire registry
`global` event carries only a filtered subset of an object's properties, and
eight the design depends on are never announced.

Design doc (v3.4 → v3.5):
- NEW §6.7 — the observation boundary. The global is an index, not a source of
  truth: bind every Node and Device, `info` props are the sole source, live
  prop tracking, one readiness obligation per unbound node, fail closed.
  Four user design calls recorded.
- §5.1 — a second, registry-visible tag carrier (`node.name` prefix) alongside
  `peerspeak.owned`, so the primary root does not rest on one mechanism.
- §6.4 — node/device props are not an optimisation to skip, they are
  unavailable from the global; the round-6 Link lesson was right and applied
  to exactly one object type.
- §6.1.0, §6.1.4 — the two corrections the impl plan owed v3.5: a
  time-dependent "hazard is LIVE" claim, and an unreachable nominated test
  case (twice over).
- §9.1 measured facts, §12 rig discipline (pw-dump binds; the registry does
  not), §14 readiness.

Impl plan:
- NEW phase 3r with a four-part exit gate, the first the direct inverse of the
  finding. Ports deliberately not bound in v1, with a revisit trigger.
- Phase 1 pins the second carrier literal as a cross-repo contract.
- Phase 5 marked GATE FAILED; matrix and O5 re-run after 3r and 1.
- Risk register: the over-exclusion row fired and worked; new row for the
  observation boundary class.

Architecture is unchanged and vindicated: fed correct properties, the engine
decided correctly in every fixture. The §5.1 exact-partition requirement is
what caught this — every exclusion was defensible and the eligible half was
empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 17:04:51 -04:00
molluskandClaude Opus 5 297f4397a7 docs: phase 5 dry-run audit results — GATE FAILED, two findings
Impl plan §5's required results file. Phase 6 does not start.

F1 (fatal): the PipeWire registry `global` event delivers only a filtered
subset of node properties, and eight of the properties the phase-3 adapter
reads are not among them — peerspeak.owned, pulse.module.id, node.link-group,
application.process.id, node.passthrough, device.api, factory.name,
alsa.driver_name (plus port.exclusive on Ports). They are silently absent, so
the primary taint root never fires, the AEC identity can never validate, and
session_device is universally false. Measured on PipeWire 1.6.8 /
WirePlumber 0.5.15, with the full announced key set for all five object types
recorded. Links and Clients are unaffected; pulse-PID derivation works.

F2: with F1 in force no node has a strong owner key, so any tainted capture
stream is an unbounded tainted reader and phase 2's fail-closed backstop
excludes every Stream/Output/Audio on the machine. Fail-closed, so silence
rather than echo — but entirely non-functional, and non-functional in a way an
exclusion-only checklist would have scored as passing. The eligible half of
the §5.1 partition is what caught it, exactly as the plan argued it would.

The fix direction is measured and recorded: binding each Node and reading its
info props recovers every missing property, which is the pattern phase 3
already built for Links. factory.id is not a shortcut — it resolves to
"adapter", not api.alsa.pcm.sink.

O5 is closed with ~4 orders of magnitude of headroom: 308 graph events in
6.5s under churn, every recompute under 50us (max 15us), busy fraction 0.0004.
Caveat recorded — measured on the degraded graph, and the F1 fix adds
per-node bind I/O this run did not measure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 15:43:23 -04:00
molluskandClaude Opus 4.8 283d938b79 docs: sequenced implementation plan for screenshare audio exclusion
Turns the converged v3.4 design into ordered phases with falsifiable exit
gates. Three adversarial review rounds with Codex (gpt-5.6-sol, xhigh);
findings adjudicated rather than accepted wholesale, with reachability
verified against source on both sides.

Structural decisions:
- Phase 0d closes BOTH unsafe paths into the capture (source string and
  capture-sink inputs) before any machinery that could take them exists.
- Phase 5 dry-run audit mode is a hard gate: the taint engine runs against
  the live graph, creating no links, asserting exact eligible/excluded
  partitions with reason codes.
- Link manager is deliberately last among the pixelpass components.

Two measured corrections owed back to v3.4 (plan §11): §6.1.0's "hazard is
LIVE right now" has already flipped and must not be gated on, and §6.1.4
nominates an unreachable test case (as did my first replacement for it).

Design approval only. No code, nothing approved for merge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 15:34:27 -04:00
molluskandClaude Opus 4.8 fd72e6f018 docs: close the AEC default-sink question raised by §6.1.0
Checked ~/.config/peerspeak/config.json: output_device and input_device are
both pinned to the Arctis, so echo_cancel::enable always passes sink_master
explicitly and the AEC binds to real hardware regardless of Sunshine owning
the default sink. Not live for this user.

Kept as a low-priority general defect: on "system default", the master args
are omitted (echo_cancel.rs:89-94) and module-echo-cancel binds to whatever
the default is, which on a box like this one is a null sink. Hardening would
be to resolve and validate the default before load. Own task, not this
feature.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 05:22:32 -04:00
molluskandClaude Opus 4.8 8768cd242c docs: v3.4 audio-exclusion — CONVERGED, ready for implementation planning
Round 7. Codex ratifies: v3.3 is ready to become the implementation plan.
Seven rounds, every blocker fixed or refuted with evidence. Design approval
only — nothing approved for merge, no code written.

Two subtle catches from the ratification round, both applied:

- The owner-key union had a wording trap that would have preserved the exact
  bug it was written to fix. "Resolves" must mean "yields a MATCH between the
  two legs", not "first property present on the node" — client.id IS present
  on both gst-launch legs but differs, so a first-present implementation stops
  at key 3, sees a mismatch, concludes "different owners" and leaks. Now
  specified as try-in-order-until-equal, with a dedicated test.
- Sticky taint must be lifetime-aware, not keyed on raw ids. client.id, node
  ids, module indices, link-groups and PIDs all recycle on this stack, so a
  bare key would hand an unrelated future app permanent inherited taint.
  Stored against live owner components, cleared only when all members vanish.

Also added: how pixelpass learns the pipewire-pulse PID itself (consistent
pipewire.sec.pid across Pulse clients, validated against /proc/<pid>/comm),
with the failure modes in both directions — safe only because unresolved
ancestry is fail-closed, which is the invariant the section rests on.

NEW LIVE FINDING (§6.1.0), the strongest reachability evidence yet and one
Codex's sandbox could not have seen: the user's CURRENT DEFAULT SINK is
sink-sunshine-stereo, a support.null-audio-sink. Every hardware sink is
SUSPENDED; the only RUNNING sink is Sunshine's virtual one, with Firefox
playing into it and sunshine reading its monitor. The hazardous forwarder
topology is live in the default audio path full time, with no EasyEffects
involved. It also means the rejected hardware-sink-only shortcut would have
captured NOTHING on this machine. Flagged separately, explicitly UNVERIFIED:
what module-echo-cancel binds to when the default sink is an app-owned null
sink.

§12 expanded with a graph-engine test surface (node-local tests cannot catch
C2/C3-class defects). §14 rewritten: convergence table, agreed v1 scope, and
what is deliberately out.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 04:14:34 -04:00
molluskandClaude Opus 4.8 da72541e18 docs: v3.3 audio-exclusion — owner-key union + sticky taint
Round 6. Codex disagreed with two of my three round-5 claims and was right
about both; I had independently refuted one of them with a sharper test.

C2 REFUTED (by my own measurement): client.id is NOT an owner bridge. One
gst-launch process doing capture+playback produced TWO client objects (209
input, 210 output), no link-group, same application.process.id 20172. So
client.id bridges a *connection*, not an owner, and GStreamer — the same
framework pixelpass uses — splits them by default. Replaced with a
conservative union, strongest first: node.link-group, owned pulse.module.id,
client.id, node application.process.id, else fail closed.

With a trap Codex did not flag: application.process.id is pipewire-pulse's
PID for module-created streams, so bridging on it would fuse every Pulse
module's legs into one owner and mass-exclude tunnel/RTP/loopback audio the
user may legitimately want shared. Never bridge on that key when it equals
the pipewire-pulse PID; keys 1-2 already cover those precisely. PID thus
returns to the design in the CORRELATION role while remaining unusable in
the IDENTITY role — and in that role a wrong answer fails closed.

C3 CONCEDED: taint must be STICKY. Current-topology taint forgets buffered
audio — an app that reads a tainted monitor, buffers, then closes its input
leg would be relinked while still emitting peerspeak audio from the buffer,
and no graph event marks the drain. Taint now persists per owner until its
nodes disappear. Added §6.1.4 quantifying the arrival-side window (~10.6-21.3
ms quantum plus scheduling) and noting it is zero when taint roots already
exist, which is the common case.

C1 SUSTAINED with Codex's caveat: node-granular traversal is free for the
monitor boundary, but over-taints Audio/Duplex nodes. Fail-closed, accepted
for v1, documented as a known contradiction of the "Firefox with a mic stays
shareable" promise on duplex devices.

S1: endpoint props demoted to an optimization; bind-LinkInfo fallback is the
correctness path. S2: readiness epoch + revalidate before each link creation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 04:01:37 -04:00
molluskandClaude Opus 4.8 8610ab2eb6 docs: v3.2 audio-exclusion — signal-graph taint, measured
Round 5. Codex and I converged independently on the same conclusion — a
Link-only ancestry walk does not catch the leak — it from the crate/header/
WirePlumber sources, me from the live graph. Its sandbox could not reach the
daemon (pw-dump: Operation not permitted), so the measurements are mine.

Reproduced the EasyEffects topology with module-null-sink + module-loopback
(same shape, no EasyEffects needed). Result: there is NO Link object between
a forwarder's input leg and its output leg. Walking upstream from the leaking
node over Links alone finds no inbound links at all — a dead end that reads
as "clean". The legs are related only by shared node.link-group / client.id /
pulse.module.id.

So the signal graph needs three edge types:
1. Link edges — measured: registry Links carry all four endpoint props.
2. Sink-monitor — measured FREE at node granularity: the monitor connection
   IS a real Link whose output node is the sink itself. Codex held that this
   must be modelled explicitly; that is true only for a port-granular walk.
   Taint walks at node granularity, links are created per port.
3. Owner bridge — node.link-group when present, else client.id (measured
   shared across the forwarder's legs, distinct per app). Only modules set
   link-group, so client.id is what covers ordinary apps.

New §6.1.1: bridge taint must be CONDITIONAL on the input leg being tainted.
"Client has both legs ⇒ exclude" would exclude every app using a microphone.
Firefox in a Meet call stays shareable; Firefox sharing desktop audio does not.

Also: §6.5 rejects the cheap "hardware-sink-only" predicate with a measurement
— the forwarder's output leg links directly to alsa_output, so the shortcut
passes the leak and excludes the innocent app, backwards on both halves.
§6.3 barrier corrected: core sync/done is a previous-work roundtrip, not graph
quiescence. §6.4 adds crate version, endpoint fast path + bind fallback, and
full-recompute cost. §5.2 correction 5 rewritten: application.process.id lives
on the Node and is the app's own PID; pipewire.sec.pid lives on the Client and
is pipewire-pulse's for every Pulse client. That resolves four rounds of
contradictory PID claims.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 03:49:09 -04:00
molluskandClaude Opus 4.8 100117085d docs: v3.1 audio-exclusion — apply Codex round-4 findings
Round 4 (review-2026-07-21-design-v3-round4.md) returned 5 findings, 3 of
them blocking. All claims re-verified against source before acceptance.

Biggest correction: eligibility is a GRAPH property, not a node property.
Exclusion does not propagate downstream — a filter-chain/loopback/combine-sink
re-emits the mix as a fresh untagged Stream/Output/Audio that passes both the
peerspeak.owned and pulse.module.id checks, re-injecting the whole call into
the share. Reachability confirmed: easyeffects IS installed on this machine
(it merely wasn't running during the fan-out spike, which is why the spike
missed it). §6 rewritten around transitive upstream reachability, tracking
Node/Port/Link globals, with a registry sync barrier and revalidation
immediately before each link creation.

Also applied:
- §5.3 is now a bounded validation state machine, not a one-shot check.
  wait_for_nodes only waits for the virtual source/sink, never the playback
  hazard leg, and pixelpass capture spawns lazily on first viewer, so the
  one-shot check raced in both directions. Revocation redefined as loss of
  the module identity, not transient absence of one leg.
- §7.2: reordering ActiveSession fields is NOT sufficient — kill_on_drop
  sends SIGKILL without waiting, so AEC can still unload while pixelpass
  lives. Fix is explicit shutdown().await at both channel-close breaks,
  field order as defence in depth, plus a fake-resource ordering test.
- §5.1 relabelled implementation sites; none of them tag anything today.
- Stop Share citation corrected to :699/:3480.
- D1-D7 resolved; readiness section added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 03:31:12 -04:00
molluskandClaude Opus 4.8 cab6bafce5 docs: v3 audio-exclusion design — rewrite around Option C + AEC gate
v1/v2 described a move-based design that Option C superseded on 2026-07-20,
and the AEC playback-leg identity gate has since passed. Roughly two thirds
of v2 documented problems Option C does not have, so this is a rewrite rather
than a patch (v1/v2 remain at 88ad5a0 / 10203e1).

Folds in: the four AEC gate results, the five corrections that constrain them
(observed correlation not a contract; exact-equality only; index/link-group
reuse and node-id recycling; group prefix = hazard detection not ownership;
application.process.id == pipewire-pulse for module-created streams), the
verified implicit-drop ordering defect in ActiveSession, fail-closed
validation/revocation, the IPC shape, and the split-out prerequisites.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 03:16:30 -04:00
molluskandClaude Opus 4.8 10203e1edb docs: adopt fan-out (Option C) after feasibility spike
Ran the direct-link spike on the live graph (PipeWire 1.6.8,
WirePlumber 0.5.15). Fan-out carries full-level audio for paplay, mpv
and VLC while the application keeps its existing speaker link;
WirePlumber does not reap foreign links across default-sink switch,
suspend/resume or 100s steady state; and non-lingering links are
destroyed automatically when their owning connection is SIGKILLed.

The decisive result is that destroying the capture sink mid-share left
the application playing to its speakers undisturbed, so capture-side
failure degrades to "not captured" rather than breaking the user's
audio. That is the property the move-based design had to work hard to
approximate.

Records what the spike does not prove: fidelity beyond signal presence,
daemon restart, quantum perturbation, and exclusive/passthrough streams.
The capture null sink is still pactl-owned, so Stop Share continues to
leak a module every time and the graceful-stop work is still owed.

Eligibility becomes a broad guarded selector rather than a narrow
allowlist, since copying no longer risks disturbing the source.

Option A and its attendant cleanup, restore and output-switch machinery
are retained for the record but are no longer the plan. A v3 rewrite is
owed once the AEC playback-leg identity is settled.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 05:04:25 -04:00
molluskandClaude Opus 4.8 88ad5a0807 docs: screenshare audio exclusion design v2
Rewrite after Codex's adversarial review of v1 found four release
blockers, all independently verified against source.

v1's premise was wrong: whole-desktop capture bypasses Routing::start
entirely (pipeline.rs:121), so this needs a new capture mode rather than
an inverted predicate.

v2 replaces PID-based identity with ownership by inherited tag, and makes
the router an allowlist so unrecognized infrastructure is left alone
rather than optimistically moved. Graceful stop becomes a prerequisite:
Stop Share is currently SIGKILL, so cleanup never runs on the normal path.

Records live measurements taken 2026-07-20: PULSE_PROP tagging reaches
the graph for paplay, mpv and VLC, and application.process.id is the
client's own PID, not pipewire-pulse's — correcting a claim both the
review and v1 relied on.

Adds Option C (fan out a second owned link instead of moving streams),
which deletes most of the cleanup, latency and multi-host problems the
move-based design has to solve. Not yet implemented; gated on a
feasibility spike.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 02:54:53 -04:00
molluskandClaude Opus 4.8 0588d92537 release: 0.6.6
CI / check (push) Failing after 5m35s
The live-edge catch-up (8c4f4a0, b4a4c00) landed after the v0.6.5 tag, so
the 0.6.5 artifacts do not contain it — the same gap that left the fix out
of v0.6.4. Cut 0.6.6 so the published build actually carries it.

Local-only changes (no wire change; PROTO planes unchanged), so this is a
PATCH bump per VERSIONING.md.

601 lib tests green, clippy clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:57:05 -04:00
molluskandClaude Opus 4.8 b4a4c00711 fix(screenshare): make live-edge catch-up actually recover
CI / check (push) Failing after 2m37s
The first cut used a fixed 1.05x drain, which measurement showed was too
gentle to matter: clearing a 6 s backlog would take two minutes, which a
viewer experiences as still broken.

Two changes, both measured on the netem satellite rig (loopback
impairment, gst -> ffmpeg HTTP relay -> mpv, matching the http:// URL
production actually serves):

1. Proportional drain. Speed now scales with buffer depth,
   1 + 0.05*(cache - 0.5), clamped to 1.15x, keeping the hysteresis band
   so it cannot oscillate. Deep backlogs recover in tens of seconds;
   small excursions still get an inaudible nudge.

2. Bound the byte cache in Low latency. The demuxer cache is a *byte*
   budget, so at a given bitrate it sets the worst-case backlog: 2 MiB
   held ~6 s of a 2.5 Mbps share. Capping Low latency at 1 MiB halved the
   standing buffer, 6.0 s -> 2.8 s, on its own. Smooth keeps the user's
   value, since a deep buffer is that posture's whole point.

Measured effect with both: playback consumes 11.6% faster than realtime
while behind (ratio 1.1157 vs 0.9988 with catch-up off), i.e. ~9 s of
backlog cleared in 80 s where before it recovered nothing at all and the
viewer stayed behind for the rest of the call.

Rig caveat: its upstream queues hold an unbounded backlog, so the cache
never drops back through the low mark and the return-to-1x transition is
only covered by unit tests, not the rig.

601 lib tests green, clippy clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:47:32 -04:00
molluskandClaude Opus 4.8 8c4f4a0b8b feat(screenshare): drain a lagging viewer back to the live edge
CI / check (push) Failing after 2m12s
On a lossy link the reliable PixelPass transport turns every loss burst
into buffered latency that nothing trims back, so the viewer settles
seconds behind the host and stays there. Measured on a tc netem satellite
simulation: a viewer parks at a ~6 s standing buffer indefinitely.

--untimed (0.6.5) does NOT fix this and measured marginally worse (+1.38 s
vs +1.24 s): it only unpaces presentation, while audio still drains at 1x
the DAC rate, so an accumulated backlog never shrinks. Drop it.

Instead give mpv a JSON IPC socket in the Low latency posture and drive
playback slightly fast while the buffer is deep, returning to 1x once it
drains. Pitch correction keeps it inaudible and A/V sync is preserved,
because audio and video speed up together.

The control law and IPC message handling are pure functions with unit
tests; the only I/O is livesync::drive, which ends by itself when the
player exits. Smooth is deliberately excluded — its ~2 s readahead is the
point of that posture, and catch-up would fight it every poll.

Known limitation: 1.05x needs ~120 s to clear a 6 s backlog, so recovery
is slower than ideal. Tuning (a proportional law, or a seek-to-live for
large backlogs) is the follow-up.

598 lib tests green (+11), clippy clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:36:06 -04:00
molluskandClaude Opus 4.8 76c4f68bb3 release: 0.6.5
CI / check (push) Successful in 2m54s
Local-only changes since 0.6.4 (no wire change; PROTO planes unchanged),
so this is a PATCH bump per VERSIONING.md.

Ships the low-latency screen-share live-edge fix (4bfc184), which landed
three hours after the v0.6.4 tag and was therefore never released.

Also adds the missing CHANGELOG entry for the participant "Advanced audio"
foldout (26d6600), which shipped without one.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:45:16 -04:00
mollusk c427231858 feat(notifications): add chat and contact sounds
CI / check (push) Successful in 2m39s
2026-07-19 02:02:03 -04:00
mollusk 4bfc18463b fix(screenshare): keep low-latency playback live 2026-07-18 22:22:24 -04:00
mollusk 26d66007de ui: fold participant audio controls 2026-07-18 20:14:09 -04:00
molluskandClaude Fable 5 3d7b01c8a2 release: 0.6.4
CI / check (push) Failing after 3m17s
Wire-compatible refinement release (GOSSIP_PROTO stays 5). Highlights:
honest chat send status + sender-side pacing (chat-hardening Phase 5),
completing the chat-hardening plan; playlist drawer de-clutter + auto-resize;
plus the FEC-gap and network-restart fixes already on the branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 19:16:39 -04:00
molluskandClaude Fable 5 5f4eba1815 chat: honest local send status + sender-side pacing (Phase 5)
CI / check (push) Successful in 4m17s
Closes the final phase of docs/chat-hardening-plan.md. Two problems: a
locally echoed message always looked sent even when the core had no active
session or the gossip broadcast failed; and a fast burst could broadcast
'successfully' yet be silently dropped by every receiver's per-author rate
bucket (8 burst, then 1/s) with no sender feedback.

Send status: CoreCommand::SendChat/SendChatFile carry a local-only id (never
on the wire); the core replies with UiEvent::ChatSendResult after the gossip
broadcast succeeds or fails, and a no-active-session is now an explicit
failure rather than a silent no-op. gossip send_chat, which previously
returned Ok on a missing sender/topic or an encode failure, now returns Err.
ChatEntry gains local_send: Option<LocalSend>; failed sends render a red
'Not sent — {reason}  [Retry]' line, Broadcast/Pending render nothing
(there are no delivery receipts, so silence is the honest success state).

Sender-side pacing (new src/app/sendqueue.rs): sends past the burst queue
locally as 'queued…' and trickle out at the receivers' sustained rate, so
nothing is lost and typing is never blocked (user chose queue-and-trickle
over input throttling). The pacer reuses the gossip gate's own TokenBucket +
per-author constants (now pub(crate)) so the two sides of the policy can't
drift. A 250ms drain subscription runs only while the queue is non-empty.
Retry re-dispatches the retained payload; re-serving the same attachment id
replaces the ServeStore entry rather than double-counting bytes. The pacer
and monotonic send-id counter survive a room reset (receivers' buckets
persist; ids never alias a late result); queue and retry payloads are cleared.

582 lib tests (+11: 4 pacer/queue seam, 7 app-level transition/retry/reset);
all-targets green, clippy -D warnings clean, fmt clean, smoke launch OK. No
wire change (GOSSIP_PROTO stays 5). Tests-green-only — the two owed
two-machine field-test items are logged in the plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 19:03:25 -04:00
molluskandClaude Fable 5 77d2bf2992 style(music): de-duplicate drawer transport controls, let playlist fill drawer height
CI / check (push) Successful in 2m43s
The playlist drawer duplicated the player bar's |prev/play/next| transport
row even though the drawer can only be open while the bar is visible
(drawer_open gates on show_player_bar), so the drawer copy is removed;
seek, music volume, Browse, and the tune-in checkbox remain drawer-only.

The track list (and the Public tab's broadcast list) was a 160px-fixed
scrollable nested inside a second full-height scrollable, showing only a
few entries. The outer scrollable is gone and both lists now fill the
drawer's remaining height, resizing with the window.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 18:25:34 -04:00
molluskandClaude Fable 5 c8d0053431 docs(plan): add Phase 4 items to the two-machine field-test checklist
CI / check (push) Successful in 3m31s
Also re-triggers CI: run 165 on 1d038be died to rust-lld crashes from disk
exhaustion on the runner host (12G free vs ~12G cold-build transient), not a
code failure; 18G of local build artifacts have been swept (30G free now).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 05:04:25 -04:00
molluskandClaude Fable 5 1d038be9a0 chat: parsed-URL link policy, cached link ranges, history byte budget (Phase 4)
CI / check (push) Failing after 3m10s
Phase 4 of docs/chat-hardening-plan.md — URL and rendering resilience.
Closes the chat-body half of S14 (bidi override strip).

- sanitize: new is_safe_web_url shared link policy (url crate, promoted to a
  direct dependency): http/https scheme + non-empty host + no userinfo;
  candidates failing it stay plain text (their whole whitespace run, interior
  not re-scanned). Scheme detection is now ASCII-case-insensitive.
- sanitize: linkify() -> link_ranges()/segments(): validated byte ranges
  computed once, exact-roundtrip slicing, at most CHAT_MSG_MAX_LINKS (8)
  clickable links per message; the rest stays selectable plain text.
- sanitize_chat: strips bidi overrides/isolates (U+202A-202E, U+2066-2069)
  from message bodies while keeping ZWJ/ZWNJ/LRM/RLM (S14 chat-body half).
- app: ChatEntry caches its link ranges (filled in push_chat), so redraws
  slice instead of rescanning/re-validating; only link spans allocate.
- app: chat history now also bounded by 512 KiB total sanitized text
  (CHAT_HISTORY_MAX_TEXT_BYTES) alongside the 300-entry cap; the attachment
  byte cache is deliberately untouched by history eviction (own budgets).
- app: AppMessage::OpenUrl re-checks the same parsed policy (defence in
  depth) instead of prefix checks - non-web schemes can never reach the
  opener even if the handler is invoked directly.

571 lib tests green (+3 net); clippy -D warnings + fmt clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 04:57:15 -04:00
molluskandClaude Fable 5 554b613466 chat: attachment cache, download, and transfer hardening (Phase 3)
CI / check (push) Successful in 2m33s
Phase 3 of docs/chat-hardening-plan.md — attachments can no longer turn
into unbounded memory, bandwidth, decoder, or task pressure (S15 closed;
S14's filename half closed).

Cache and image cost (3A): AttachmentCache now carries encoded- and
decoded-byte budgets (96 MiB / 64 MiB) on top of the count cap, with
per-entry weights, replacement accounting, and oldest-first eviction; an
individually over-budget fetch services any pending Save/Play from the
bytes in hand and is exposed as Evicted instead of retained.
validate_image_bytes prechecks header dimensions (per-side AND a new
14 MP total-pixel limit) before any decode; the renderer only ever
receives a ≤1600 px downscaled RGBA preview whose w*h*4 cost counts
against the decoded budget — originals stay encoded-only for Save.
sanitize_filename strips the bidi/zero-width spoofing set (RTL-override
extension spoof).

Download policy and state (3B): images auto-fetch only when roster-
authored AND declared ≤4 MiB, gated by a new deterministic
AutoFetchBudget (per-author and session request+byte token buckets,
check-then-take, bounded author map) alongside the existing dedup and
four-permit bound. Attachment state is now explicit — absence/Loading/
Ready/Failed/Evicted — driven by a new AttachmentFetchStarted event, so
skipped or evicted images render a "Load image" button instead of an
indefinite "loading…", and repeated clicks can never spawn duplicate
fetch tasks.

Exact transfers and serve store (3C): fetch_blob requires the received
length to equal the declared size (short = local error, overlong =
bounded-read reject, empty keeps meaning "sender no longer has it");
the file picker's unbounded read is replaced by a metadata-prechecked
cap+1 bounded reader; one Arc<Vec<u8>> now backs the UI cache, command
queue, and serve store; served_files is a count- and byte-budgeted FIFO
ServeStore (16 entries / 128 MiB).

37 new tests (568 lib total) including a real two-endpoint loopback
exercising exact/short/overlong/unknown-id transfers. Plan checkboxes
ticked and constant deviations decision-logged. Tests-green-only: the
plan's two-machine field-test section remains open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 02:09:30 -04:00
molluskandClaude Fable 5 8898652349 chat: roster-bound authorship, replay dedup, and rate limits (Phase 2)
CI / check (push) Successful in 2m59s
Chat-hardening plan Phase 2 — only current authenticated room members can
create chat UI work, impersonation via the wire name is structurally closed,
and no member can monopolize the event channel:

- core: new ChatRoster (bounded id -> sanitized-name map, shared) replaces the
  event task's bare HashSet; upserted on PeerJoined/PeerUpdated, removed on
  graceful PeerLeft AND terminal grace-expiry eviction (both timer paths).
  Non-roster chat is dropped before attachment handling; the rendered author
  label is the roster-bound name — the sender-claimed wire name is never read.
- gossip: ChatIngressGate after verify_gossip, before any sanitize work or
  event send: early known-author gate (live + mid-reconnect peers), exact-
  replay suppression keyed on the deterministic Ed25519 signature (1024-entry
  cap + freshness-window TTL, zero new deps vs the plan's BLAKE3 option), then
  per-author (8 burst, 1/s) and room-wide (32 burst, 8/s) token buckets.
  Replays are detected before tokens are consumed; a room-bucket reject
  refunds the author token; rejection logging is squelched per author.
- The inner Chat.ts is now ignored entirely; RoomEvent carries the signed
  envelope timestamp.

550 lib tests (+18), reconnect_eviction +1 (grace keeps chat authority,
terminal eviction revokes it), clippy --all-targets -D warnings clean.
Tests-green-only: the plan's two-machine field-test section remains open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 20:37:13 -04:00
molluskandClaude Fable 5 5927148ee4 chat: enforce shared text policy at UI, sign point, and gossip ingress
CI / check (push) Successful in 3m4s
Chat-hardening plan Phase 1. The chat body policy (2,000-char + 8 KiB
ceilings, single-pass control/whitespace normalization) moves from the UI
layer into src/sanitize.rs and is now enforced at every trust boundary:
cap_chat_input bounds the live input (oversized paste), the gossip sign
point re-sanitizes so non-UI callers can't bypass policy, and gossip
ingress rejects oversized raw text before sanitizing (admit_chat_text)
and drops messages with neither visible text nor an attachment. The
incoming chat author label now uses the strict name sanitizer until
Phase 2 roster-binds it. +8 tests (532 lib green).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 18:36:21 -04:00
molluskandClaude Fable 5 93f4954653 docs: changelog for the jitter FEC and net-rebuild resilience fixes
CI / check (push) Successful in 2m26s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:42:08 -04:00
molluskandClaude Fable 5 e6eb490939 audio: only FEC-recover a gap from its immediate successor packet
CI / check (push) Failing after 10m54s
The jitter buffer's gap path fed the LOWEST buffered packet to
decode_fec regardless of position. Opus in-band FEC in packet N carries
a copy of frame N-1 and nothing else, so that reconstruction is only
correct when the smallest survivor is exactly next+1 (single loss).
On burst loss it spliced a later frame's audio into the wrong slot —
worse than concealment. Gate FEC on adjacency (new fec_covers_gap(),
wraparound-aware); everything else falls back to plain PLC.

Two new tests: the gate itself, and a burst-loss test proven to bite —
it compares bit-exact against a twin decoder and fails against the old
unconditional-FEC behavior (checked by mutation).

Fixes finding 3 of the 2026-07-16 full-codebase review.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:31:47 -04:00
molluskandClaude Fable 5 e724167b03 docs: add chat-hardening plan as scope contract
GPT-5.6's 5-phase plan for the chat identity/replay/rate-limit cluster
(2026-07-16 review findings 5-8): roster-bound display names, replay
dedup, quiet rate limiting, bidi-aware sanitization, attachment size
checks. Self-describes as temporary — delete when the work completes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:27:47 -04:00
molluskandClaude Fable 5 7b9cb57003 core: survive a failed net-stack rebuild instead of silently dying
CI / check (push) Failing after 13m37s
The four live-rebuild sites (deferred rebuild on Join/Leave, idle
SetNetworkMode, idle RegenerateIdentity) all did
`net.shutdown().await` then `build_net_stack(...).await?` — a build
failure propagated out of run_core_loop, which its supervisor only
logs. Every subsequent command went nowhere: window alive, app dead,
user told nothing. (The initial startup build already reported.)

New replace_net_stack() helper: tear down the old stack, build for the
requested posture, and on failure fall back to the posture the old
stack was actually running (tracked in the new `net_mode` local; when
the postures are equal the fallback is a plain retry — e.g. identity
regeneration, where reverting the already-persisted key would be
wrong). If the fallback lands, the UI is told the change didn't stick
and `network_mode` reverts so state stays honest and the change stays
re-attemptable. If both builds fail the UI gets a fatal 'Networking
lost … restart' error before the loop exits — informed, not a zombie.

Retry policy isolated in rebuild_with_fallback(), generic over the
builder: 4 new unit tests cover first-try success, fall-back, plain
retry, and double failure without binding sockets.

Fixes finding 2 of the 2026-07-16 full-codebase review.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 23:58:58 -04:00
molluskandClaude Fable 5 af7a42a049 ci: remove zombie workflows from the per-push pipeline
cargo-deny.yml (runs-on: ubuntu-latest) and windows-build.yml (runs-on:
windows-latest) target runner labels no registered runner advertises, so
every push queued two runs Gitea auto-cancelled ~24h later — the Actions
page has shown 2 cancelled runs per push since the runner went live.

- cargo-deny.yml: deleted; redundant with ci.yml's deny step, which now
  runs `cargo deny --locked check` to preserve the locked-tree stance.
- windows-build.yml: kept but workflow_dispatch-only until a Windows
  runner exists; restore instructions in the header comment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 23:46:56 -04:00
molluskandClaude Fable 5 e78e7bc2a5 supply-chain: ignore quick-xml build-time DoS advisories + ttf-parser unmaintained
RUSTSEC-2026-0194/0195 (quick-xml 0.39.4, published 2026-06-29) broke the
deny/audit CI gates on every push since June 29. quick-xml is reached only
via the wayland-scanner proc-macro parsing vendored protocol XML at compile
time — attacker input never touches it and it is absent from the shipped
binary. The fixed 0.41.0 is semver-incompatible with wayland-scanner's
`^0.39` req (no upstream bump yet); documented ignores until one exists.

RUSTSEC-2026-0192 (ttf-parser unmaintained, via iced/cosmic-text) joins the
existing unmaintained ignores (paste, audiopus_sys) — same class, same
lockfile-pinning protection.

New .cargo/audit.toml keeps cargo-audit in sync with deny.toml.

Known leftover warning (allowed, non-failing): spin 0.10.0 is yanked but
futures-buffered (via iroh) requires ^0.10 and no unyanked 0.10.x exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 23:46:13 -04:00
molluskandClaude Fable 5 52ab374b74 style: cargo fmt under rustfmt 1.9.0 (toolchain update 2026-07-08)
Six diffs across four files: the 2026-07-08 stable toolchain update
(rustc 1.96.1 / rustfmt 1.9.0) re-flags code that was fmt-clean when
committed under the previous rustfmt. No semantic change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 23:45:18 -04:00
mollusk 8825707c17 chore: patch crossbeam-epoch RustSec advisory
CI / check (push) Failing after 7s
cargo-deny / cargo-deny (push) Has been cancelled
windows-build / windows-build (push) Has been cancelled
2026-07-15 06:31:09 -04:00
molluskandClaude Fable 5 76c62e5ac3 docs: mark connection badge field-verified (2-machine call 2026-07-08)
CI / check (push) Failing after 5s
cargo-deny / cargo-deny (push) Has been cancelled
windows-build / windows-build (push) Has been cancelled
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 17:07:54 -04:00
molluskandClaude Fable 5 d2432740c1 network: per-peer connection badge (direct/relay, RTT, loss, bitrate)
Answer "am I actually P2P right now?" per peer. A 1 Hz session task
snapshots the selected QUIC path of every live audio connection
(IrohTransport::connection_stats), core::connstats::derive turns
consecutive snapshots into RTT/loss/bitrate (path switches and counter
resets invalidate the rate window), and the peer card shows a
Direct/Relay badge with a hover tooltip for address, loss, and up/down
bitrate. No new dependencies, no wire change.

Loopback-integration-tested against real iroh endpoints; not yet
field-verified on a 2-machine call (FEATURES.md row marked 🧪).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 15:15:22 -04:00
molluskandClaude Opus 4.8 99a4a336ad Release 0.6.3 — in-app screen-sharing controls + hwdec fixes
CI / check (push) Failing after 4s
cargo-deny / cargo-deny (push) Has been cancelled
windows-build / windows-build (push) Has been cancelled
Bump to 0.6.3 and document the screen-share work merged on this branch:
the advanced Settings section + per-call quality picker (96e41de), the
hardware-decode-defaults-off frame-1 freeze fix (96e41de), the per-call
quality override fix (e378b2e), and the VLC-honors-viewer-settings fix
(faad8ce). All local-only — no wire-protocol change, old configs load
unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 17:20:00 -04:00
molluskandClaude Opus 4.8 df45c0bfeb screenshare: log the pixelpass-host and player argv on spawn
The screen-share code only logged pixelpass's high-level JSON events, never
the argv it spawned children with, so a field log couldn't confirm which
encode/viewer settings actually reached the helpers — e.g. the per-call
quality's --bitrate (host) or the hardware-decode --avcodec-hw/--hwdec flag
(player). Log both verbatim at spawn: host args carry no secret, and the
player line omits the local stream URL. Logged per attempt so a player
fallback is visible too.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 16:47:06 -04:00
molluskandClaude Opus 4.8 faad8ce26a screenshare: honor viewer settings for the VLC player too
The viewer playback settings (hardware decode + buffering) only shaped
mpv's argv; vlc_args() was fixed, so a VLC viewer silently ignored them.
The load-bearing case is hardware decode: mpv defaults to software decode
(the A-bug fix), but VLC hardware-decodes by default, so a VLC viewer with
the default hardware_decode=false still got GPU decode and could hit the
frame-1 freeze the default exists to avoid — the toggle did nothing.

vlc_args() now takes the settings and maps the knobs that translate
cleanly to VLC: hardware decode (--avcodec-hw=none/any) and buffering
posture (network/live caching ms). The genuinely mpv-specific knobs
(cache_mb byte-cache, extra_mpv_args) stay mpv-only; the Settings UI
hints are reworded to say which knobs are mpv-only vs universal. +2 tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 13:22:31 -04:00
molluskandClaude Opus 4.8 e378b2e33b screenshare: fix inline per-call quality override being discarded
The Share control's inline quality dropdown sets a session-only
`share_quality_selection`, but ToggleScreenShare (which opens the audio
picker on the only real path to a share) unconditionally reset it back to
the saved config default before ConfirmShareScreen read it. The picker has
no quality control of its own, so the user's per-call pick was silently
dropped 100% of the time and every share used the persisted default.

Drop the reset; add a regression test asserting the override survives
picker-open and reaches the confirm.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 13:09:22 -04:00
molluskandClaude Opus 4.8 96e41de1b1 screenshare: advanced in-app streaming controls + hwdec toggle
Add a local-only "Screen sharing" section to Settings plus a per-call quality
picker on the Share control: in-app control over how a share is encoded
(quality/bitrate/framerate/max-height/max-viewers/software-x264, + extra
pixelpass args) and how it's played back (mpv/vlc, hardware decode, buffering,
cache, + extra mpv args). Settings live in AppConfig.screen_share (all
serde-defaulted, so old configs load unchanged) and become pixelpass host CLI
flags / mpv args at share/view launch.

Hardware decode defaults OFF, which also fixes the frozen-frame-with-audio bug:
forcing --hwdec=auto stalled some viewers' HW decoder on frame 1 while audio
kept playing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:15:11 -04:00
molluskandClaude Opus 4.8 5c888f8357 context_input: middle-click pastes the X11 PRIMARY selection
Joe (X11) reported that middle-click paste did nothing in the ticket and
node-ID fields. iced's base text_input only binds Ctrl+V to the Standard
(CLIPBOARD) selection and never reads PRIMARY or binds mouse button 2, so
the "select text, middle-click to paste" workflow was dead.

Add a Button::Middle branch to ContextInput::update that reads
clipboard::Kind::Primary, sanitizes it, and pastes at the cursor (reusing
the already-tested pure paste()). Factor the control-char stripping into a
shared, unit-tested sanitize_clip() helper also used by the menu Paste, so
a trailing newline on the PRIMARY selection is dropped. Respects `locked`
so read-only display fields still reject paste.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 01:25:01 -04:00
molluskandClaude Opus 4.8 074f004227 core: one screen-share player per share — re-watch replaces, not stacks
Every click of Watch (`CoreCommand::ViewShare`) spawned a fresh pixelpass
viewer + mpv and pushed it onto an untracked Vec. A field test hit the
consequence: the first click gave a frozen player (the host's capture was
stalling), so the viewer clicked again to retry — and got a SECOND mpv,
doubling the shared audio.

Track viewers paired with their share ticket. On ViewShare, reap players
whose window already closed (try_wait), then if a live player for the same
ticket exists, kill it before spawning the replacement. Re-watching a
share now swaps its player instead of stacking a second one. Pure
`replace_viewer_index` seam + test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 20:48:52 -04:00
molluskandClaude Opus 4.8 2d22036930 screenshare: drop mpv --untimed so shared video stays A/V-synced
The viewer launched mpv with `--untimed`, which displays each video
frame the instant it decodes and ignores audio timestamps. Sharing a
desktop (no audio) that just minimizes latency, but sharing a *video*
made its audio drift progressively out of sync — confirmed in a field
test watching a video together. Remove the flag so mpv paces video to
the audio clock; the remaining low-latency flags keep lag negligible for
desktop pointing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 20:29:59 -04:00
48 changed files with 12326 additions and 650 deletions
+11
View File
@@ -0,0 +1,11 @@
# cargo-audit configuration. Keep the ignore list in sync with deny.toml,
# which carries the full justification for each entry.
[advisories]
ignore = [
# quick-xml DoS advisories: build-time only, reached solely via the
# wayland-scanner proc-macro parsing trusted vendored protocol XML.
# Fix (0.41.0) is semver-incompatible with wayland-scanner's `^0.39`;
# drop once wayland-scanner bumps. See deny.toml.
"RUSTSEC-2026-0194",
"RUSTSEC-2026-0195",
]
-34
View File
@@ -1,34 +0,0 @@
name: cargo-deny
# Enforce the supply-chain policy in deny.toml (advisories / bans / licenses /
# sources) on every push to main and every PR. Runs on a *locked* tree so the
# pinned, vetted versions in Cargo.lock are exactly what get audited — see the
# deny.toml header and VERSIONING.md. A new poisoned release of a dependency
# cannot reach CI until Cargo.lock is deliberately updated.
on:
push:
branches: [main]
pull_request:
jobs:
cargo-deny:
runs-on: ubuntu-latest
# rust:1 provides the cargo toolchain that cargo-deny shells out to for
# `cargo metadata`. Adjust the runner label if your act_runner uses a
# different one.
container: rust:1
steps:
- uses: actions/checkout@v4
- name: Install cargo-deny (pinned prebuilt)
run: |
set -euo pipefail
version=0.19.9
curl -sSfL \
"https://github.com/EmbarkStudios/cargo-deny/releases/download/${version}/cargo-deny-${version}-x86_64-unknown-linux-musl.tar.gz" \
| tar -xz -C /usr/local/bin --strip-components=1 --wildcards '*/cargo-deny'
cargo-deny --version
- name: cargo deny check
run: cargo deny --locked check
+3 -1
View File
@@ -36,7 +36,9 @@ jobs:
run: cargo test --doc
- name: cargo-deny (advisories, bans, licenses, sources)
run: cargo deny check
# --locked so the pinned, vetted versions in Cargo.lock are exactly
# what get audited (the lockfile-as-review-checkpoint model).
run: cargo deny --locked check
- name: cargo-audit
run: cargo audit
+15 -11
View File
@@ -7,11 +7,20 @@ name: windows-build
# alias) so a Unix-only assumption can't sneak back in and break Windows.
#
# RUNNER REQUIREMENT: this needs a Windows act_runner registered with the
# `windows-latest` label (the Linux `cargo-deny` job's container approach does
# NOT apply here — Windows jobs run on the host, not a Linux container). If your
# runner advertises a different label, change `runs-on` below. Until a Windows
# runner exists this workflow is simply skipped/queued, not a failure of the
# Linux CI.
# `windows-latest` label (a Linux-container approach does NOT apply here —
# Windows jobs run on the host, not a Linux container). If your runner
# advertises a different label, change `runs-on` below.
#
# MANUAL-ONLY until that runner exists: with push/PR triggers enabled, every
# push queued a run no runner could claim and Gitea auto-cancelled it ~24h
# later, littering the Actions page with cancelled runs. Restore the push/PR
# triggers when a Windows runner is registered:
#
# on:
# push:
# branches: [main, "windows-port-**"]
# pull_request:
# workflow_dispatch:
#
# BUILD-HOST REQUIREMENTS (validated by the opus spike, see
# peerspeak-windows-opus-spike.md):
@@ -23,12 +32,7 @@ name: windows-build
# must provide both.
on:
push:
# `main` plus the in-progress port branches, so the Windows path is exercised
# before merge rather than only after.
branches: [main, "windows-port-**"]
pull_request:
# Allow manual runs from the Gitea Actions UI.
# Manual runs from the Gitea Actions UI only — see the header comment.
workflow_dispatch:
permissions:
+102
View File
@@ -2,6 +2,108 @@
All notable changes to PeerSpeak are documented here.
## [Unreleased]
## [0.6.6] — 2026-07-19
### Fixed
- **A screen share that falls behind now catches back up.** On a lossy
connection (satellite links are the worst case) the share could settle several
seconds behind the host and simply stay there for the rest of the call. The
viewer now notices a deep buffer and plays imperceptibly fast until it is back
at the live edge — the audio stays in tune and in sync while it does. This
replaces the previous attempt at the problem, which measurement showed did not
help. Applies to the Low latency setting; Smooth intentionally keeps its
larger buffer.
### Changed
- **Low latency now keeps a tighter viewer buffer.** The screen-share cache
setting is a size in megabytes, which at a given bitrate quietly decides how
many *seconds* behind a viewer can drift — a 2 MB buffer turned out to hold
about six seconds of a typical share. Low latency now caps that buffer at 1 MB
regardless of the setting, which halved how far behind a share fell on a bad
connection before anything else kicked in. Smooth still honors the value you
choose, since a deep buffer is the point of that mode.
## [0.6.5] — 2026-07-19
### Added
- **Chat message sounds.** Successful outgoing messages and admitted incoming
messages now have distinct notification chimes, each with its own enable
toggle and optional custom WAV path in Notifications settings.
- **Contact presence sounds.** The home-screen contacts list now announces a
contact becoming online or offline. Initial online contacts are announced;
initial offline results stay silent. Both events have independent toggles and
optional custom WAV paths.
- **Notification sound browser.** Every notification event now has a native
Browse button for choosing a custom WAV instead of typing its path manually.
### Changed
- **Tidier per-participant audio controls.** The equalizer bands and noise gate
for each participant now live behind an **"Advanced audio"** foldout instead
of being expanded all the time, so a call with several people no longer fills
the panel with sliders. The controls themselves are unchanged.
### Fixed
- **Low-latency screen sharing stays near the live edge again.** mpv's
timestamp pacing could let stale frames accumulate across the reliable
PixelPass transport until a share was 710 seconds behind. Low-latency mode
now presents decoded frames immediately; Smooth mode retains timestamp pacing
when keeping shared-video audio and video synchronized matters more.
## [0.6.4] — 2026-07-18
### Added
- **Chat now tells you when a message didn't send.** A message that couldn't go
out — because you weren't in a room, or the broadcast failed — is marked
**"⚠ Not sent"** with a **Retry** button, instead of sitting in the transcript
looking delivered. A successful send shows nothing (PeerSpeak has no
delivery/read receipts, so anything else would be a false promise).
- **Fast typing no longer loses messages.** When you fire off a quick burst,
messages past the first few are held as **"queued…"** and sent a moment apart,
matching the rate other people's clients accept. Previously a fast burst could
look sent on your end while some messages silently never reached the room.
### Changed
- **Tidier music playlist drawer.** The slide-out playlist no longer repeats the
play/skip controls already on the player bar, and the track list now grows to
fill the drawer instead of being boxed into a short scroll area, so you can see
more of your playlist at once.
- **Safer chat under the hood.** A round of chat hardening tightened how incoming
messages, display names, links, and file/image attachments are validated and
bounded, so a malformed or hostile message from a peer can't spoof a name,
replay, flood, or run the app out of memory. No change to how normal chat looks
or works.
### Fixed
- **Burst packet loss no longer splices the wrong audio into the gap.** Loss
concealment used Opus in-band FEC even when the next packet to arrive wasn't
the one immediately after the gap, so losing several packets in a row could
briefly play a later frame's audio in the wrong position. FEC now only
reconstructs a gap from its immediate successor packet; larger gaps are
concealed normally.
- **A failed network restart no longer silently kills the app.** Changing the
network mode (or regenerating your identity) rebuilds the connection stack;
if that rebuild failed — rare, but possible when the local socket can't
bind — PeerSpeak kept its window open but silently stopped responding to
every command. It now falls back to your previous network settings and says
so, and only gives up (with a clear error telling you to restart) if even
the fallback fails.
[0.6.4]: https://gitbutter.xyz/mollusk/peerspeak/releases/tag/v0.6.4
## [0.6.3] — 2026-07-06
### Added
- **In-app screen-sharing controls.** A new **Screen sharing** section in Settings, plus a per-call **quality picker** on the Share control, put the whole share pipeline under your control without editing config files. Encode side: quality preset, bitrate, framerate, maximum resolution, maximum viewers, a force-software-encode switch, and an escape hatch for extra pixelpass arguments. Playback side: choose **mpv or VLC**, toggle **hardware decoding**, pick a buffering posture (low-latency vs. smooth), set the demuxer cache, and pass extra mpv arguments. Everything is stored locally in your config and defaults are unchanged, so existing setups keep working as-is.
### Fixed
- **Shared video no longer freezes on the first frame while audio keeps playing.** Hardware decoding now defaults **off**; forcing `--hwdec=auto` stalled some viewers' hardware decoder on frame 1. You can re-enable hardware decoding from the new Screen sharing settings if your machine handles it well.
- **The per-call quality picker is now honored.** The inline quality dropdown next to the Share button was being reset to the saved default before a share started, so every share silently used the default quality regardless of what you picked.
- **VLC now respects your playback settings.** VLC hardware-decodes by default, so a VLC viewer previously ignored the hardware-decode toggle (and could hit the same frame-1 freeze) and the buffering posture. VLC viewers now map both settings onto VLC's own options.
[0.6.3]: https://gitbutter.xyz/mollusk/peerspeak/releases/tag/v0.6.3
## [0.6.2] — 2026-07-03
### Fixed
Generated
+5 -3
View File
@@ -1207,9 +1207,9 @@ dependencies = [
[[package]]
name = "crossbeam-epoch"
version = "0.9.18"
version = "0.9.20"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5b82ac4a3c2ca9c3460964f020e1402edd5753411d7737aa39c3714ad1b5420e"
checksum = "2d6914041f254d6e9176c01941b21115dcfb7089e55135a35411081bd106ef3f"
dependencies = [
"crossbeam-utils",
]
@@ -4871,7 +4871,7 @@ checksum = "35fb2e5f958ec131621fdd531e9fc186ed768cbe395337403ae56c17a74c68ec"
[[package]]
name = "peerspeak"
version = "0.6.2"
version = "0.6.6"
dependencies = [
"anyhow",
"async-trait",
@@ -4883,6 +4883,7 @@ dependencies = [
"image",
"iroh",
"iroh-gossip",
"libc",
"opus",
"pipewire",
"rand 0.10.1",
@@ -4894,6 +4895,7 @@ dependencies = [
"thiserror 2.0.18",
"tokio",
"tokio-stream",
"url",
"windows-sys 0.61.2",
]
+12 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "peerspeak"
version = "0.6.2"
version = "0.6.6"
edition = "2024"
description = "Decentralized peer-to-peer voice chat (Rust/iroh/PipeWire/Opus/iced)"
license = "MIT"
@@ -75,6 +75,10 @@ serde_json = "1.0.150"
thiserror = "2.0.18"
tokio = { version = "1.52.3", features = ["full"] }
tokio-stream = "0.1.18"
# Chat link policy: parse + validate clickable URL candidates (scheme/host/
# userinfo checks in `sanitize::is_safe_web_url`). Already in the tree
# transitively via iroh — this only promotes it to a direct dependency.
url = "2.5"
# --- Platform-specific dependencies -----------------------------------------
# Audio and the native file-picker backends differ per OS. Everything else in the
@@ -105,3 +109,10 @@ windows-sys = { version = "0.61", features = [
"Win32_System_Diagnostics_ToolHelp",
"Win32_System_Threading",
] }
# Unix-only. Used for exactly one thing: sending SIGINT to our own
# pixelpass child so it can run its cleanup before we resort to SIGKILL
# (src/core/teardown.rs). Already in the tree via alsa/cpal/tokio, so
# declaring it directly adds no new code to the build.
[target.'cfg(unix)'.dependencies]
libc = "0.2.186"
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+8
View File
@@ -76,6 +76,14 @@ CHIMES = {
"mic-toggle.wav": [(E5, 0.08)],
# Reconnect gave up: disappointing low two-note fall.
"reconnect-failed.wav": [(C5, 0.15), (349.23, 0.30)],
# Our chat message entered the room: a tiny bright acknowledgement.
"chat-sent.wav": [(1046.50, 0.06)],
# A peer message arrived: a soft two-note lift, distinct but unobtrusive.
"chat-received.wav": [(E5, 0.07), (G5, 0.11)],
# A saved contact came online: a light, higher two-note arrival.
"contact-online.wav": [(E5, 0.09), (880.00, 0.18)],
# A saved contact went offline: the same tonal family falling away.
"contact-offline.wav": [(E5, 0.09), (440.00, 0.18)],
}
+13
View File
@@ -24,6 +24,19 @@ ignore = [
# audiopus_sys: unmaintained FFI bindings to the stable libopus C library,
# pulled in via our direct `opus 0.3.1` dep. No drop-in replacement.
"RUSTSEC-2026-0150",
# ttf-parser: unmaintained, transitive via iced/cosmic-text (font parsing
# for the GUI). Inputs are system + embedded fonts, not network data. No
# upstream migration yet; revisit when iced moves off it.
"RUSTSEC-2026-0192",
# quick-xml 0.39.4 DoS advisories (quadratic dup-attr check; unbounded
# namespace allocation). Build-time only: quick-xml is reached solely via
# the wayland-scanner PROC-MACRO, which parses the wayland protocol XML
# files vendored inside the wayland-* crates at compile time. Attacker
# input never reaches it and it is not in the shipped binary. The fix
# (0.41.0) is semver-incompatible with wayland-scanner 0.31.x's `^0.39`
# requirement; drop both ignores once wayland-scanner releases a bump.
"RUSTSEC-2026-0194",
"RUSTSEC-2026-0195",
]
# ---------------------------------------------------------------------------
+1
View File
@@ -102,6 +102,7 @@ covers internals). When you ship a feature, add it here.
| iroh QUIC transport | ✅ | |
| Network mode picker | ✅ | `RelayNoDiscovery` (default), `N0Full`, `DirectOnly`. Takes effect next join. |
| Retained-address reconnect | ✅ | Dials last-known full addr before falling back to bare id. |
| Per-peer connection badge (direct/relay + RTT, hover for addr/loss/bitrate) | ✅ | Peer-card badge fed by a 1 Hz poll of the live audio link's selected QUIC path (`connection_stats``core::connstats::derive`). Field-verified on a real 2-machine call 2026-07-08. |
| Reconnect + eviction model | ✅ | Incl. two-outage reconnect-eviction fix + regression test. |
| Self-hosted relay | ❌ | Decided against — rely on n0 relays, `RelayNoDiscovery` default. |
+584
View File
@@ -0,0 +1,584 @@
# Chat hardening — ephemeral implementation plan
**Status (2026-07-18):** Phases 15 COMPLETE (all plan phases done). Phase 1 =
shared text policy in `src/sanitize.rs`, ceilings enforced at UI input, sign
point, and gossip ingress. Phase 2 = roster-bound authorship
(`src/core/chatroster.rs`), replay dedup + rate limits (`ChatIngressGate` in
`src/network/gossip.rs`). Phase 3 = attachment cache/serve-store budgets,
downscaled previews, auto-fetch byte/request budgets (`src/core/fetchbudget.rs`),
exact transfers, bounded local reads. Phase 4 = parsed-URL link policy
(`is_safe_web_url`/`link_ranges` in `src/sanitize.rs`, `url` crate), 8-link cap,
cached link ranges in `ChatEntry`, 512 KiB history text budget, chat-body
bidi-override strip (closes S14). Phase 5 = honest local send status
(`CoreCommand::SendChat`/`SendChatFile` carry a local id, `UiEvent::ChatSendResult`,
`SendStatus` on own echoes) PLUS sender-side pacing (`src/app/sendqueue.rs`
mirrors the receivers' per-author budget so fast bursts trickle instead of being
silently dropped downstream). All gates green each phase. This is a temporary
scope contract for hardening the existing room chat; with every phase complete
and the two-machine field test done, delete this file (see the completion note
at the end). The two-machine field-test section below is still owed before that
deletion. Do not add link previews as part of this effort.
## Goal
Strengthen the current encrypted, signed, session-only room chat without changing
its product model: plain selectable text, clickable web links, and peer-to-peer
attachments over the existing gossip and files planes. The work should make chat
resistant to identity spoofing, replay, spam, oversized input, expensive rendering,
and attachment-driven memory/bandwidth pressure while preserving normal Unicode
conversation and the existing full-mesh architecture.
## Existing foundation to preserve
- Gossip payloads are signed by the claimed `EndpointId`, bound to the raw room
topic and protocol domain, and checked before dispatch.
- The signed envelope timestamp is admitted only within the two-minute gossip
freshness window.
- Inbound gossip frames are capped at 128 KiB before JSON deserialization. This
larger plane-wide cap must remain because `Announce` may contain a custom avatar.
- Chat history is session-only and capped at 300 entries.
- Only `http://` and `https://` links are opened, as a single process argument
without a shell.
- Attachment descriptors are signed with the chat payload; attachment bytes use
the encrypted files plane, have a 25 MiB per-file cap, and are keyed by both
author and attachment id.
- Image bytes are decoded defensively and automatic image fetches already have a
four-task concurrency limit.
## Working design decisions
These are the implementation defaults unless code inspection or tests reveal a
concrete reason to adjust them. Record any adjustment in the decision log.
1. **No wire change.** Keep `GossipMessage::Chat` unchanged and do not bump
`GOSSIP_PROTO`. The redundant wire `name` and inner `Chat.ts` remain serialized
for compatibility but are not trusted. Remove them only during a future planned
gossip-version bump.
2. **Roster identity is authoritative.** A chat line is admitted only for an
authenticated identity already known to the current room (including the
reconnect grace state). Its displayed name comes from the sanitized roster
state, never from `GossipMessage::Chat.name`.
3. **Body Unicode remains expressive.** Do not apply the short-label sanitizer to
the message body; it strips format characters used by some languages and emoji.
Continue neutralizing controls and whitespace, while treating author labels,
filenames, and URLs more strictly because those are spoof-sensitive surfaces.
4. **Bounds apply at every trust boundary.** UI input is bounded while editing,
outgoing text is normalized before signing, and incoming text is byte-checked
and normalized before it leaves the gossip layer. UI-only truncation is not an
adequate ingress defense.
5. **Automatic network work is stricter than manual work.** Keep the 25 MiB manual
attachment ceiling, but auto-fetch only small images. Larger images remain
available behind an explicit Load/Download action.
6. **Caches are bounded by cost, not only entry count.** Count encoded bytes and
estimated decoded image bytes. A count cap remains as a secondary bound.
7. **Rate limiting degrades quietly.** Drop excess/replayed peer messages with a
rate-limited log entry. Do not let a spammer produce a second UI-notification
flood.
## Proposed policy constants
Keep these together near the code that enforces them and cover them with boundary
tests. Values are starting points, not a compatibility contract.
| Policy | Initial value | Reason |
| --- | ---: | --- |
| Chat body characters | 2,000 | Preserves current UI behavior |
| Chat body UTF-8 bytes | 8 KiB | Covers 2,000 four-byte scalars with small headroom |
| Live input characters/bytes | Same as body | Prevent oversized paste/edit state |
| Clickable links per message | 8 | Bounds spans and opener targets |
| Retained chat text | 512 KiB plus 300 entries | Bounds redraw and selection work |
| Per-author chat limiter | Burst 8, refill 1/second | Allows normal bursts, stops sustained spam |
| Room-wide chat limiter | Burst 32, refill 8/second | Protects shared event/UI queues |
| Exact-chat replay cache | 1,024 digests, 2-minute TTL | Covers freshness window with a hard bound |
| Auto-fetch image encoded size | 4 MiB | Limits unsolicited bandwidth and allocations |
| Attachment cache encoded budget | 128 MiB | Allows several ordinary files without GiB growth |
| Attachment cache decoded-preview budget | 64 MiB | Bounds renderer-side image pressure |
| Served attachment budget | 256 MiB plus a count cap | Bounds sender memory for a long session |
| Inline preview longest side | 1,600 px | Chat renders near 260 px; full 4K decode is wasteful |
| Decoded source image pixels | 16 megapixels maximum | Adds a total-pixel bound to per-side bounds |
## Phase 1 — Shared text policy and live-input bounds
**Target:** downstream layers never receive or retain an unexpectedly large or
unsafe chat string.
- [x] Move chat constants and `sanitize_chat` from `src/app/mod.rs` into
`src/sanitize.rs` (or a narrowly scoped shared chat-policy module if that keeps
the API clearer).
- [x] Implement a single-pass sanitizer that:
- maps control characters to spaces;
- collapses whitespace and trims ends;
- enforces both the character and UTF-8 byte ceilings without splitting a scalar;
- returns empty for content with no visible text.
- [x] Add `cap_chat_input` for live editing. It must preserve the user's current
whitespace while enforcing character and byte ceilings; normalization remains a
submit/ingress operation so typing does not visibly jump.
- [x] Apply `cap_chat_input` in `AppMessage::ChatInputChanged`, covering keyboard,
clipboard, primary-selection, and context-menu paste paths through the controlled
input widget.
- [x] Sanitize outgoing text immediately before local echo and `CoreCommand` send.
- [x] Sanitize again before `GossipMessage::Chat` is signed, so a future non-UI
caller cannot bypass policy.
- [x] At gossip ingress, reject raw chat text over the byte ceiling before doing
downstream sanitization; sanitize accepted text before creating `RoomEvent`.
- [x] Keep attachment-only messages when the sanitized caption is empty; drop a
chat with neither visible text nor a valid attachment.
- [x] Stop sanitizing an incoming chat `name` with the body sanitizer. Phase 2
replaces it with the roster-bound name.
### Phase 1 tests
- [x] ASCII, multibyte Unicode, emoji, whitespace, NUL/CR/LF/TAB/ESC, empty input.
- [x] Exact character and byte boundaries, including a four-byte scalar at the
cutoff.
- [x] Oversized paste never makes `state.chat_input` exceed either ceiling.
- [x] Outgoing, incoming, and direct core/network paths converge on the same
normalized result.
- [x] Empty captions are retained only when a valid attachment remains.
## Phase 2 — Admission, identity binding, replay, and spam control
**Target:** only current authenticated room members can create chat UI work, and a
member cannot impersonate another participant or monopolize the control/UI queues.
- [x] Change the core event task's chat roster from a bare `HashSet<EndpointId>` to
a bounded map containing each member's latest sanitized display name (or retain a
parallel name map if less invasive).
- [x] Insert/update the map on `PeerJoined`/`PeerUpdated`, retain it during transient
reconnect grace, and remove it on graceful or terminal eviction.
- [x] Before attachment handling or UI forwarding, reject `RoomEvent::ChatMessage`
whose author is not present in that authoritative roster.
- [x] Replace the embedded wire name with the roster map's name before constructing
`UiEvent::ChatMessage`. The UI may keep storing a name snapshot so old chat lines
remain labeled after a peer leaves.
- [x] Add a lightweight early known-author gate in the gossip loop using its live
and disconnected-peer sets. Keep the core roster gate as defense in depth and as
the final authority.
- [x] Validate that the inner `Chat.ts` equals the signed envelope timestamp, or
ignore it entirely. Do not use the inner timestamp for replay or ordering.
- [x] Add exact-chat replay suppression after signature verification and before
event-channel send:
- hash the canonical signed bytes, not raw JSON formatting;
- use BLAKE3 (make it a direct dependency if needed; it is already in the iroh
dependency graph) or an equally collision-resistant existing primitive;
- store a `HashSet` plus FIFO/TTL order for bounded lookup and eviction;
- prune by both the gossip freshness window and the hard entry cap.
- [x] Add a bounded token bucket per admitted author and a room-wide bucket before
awaiting `event_tx.send`. Limiter state must be removed with roster eviction and
remain bounded by the roster cap.
- [x] Ensure duplicate messages are dropped before consuming rate-limit tokens, so
a replay cannot starve a legitimate new message from that author.
- [x] Rate-limit rejection logging per author/reason.
- [ ] Consider applying the same local submit policy to accidental rapid Enter or
button activation, without routing chat through the coalescing command path.
### Phase 2 tests
- [x] Valid roster author is admitted; never-announced, post-leave, forged, and
stale authors are rejected.
- [x] A peer sending `name = "Victim"` renders under its own roster name.
- [x] A name update affects future messages without rewriting history.
- [x] Reconnect grace continues accepting the known author; terminal eviction does
not.
- [x] The same signed chat is displayed once; distinct chats created in the same
millisecond are both admitted.
- [x] Replay-cache TTL/cap pruning cannot grow without bound.
- [x] Per-author burst/refill and room-wide burst/refill boundaries.
- [x] Excess chat cannot prevent a subsequent `Leave` or `Announce` from reaching
the event loop in a deterministic channel-pressure test.
## Phase 3 — Attachment transfer and memory hardening
**Target:** neither peers nor long local sessions can turn chat attachments into
unbounded memory, bandwidth, decoder, or task pressure.
### 3A. Cache and image cost
- [x] Extend `AttachmentCache` with encoded-byte and decoded-preview-byte counters.
Preserve the count cap, but evict oldest entries until all three budgets fit.
- [x] Give every entry an explicit weight. Replacement must subtract the old
weight before checking/inserting the new one.
- [x] Decide behavior for a single entry larger than the cache budget: service an
immediate pending Save/Play request without retaining it, then expose it as
evicted/unavailable rather than exceeding the budget.
- [x] Add a total-pixel limit to `validate_image_bytes` in addition to the existing
width/height limit.
- [x] Build a downscaled inline preview handle with a maximum 1,600 px side. Keep
original bytes only for Save; do not hand a full-resolution 4K image to the
renderer merely to display it at chat width.
- [x] Count estimated RGBA preview cost (`width * height * 4`) against the decoded
budget even if iced internally copies or uploads it.
- [x] Strip the same bidi/zero-width spoofing characters used for display labels
from attachment filenames, while preserving ordinary Unicode filenames.
### 3B. Automatic download policy and state
- [x] Auto-fetch only roster-authored images whose declared size is at or below
`MAX_AUTO_IMAGE_BYTES`; keep the existing `(author,id)` dedup and four-permit
concurrency bound.
- [x] Add per-author and session byte/request budgets for automatic fetches so a
peer cannot drain bandwidth sequentially after each permit is released.
- [x] Represent `NotFetched`, `Loading`, `Ready`, `Failed`, and `Evicted` distinctly
enough for the UI to avoid an indefinite “loading…” label when auto-fetch was
skipped or the cache evicted an item.
- [x] Render a Load image button for large/skipped images. A manual click may use
the 25 MiB file cap but still observes cache/decoder budgets.
- [x] Ensure a repeated click cannot create duplicate unguarded fetch tasks.
- [x] Keep non-image attachments manual-only.
### 3C. Exact transfers, local reads, and served files
- [x] In `IrohTransport::fetch_blob`, require `bytes.len() as u64 == declared_size`.
Reject empty, short, and overlong transfers with a concise local error.
- [x] Replace the file picker's unbounded `FileHandle::read()` with a helper that
reads at most `MAX_ATTACHMENT_BYTES + 1`. Check metadata first where available,
but retain the bounded read because metadata can race or be unavailable through
a portal.
- [x] Avoid duplicating a full attachment across UI, command queue, and serve store.
Prefer `Arc<Vec<u8>>`/`Arc<[u8]>` through `AttachmentState`, `CoreCommand`, and
`serve_attachment`, subject to iced handle API constraints.
- [x] Replace the unbounded session `served_files` map with a count- and byte-
budgeted FIFO store. Evicted ids should produce the existing “sender no longer
has the file” response rather than stale or aliased data.
- [x] Keep attachment ids keyed by author on receipt and preserve all existing
request-length, timeout, filename, and decoder checks.
### Phase 3 tests
- [x] Byte-budget eviction, count eviction, replacement accounting, clear/reset,
and an individually overweight entry.
- [x] Decoded-preview budget and downscale dimensions for wide, tall, square, and
boundary images.
- [x] Image with valid per-side dimensions but excessive total pixels is rejected.
- [x] A declared 4 MiB image auto-fetches; the first byte over the limit requires a
click.
- [x] Per-author/session auto-fetch budgets recover according to their policy and
never exceed task concurrency.
- [x] Short, exact, and overlong file responses.
- [x] Local file reader stops at cap + 1 instead of allocating the full source.
- [x] Served-file FIFO/byte eviction and replacement accounting.
- [x] Same attachment id from two authors remains isolated throughout fetch, cache,
save, and display.
## Phase 4 — URL and rendering resilience
**Target:** keep clickable links without making malformed/deceptive input or many
small spans an unnecessary UI/launcher surface.
- [x] Make `url` a direct dependency (already present transitively) and validate
link candidates with `url::Url`.
- [x] A clickable URL must have an `http` or `https` scheme and a valid host.
- [x] Treat URLs containing username/password syntax as plain text, or require an
explicit confirmation that shows the parsed destination host. Prefer plain text
for the first implementation.
- [x] Preserve the existing defense-in-depth validation in `AppMessage::OpenUrl`;
replace prefix checks with the shared parsed-URL policy.
- [x] Cap clickable candidates at eight per message. Remaining content stays
selectable plain text and must still round-trip exactly.
- [x] Refactor linkification to return borrowed ranges/offsets or cache link ranges
in `ChatEntry`, avoiding allocation and rescanning on every redraw.
- [x] Bound retained history by total sanitized text bytes as well as 300 entries.
Eviction must keep attachment bookkeeping coherent and should not invalidate an
open Save/Play operation.
- [x] Do not add metadata fetching, remote images, Markdown, or link previews.
- [x] (Folded in from S14, per the security handoff) Strip bidi
overrides/isolates from the chat BODY in `sanitize_chat`, keeping the other
expressive format characters (ZWJ/ZWNJ/LRM/RLM).
### Phase 4 tests
- [x] Valid HTTP/HTTPS, malformed host, empty host, mixed case, Unicode path/query,
punctuation, credentials/userinfo, and non-web schemes.
- [x] Eight-link boundary and many-link adversarial input.
- [x] Segment/range reconstruction exactly reproduces the sanitized message.
- [x] Entry-count and total-text-budget history eviction.
- [x] Opener policy cannot launch a non-web scheme even if called directly.
## Phase 5 — Honest local send status
**Target:** never present a locally echoed message as successfully broadcast when
the core rejected it or gossip broadcast failed.
- [x] Add a local-only message id and `Pending`/`Broadcast`/`Failed` state to local
chat entries. Do not put this id or state on the wire. (`ChatEntry.local_send:
Option<LocalSend>`; `SendStatus` also has `Queued` for the paced-but-not-yet-sent
state — see the pacing decision-log entry.)
- [x] Carry the local id through `CoreCommand::SendChat`/`SendChatFile` and return a
`UiEvent` result after the local gossip broadcast call succeeds or fails.
(`SendChat`/`SendChatFile` gained `local_id`; new `UiEvent::ChatSendResult { local_id,
error }`.)
- [x] If the core is not in an active session, return failure instead of silently
doing nothing. (`send_chat` now `Err`s on missing sender/topic and on encode
failure; the core arm maps no-session to a `ChatSendResult` error.)
- [x] Show failure compactly with a retry action. A successful local broadcast must
not be labeled “delivered” or “read”; PeerSpeak has no peer acknowledgements.
(Failed → red "⚠ Not sent — {reason} [Retry]" line; Broadcast/Pending render
nothing — silence is the honest success state.)
- [x] Retry creates one new signed broadcast while retaining replay correctness and
attachment serving state. (`RetryChatSend(id)` re-dispatches the retained
`PendingSend`; re-serving the same attachment id REPLACES the `ServeStore`
entry, never double-counts — see `serve_store_replacement_accounting_and_remove_clear`.)
### Phase 5 tests
- [x] Local echo starts pending, becomes broadcast on success, and becomes failed
on no-session/channel/gossip error. (`send_status_pending_then_broadcast_on_success`,
`send_status_failed_keeps_payload_for_retry`.)
- [x] Results update only the matching local entry, including after history
eviction or room reset. (`send_result_updates_only_the_matching_entry`,
`send_result_after_eviction_drops_orphan_payload`, `send_result_after_room_reset_is_a_noop`.)
- [x] Retry does not duplicate served bytes or mutate an unrelated entry.
(`retry_redispatches_only_the_targeted_send`; served-byte dedup =
`serve_store_replacement_accounting_and_remove_clear` in `files.rs`.)
## Compatibility and versioning
- The planned implementation changes validation, local data structures, and
internal `CoreCommand`/`UiEvent` shapes only. Keep the serialized
`GossipMessage::Chat` and file request/response formats unchanged.
- Therefore do **not** bump `GOSSIP_PROTO`, `FILES_PROTO`, or the pre-1.0 MINOR
solely for this plan. The eventual release is a compatible PATCH unless scope
expands into a wire change.
- If implementation requires removing/adding serialized fields, changing
attachment request framing, or introducing acknowledgements on the wire, stop
and revise this section before coding that part. Follow `VERSIONING.md` and use
the appropriate protocol plus release MINOR bump.
## Verification gates
Run after each phase, with focused tests first and the full gates before handoff:
```text
cargo fmt --check
cargo test --lib
cargo test --all-targets
cargo clippy --all-targets -- -D warnings
```
Also retain the existing ignored/loopback coverage where the environment supports
it; do not make ordinary unit tests depend on external network access.
### Two-machine field test
- [ ] Ordinary ASCII/Unicode conversation, rapid short burst, long boundary text,
and oversized paste.
- [ ] Rename during a room: new lines use the new roster name; old lines retain
their snapshot.
- [ ] Disconnect/reconnect grace and post-leave chat admission behavior.
- [ ] Multiple normal images, one image above the auto threshold, a malformed
“image”, and a maximum-size manual file.
- [ ] Download/save after cache eviction; clear failure state and no runaway
memory across repeated attachments.
- [ ] Observe process RSS and UI responsiveness during a bounded spam/attachment
stress run; verify leave/reconnect controls remain responsive.
- [ ] Linux and Windows URL opening for valid links; malformed/userinfo links remain
selectable but do not launch.
- [ ] A message with more than eight URLs renders eight clickable links and the
rest as selectable plain text, with nothing dropped.
- [ ] A message attempting bidi-override display spoofing renders in send order
(the override characters are stripped, emoji/joining-script text intact).
- [ ] Send a fast burst (>8 messages in a second): all arrive at the peer in
order, none silently lost; the sender sees "queued…" on the overflow that
then clears as each goes out.
- [ ] Send with no active session (or a failing broadcast): the message shows
"⚠ Not sent" with a Retry, and Retry resends it once when connectivity is back.
## Completion criteria
The plan is complete when:
1. Only active/grace-rostered authenticated authors reach chat UI state.
2. Chat identity is roster-bound and cannot be overridden by the embedded wire
name.
3. Exact replay and sustained spam are bounded before shared event queues.
4. Live input, inbound/outbound body size, history text, attachment caches,
automatic transfers, served files, and decoded previews all have tested hard
bounds.
5. File transfer length and image decoding/display costs are validated.
6. Clickable links pass a shared parsed-URL policy and rendering work is bounded.
7. Local broadcast failure is visible without claiming peer delivery.
8. Unit/all-target/clippy gates and the two-machine field test pass.
9. Relevant durable docs (`README.md`, `docs/FEATURES.md`, `CHANGELOG.md`, security
notes, and comments) describe the final behavior.
10. This ephemeral plan is deleted after its useful status/history is transferred
to durable documentation.
## Out of scope
- Link previews, metadata fetches, or remote thumbnail requests.
- Persistent/offline chat history or server-side message storage.
- Markdown, rich embeds, reactions, editing, deletion, threads, or search.
- Read receipts or peer delivery acknowledgements.
- Moderation UI, kicking, blocking, or trust-list redesign.
- Antivirus/malware scanning of user-requested downloaded files.
- A new application-layer group-encryption protocol or a broader cryptographic
redesign. If PeerSpeak makes a formal end-to-end-encryption product claim, audit
and document the exact iroh/gossip/relay threat model as a separate project.
## Decision log
- **2026-07-15:** Chose hardening over automatic link previews because receiving a
message should not trigger third-party web requests or weaken PeerSpeak's
privacy-oriented design.
- **2026-07-15:** Initial scope keeps all wire formats stable; hardening is local
admission, validation, resource accounting, and honest UI state.
- **2026-07-17 (Phase 1):** The 8 KiB byte ceiling deliberately cannot bind on
*sanitized* output (2,000 scalars × 4 bytes = 8,000 ≤ 8,192), so inside
`sanitize_chat`/`cap_chat_input` it is defense in depth; its operative role is
the raw-ingress reject in `admit_chat_text`.
- **2026-07-17 (Phase 1):** Interim until Phase 2's roster binding: the incoming
chat `name` now goes through the strict `sanitize_name` label sanitizer at the
UI edge (was the body sanitizer), so author labels already get bidi/zero-width
stripping and the 48-char label cap.
- **2026-07-17 (Phase 1):** `send_chat` at the gossip sign point silently no-ops
(Ok) on an empty-after-sanitize body with no attachment rather than erroring;
the UI already prevents this case, and Phase 5's send-status work is where
send-path feedback gets designed.
- **2026-07-17 (Phase 2):** Replay dedup is keyed on the payload's own Ed25519
**signature bytes** instead of a BLAKE3 digest (the plan allowed "an equally
collision-resistant existing primitive"): ed25519 signing is deterministic
(RFC 8032), so the 64-byte signature is already a collision-resistant
fingerprint of the exact signed bytes — same dedup power, zero new direct
dependencies. Cache entries are stamped with the signed envelope `ts` and
pruned once it exits the freshness window, because `verify_gossip` already
rejects such a frame before the cache is consulted.
- **2026-07-17 (Phase 2):** A room-bucket reject refunds the just-consumed
author token, so a room-wide squeeze caused by other members does not also
drain an innocent author's personal budget.
- **2026-07-17 (Phase 2):** Rate-limited frames are NOT entered into the replay
cache: only fully admitted chats are. A legitimate message the room was too
busy for, redelivered later by the swarm, is then displayed once instead of
being misread as a replay of something never shown.
- **2026-07-17 (Phase 2):** The "wire name never renders" guarantee is
structural: the core event task binds the wire field as `name: _` and builds
`UiEvent::ChatMessage` exclusively from `ChatRoster::name_of`, so there is no
code path from wire name to UI. The roster map behavior is unit-tested; the
end-to-end impersonation scenario stays on the (still-open) two-machine
field-test list.
- **2026-07-17 (Phase 2):** The channel-pressure requirement is met at the seam
level: chat admission is bounded (32-burst / 8-per-s room-wide) BEFORE any
`event_tx.send`, and `Announce`/`Leave` admission is independent of the chat
gate — verified by unit tests. A full gossip-loop pressure harness was not
built; the seam bound is what protects the channel.
- **2026-07-17 (Phase 2):** An empty-after-sanitize roster name falls back to
the short node id, so a member who announces an all-control-character name
still gets a stable, non-blank chat label.
- **2026-07-17 (Phase 2):** The "Consider applying the same local submit policy
to accidental rapid Enter" item is DEFERRED: the receiving side is the
security boundary (every peer independently enforces the buckets), and a
local silent drop would be a UX regression better designed alongside Phase
5's honest send status.
- **2026-07-18 (Phase 3):** Constants that deviate from the proposed table, all
bound-tested: total decoded pixels **14 MP** (not 16 MP) so the bound clears
12 MP phone photos (4032×3024) yet actually binds inside the 4096²≈16.8 MP
per-side envelope; cache encoded budget **96 MiB** (not 128) — still several
full-size files, tighter worst case; serve store **128 MiB + 16 entries**
(not 256 MiB) — a sender's own session should not pin a quarter GiB.
- **2026-07-18 (Phase 3):** `validate_image_bytes`/`decode_preview` precheck
dimensions from the container HEADER (`into_dimensions`) before any pixel
decode, so an over-limit decode bomb is rejected without paying its decode
cost; the decode-time `image::Limits` remain as defense in depth, and the
decoded dimensions must equal the prechecked header dimensions.
- **2026-07-18 (Phase 3):** Budget-pressure evictions leave NO cache entry
(absence = NotFetched → the same Load/Download affordance), while the
explicit `Evicted` state marks only an *individually over-budget* fetch whose
bytes were used once (pending Save/Play serviced from hand) and dropped. Both
render load-on-demand; only the bookkeeping differs.
- **2026-07-18 (Phase 3):** The core still runs `validate_image_bytes` before
emitting `AttachmentReady`, and the UI decodes once more to build the ≤1600px
preview. Two bounded decodes per image were accepted over shipping decoded
RGBA across the channel (which would defeat the encoded-only Arc sharing).
- **2026-07-18 (Phase 3):** The image lightbox now enlarges the ≤1600px preview
handle, not the original bitmap — originals are retained encoded-only for
Save. At the lightbox's window-sized draw area the visual difference is nil
for the chat use case; full fidelity remains one Save away.
- **2026-07-18 (Phase 3):** `AutoFetchBudget` checks all four buckets
(author/session × requests/bytes) and only then consumes atomically, so a
rejection burns nothing (no refund path like Phase 2's room bucket needed).
Tokens ARE consumed if the four-permit semaphore then rejects the spawn —
that only happens mid-flood, when charging the author is the intent.
- **2026-07-18 (Phase 3):** The auto-fetch budget's author map prunes
least-recently-active past 64 entries instead of wiring roster eviction into
the event task: authors are roster-gated upstream (≤32 live members), so
strangers cannot churn the map, and a pruned author returning with full
buckets is within policy.
- **2026-07-18 (Phase 3):** Music-track serving shares the bounded serve store
with chat attachments. A user who sends enough large attachments during a
broadcast can evict their own current track; listeners then get the standard
"sender no longer has the file" failure. Accepted: budget honesty over a
second store, and the store comfortably fits current+next track plus a
normal chat working set.
- **2026-07-18 (Phase 3):** The clip player's command channel still takes one
owned byte copy at the moment of a Play click (small, human-initiated). The
Arc de-duplication targeted the send path (UI cache / command queue / serve
store), which now shares a single allocation.
- **2026-07-18 (Phase 3):** Overlong transfers are rejected by the transport
read itself (`read_to_end(size)` errors past the bound) rather than an
explicit length compare; short transfers get the explicit
`len == declared_size` check. Music fetches ride `fetch_blob`, so they
inherit exactness for free.
- **2026-07-18 (Phase 4):** The S14 chat-body half (bidi strip) landed here per
the security handoff: `sanitize_chat` strips ONLY bidi overrides/isolates
(U+202A202E, U+20662069) — the characters that can visually reorder a
rendered line — while ZWJ/ZWNJ (emoji sequences, joining scripts) and the
LRM/RLM direction *marks* (which cannot reorder) are kept. Labels/filenames
keep the stricter full-format-strip.
- **2026-07-18 (Phase 4):** A link's href is the exact displayed slice of the
message — validation is parse-only, no normalization on open — so what the
user sees IS the argv the opener receives. Consequence: WHATWG slash
collapsing means `http:///path` parses to host `path` (as in browsers) and is
accepted; the empty-host rejects are `http://` and friends that fail parsing.
- **2026-07-18 (Phase 4):** URLs with userinfo syntax went the plan-preferred
plain-text route (no confirmation dialog). A candidate that fails the policy
leaves its WHOLE whitespace-delimited run as plain text without re-scanning
the interior — `http://a@http://b.com` yields zero links, by design.
- **2026-07-18 (Phase 4):** Scheme detection became ASCII-case-insensitive
(`Http://…` from sentence auto-capitalization now linkifies); the policy
check is unaffected since `url` normalizes scheme/host case during parsing.
- **2026-07-18 (Phase 4):** Cached ranges in `ChatEntry.links`, filled inside
`push_chat` (the single history choke point), were chosen over
borrowed-return-per-redraw: redraws now slice cached char-boundary ranges,
and only link spans allocate (their href String).
- **2026-07-18 (Phase 4):** History byte-budget eviction (512 KiB, alongside
the 300-entry cap) deliberately does NOT touch the attachment byte cache:
that cache is bounded by its own Phase 3 budgets, and leaving it alone means
an open Save/Play on an evicted line keeps its bytes-in-hand (the save
dialog falls back to the generic "download" name). The just-pushed entry is
never evicted; a single message's 8 KiB ceiling cannot exceed the budget.
- **2026-07-18 (Phase 5):** Sender-side PACING was added to Phase 5's scope
(originally receiver-status only). The Phase 2 decision log deferred the
"apply the same local submit policy to accidental rapid Enter" item to pair
with Phase 5, and honest status alone would still let a fast burst broadcast
successfully yet be silently dropped by every receiver's per-author bucket
(8 burst, then 1/s) with no sender feedback. The user chose "queue and
trickle" over "throttle input": sends past the burst queue locally as
`SendStatus::Queued` ("queued…") and release at the receivers' sustained
rate, so nothing is lost and typing is never blocked.
- **2026-07-18 (Phase 5):** The pacer (`src/app/sendqueue.rs`) reuses the
gossip gate's OWN `TokenBucket` + `CHAT_AUTHOR_BURST`/`CHAT_AUTHOR_REFILL_PER_MS`
(made `pub(crate)`), so the two sides of the rate policy are one definition
and cannot drift. It mirrors only the PER-AUTHOR budget, not the room-wide
one — we cannot know other members' send rates, and the per-author bucket is
the one guaranteed to apply to us at every receiver.
- **2026-07-18 (Phase 5):** Send status renders as a line UNDER the message
(user pick over an inline suffix glyph); `Broadcast` and the transient
`Pending` show nothing because PeerSpeak has no delivery/read receipts, so an
unadorned message IS the honest "handed to the swarm" state. Only `Queued`
and `Failed` (with Retry) are surfaced.
- **2026-07-18 (Phase 5):** The pacer and the monotonic send-id counter
deliberately SURVIVE a room reset while the queue and retry payloads are
cleared: receivers' per-author buckets persist across our rejoin (so the
pacer should not refill to full), and never-reused ids keep a late
`ChatSendResult` from a pre-reset send from aliasing a new entry — verified by
`send_result_after_room_reset_is_a_noop`.
- **2026-07-18 (Phase 5):** The pacer clock is `Instant`-based
(`AppState.send_clock`), not wall-clock, so a system time jump can neither
rewind nor fast-forward the send budget.
## Completion
All five phases are implemented and every gate is green. Per the scope-contract
note at the top, this file should be DELETED once the owed two-machine field
test (the checklist below) has been run — that deletion is a separate,
user-gated step, not part of the Phase 5 commit. Until then the plan stays as
the record of what shipped and what remains to verify on real hardware.
@@ -0,0 +1,970 @@
# Implementation plan: whole-desktop screen-share audio without self-echo
**Status:** 🟢 **v4 — three review rounds applied. Approved to start Phase 0a.**
**Date:** 2026-07-21
**Design of record:** [`screenshare-audio-exclusion-plan.md`](screenshare-audio-exclusion-plan.md) v3.4 (`8768cd2`), converged round 7.
**Scope:** *ordering, gates and acceptance criteria only.*
**Reference convention.** `v3.4 §N` = the design doc. `plan §N` = this document. The two
numbering schemes collide (both have a §11 and a §12) and an unqualified reference in v2 sent
Phase 6's most important gate to a section that does not exist. Every cross-reference below is
qualified.
**Review history**
| Round | Findings | Outcome |
| --- | --- | --- |
| 1 | 7 P1 + 5 P2 | 12 accepted, 1 half-rejected → v2. `~/Documents/handoff-docs/Codex/peerspeak/review-2026-07-21-impl-plan-round1.md` |
| 2 | 3 P1 + 7 P2 + 1 P3 | all accepted → v3. `…/review-2026-07-21-impl-plan-round2.md` |
| 3 | verification pass: 4 of 7 edits landed, 3 partial; 2 P1 + 2 P2 + 1 P3 | all accepted → v4. Approved to start Phase 0a. `…/review-2026-07-21-impl-plan-round3.md` |
Adjudication in plan §10. Two of my own claims were refuted by Codex with source evidence and
two of its claims were refuted or narrowed by mine; both are recorded there rather than
quietly dropped.
---
## 0. What this plan is optimising for
The design is converged; the risk has moved from "is it right?" to "will it be built in an
order where each mistake is caught while it is still cheap." Three properties drive every
ordering decision:
1. **Nothing that can create an echo runs before the thing that decides eligibility has been
validated against a real graph.** The taint engine is the hull. It is built, unit-tested,
and floated empty (Phase 5, dry-run) before a single link is created.
2. **Every path to unsafe audio is closed structurally before the machinery that could take
it is written.** There are **two** such paths, not one — the capture *source* and the
capture sink's *inputs*. Both are closed in Phase 0d.
3. **Every phase ends in a state that is shippable or trivially revertible**, and every gate
is one a broken implementation can *fail*. A gate that cannot fail is not a gate; where a
gate asserts only that bad things are absent, it must also assert that good things are
present, or "captures nothing at all" passes it.
The corollary, stated plainly because it is the most likely way this goes wrong: **the
temptation will be to write the link manager early**, because it is the visible feature.
Fan-out is roughly 400 lines and demos beautifully with a hand-picked node. It is also the
component that, shipped ahead of a validated engine, produces exactly the bug this feature
exists to prevent — in front of Joe.
## 0.1 Cross-repo reality
Two repos, no Cargo dependency; the contract is pixelpass's CLI plus its `--output json`
event stream (`peerspeak/src/screenshare/mod.rs:1-14`). The bulk of the work — graph engine,
link manager, AEC state machine — is **pixelpass**. peerspeak's share is tagging, argv,
capability gating, teardown ordering and the user-visible status surface.
**Hard ship-order constraint.** peerspeak spawns whatever `pixelpass` resolves on `PATH`. If
peerspeak passes `--aec=…` to a pixelpass that predates the flag, clap rejects it and **the
share hard-fails** — the documented A23 / audit-P2 skew failure that already governs
`--strict-audio` (`screenshare/mod.rs:249-262`).
> **pixelpass ships the capability first (Phase 7), including the old-peerspeak/new-pixelpass
> golden test. peerspeak only passes the new flags to a binary that advertised support
> (Phase 8), which owns the new-peerspeak/old-pixelpass golden test. Never "flag present or
> absent" as the protocol — v3.4 D5.**
⚠️ **Capability is resolved twice, against two independently-resolved binaries.**
`ListAudioApps` resolves pixelpass and probes it at `core/mod.rs:3375-3378`; `StartScreenShare`
resolves it *again* at `:3409-3419`. Between those moments `PATH` or the override can change.
Verified in source. Phase 8 must bind the capability result to the resolved path and re-probe
if it differs, failing closed.
---
## 1. Phase map and dependency DAG
| # | Phase | Repo | Mutates graph? | Exit gate |
| --- | --- | --- | --- | --- |
| 0a | `object.serial` u64 fix | pixelpass | no | boundary parse tests |
| 0b | Explicit teardown + drop ordering | peerspeak | no | **four independent mutations** (plan §2; revised from five, §10 r14) ✅ built |
| 0c | Graceful stop + connection-owned capture sink | both | sink ownership | two-host SIGKILL live gate **+ SIGINT-first gate** |
| 0d | **Typed capture plan + internal mode input** | pixelpass | no | mode matrix; neither unsafe source nor unsafe sink input constructible |
| 1 | peerspeak ownership tagging | peerspeak | no | tag on live nodes; literal pinned in plan §3 |
| 2 | Graph model + taint engine (pure) | pixelpass | **no PipeWire at all** | v3.4 §12 fixture matrix + degenerate-snapshot case |
| 3 | Registry observer + readiness epoch | pixelpass | read-only | six-part gate incl. **PID derivation** |
| **3r** | **Observer revision — bind every Node/Device (v3.5 §6.7)** | pixelpass | read-only | **four-part gate (plan §4 "Phase 3 revision")** |
| 4 | AEC identity validation state machine | pixelpass | read-only | fake-clock transition matrix |
| 5 | **Dry-run audit mode** | pixelpass | read-only | 🚦 **MAJOR GATE** — exact decision partitions (plan §5) |
| 6 | Link manager + status events, driven through the real host path | pixelpass | **yes — first mutation** | link-manager matrix (plan §6.1) + live dynamic matrix |
| 7 | **Public mode selector** + capability advertisement | pixelpass | no | old-peerspeak/new-pixelpass golden test |
| 8 | peerspeak integration (mode flag, argv, picker, UI, status) | peerspeak | no | new/new argv golden (mode **and** `--aec`); new-peerspeak/old-pixelpass golden; causal status delivery |
| 9 | Rig upgrade + field matrix | both | yes | 🚦 **SHIP GATE** — plan §7 |
**Landing DAG** (development may be concurrent; *landing* order may not):
```
0a ──────────────────► 2 ──► 3 ──► 4 ──► 5 ──► 3r ──► 5 (re-run) ──► 6 ──► 7 ──► 8 ──► 9
0b ──────────────────────────────────────────────────────────────────┤
0c ──► 0d ───────────────────────────────────────────────────────────┘
1 (r8 carriers) ──────────────────────────────────────► 5 (re-run)
```
⚠️ **Status 2026-07-25 (evening): 3r is BUILT AND MERGED; the re-run has not happened yet.**
The phase-5 gate failed on its first live run and put 3r into the DAG; 3r's own four-part
gate now passes, including the live prop-recovery row on this host. Phase 5's machinery is
built and correct — it is the audit that found the defect, twice — so "5 (re-run)" is a
*re-run of the matrix*, not a rebuild. **Phase 6 still does not start** until a passing
results file exists. **Phase 1 is a hard prerequisite of the re-run for both carriers**
(plan §3).
⚠️ **A smoke run of the audit against the fixed observer immediately found a second defect
(design v3.6 §6.8): a fail-closed `unresolved-ancestry` mark was being promoted to permanent
sticky taint.** Fixed in the taint engine (evidence-only sticky pass, 3 new tests,
mutation-verified) and merged. Decisions were unaffected — all 57 phase-2 tests passed
untouched — so this is a change to what stickiness *remembers*, not to what it *decides*.
Note the pattern for the re-run: the matrix rows assert exact partitions, and a stale sticky
entry from enumeration would have contaminated every one of them.
- **0b strictly precedes 6.** v2/v3 drew 0b with no continuing edge. Phase 6 is the first phase
that creates objects whose lifetime is tied to pixelpass being alive, so the teardown-ordering
guarantee must exist before it: without it, the AEC can unload while a fanning-out pixelpass
still holds link proxies and a stale module index (v3.4 §7.1).
- **0a strictly precedes 2**: the taint engine's lifetime-awareness (v3.4 §6.1.3) is keyed on
`object.serial`; building the model against the current lossy `u32`
(`pixelpass/src/host/audio.rs:534-540`; `RouterState::sink_serial` at `:594-598`) means a
cross-cutting migration later.
- **0c strictly precedes 0d** (round-2): 0d's types must be built around the *final*
connection-owned bare sink, not today's `Routing`. Building the type boundary against the
legacy pactl sink means rebuilding it when 0c lands.
- **1 strictly precedes 5**: without tags the taint engine has no roots and the dry-run can
only exercise the forwarder half of the problem.
---
## 2. Phase 0 — prerequisites
### 0a. `object.serial` u32 truncation — pixelpass
Parse as `u64` throughout; audit the other `parse::<u32>` at `:369` (determine whether it is a
serial or a genuinely-32-bit value before changing it). Tests: value > `u32::MAX`, and the
boundary.
### 0b. Explicit teardown + drop ordering — peerspeak
v3.4 §7.2, decision D4. All **three** of v3.4's fixes:
1. Replace the implicit-drop path at both channel-close sites. **Citation correction:** v3.4
says `core/mod.rs:1516`; the actual `None => break` arms are at **`:1514`** and **`:1532`**.
Take the session and `shutdown().await` it.
2. Move `echo_cancel` (currently `:682`) to the last declared field, after `screenshare_host`
(`:685`), with a comment naming the invariant.
3. **Last-ditch drop wrapper**: `start_kill` + a bounded `try_wait` reap on the host before the
AEC guard unloads. This is the *only* protection on the panic/unwind path, and unwind is
reachable — the core has numerous `unwrap()` sites and no `panic=abort` profile.
> ✅ **0b IMPLEMENTED 2026-07-26** (peerspeak branch `phase-0b-teardown`). The ordering
> defect was live: `echo_cancel` was declared *ahead* of `screenshare_host`, so any unwind
> unloaded the AEC while the host was still fanning out. Fields moved into
> `src/core/teardown.rs` with `echo_cancel` declared last, and `ReapOnDrop` added because
> `kill_on_drop(true)` only *signals* — it hands the child to the runtime's orphan queue,
> which an unwinding runtime may never drain. Matrix revised to four mutations; see §10
> round 14, and round 15 for the two blocking review findings that followed.
>
> **Owed to phase 9:** an explicit lifecycle row — *drop the controller / close the command
> channel while sharing* — which is the live proof for the hoisted teardown call site.
⚠️ **Mutation testing: five mutations, each independently breaking a named test.**
⚠️ **SUPERSEDED by §10 round 14 — mutation 2 is vacuous and the matrix is now four.** v1 demanded
a mutation that targeted the wrong defense; v2 fixed that but bundled two defenses into one
combined mutant, which proves neither. Final form:
| # | Mutation | Must break |
| --- | --- | --- |
| 1 | remove `shutdown().await` at `:1514` | close-arm-A teardown test |
| 2 | remove `shutdown().await` at `:1532` | close-arm-B teardown test |
| 3 | remove the explicit **wait** after host kill | explicit-ordering test (host kill+wait strictly precedes AEC unload) |
| 4 | reverse the field order | panic/unwind ordering test |
| 5 | remove the wrapper reap | panic/unwind ordering test (distinct assertion from #4) |
Both channel-close arms get their own test; a single "closes the command channel" test can
exercise one arm and leave the other unsafe.
### 0c. Graceful stop + connection-owned capture sink — both
> ✅ **MECHANISM PROBE PASSED on this host, 2026-07-26** (PipeWire 1.6.8). Run *before* any
> structural work, on the reviewer's insistence, because a single unverified assumption could
> have invalidated the entire approach: whether a hand-created adapter is visible to
> pipewire-pulse under the name pixelpass's capture path depends on. It is.
>
> ```
> pw-cli> create-node adapter factory.name=support.null-audio-sink \
> node.name=pixelpass_probe_<pid> media.class=Audio/Sink \
> audio.channels=2 audio.position=[FL,FR] node.virtual=true \
> monitor.channel-volumes=true object.linger=false
> ```
>
> Five gates, all green:
> 1. `pactl list short sinks` shows the sink under the **exact** `node.name`.
> 2. `pactl list short sources` shows **`<node.name>.monitor`** — the derived monitor name is
> a pipewire-pulse contract, not a property of Pulse-created sinks. This was the one that
> could have sunk the approach.
> 3. `gst-launch-1.0 pulsesrc device=<node.name>.monitor num-buffers=40 ! fakesink` pulled its
> buffers and exited clean, and a real recording stream attached — so the pixelpass capture
> path works against it unchanged.
> 4. **No null-sink module was loaded** (`pactl list short modules | grep -c null-sink` stayed
> at its baseline of 3). It is genuinely not a Pulse module.
> 5. **SIGKILL of the owning connection removed both Pulse-visible names**, with zero residue
> anywhere in `pw-dump`. That is the entire point of 0c, demonstrated on the real graph.
>
> The default sink never moved, so this is also safe to run on a live desktop.
> **Every O1 stop condition listed for 0c is retired.** `object.linger=false` is load-bearing:
> the bundled pipewire-rs example sets `linger=1` for the opposite behaviour.
>
> ⚠️ **`--repair`'s job does not shrink — it BREAKS.** Discovery derives dead host PIDs
> **only** from `module-null-sink` entries (`pixelpass/src/repair.rs`), and only then matches
> loopbacks against that PID set. The native sink is scoped to **every mode that owns a
> capture sink**, not just `DesktopExcluding`, so legacy Pulse loopbacks will coexist with a
> connection-owned sink; when that host dies the sink vanishes automatically and its loopbacks
> become **undiscoverable orphans**. Candidate PIDs must be derived independently from all
> three module shapes (`null-sink sink_name=`, `loopback sink=`, `loopback source=…monitor`),
> with a liveness recheck immediately before each destructive unload. This makes the repair
> rework **load-bearing, not defensive**.
v3.4 §7.4. The **largest hidden cost in Phase 0**: moving the null sink off `pactl load-module`
(`pixelpass/src/host/audio.rs:69`, cleaned up only in `Routing::cleanup` at `:259-260`, which
SIGKILL skips) onto a connection-owned PipeWire object.
- peerspeak: SIGINT (**not** SIGTERM — pixelpass installs only `ctrl_c()`,
`pixelpass/src/common/signal.rs:6`), bounded wait, SIGKILL fallback, at `core/mod.rs:699`
and `:3480`.
- pixelpass: connection-owned sink; `--repair` (`src/repair.rs:15-63`) extended and proven safe
with a second live host.
**Exit gate — two halves. The round-2 finding was that v2 gated only the first.**
*(i) Ownership,* a live two-host test — connection-ownership is a runtime property no unit test
can establish:
```bash
pactl list short sinks | rg 'pixelpass_capture_'
pw-dump | jq -r '.[] | select(.type=="PipeWire:Interface:Node") | .info.props as $p
| select(($p["node.name"] // "") | startswith("pixelpass_capture_"))
| [$p["object.serial"], $p["node.name"]] | @tsv'
kill -KILL <first-pixelpass-pid>
# re-run both: the killed host's object GONE, the second host's REMAINS
pixelpass --repair # the live host must be untouched
```
*(ii) Graceful stop,* which the above does not touch at all — it exercises only external
SIGKILL, so Stop Share could remain `child.kill().await` (`core/mod.rs:3480`, still true today)
and every command above would pass:
- a fake-child signal-order test: SIGINT first, SIGKILL only after the bound expires;
- a live Stop Share run: SIGINT sent, child exits within the declared bound, **no fallback
kill on the normal path**.
**O1 is closed: no demotion path.** v1 pre-authorised moving 0c after Phase 6 if it ballooned.
That is a waiver of settled decision v3.4 D6 hidden in a sequencing document, which is how a
converged design quietly decays. If 0c balloons, **stop and reopen D6 as design round 8.**
### 0d. Typed capture plan + internal mode input — pixelpass
Today `setup_audio` returns `(Option<Routing>, String)` (`pipeline.rs:123-142`) and the `String`
flows untyped into `build_args` (`:151-157`). There are **two** unsafe paths into the capture,
and v2 closed only the first:
**Path 1 — the source string.** `default_audio_monitor()` has exactly one call site,
`pipeline.rs:138`. That single line hands `pulsesrc` the real default monitor.
**Path 2 — the sink's inputs (round-2 P1, the defect v2 missed).** Even with a type-safe source,
`Routing::start` loads `module-loopback source=@DEFAULT_SINK@.monitor → pixelpass_capture_*`
whenever it runs outside strict-app mode (`audio.rs:80-92`), and it runs whenever
`PIXELPASS_AUDIO_VIA_NULL_SINK` is set (`pipeline.rs:124-125`). So `DesktopExcluding` could
correctly read *its own* sink's monitor while the legacy loopback has already filled that sink
with the whole-desktop mix — **full echo, with no source switch anywhere.** v3.4 §3 states this
loopback must never load in the new mode; nothing structurally enforced it.
Both are closed by construction:
- `LegacyDesktop` — the only variant that can produce `DefaultMonitor`.
- `PerApp { routing }` — legacy `Routing`, unchanged.
- `DesktopExcluding { capture_sink }` — owns a **bare** connection-owned sink type (0c) whose
API **cannot construct the legacy loopback at all**. Not "does not call it": the constructor
is not reachable from this variant.
- **Conflict policy, pinned here rather than discovered later:** the new mode combined with
`--app` or `PIXELPASS_AUDIO_VIA_NULL_SINK` **rejects at CLI parse time**. It must never fall
through to legacy `Routing`, and it must never silently ignore an input the user set.
- **Internal mode input lands here too** (round-2 P1): a non-advertised `HostOpts` field plus a
hidden trigger, so the variant is reachable through the real host/spawn path before Phase 6
needs to measure through it. `HostOpts` has no mode field today and `setup_audio` selects
solely on `app` + the env override.
**Why 0d is a prerequisite rather than a Phase 6 deliverable** (my divergence from Codex's
round-1 suggestion; it agreed in round 2): this is a pure non-mutating refactor of one function,
and landing it in Phase 6 means the link manager is written against the untyped API and then
refactored underneath itself, while the two most dangerous paths in the codebase stay unguarded
through four phases of active work around them. Guardrails go up before the scaffolding.
**Exit gate:**
- mode matrix over every `(app, strict_audio, mode, env-override)` combination asserting the
resulting capture plan, including every conflict combination rejecting;
- type-level: `DesktopExcluding` can name neither the default monitor nor the legacy loopback;
- **graph assertion**: with the new mode active, no default-monitor link or module feeds the
capture sink;
- legacy behaviour byte-identical.
A constructible-but-not-yet-public variant is acceptable for the interval between 0d and Phase
6 provided it is unit-tested and reachable by the hidden trigger.
---
## 3. Phase 1 — peerspeak ownership tagging (v3.4 §5.1)
Zero behaviour change; it is what makes Phase 5 observable.
- Native playback: prop on the stream dict, `src/audio/pipewire_impl.rs:374-388`.
- mpv/VLC spawn (`src/screenshare/mod.rs:768-775`), notification spawn (`src/notify.rs:265-272`):
`PULSE_PROP` + `PIPEWIRE_PROPS` on the `Command`.
⚠️ **The literal is a cross-repo wire contract and is pinned HERE, before Phase 1 starts** — not
deferred with the v3.4 §11 product naming, which is a separate and genuinely user-facing question.
```
key: peerspeak.owned
value: 1
```
⚠️ **Round 8 — a SECOND carrier is required, and its literal is pinned here too** (v3.5 §5.1).
`peerspeak.owned` is invisible to the registry `global` event and readable only via a node
bind (v3.5 §6.7); the prefix below is announced by the registry and needs no bind, so the
primary taint root no longer rests on a single observation mechanism.
```
key: node.name
format: peerspeak_owned_<role>_<pid> e.g. peerspeak_owned_mpv_31284
prefix: peerspeak_owned_ ← the matched literal
```
- **Both carriers are set at every tagging site.** A node is owned if **either** matches —
union, the fail-closed direction. The engine's tag root is `peerspeak.owned == 1` **OR**
`node.name` starts with `peerspeak_owned_`.
- **`node.description` is NOT touched**, so mixers still show "mpv". Only `node.name`, which
is the internal identifier, carries the prefix.
- The prefix mechanism is already proven here: `pixelpass_capture_*` is matched on
`node.name` and was the only root that kept working under the F1 defect.
- Same three requirements as the property literal: one named constant per repo, the
black-box cross-repo test driven from a shared fixture, and phase 5 as the real proof.
- ⚠️ Native call playback sets both on its own stream dict. The child spawns set the prefix
through the same `PULSE_PROP` / `PIPEWIRE_PROPS` env that carries the property —
`node.name` is settable there, and **the phase-1 exit gate must show it landing on a live
mpv node**, not just in the env.
**A per-repo literal test is not a contract test.** Two tests, one per repo, each maintained
beside its own implementation, get updated in lockstep with a rename and prove nothing. Required:
1. The literal appears **once** per repo as a named constant, commented with a pointer to this
section and to the other repo's constant.
2. A **black-box cross-repo test**: peerspeak constructs the child `Command`, the test reads the
env it would set, and asserts it produces the exact property string pixelpass's engine
matches on — driven from a single shared fixture string committed in both repos.
3. The real proof is the **Phase 5 dry-run**, which requires pixelpass to classify all three
live peerspeak playback paths as `NotEligible` *for the tag reason*. Emission alone proves
only that peerspeak talks, not that pixelpass listens.
**Exit gate:** `pw-dump` shows the tag on a native call playback node, an mpv node and a
notification node on this box. (Consumption is gated in Phase 5.)
**Non-goal:** the known grandchild-inheritance leak (v3.4 §5.1) stays accepted in v1.
---
## 4. Phases 24 — the engine (pixelpass)
### Phase 2 — graph model + taint engine, pure
All of v3.4 §6.1–§6.1.3, **with no PipeWire types in any signature**:
```
fn evaluate(snapshot: &GraphSnapshot, ctx: &ExclusionCtx, prior: &StickyState)
-> (Decisions, StickyState)
```
- `GraphSnapshot` = plain owned Node/Port/Link/Client structs keyed on `u64` serial, with the
recyclable id retained only as a lookup key, never as identity (v3.4 §6.1.3).
- `Decisions` carries `Eligibility::NotEligible { reason }` with **stable reason codes**, not
prose. That code is the Phase 5 dry-run output, the Phase 6 JSON status event, and the
eventual "why isn't this app shared" answer. Design it once, here.
- `StickyState` threaded explicitly, so stickiness is testable as a snapshot sequence.
**Every row of the v3.4 §12 taint-engine table is a required deliverable**, plus one addition:
an **empty/degenerate snapshot must yield "nothing eligible", not "everything eligible"** — the
fail-closed default asserted at the boundary.
**Exit gate:** fixture matrix green; the engine has never been linked against libpipewire.
### Phase 3 — registry observer + readiness epoch, read-only
v3.4 §6.3 and §6.4. Replaces (not extends) the existing router, which watches Node and Metadata
adds, forwards raw removals, and binds no graph (`src/host/audio.rs:523-585`).
> ⚠️ **Built and merged, then superseded in part by "Phase 3 revision (round 8)" below.** This
> section's node-property requirements assume the registry `global` event carries them. It does
> not (v3.5 §6.7). Everything here about removals, the readiness epoch, PID derivation and the
> Link path is unaffected and still holds.
- Node, Port, Link **and Client** globals; adds **and removes**.
- Link endpoint props from the global are the **optimisation**; the bind-`LinkInfoRef` fallback
is the correctness path.
- Readiness = `core.sync()`/`done` **plus** no outstanding required observations, fail-closed
timeout. Log which condition released the epoch.
- pipewire-pulse PID derivation (v3.4 §6.1.2): consistent `pipewire.sec.pid` across Pulse
clients, validated against `/proc/<pid>/comm`.
**Exit gate — six parts.** A one-time `pw-dump` diff passes while Port observation is absent,
removals are ignored, the fallback is dead code, and readiness releases early:
| gate | proves |
| --- | --- |
| adapter tests: add **and remove** of all four object types | removal handling exists |
| forced-absent Link endpoint props | the bind fallback is live, not decorative |
| unresolved-observer timeout test | readiness fails closed |
| readiness does not release with an observation outstanding | the epoch means something |
| **PID derivation matrix** (round-2): consistent valid PID · inconsistent PIDs · missing client property · `/proc` entry missing · `comm` mismatch · PID reuse — **every failure makes owner-bridge key 4 unusable** | the pure engine can be correct on a wrong context; this is where the context is built |
| **live**: create and destroy a controlled node/link topology; diff Nodes, **Ports**, Links and Clients before/during/after | the adapter tracks a *changing* graph, not a static one |
### Phase 3 revision (round 8) — bind every Node and Device ✅ BUILT AND MERGED 2026-07-25
v3.5 §6.7. Phase 3 shipped reading node properties off the registry `global` event, where
**eight of them are never announced**. This is the fix. Scope is the observer only — phases 2
and 4 are unaffected, and the phase-5 audit machinery is already correct.
> **🟢 Done.** Pure core + adapter, split-seam with mutual review as in phase 3 (mine and
> Codex's respectively, each reviewing the other). All four gate rows below pass, the live
> row on this host. Codex's review of the core found no certain P1; two findings taken and
> mutation-verified (a `device_props` ambiguity test that checked for one live *Device*
> rather than one live *global*, and `device.api` corroborating by presence). Two findings
> left open as design items, both pre-existing — hardware playback-to-capture paths and the
> readiness-budget calibration, both recorded in design §6.8.
>
> **Added beyond the spec: a second live gate for the Device-side path.** Row 1's
> `session_device` assertion is satisfied by a union, and WirePlumber 0.5.15 copies
> `device.api`/`alsa.driver_name` onto ALSA nodes on this host — so row 1 passes through the
> node fallback and would keep passing if the Device bind delivered nothing at all, leaving
> §6.7 decision 4 ungated on the development machine. Verified by mutation: breaking the
> Device-side driver read fails the new test while row 1 still passes.
**Requirements.**
1. **Bind every `Node` global**, unconditionally, no `media.class` filter. Retain the proxy
and its `info` listener in that global's slot in the existing per-id FIFO
(`LiveGlobal.bound_link` generalises to a bound-proxy slot).
⚠️ The phase-3 review's finding 3 — record the id and apply the add as **one** step, so
the proxy FIFO stays lockstep with the model's `live_ids` — now applies on the **hottest**
path in the observer. A recycled Node id must not pop another generation's proxy.
2. **The global is an index; `info` is the source of truth.** Read from the global only what
must exist before the bind resolves: `object.serial` (identity), the object's id, and
`device.id`/`node.id` linkage. **Every** taint-relevant property — including `node.name`
and `media.class`, so there is exactly one source — comes from the bound `info` props.
3. **A node with no `info` yet is WITHHELD from the snapshot and is a readiness obligation**
(`pending_nodes`, beside `withheld` and `pending_links`). `graph_ready` false while any is
outstanding; the existing bounded deadline makes an unresolvable bind sticky-`TimedOut`,
fail closed. No provisional-ownership admission, ever (v3.4 §6.1.3).
4. **Track props for the node's lifetime.** On a later `info` with `PROPS` in `change_mask`,
re-read, re-classify, and apply a `NodePropsUpdated` event.
⚠️ **Suppression rule:** a prop update may be dropped **only** when the resulting
`Projection` is identical to the current one. Anything looser breaks phase 4's
no-coalescing contract; anything stricter (emitting on every `info`, including
state-only changes) inflates the O5 event rate with non-events.
5. **Bind every `Device` global** and read `device.api` **and** `alsa.driver_name` from its
`info` props — authoritative, and the phase-3 review's owed fix (on PipeWire ≥ 1.2.6 with
WirePlumber < 0.5.13 the driver name is not copied to the node, and the fail-closed
absent-driver rule would over-exclude real cards). `factory.name` exists only on the node.
`classify` takes both sides; node values are the fallback, Device values win.
6. **Ports are NOT bound in v1 — an explicit accepted limitation.** `port.exclusive` is the
only port property missing from the global, and it guards a *mutation* (don't fan out into
an exclusive port), not echo: an exclusive port rejects the second link, so phase 6 sees a
clean link-create failure it must handle correctly anyway. Binding ~21 more objects at
rest to pre-empt an error that surfaces safely is not worth the obligation surface in v1.
**Revisit trigger:** any phase-6 link-matrix row where an exclusive-port link failure is
not cleanly recoverable. (`node.passthrough`, the *other* half of that §6.2 row, is a node
property and **is** recovered by this revision.)
**Exit gate — four parts.** The first is the direct inverse of the F1 finding.
| gate | proves |
| --- | --- |
| **live prop recovery**: a `module-null-sink` tagged `peerspeak.owned=true` plus a `module-loopback` reading its monitor — assert the projection carries `peerspeak.owned`, `pulse.module.id`, `node.link-group` **and** `factory.name`/`device.api`/`alsa.driver_name` on a real ALSA node | the eight properties actually arrive — F1 cannot recur silently |
| **pure-model prop-update matrix**: props-changed → re-classified; identical props → suppressed; a `session_device`-relevant change flips classification. (A *live* prop mutation has no reliable CLI trigger — the pure test is the gate, a live sighting is opportunistic) | the lifetime-tracking path exists and its suppression rule is exact |
| **readiness with node binds**: no projection reports `graph_ready` while a node bind is outstanding; an `info` that never arrives ends in sticky `TimedOut` | withholding and fail-closed timeout still hold with the new obligation class |
| **recycled Node id under churn**: repeated add/remove of the same id; no proxy leak, no cross-generation misattribution | the FIFO lockstep rule survives being moved to the hot path |
**Then re-run the whole phase-5 §5.1 matrix and re-measure O5** with bind I/O included — the
existing numbers were taken on the degraded graph and inherit nothing.
### Phase 4 — AEC identity validation state machine, read-only
v3.4 §5.3 verbatim: `NotConfigured / Validating / Validated / Failed / Revoked`;
`--aec=off|pulse-module:<idx>` parsing (D5); bounded deadline; **no fan-out while `Validating`**;
revocation = loss of *all* nodes bearing the index, never one leg corking. Foreign
`echo-cancel-*` groups: warn and exclude (D3).
**Moved ahead of the dry-run gate (round-1 P1).** v1 put this after the dry-run while the
dry-run checklist required AEC validation and revocation semantics — a circular dependency that
made the major gate uncompletable as written.
**Exit gate — a fake-clock/event-sequence transition matrix**, because these are timing
semantics a live poke cannot cover: `Validating → Failed` on deadline expiry; `Validating →
Validated` on first matching node; partial-node disappearance ⇒ **stays `Validated`**; all
nodes gone ⇒ `Revoked`; `Revoked` stops fan-out and drops proxies; a retained stale index does
not alias onto a reloaded module (v3.4 §5.2 correction 3 — indices *are* reused). Parsing: JSON
number and string forms, `> u32::MAX`, absent, malformed.
Then wire the state machine's output into the dry-run so `Validating`/`Failed`/`Revoked` are
observable in Phase 5 before they gate anything real.
---
## 5. Phase 5 — dry-run audit mode 🚦 MAJOR GATE
> **🚦 STATUS 2026-07-26: GATE PASSED on run 2.** Results:
> `docs/screenshare-audio-exclusion-phase5-results.md`. All 13 rows completed, the eligible
> half of every row is non-empty, and O5 is re-measured on the fixed graph (worst recompute
> 67 µs; readiness 12 ms with 18 binds). Three rows carry recorded substitutions (8, 9, 13)
> and three findings are recorded as non-blocking.
>
> **Run 2 found and fixed a third defect of the F2 class, F13-1:** pipewire-pulse's PID was
> unresolvable on this host *permanently*, because stage 1 of the derivation required exactly
> one repeated `sec_pid` and **WirePlumber repeats one too** (two Clients, one PID). Key 4's
> suppression therefore never fired and every Pulse-emulated node fused into one owner. Fixed
> in pixelpass `91c4ded`: probe every distinct `sec_pid` and let `/proc/<pid>/comm` decide.
> **The eligible half of row 1 is the only thing that exposed it** — the verdict was
> fail-closed and silent.
>
> ⚠️ **Phase 6 is NOT unblocked by this file alone.** F11-1 was the other gate and is now
> **closed** (2026-07-26, pixelpass `c78eb2d`: key 4 bounds an owner only when the node's
> Client resolves; measured cost on the live graph, zero — see the results file). Phases
> 0b/0c/0d and the "Stereo Mix" design call still precede phase 6.
>
> Two things to keep when re-running: **every partition row must run with `AEC=off`** (a
> configured-but-unvalidated AEC shuts the fan-out gate and empties the eligible half of every
> row, which reads as a failure that is really a harness error), and **start the audit BEFORE
> building the fixture**. Fixture-first makes the whole graph arrive as one enumeration burst,
> so every node is first tainted while `graph_ready` is false; that partial-graph taint enters
> sticky state and the keyless sticky reason then wins over the evidence-derived one, so a row
> cannot assert its own key. Read keys at *derivation* (first non-sticky appearance).
**Adds no capability. Its entire purpose is to be wrong loudly and safely.**
A hidden trigger (`PIXELPASS_AUDIO_AUDIT=1`) running Phases 24 against the live graph on every
graph event, emitting per `Stream/Output/Audio` node: serial, name, decision, stable reason
code, graph epoch. It creates **no links**. Output goes to **stderr or a defined JSON event**
never unstructured prose into `--output json`, which peerspeak parses
(`screenshare/mod.rs:92`).
Why this is the gate: the C2/C3-class defects are graph-*reasoning* defects. A fixture proves
the code matches my model of PipeWire; only a live run proves my model matches PipeWire. A wrong
answer here costs a log line; the same wrong answer in Phase 6 costs an echo.
### 5.1 Every row asserts an exact partition, not a spot check
Round 2's sharpest structural point: checking only named targets constrains nothing about
everything else, so **each row must assert the complete candidate universe partitioned into
exact eligible and excluded sets, with reason codes on the excluded side.** That single
requirement is also the answer to O7 — it is the over-exclusion gate, because an
exclude-everything implementation fails the eligible half of every row.
| # | Scenario | Excluded (with reason code) | Eligible |
| --- | --- | --- | --- |
| 1 | `module-null-sink` + `module-loopback` forwarder (the v3.4 §6.1 measured shape) | output leg, reason = **owner bridge**, naming the key — *not* a Link walk | same forwarder shape with **no** tainted input |
| 1b | *opportunistic, non-gating:* Sunshine's null-sink topology while it is routing desktop audio | its forwarder leg, if a re-emitting leg exists | — |
| 2 | `gst-launch pulsesrc ! pulsesink` split clients, input **explicitly rooted on a tainted monitor** | output leg via key 4 | the same process reading an **untainted** source |
| 3 | **two** Pulse modules; **one** tainted input | the tainted module's output only | **the other module's output must be ELIGIBLE** — this is what makes wrong pipewire-pulse-PID fusion observable |
| 4 | peerspeak native call playback | that node, reason = tag | — |
| 5 | peerspeak-spawned **mpv** (watched share) | that node, reason = tag | mpv launched by hand |
| 6 | peerspeak **notification** sound | that node, reason = tag | — |
| 7 | a **second** pixelpass host's capture sink, **plus a controlled forwarder reading that sink's monitor** | the forwarder's **named output serial** (cycle prevention, v3.4 §6.2) | — |
| 8 | EasyEffects running | combined output leg | EasyEffects stopped ⇒ ordinary streams |
| 9 | Firefox: music only / mic on untainted source / capturing a tainted monitor | the third only (v3.4 §6.1.1) | the first two |
| 10 | sticky taint: tainted input leg removed, output leg lives | still excluded | after full owner teardown + restart |
| 11 | recycled serial/index/link-group after teardown | — | must **not** inherit taint |
| 12 | AEC loaded, then unloaded | four nodes; then `Revoked` | — |
| 13 | `Audio/Duplex` device | over-taints, recorded as **known accepted** (v3.4 §6.1 caveat) | — |
Rows 3 and 7 were vacuous in v2: row 3 had no tainted module, so incorrect fusion of all
pipewire-pulse modules changed no emitted decision; row 7 observed a capture sink without naming
a downstream candidate, so recognising `pixelpass_capture_*` as a mere sink name would pass
without any transitive propagation.
### 5.2 Also record, per O5
Graph-event rate, recompute duration **distribution and maximum**, and whether events queue
behind recompute/logging. v3.4 §6.4's "full recompute is fine for v1" then rests on measured
headroom and epoch lag rather than on a node count.
### 5.3 ⚠️ Do not build a gate on a transient topology
v1 leaned on v3.4 §6.1.0's "the hazard is LIVE on this machine right now." **Measured
2026-07-21 ~14:55 — no longer true**, six hours after it was written: `Default Sink` is
`alsa_output.pci-0000_10_00.6.analog-stereo` (IDLE), all three `sink-sunshine-*` null sinks
SUSPENDED. Sunshine is still running (pid 4104) and still reads a monitor — but the hardware
sink's, via active link `56 → 95`, not a null sink's. So Sunshine running is **not** sufficient
for the topology to be present; see plan §11.
The **controlled fixture (row 1) is authoritative** — deterministic and always available. But
the wild sample is not therefore unnecessary: a fixture I build tests my model against my own
assumptions, whereas Sunshine is an uncontrived third-party forwarder nobody designed for this
test. It stays as row 1b, **opportunistic and non-gating**, because it cannot be relied on to
be present.
**Any surprise here goes back to the design doc as round 8. Phase 6 does not start until this
results file exists.**
---
## 6. Phases 68 — mutation, capability, integration
### Phase 6 — link manager + status events, driven through the real host path
v3.4 §4.2 + §6.2 + §6.3 items 34. Non-lingering links (rig gotcha: `object.linger=false` is
*ignored* by `pw-link --props` and `pw-cli create-link`; only `pw-link -m` yields one), proxies
retained for the life of the share, per-port link sets, "captured" only when **every** required
link is `ACTIVE`, same-epoch revalidation immediately before each creation, proxy drop on
ancestry becoming unsafe.
Failure ⇒ report the stream unsupported. **Never** fall back to the default monitor — and after
0d that fallback is unconstructible in this mode, by either path.
**Everything here is measured through the real selector → sink → link manager → `pulsesrc`
path**, using 0d's hidden trigger. A harness-only measurement would pass while the production
CLI still reaches only legacy branches.
**Status events land here** (round-1/2: no phase owned them). pixelpass's event enum
(`src/common/output.rs:36-66`) has nothing for exclusion status, and capture-spawn failure
(`host/mod.rs:309-312`) replies to the viewer while emitting no event at all. Required as
**versioned wire-shaped events**, not stderr lines. Without them the safe failure mode is
unexplained silence after the first viewer connects — and a sharer who cannot see why will
switch back to unsafe whole-desktop audio.
There are **four** production causes, and each needs an **exact JSON golden plus a cause →
emission test** — not a shared "an event is emitted" assertion, which passes while three of the
four remain unwired:
| cause | event | trigger under test |
| --- | --- | --- |
| per-stream link failure | `stream_unsupported` | link-matrix row 8c |
| AEC validation deadline | `aec_failed` | Phase 4 `Validating → Failed` |
| AEC identity lost mid-share | `aec_revoked` | Phase 4 `Validated → Revoked` |
| foreign `echo-cancel-*` present (D3) | `foreign_aec_warning` | a second AEC module loaded |
Phase 8 owns the other half of each: parse, traverse the **new mode's** notice channel, and
reach the intended UI state. The channel is currently created only for `audio_app`
(`core/mod.rs:3424`) and only the two `AppAudio` events are translated (`:3431-3434`).
#### 6.1 Link-manager matrix (local anchor — v3.4 §12 has only a one-line bullet)
v2 pointed its most important gate at a "v3.4 §12 bookkeeping matrix" that does not exist.
Here it is. Each row is a deterministic test with an injected graph, not a live observation:
| # | Case | Assertion |
| --- | --- | --- |
| 1 | graph mutated to tainted **between evaluation and `create_link`** | **zero unsafe `create_link` calls** — not "eventually cleaned up" |
| 2 | per-port enumeration | exact set of attempted links and their states |
| 3 | partial activation (FL `ACTIVE`, FR not) | **not** reported captured |
| 4 | duplicate enumeration of the same node | idempotent; no second link set |
| 5 | ancestry becomes unsafe after `ACTIVE` | owned proxies dropped |
| 6a | `port.exclusive` port | refused, reason code emitted |
| 6b | encoded stream | refused, reason code emitted |
| 6c | passthrough (IEC958) stream | refused, reason code emitted |
| 7 | capture sink replaced | relink **succeeds** — every required port back to `ACTIVE` and the node reported captured again; stale proxies dropped |
| 8a | AEC `Failed` (validation deadline) | plan stays `DesktopExcluding`; no capture |
| 8b | AEC `Revoked` mid-share | plan stays `DesktopExcluding`; fan-out stops |
| 8c | link creation error | plan stays `DesktopExcluding`; that stream reported unsupported |
| 8d | capture-sink creation failure | plan stays `DesktopExcluding`; mode fails, does not degrade |
| 8e | readiness-epoch timeout | plan stays `DesktopExcluding`; fail closed |
| 9 | an **eligible** late-arriving node | **positively captured** — the over-exclusion counterpart to row 1 |
Rows 6a6c were one combined fixture in v3: a single working refusal predicate would have
masked two missing ones. Row 7 required only "relink attempted", which a permanently-failing
attempt satisfies while v3.4 §4.2 requires a live owner to actually restore links after sink
recreation. Rows 8a8e replace an unenumerated "any failure path", under which testing one
handler passes while another silently swaps the plan to `LegacyDesktop`.
Row 1 is the one v2 could not falsify: "clean → tainted mid-share ⇒ links dropped" can pass by
observing eventual removal, while an unsafe link genuinely existed for a window. Row 8 is O8's
answer: 0d's enum prevents a `DesktopExcluding` value from *containing* `DefaultMonitor`, but
not a failure handler from replacing the whole plan with `LegacyDesktop`, so this needs a
release-mode integration test per failure transition. A `debug_assert!` is cheap and worth
adding, but it is not a gate.
Plus the live dynamic matrix (SIGKILL removes owned links; node appearing after share start;
sink recreation) and the three-arm leak measurement re-run through the production path with
v3.4 §12's rig discipline (`media.class` filter first, never drop stderr, verify the link is
in-graph, `parec -d <sink>.monitor`). The deliberately-naive predicate used as that
measurement's positive control lives in a **test-only injected implementation**, never a
shippable runtime override.
### Phase 7 — public mode selector + capability advertisement (pixelpass ships first)
⚠️ **Round-3 P1: nothing in v3 ever promoted the hidden trigger to a public flag.** 0d added an
internal mode input; Phase 7 advertised capability and naming; Phase 8 added `--aec`, the picker
and status. No phase required the actual **mode selector** to exist publicly or to be passed.
The result would be a capability-gated picker entry that, when chosen, still spawns legacy
whole-desktop capture — the feature appearing to ship while doing nothing. Reachable: peerspeak's
host argv has no mode parameter (`screenshare/mod.rs:152`) and pixelpass's `HostOpts` has no mode
field (`cli.rs:153`); v3.4 §11 requires a distinct mode selector.
So Phase 7 lands **both**:
1. the public mode flag (naming per v3.4 §11), replacing the 0d hidden trigger as the production
entry point — the hidden trigger may remain for testing;
2. D2's **versioned machine-readable** capability response or bitset. Must **not** overload
`app_audio_supported: bool` — per-app-strict and desktop-excluding are independent
capabilities. `--help` substring probing survives only as the legacy fallback.
**Old-peerspeak + new-pixelpass golden test lands here, before pixelpass ships**: behaviour
byte-identical, absent `--aec` still accepted.
v3.4 §11 public naming is a **blocking user input at the start of this phase**. Internal typed
variant names (0d) do not block on it.
### Phase 8 — peerspeak integration
- `EchoCancelGuard::module_index()` accessor (currently private; only `source_name()` /
`sink_name()` exist).
- **Emit the public mode flag** when the new picker choice is selected. Gated by an **exact
new/new argv golden** that requires *both* the mode flag and `--aec=…` to be present — the
round-3 P1. An argv test that checks only `--aec` passes while the mode flag is never sent and
pixelpass silently runs legacy capture.
- Always pass `--aec=off|pulse-module:<idx>` — absence is not a protocol state (D5).
- **Bind capability to the resolved binary path**; re-probe immediately before constructing
new-mode argv if resolution differs from the probe's; fail closed. Closes the `:3375-3378`
vs `:3409-3419` double-resolution gap.
- **New-peerspeak + old-pixelpass golden test** (round-2: the phase map promised "both
directions" and only one was specified) — against an old-capability response and an old fake
binary: the new picker entry stays absent and **no new flags are emitted**. Path rebinding
alone does not test the failure policy.
- Parse and surface the Phase 6 status events, with a **causal** test: an event emitted by
pixelpass must reach the UI. Matching enums defined independently in both repos would
otherwise pass. The notice channel is currently created only when `audio_app` is set
(`core/mod.rs:3424`) and only the two `AppAudio` events are translated (`:3431-3434`);
everything else is logged and lost — so the new mode needs its own channel creation path.
- Capability-gated picker entry; wording per v3.4 §11.
- Regression: existing `--app` / `--strict-audio` argv byte-identical to today.
---
## 7. Phase 9 — rig upgrade and field tests 🚦 SHIP GATE
v3.4 §9.2's rig upgrade is **owed before any exclusion claim is published**: two orthogonal
PN/MLS probes, windowed per-channel normalised cross-correlation reporting max per-window
correlation, plus xrun telemetry. Until it exists the only defensible claim is the gross-leak
distinction, in v3.4 §9.2's exact wording.
**Every row gets a declared pass/fail threshold before the run, not after.** Baseline for all
rows: excluded probe ≤ the declared rig criterion; **eligible control audio present**; original
playback routes intact; zero surviving owned links or capture sinks after teardown; xrun and CPU
within recorded bounds.
The **full** v3.4 §12 matrix — v1 silently dropped rows 2 and 7:
1. Sharer in a call while sharing, AEC on **and** off.
2. **Sharer simultaneously viewing another share while sharing** (restored). Highest-value test
of the child-tag path: mpv playing a watched share while hosting. Reachable —
`StartScreenShare` stores a host at `core/mod.rs:3458`, `ViewShare` stores viewer children at
`:3525`, no mutual exclusion.
3. Lifecycle, each separately: Stop, room leave, UI crash, pixelpass panic, SIGINT, SIGTERM,
SIGKILL, last viewer, pipewire-pulse restart, PipeWire daemon restart, `--repair`.
4. EasyEffects running for the whole share.
5. Output-device switch mid-share via the real `Ctrl+Meta+F` / `Ctrl+Meta+S` scripts.
6. Two concurrent hosts; notification mid-share; app that starts playing after the share.
7. **Sample-rate / channel / passthrough behaviour on real sinks, and CPU cost** (restored).
⚠️ **Mid-share taint-root arrival — v3.4 §6.1.4 names an unreachable case, and so did my first
replacement.** v3.4 says "AEC-load-mid-share is the case to test": unreachable, because there is
exactly one `echo_cancel::enable` site at session join (`core/mod.rs:1850`), the guard moves
into `ActiveSession` at `:2729`, and `StartScreenShare` rejects `active_session == None` at
`:3397-3404` ("Join a call before sharing your screen") — so the AEC always predates the share.
My proposed replacement, "a peer joining creates their playback node," is **also wrong**:
peerspeak starts **one mixed playback stream** at session construction (the sole core
`start_playback`, `core/mod.rs:1900`), and `PeerJoined` (`:2396-2409`) only admits and connects
the sender. No per-peer node is ever created.
The reachable newly-created mid-share taint roots are: **a notification sound played mid-share**
(`notify.rs:265-272`), and **starting to view another share mid-share**, which spawns a tagged
mpv/VLC (`screenshare/mod.rs:768-775`). Those are the transition-window field tests. Owned-AEC
mid-share load stays a **synthetic** test until a second `enable` site or hot reload arms it.
---
## 8. Open questions — final status
| # | Question | Status |
| --- | --- | --- |
| O1 | Is 0c a true blocker? | **CLOSED — yes, no demotion path.** The *echo* argument for demoting it is sound and irrelevant: D6 settled it. Balloon ⇒ round 8. |
| O2 | Ship the dry-run mode? | **CLOSED — keep**, env-gated, stable reason codes + epoch + serial, stderr or defined JSON event so `--output json` stays clean. |
| O3 | Naming | **SPLIT.** The `peerspeak.owned` wire literal is pinned in plan §3 now (a contract, not product wording). Public mode/picker wording remains the **user's call**, blocking at the start of Phase 7 only. |
| O4 | Sticky state: engine or observer? | **CLOSED — pure engine.** Observer supplies lifetime-bearing membership/removal facts; the engine decides. |
| O5 | Is full recompute really fine? | **CLOSED — measure it in Phase 5**: duration distribution, maximum, and queueing, not a recompute count. |
| O6 | Can the default-monitor fallback be made structurally impossible? | **CLOSED — yes, but it took two closures, not one.** Source path *and* sink-input path, both in 0d. Placement 0d rather than Phase 6 was my divergence; Codex agreed in round 2 with the added constraint `0c → 0d`. |
| O7 | Does over-exclusion need its own gate? | **CLOSED — subsumed.** Codex correctly narrowed my premise: an exclude-everything build already fails the eligible controls in six Phase 5 rows *provided they are asserted*. Fix is the exact-partition requirement (plan §5.1) plus link-matrix row 9's positive capture assertion. |
| O8 | Runtime assertion for "no source switch on failure"? | **CLOSED — 0d's enum is insufficient.** It stops a `DesktopExcluding` value containing `DefaultMonitor`, not a failure handler swapping the whole plan for `LegacyDesktop`. Release-mode integration test per failure transition (link-matrix row 8); `debug_assert!` in addition, but it is not the gate. |
---
## 9. Risk register
| Risk | Where it bites | Mitigation |
| --- | --- | --- |
| Engine correct, **source** wrong | full echo, engine bypassed | 0d path 1 — typed capture plan |
| Engine correct, **sink inputs** poisoned | full echo, no source switch anywhere | 0d path 2 — bare sink type + conflict rejection + graph assertion |
| Taint engine subtly wrong about real PipeWire | Phase 6 leaks the call into the share | Phase 5 exact-partition gate with negative controls |
| Unsafe link exists briefly, then is cleaned up | a real leak that "eventual cleanup" tests score as a pass | link-matrix row 1: zero unsafe `create_link` calls |
| Link manager built before the engine is validated | same, discovered in front of a viewer | strict DAG; plan §0's stated temptation |
| Tag literal mismatch across repos | v3.4 §5.1 silently does nothing, *quietly* | literal pinned in plan §3; cross-repo black-box test; consumption gated in Phase 5 rows 46 |
| Cross-repo skew | share hard-fails on spawn | pixelpass-first; a golden test in **each** direction; capability bound to resolved path |
| Fail-closed with no explanation | user switches back to unsafe whole-desktop audio | causal status-delivery test, Phase 6 → Phase 8 |
| Over-exclusion ships as "working" | mode captures silence, all gates pass | exact partitions (plan §5.1) + link-matrix row 9 — **🟢 FIRED 2026-07-25 and worked**: the build *was* the exclude-everything degenerate case, and the empty eligible half is what exposed it |
| **A property the engine reads is silently absent at the observation boundary** | engine correct, context permanently `None`; fails in *both* directions at once (F1: no taint root ⇒ echo; F2: no owner key ⇒ exclude everything) | **v3.5 §6.7 — never read node/device props off a registry global.** Phase 3r's live prop-recovery gate asserts each one arrives. General form: `pw-dump` is a **bound** view; the registry is not, and the difference is silent |
| Wrong pipewire-pulse PID | mass over-exclusion from a correct engine on a wrong context | Phase 3 PID-derivation matrix |
| 0c balloons | prerequisites eat the schedule | reopen D6 as round 8 — no silent waiver |
| "It works on my box" | the only box is this box | two-machine field test is the ship gate |
## 10. Adjudication record
**Round 14 (2026-07-26) — 0b's five-mutation matrix is revised to four, and one pinned
mutation is retired as vacuous.** Reached independently by both reviewers, then agreed.
- **Mutation 2 cannot be killed by any test, because its site cannot execute.** The
best-effort wake arm (`core/mod.rs`, the `besteffort_wake_rx` close arm) is unreachable
**by construction, twice over**: (i) `run_core_loop` owns a clone of `besteffort_wake_tx`
— created at `CoreController::new` and used for the `has_more` re-arm inside the loop —
and a tokio `Receiver::recv()` yields `None` only once *every* sender is dropped; (ii)
even without that clone, both `CoreController` and `CoreCommandSender` hold `reliable_tx`
alongside the wake sender, and the `select!` is `biased` with the reliable arm first, so
the reliable arm always wins the race to exit. Writing teardown there would be code that
provably never runs, dressed as a tested path.
- **Mutation 1's site is reachable but not unit-testable.** It sits inside `run_core_loop`,
which builds a real iroh endpoint and loads identity; no unit test can drive it.
- **Decision: (b) + (c).** Teardown is **hoisted to one unconditional site after the loop**,
so every `break` is covered structurally, including any added later — strictly better than
duplicating teardown across one live arm and one dead one. The **seam-level mutation gates
are the real ordering proof**, and the call site's live proof is owed to **phase 9**, which
gains an explicit row: *drop the controller / close the command channel while sharing*.
"UI crash" is not precise enough to serve as that row.
- **Rejected: an integration test built to preserve the number five.** It would pay for a
full iroh core plus test-only observability and prove only that a method was called — not
the ordering invariant, which is the thing that actually breaks.
- The 0b gate is therefore **four mutations**, enumerated exactly (round 16 P3-5 — the earlier
wording said "three plus the 0c pair", which reads as five and blurred what 0b owns):
| # | mutation | killed by | status |
|---|----------|-----------|--------|
| 1 (old gate 3) | remove the wait after the host kill | `explicit_shutdown_reaps_the_host_before_the_aec_can_unload` | killed now |
| 2 (old gate 4) | reverse `ScreenshareTeardown`'s field declaration order | `the_aec_unloads_after_the_children_on_the_drop_path` | killed now |
| 3 (old gate 5) | remove the reap loop from `ReapOnDrop::drop` | `dropping_a_guard_kills_and_then_reaps_the_child` | killed now |
| 4 (old 1) | remove the teardown at the hoisted post-loop call site | — | **deferred to the phase-9 row** *drop the controller / close the command channel while sharing* |
4-vs-5 separation verified: reversing the field order leaves the reap test green, and
removing the reap loop leaves the ordering test green. **0c's own pair (no SIGINT · no
SIGKILL fallback) is counted under 0c, not here**, along with the round-15/16 additions
(disarm the wrapper at entry · disarm it between the waits · treat a wait error as a reap ·
report an unconfirmed stop as clean · zero the grace).
**Round 15 (2026-07-26) — review of the 0b/0c-peerspeak implementation returned two blocking
findings, both accepted.** Recorded because both are the same shape: a defence that existed
but was disarmed exactly when it was needed.
- **The drop fallback was disarmed across its own wait.** `shutdown` took the child out of
the wrapper before the first `.await`; a cancellation or unwind during the wait left the
raw child to drop with `kill_on_drop` (which signals without reaping) while `Drop` found
`None`. The child now stays owned until the reap is **confirmed**.
- **A failed wait was reported as a reap, and the hard-kill wait was unbounded.** The
`io::Result` was discarded, so a wait error returned "reaped"; and a process in
uninterruptible sleep after SIGKILL could wedge the core loop forever. Both waits are now
bounded and the conflict case has a written policy: availability wins, the child stays
owned so the bounded `Drop` retry stays armed, and the residual risk is logged.
- **A gate of mine was vacuous and the review's fourth test-double point caught it.** The
elapsed-time assertion compared against `STOP_GRACE` itself, so zeroing the constant left
it trivially true. `the_grace_is_a_real_interval` now pins the constant to a band.
**Round 16 (2026-07-26) — the re-review of the 0b/0c-peerspeak fixes returned *approve with
follow-ups*: no blocking findings, five P3s, all five applied before the merge.** The two that
carry design content:
- **An unconfirmed stop was reported to the user as a clean one.** `stop_host` returned a bare
"was sharing" bool, so the one case where availability-first gives up (SIGKILL queued, reap
never confirmed) still emitted `ScreenShareStopped` with no warning — the UI would say
sharing ended while pixelpass might still be fanning out. `ReapOnDrop::shutdown` now returns
`StopOutcome`, `stop_host` returns `Option<StopOutcome>`, and an `Unconfirmed` user-initiated
stop raises a UI error naming the stray process. Session/viewer teardown discards the outcome
on purpose: no user is waiting on an answer there and the risk is already logged.
- **Cancellation coverage only reached the graceful wait.** The mid-wait test could not kill a
mutant that disarmed the wrapper *between* the two waits. Verified: the naive form of that
mutant does not compile (the child is borrowed from `self`), but the restructured form —
`self.child.take()` once cooperation has failed — compiles, and the pre-existing test passes
it. `cancelling_shutdown_after_the_kill_leaves_the_fallback_armed` kills it.
**Deferred item — aggregate teardown latency (round 16 P3-4).** Bounds are per child, not per
teardown. Sequential drain gives `2 × STOP_GRACE` per unconfirmed child inline (≈4 s), plus
`REAP_BUDGET` (250 ms) per child on the `Drop` path: three wedged children ≈6 s of command-loop
stall, ≈12.75 s worst case including drop retries. Accepted as-is for 0b — one host plus one or
two viewers is the real shape, and concurrency here would mean detaching children from the
session that owns the AEC's lifetime. **Trigger to revisit: a fourth tracked child becomes
routine, or a measured teardown exceeds 5 s.** The fix, when triggered, is to drain viewers
concurrently while still owned by `shutdown_children` — not to detach them.
**Round 1 — 13 items, 12 accepted.** Phase reorder (AEC machine before dry-run); typed capture
plan (accepted, moved *earlier* than proposed); Phase 3 five-part gate; tag-consumption gating;
Phase 6 matrix mandatory; 0b unwind backstop restored **and my mutation test corrected — it
targeted the wrong mutation**; Phase 9 rows restored with pre-declared thresholds; DAG stated;
O1 demotion language removed; 0c two-host gate; skew tests moved to Phase 7; status events
assigned. **Half-rejected:** "the Sunshine topology is unverified and unnecessary" — *unverified*
was right and it has since flipped; *unnecessary* rejected, retained as non-gating row 1b.
**Round 2 — 11 items, all accepted.** The three that mattered:
- **P1, the sink-input path.** My 0d closed the source and left the capture sink's inputs open;
`PIXELPASS_AUDIO_VIA_NULL_SINK` + `Routing::start` would have filled the owned sink with the
whole-desktop mix and produced full echo with no source switch. This is the "at least one
comparable error" I asked round 2 to find, and it was in the fix for round 1's headline P1.
- **P2, my v3.4 §6.1.4 replacement was also unreachable.** I claimed a peer joining creates their
playback node; verified false — one mixed playback stream at session construction
(`core/mod.rs:1900`), `PeerJoined` only admits the sender. Replaced with notification sound and
watched-share start, both reachable.
- **P1, no invocable production selector**, so Phase 6's "production path" measurement would have
run through a harness. Internal mode input moved into 0d.
**Round 3 — verification pass, 5 items, all accepted.** It confirmed 4 of the 7 round-2 edits
landed and 3 were partial, which is the reason to run a verification round at all rather than
declaring the fixes done. The one that mattered:
- **P1, the public mode selector was never assigned to any phase.** 0d added a hidden trigger,
Phase 7 added capability + naming, Phase 8 added `--aec` — and nothing required the mode flag
itself to exist publicly or be passed. A capability-gated picker entry would have appeared and,
when chosen, spawned legacy whole-desktop capture: the feature shipping while doing nothing,
with the echo intact. Fixed in Phase 7 (flag) and Phase 8 (emission + new/new argv golden).
- Three link-matrix rows I had just written were insufficiently falsifiable — a combined
exclusive/encoded/passthrough fixture (one working predicate masks two missing), "relink
attempted" (a permanently-failing attempt passes), and an unenumerated "any failure path".
Split into 6a6c, a success assertion, and 8a8e.
- Status delivery was gated by one generic causal test that passes while three of the four
events stay unwired. Now four exact JSON goldens with named triggers.
- `0b` was drawn in the DAG with no outgoing edge; it now explicitly precedes Phase 6.
**Codex's positions I narrowed:** it agreed the Sunshine sample is worth keeping as non-gating,
and corrected my wording — the topology appears when Sunshine *routes desktop audio through its
null-sink topology*, not merely whenever Sunshine is running, since I measured it running
without that topology. It also correctly narrowed O7's premise (an exclude-everything build does
already fail six rows' eligible controls, *if* asserted) while agreeing the exact-partition fix
is right.
**Verified by me before accepting:** the single `default_audio_monitor` call site
(`pipeline.rs:138`); the double binary resolution (`core/mod.rs:3375-3378` vs `:3409-3419`); the
no-session guard on `StartScreenShare` (`:3397-3404`); the sole core `start_playback` (`:1900`)
and `PeerJoined`'s scope (`:2396-2409`); and the current default-sink/Sunshine graph state.
## 11. Corrections owed to the design doc — ✅ APPLIED in v3.5 (round 8, 2026-07-25)
Both are now in the design doc (§6.1.0 and §6.1.4 respectively), alongside round 8's own
finding (§6.7, the observation boundary). Kept here as the record of what was owed and why:
1. **v3.4 §6.1.0's "🔴 the hazard is LIVE on this machine right now" is time-dependent and has
already flipped.** Measured 2026-07-21 ~14:55 (details in plan §5.3). The reachability
argument is unaffected — the topology appears **when Sunshine routes desktop audio through
its null-sink topology**, which is narrower than "whenever Sunshine is running," since it was
measured running without it. Nothing should gate on its presence.
2. **v3.4 §6.1.4's nominated test case is unreachable, and so was my first replacement.**
Details in plan §7. The conclusion (the transition window exists only for newly-created
roots) stands; the example must become the notification sound or watched-share start.
## 12. Not in this plan
**Round 8 additions:** **port binding** (so `port.exclusive` is never observed — plan §4 "Phase
3 revision" item 6, with its revisit trigger), **per-node quarantine** (an unresolvable node
bind fails the whole readiness epoch closed instead of isolating that one node — v3.5 §6.7
decision 3), and the **serial-continuity signal** for the AEC validator's no-coalescing
contract (phase 4's owed F4 hardening).
Everything v3.4 §14 lists as out of v1 — port-granular taint, timed drain, hot-AEC-reload epoch
protocol, native PipeWire AEC, incremental dirty-set, seamless daemon-restart recovery — plus
v3.4 §10 items 2 and 3 (per-app debt; D6 says they do not block Option C), the v3.4 §5.1
grandchild leak, and the general "AEC binds to the default sink when no device is pinned" defect
(v3.4 §6.1.0, resolved for this user, own task).
@@ -0,0 +1,472 @@
# Phase 5 — dry-run audit gate: results
**Status: 🟢 GATE PASSED (run 2, 2026-07-26). All 13 §5.1 rows completed; the
eligible half of every row is non-empty. O5 re-measured on the fixed graph and
stays closed.** One new defect was found and fixed during the run (F13-1); three
findings are recorded as non-blocking, and three rows carry recorded
substitutions. Phase 6 is unblocked **by this file**, and F11-1 — the other gate —
was closed with this data on 2026-07-26 (see "What still blocks phase 6").
- **Run date:** 2026-07-26 (run 1: 2026-07-25, gate FAILED — see history below)
- **Host:** `cazen` — PipeWire 1.6.8, WirePlumber 0.5.15, CachyOS
- **Audit build:** pixelpass `main` @ `91c4ded`, release profile
- **peerspeak build:** `main` @ `b68fca6` (phase 1 merged)
- **Ambient load:** Firefox playing audio throughout (a live, uncontrived
candidate); Sunshine running (pid 3838); Arctis 1 Wireless as active sink
- **Graph size:** 14 Nodes, 4 Devices, 57 Ports, 4 Links, 24 Clients
---
## What changed since run 1
Run 1 failed on two defects, both fixed before this run:
- **F1** (fatal): the registry `global` event delivers only a filtered subset of
node properties, so eight properties the engine depends on were permanently
absent. Fixed by design round 8 / **phase 3r** — bind every Node and Device
and read properties from `info`.
- **F2**: a machine-wide over-exclusion cascade downstream of F1.
Both are gone: the baseline run (no fixture at all) reports **1 candidate,
eligible, empty taint set**.
### 🔴 F13-1 — FOUND AND FIXED DURING THIS RUN
**Row 1 failed on its first attempt, and the cause was a third defect of exactly
the F2 class from a new source: pipewire-pulse's PID was unresolvable on this
host, permanently.**
`pulse_pid::candidate` returned the single `pipewire.sec.pid` shared by two or
more Clients, on the stated reasoning that "native PipeWire clients carry their
own distinct PID; only the Pulse shim repeats one value". Measured: **WirePlumber
repeats one too.** It holds two Clients — `WirePlumber` and
`WirePlumber [export]` — both `sec_pid` 1747. Two values repeated (1747 and
pipewire-pulse's 2528), the rule called that ambiguous, and returned `None`.
With the daemon PID unknown, `owner::keys_of`'s documented fail-closed asymmetry
takes over: key 4's suppression never fires, every Pulse-emulated node fuses into
one owner, and the cascade follows. Row 1's observed failure:
```
ELIGIBLE (1): r1_plain_app
EXCLUDED: Firefox tainted-owner-bridge key=application.process.id
r1_c_play tainted-owner-bridge <- the CLEAN control half
TAINT: ... + both sound cards, all three sunshine sinks, sunshine itself
```
The rule was wrong in **both** directions, so the prefilter was removed rather
than patched:
- **False ambiguity** — any second process holding two Clients defeats it.
WirePlumber always does, so this was permanent, not a corner case.
- **False absence** — a session where pipewire-pulse holds exactly one Client
(one Pulse app running) repeats nothing, so the candidate is missed and the
same cascade follows.
`comm` was always the authoritative check; repetition was a heuristic standing in
front of it, and it was a guess about other processes' Client counts. Fixed in
pixelpass `91c4ded`: `candidates()` lists every distinct `sec_pid`, `resolve()`
picks the unique one whose `/proc/<pid>/comm` is exactly `pipewire-pulse`, and
several matches still fail closed (a single `Option<u32>` cannot suppress two
daemons — recorded, not approximated). The adapter probes only PIDs *entering*
the candidate set, and `retain_probed_comms` bounds the map to live PIDs so a PID
that leaves and returns is re-probed instead of answered from a stale `comm`.
**This is the §5.1 exact-partition requirement earning its keep for the second
time.** The verdict was fail-closed and silent; only the asserted *eligible* half
exposed it. An exclusion-only checklist would have passed this build too.
---
## §5.1 — the matrix
Every row ran with `PIXELPASS_AUDIO_AUDIT_AEC=off` except row 12. Every row ran
in its **own** audit process, so nothing carries over (sticky taint is
per-process state).
⚠️ **Methodology change from run 1, and it is load-bearing.** Run 1 built each
fixture *before* starting the audit. On this host the entire graph then arrives
as one enumeration burst (~122 events in 12 ms), so every node is first tainted
while `graph_ready` is still false, that partial-graph taint is recorded into
sticky state, and on the single ready record the sticky pass raises
`TaintedOwnerBridge { key: None }` before the evidence pass can name a key —
`raise` will not replace a same-rank reason. Verdicts were still correct but rows
could not assert their key. This run starts the audit first, waits for readiness,
then builds the fixture, so taint is derived from real topology *changes* against
a ready graph — which is also the dynamic path §6.3 cares about. Keys are read at
**derivation** (first non-sticky appearance), not from the final record.
| # | scenario | status |
| --- | --- | --- |
| 1 | null-sink + loopback forwarder, owner bridge | ✅ **pass** (after F13-1 fixed) |
| 1b | Sunshine's topology (opportunistic, non-gating) | 🟡 observed, nothing to exclude — see below |
| 2 | gst split clients, tainted input | ✅ **pass**, key 4 named at derivation |
| 3 | two Pulse modules, one tainted | ✅ **pass** |
| 4 | peerspeak native call playback | ✅ **pass** — real tagging site |
| 5 | peerspeak-spawned mpv | ✅ **pass** — real tagging site, hand-launched mpv eligible |
| 6 | peerspeak notification sound | ✅ **pass** — real tagging site |
| 7 | second host's capture sink + forwarder | ✅ **pass**, eligible half non-empty |
| 8 | EasyEffects | 🟡 **pass with substitution** — echo-cancel stood in |
| 9 | Firefox three cases | ✅ **pass** (cases 23 via gst; see substitution) |
| 10 | sticky taint across teardown | ✅ **pass**, all four phases incl. retirement |
| 11 | recycled serial / index / link-group | ✅ **pass**, and provably non-vacuous |
| 12 | AEC loaded → unloaded → Revoked | ✅ **pass** |
| 13 | `Audio/Duplex` device | 🟡 **pass with synthetic node** — over-taint confirmed |
### Row 1 — owner bridge, key named
```
ELIGIBLE (3): Firefox · r1_c_play · r1_plain_app
EXCLUDED (2): peerspeak_owned_call_4242 peerspeak-owned
r1_t_play tainted-owner-bridge key=node.link-group
TAINT (5): the tagged producer, r1_t_src, r1_t_cap, r1_t_play, r1_t_dest
```
The clean half is an **identically shaped** forwarder — same module type, same
monitor-read, same re-emit — differing only in whether anything tainted feeds it.
`r1_c_play` eligible is the assertion an exclude-everything build cannot satisfy.
The key is `node.link-group`, a strong key, not a link walk.
### Row 2 — GStreamer split clients, key 4
Measured props confirm the shape is the real refutation: `r2_gst_tainted_src`
(client 188) and `r2_gst_tainted_sink` (client 191) are **different Clients** of
**one process**, pid 235628, with no `link-group` and no `pulse.module.id`. So
`application.process.id` is the only key that can relate them.
Derivation record (seq 209): `r2_gst_tainted_sink``tainted-owner-bridge`,
**`owner_key=application.process.id`**. `r2_gst_clean_sink`, reading an untainted
monitor in a second process, is eligible.
### Rows 46 — peerspeak's own paths, through the real call sites
Driven by peerspeak's phase-1 live gate tests (`--ignored`), i.e. the real
tagging sites, not a hand-rolled env: "emission alone proves only that peerspeak
talks, not that pixelpass listens" (impl plan §3).
| node | verdict |
| --- | --- |
| `peerspeak_owned_call_238172` | EXCLUDED `peerspeak-owned` |
| `peerspeak_owned_mpv_238196` | EXCLUDED `peerspeak-owned` |
| `peerspeak_owned_notify_238231` | EXCLUDED `peerspeak-owned` |
| `peerspeak_owned_clip_238249` | EXCLUDED `peerspeak-owned` (bonus — chat clips) |
| `mpv` (launched by hand, untagged) | **ELIGIBLE** |
This is the cross-repo contract closed end to end on live nodes.
### Row 9 — the over-exclusion promise
```
ELIGIBLE: Firefox (music only) · r9_mic_out (captures an untainted real device)
EXCLUDED: r9_mon_out tainted-owner-bridge key=application.process.id
```
`r9_mic_out` is the row that defends §6.1.1: an app that captures a real
`session_device` source and also plays audio stays shareable. The device source
itself never entered the taint set.
### Row 10 — the full sticky lifecycle
| phase | topology | verdict |
| --- | --- | --- |
| A | tainted producer + forwarder | `r10_play_out` EXCLUDED, key `node.link-group` |
| B | **tagged producer killed**, forwarder lives | **still EXCLUDED** (sticky) — current topology alone no longer justifies it |
| C | forwarder owner replaced, tainted sink kept | fresh forwarder EXCLUDED — correct: a sink that received call audio is still a hazard while it lives |
| D | **every** tainted object torn down, then restart | taint set **empty** at 16.3 s; `r10_new_out` **ELIGIBLE** at 20.3 s |
Phase B proves stickiness works; phase D proves it is not permanent. Phase C is
worth keeping in mind when reading any future report: partial teardown legitimately
does *not* retire taint, and that is easy to mistake for over-exclusion.
### Row 11 — recycled identifiers, provably non-vacuous
| generation | `node.link-group` | global id (`r11_src`) | `object.serial` (`r11_play`) | pulse module |
| --- | --- | --- | --- | --- |
| 1 (tainted) | `loopback-2528-14` | 168 | 4702 | 536870919 |
| 2 (after teardown) | **`loopback-2528-14`** | **168** | 4746 | 536870920 |
The `node.link-group` came back **byte-identical** — and it is the very key that
carried the taint in generation 1 — and the global id was reused. Generation 2's
`r11_play` is **ELIGIBLE** with an empty taint set. `object.serial` correctly did
not recycle, which is why the model keys everything by it.
### Row 12 — AEC lifecycle
| stage | `aec_state` | `fan_out_permitted` | candidates |
| --- | --- | --- | --- |
| module live, configured | `validated` | `true` | Firefox + `r12_plain_app` ELIGIBLE; `echo-cancel-playback` EXCLUDED `aec-identity` |
| module unloaded | `revoked` | `false` (`gate_reason=aec-revoked`) | every candidate EXCLUDED `aec-revoked` |
All **four** link-group siblings (`sink`, `source`, `capture`, `playback`) carry
`aec-identity`; only `echo-cancel-playback` is a candidate, so it is the only one
in the excluded partition. Ordinary apps staying eligible *while validated* is
what makes "the gate is open" observable rather than inferred.
### Row 13 — `Audio/Duplex` over-taint (known accepted)
No real duplex device exists on this host, so one was synthesised by overriding
`media.class=Audio/Duplex` on a null sink. Its playback side was tainted and its
capture-side consumer was dragged down with it (`r13_dup_play` EXCLUDED), with
the eligible half intact. **Fixture limit, stated plainly:** on a null sink the
capture side *is* the monitor, so this cannot separate the duplex smear from the
ordinary sink→monitor edge. The accepted over-taint is confirmed as *behaviour*;
a real duplex device is still the only way to isolate the mechanism.
### Row 1b — Sunshine (opportunistic, non-gating)
Sunshine ran throughout. Its three null sinks stayed SUSPENDED and it read the
**hardware** monitor instead, exactly as §5.3 warned. It appears consistently and
correctly as `sunshine` / `tainted-upstream` whenever the monitor it reads is
tainted (rows 8, 12, o5). It has **no re-emitting output leg** — it sends over
the network — so it is never a candidate and there is nothing to exclude. Recorded
as observed; the "if a re-emitting leg exists" clause did not apply. A real
third-party forwarder sample remains owed.
---
## §5.2 — O5 re-measured
The run-1 numbers do not carry over: they were measured on the graph F1 degraded,
and phase 3r adds a bind plus an `info` round-trip **per node**, which is new I/O
that run never exercised.
Per-run, across all 13 rows (`recompute` in µs):
| run | events | ev/s | max | mean | emit max | busy fraction | ready@ms |
| --- | --- | --- | --- | --- | --- | --- | --- |
| baseline | 123 | 21.4 | 20 | 3 | 6 | 0.0001 | 1 |
| o5 (churn) | 407 | 44.0 | 32 | 10 | 9 | 0.0006 | 1 |
| row01 | 219 | 41.7 | 53 | 10 | 9 | 0.0006 | 1 |
| row02 | 241 | 45.9 | **67** | 11 | 10 | 0.0006 | 1 |
| row03 | 206 | 48.5 | 54 | 8 | 8 | 0.0005 | 1 |
| row0456 | 185 | 20.0 | 38 | 7 | 10 | 0.0002 | 2 |
| row07 | 184 | 43.3 | 40 | 7 | 9 | 0.0004 | 1 |
| row08 | 172 | 32.6 | 44 | 6 | 7 | 0.0003 | 2 |
| row09 | 224 | 30.9 | 52 | 9 | 9 | 0.0004 | 1 |
| row10 | 332 | 14.3 | 41 | 11 | 15 | 0.0002 | 1 |
| row11 | 298 | 24.1 | 41 | 9 | 11 | 0.0003 | 1 |
| row12 | 188 | 25.9 | 39 | 7 | 8 | 0.0003 | 1 |
| row13 | 193 | 36.8 | 43 | 8 | 7 | 0.0004 | 1 |
The dedicated churn run (five load/unload cycles of null-sink + loopback, the
same shape as run 1's measurement):
```json
{"kind":"metrics","graph_events":407,"tick_events":37,"emitted_records":407,
"span_us":9249639,"graph_events_per_sec":44.0,
"recompute_max_us":32,"recompute_mean_us":10,
"recompute_p50":"<50us","recompute_p90":"<50us","recompute_p99":"<50us",
"recompute_distribution":[["<50us",444]],
"emit_max_us":9,"emit_mean_us":1,
"busy_us":5240,"busy_fraction":0.0006,
"queued_events":292,"queue_threshold_us":100}
```
**O5 stays closed on the real graph.** Worst recompute across every run is
**67 µs**; every single recompute in the churn run finished under 50 µs, against
a 44 Hz event rate under churn heavier than a desktop produces at rest. The
observer thread spent **0.06 %** of wall time working. Node binding roughly
doubled the per-event cost (run 1: 15 µs max / 4 µs mean; now 32 µs / 10 µs on
the same churn shape) and that is the honest cost of the F1 fix — it buys three
orders of magnitude of remaining headroom, not one.
**Readiness with node binds: 12 ms**, with ~122 enumeration events and 18 binds
(14 Nodes + 4 Devices), against the 2000 ms budget. `queued_events` is high
(292) for the same benign reason as run 1: PipeWire delivers enumeration and
teardown in bursts, and a 32 µs recompute drains a burst faster than it forms.
`busy_fraction` is the number to trust.
⚠️ **The readiness budget still has no calibration argument.** 12 ms against
2000 ms is three orders of magnitude of slack on *this* host with 18 binds; it is
not an argument about a host with a large USB interface, many virtual devices, or
a cold cache. Carried forward as open, unchanged.
---
## Findings recorded, not blocking
### R2-1 — the audit's `sticky` flag is nearly always true, so it says little
As emitted, `sticky` means "this node is in the remembered set", which
`seed_sticky` populates for any node whose current reason the sticky pass agrees
with — i.e. essentially every currently-tainted node. It does **not** mean
"excluded *only* because remembered", which is what its doc comment implies and
what a reader diagnosing "why is this still excluded?" wants.
The information exists: round 9 already computes a second, **evidence-only** pass
(that is the whole provenance mechanism). Emitting "excluded by memory alone"
would make row 10 phase B assertable from a single record instead of from a
sequence. Not fixed here — it is a reporting change to a merged phase in the
middle of a gate run. Row 10 was asserted behaviourally instead, which is
stronger anyway.
### R2-2 — a bridge key is lost when a leg reappears under a new serial
Row 2 named `application.process.id` at derivation (seq 209), then gst re-created
that node; the sticky owner re-seeded the new serial through `reason_for`, whose
documented fallback is `TaintedOwnerBridge { key: None }`, and `raise` will not
replace a same-rank reason with a better-informed one. The verdict is unaffected;
only the diagnosis degrades. The fallback is honest when the owner has no live
tainted receiver, and stale when it does — which is the case worth improving.
### R2-3 — `owner_key` had to be added to the record to run row 1 at all
Row 1 asserts "reason = owner bridge, **naming the key**", and the record could
not express it: `Reason::code` collapses `TaintedOwnerBridge { key }` to one
string. `OwnerKey::code` already documented itself as ending up in the phase 5
audit output; it was simply never wired to it. Added in pixelpass `d462754`
(read-only, diagnostic-only, mutation-verified test). Worth noting as a gate-spec
lesson: the row could not have been asserted from any previous build's output.
---
## Substitutions, stated so they are not mistaken for passes
| row | asked for | used instead | why |
| --- | --- | --- | --- |
| 8 | EasyEffects | `module-echo-cancel` with `AEC=off` | EasyEffects makes itself the default sink on start and the user had live audio playing. `module-filter-chain` cannot stand in either — it is a PipeWire module, so `pactl load-module` answers "No such entity" (measured). The stand-in produces the same shape (four nodes, one `node.link-group`) and exercises `foreign-echo-cancel` (decision D3), a reason code no other row reaches. |
| 9 | Firefox's mic + monitor capture | `gst-launch` pipelines | Firefox's mic and monitor-capture paths need interactive GUI permission grants. Firefox is present live as case 1 in every row. Case 2 captures the motherboard's **analog input**, not the headset mic the user is wearing — identical to the engine (both `session_device` sources), and nothing of the user is recorded. |
| 13 | a real `Audio/Duplex` device | synthetic `media.class` override | None on this host. See row 13 above for what the fixture cannot show. |
---
## What still blocks phase 6
This file passing removes **one** of the two gates. F11-1, the other, is now
closed. Still outstanding:
1. **Hardware playback-to-capture paths ("Stereo Mix")** defeat `session_device`
and are a real echo path — needs ALSA control inspection; user design call owed.
2. **Phases 0b / 0c / 0d** are untouched and all precede phase 6.
3. **The readiness budget calibration argument** (above).
4. **Owed samples:** a real third-party forwarder (row 1b), EasyEffects (row 8),
a real `Audio/Duplex` device (row 13).
### ✅ F11-1 — closed 2026-07-26, with this matrix's data
The rule now implemented (pixelpass `c78eb2d`, §6.1.2's round-13 box): **key 4 bounds an
owner only when the node's Client resolves** — an unambiguous Client yielding
`Some(pipewire.sec.pid)`, read *before* pipewire-pulse suppression — so a node can no
longer bound itself, and escape `propagate_unresolved_owner`'s sweep, with an
`application.process.id` it invented. Bridging still uses the full union.
Codex's round-12 sharpening was the decisive part: "resolved" must mean a `sec_pid`, not
"a unique Client object exists", and the **unique-but-pid-less** row is the only one that
tells the two apart. All five Client cases are unit tests (absent · ambiguous ·
unique-but-pid-less · resolved-native · resolved-to-pipewire-pulse), plus the recorded
three-step leak path end to end. Mutation-verified: dropping the provenance test fails
four of the six rows and leaves the two no-over-exclusion rows green.
**The cost question the deferral was waiting on, measured on this host:** the before- and
after-binaries audited the *same* live graph simultaneously (both are read-only observers)
— tagged producer into the default sink, `parec` on its monitor as a live tainted reader
so the sweep was genuinely armed, Firefox + `aplay` + `pacat` as bystanders. **181 records
each, the same 14 distinct decision states, none exclusive to either side, no
`unresolved-owner` on either, eligible half non-empty throughout.** O5 unmoved (identical
p50 15 µs and busy fraction 0.0012). Every real app here is native or Pulse-emulated and
**both resolve**; sweeping all 18 live nodes, the only unresolved-Client ones were
`Dummy-Driver` and `Freewheel-Driver`, which carry no pid key to lose.
---
## Reproducing this run
Scripts live in the session scratchpad (not committed — they hard-code paths):
one per row, plus `lib.sh`, `summarize.py` and `keys.py`. The shape of every row:
```sh
audit_start out.jsonl off # start FIRST, wait for graph_ready
... build fixture ... # taint arrives as topology CHANGES
audit_stop # SIGTERM: flushes the O5 summary
python3 summarize.py out.jsonl # final partition + derivations + metrics
```
```
env PIXELPASS_AUDIO_AUDIT_FILE=/path/out.jsonl PIXELPASS_AUDIO_AUDIT_AEC=off \
./target/release/pixelpass --audit-audio
```
Rig notes that cost time:
- A tagged producer: `env PIPEWIRE_ALSA='{ "peerspeak.owned": "1", "node.name":
"peerspeak_owned_call_4242", "target.object": "<sink>" }' aplay -c 2 -r 48000
-f S16_LE -t raw -d 30 /dev/zero`. Both carriers land, and `target.object`
routes it.
- ⚠️ `pactl load-module module-echo-cancel --help` **loads the module** with
`--help` as its argument instead of printing help. It was loaded accidentally
during this session and unloaded again; check `pactl list short modules` after
any such probe.
- ⚠️ `pkill -f <pattern>` matches the harness's own shell command line and kills
the script. Use `pkill -x` or an exact pid.
- ⚠️ Under `set -e`, `kill` on an already-exited pid aborts the row before its
modules are unloaded; and `timeout` exiting 124 is *success* for the audit.
---
## History — run 1 (2026-07-25): GATE FAILED
Kept because the reasoning is still the record of why the observation boundary
was redesigned.
### F1 🔴 FATAL — the registry `global` event delivers only a filtered subset of node properties
The phase-3 adapter read eight node properties the registry never announces.
Parsed off `obj.props` in the registry `global` callback, they were silently
absent, so every one was permanently `None`/`false`.
The complete set the registry announces for a `Node` on this host:
```
application.name client.api client.id device.id factory.id media.class
node.description node.name node.nick object.path object.serial
priority.driver priority.session
```
| property | announced? | what died without it |
| --- | --- | --- |
| `object.serial`, `node.name`, `media.class`, `client.id`, `device.id` | ✅ | — |
| **`peerspeak.owned`** | ❌ | **the primary taint root (all of phase 1)** |
| **`pulse.module.id`** | ❌ | **AEC identity exclusion + phase 4 validation** |
| **`node.link-group`** | ❌ | the link-group owner key |
| **`application.process.id`** | ❌ | the process owner key |
| **`node.passthrough`** | ❌ | the passthrough local exclusion |
| **`device.api`**, **`factory.name`**, **`alsa.driver_name`** | ❌ | `session_device` classification |
Ports lost `port.exclusive`; Links and Clients were fine — notably
`pipewire.sec.pid` **is** announced, so pulse-PID derivation was reachable.
Demonstrated end to end: a null sink carrying `peerspeak.owned=true` whose
monitor a `module-loopback` re-emitted was reported **eligible** with an **empty
taint set**. In phase 6 that is an echo.
The fix became design round 8 (v3.5 §6.7) and phase 3r: bind each Node and read
props off its `info`, exactly how `pw-dump` obtains them. `factory.id` is not a
shortcut (`factory.id=19` resolves to `factory.name = "adapter"`), and
`device.api` is on the *Device* global.
### F2 🟠 Machine-wide over-exclusion cascade, downstream of F1
With F1 in force, `pixelpass_capture_*` (matched on `node.name`, which *is*
announced) was the only surviving taint root. Row 7 then excluded every
`Stream/Output/Audio` on the machine: with no strong owner keys, every tainted
capture stream was an **unbounded tainted reader**, tripping phase 2's
fail-closed backstop, while WirePlumber's shared `client.id = 42` fused the
device layer into one owner.
Net live behaviour: exclude everything, always, as soon as pixelpass's own
capture sink existed. Fail-closed, so silence rather than echo — but entirely
non-functional, and non-functional in a way that would have looked like "working
safely" to any test that asserted only exclusions.
### What run 1's machinery got right
None of this needed revisiting:
- Running the recompute **inline on the observer thread**, once per applied
registry event, upheld phase 4's no-coalescing contract and put the cost where
O5 could measure it.
- The **complete-partition record** is what caught F2 — and, in run 2, F13-1.
- **Reason codes survived the trip** and were immediately diagnostic.
- The **`peerspeak.owned` / `pulse.module.id` fixtures were right**: the engine
does the correct thing when handed correct properties. Both failures were at
the observation boundary, which is where phase 5 was designed to look.
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,7 +1,7 @@
# Maintainer: mollusk <jitty+lc1iz0dc@protonmail.com>
pkgname=peerspeak-git
_pkgname=peerspeak
pkgver=0.6.1.r315.ga78860d
pkgver=0.6.2.r319.g8014edf
pkgrel=1
pkgdesc="Decentralized peer-to-peer voice chat (Rust/iroh/PipeWire/Opus/iced)"
arch=('x86_64')
+1 -1
View File
@@ -12,7 +12,7 @@
; (x86_64-pc-windows-gnu, statically linked -- no extra DLLs needed).
#define MyAppName "PeerSpeak"
#define MyAppVersion "0.6.2"
#define MyAppVersion "0.6.6"
#define MyAppPublisher "mollusk"
#define MyAppExeName "peerspeak.exe"
+1986 -271
View File
File diff suppressed because it is too large Load Diff
+129
View File
@@ -0,0 +1,129 @@
//! Sender-side chat send status and pacing (chat-hardening Phase 5).
//!
//! Every RECEIVER admits our chat through a per-author token bucket
//! ([`CHAT_AUTHOR_BURST`] then 1/s) and silently drops what exceeds it, with no
//! acknowledgement wire. The only way the sender can be honest about fast
//! bursts is to never exceed that budget in the first place: sends past the
//! burst are queued locally (shown as "queued…") and trickled out at the
//! receivers' sustained rate. The pacer deliberately reuses the receiver
//! gate's own [`TokenBucket`] and constants so the two sides of the policy
//! cannot drift apart.
//!
//! Everything here is pure — `now_ms` is passed in, never read from a clock —
//! so every boundary is unit-testable.
use std::collections::VecDeque;
use crate::network::gossip::{CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, TokenBucket};
/// Send lifecycle of one locally authored chat message. Success is
/// [`SendStatus::Broadcast`] — "our signed frame was handed to the gossip
/// swarm" — deliberately NOT "delivered": PeerSpeak has no peer
/// acknowledgements, so the honest success presentation is no label at all.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum SendStatus {
/// Waiting in the local outbound queue for a pacer token.
Queued,
/// Handed to the core; the broadcast result has not come back yet.
Pending,
/// The signed broadcast reached the gossip swarm.
Broadcast,
/// The send failed; carries a short reason. The entry offers a Retry.
Failed(String),
}
/// Local-only send bookkeeping attached to our own chat entries. The id never
/// goes on the wire; it ties a `ChatSendResult` back to the matching echo.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct LocalSend {
pub id: u64,
pub status: SendStatus,
}
/// Sender-side pacer mirroring the receiver's per-author admission budget.
#[derive(Debug, Clone, Copy)]
pub struct SendPacer {
bucket: TokenBucket,
}
impl SendPacer {
pub fn new(now_ms: u64) -> Self {
Self {
bucket: TokenBucket::full(CHAT_AUTHOR_BURST, now_ms),
}
}
/// Take one send token if the mirrored per-author budget allows it now.
pub fn try_send(&mut self, now_ms: u64) -> bool {
self.bucket
.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, now_ms)
}
}
/// Pop the queued ids that may be dispatched now: strict front-of-queue order,
/// one pacer token each, stopping at the first refusal so a message can never
/// overtake an earlier one.
pub fn release_ready(queue: &mut VecDeque<u64>, pacer: &mut SendPacer, now_ms: u64) -> Vec<u64> {
let mut ready = Vec::new();
while !queue.is_empty() && pacer.try_send(now_ms) {
// The unwrap is safe: the loop condition just checked non-empty.
ready.push(queue.pop_front().unwrap());
}
ready
}
#[cfg(test)]
mod tests {
use super::*;
const T0: u64 = 1_000_000;
#[test]
fn pacer_allows_the_full_burst_then_refuses() {
let mut pacer = SendPacer::new(T0);
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert!(pacer.try_send(T0));
}
assert!(!pacer.try_send(T0));
}
#[test]
fn pacer_refills_at_one_per_second() {
let mut pacer = SendPacer::new(T0);
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert!(pacer.try_send(T0));
}
// 999ms is just under one token; 1000ms grants exactly one.
assert!(!pacer.try_send(T0 + 999));
assert!(pacer.try_send(T0 + 1000));
assert!(!pacer.try_send(T0 + 1000));
}
#[test]
fn release_ready_preserves_order_and_stops_at_refusal() {
let mut pacer = SendPacer::new(T0);
// Drain the burst so only refill tokens remain.
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert!(pacer.try_send(T0));
}
let mut queue: VecDeque<u64> = [10, 11, 12].into_iter().collect();
// 2 seconds of refill = 2 tokens: exactly the first two, in order.
let ready = release_ready(&mut queue, &mut pacer, T0 + 2000);
assert_eq!(ready, vec![10, 11]);
assert_eq!(queue, VecDeque::from([12]));
// No tokens left at the same instant.
assert!(release_ready(&mut queue, &mut pacer, T0 + 2000).is_empty());
assert_eq!(queue, VecDeque::from([12]));
}
#[test]
fn release_ready_empty_queue_consumes_no_tokens() {
let mut pacer = SendPacer::new(T0);
let mut queue = VecDeque::new();
assert!(release_ready(&mut queue, &mut pacer, T0).is_empty());
// The full burst must still be available.
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert!(pacer.try_send(T0));
}
}
}
+50
View File
@@ -328,4 +328,54 @@ mod tests {
assert_eq!(seek_target(-1.0, total), Duration::ZERO);
assert_eq!(seek_target(2.0, total), total);
}
/// **The fourth playback path's exit gate (round 10, R10-2).** Drives a
/// real [`ClipPlayer`] — the same object the app uses for chat clips, peer
/// music and the local playlist — and asserts the node it puts on the
/// graph carries both ownership carriers.
///
/// This path was untagged through all of phase 1, which is a real echo:
/// B broadcasts music, A tunes in, A shares their desktop, B hears their
/// own track. It was missed because phase 1 worked from the impl plan's
/// list of three playback sites and that list was incomplete — so this
/// gate drives the *player*, not the tagging helper.
///
/// ⚠️ **Run alone**: it sets a process-wide environment variable, which is
/// only sound single-threaded. In production `main` does this before
/// anything is spawned; a test binary has no such guarantee, hence
/// `--test-threads=1`.
///
/// `cargo test --lib -- --ignored --test-threads=1 clip_player_node`
#[test]
#[ignore = "live: requires a running PipeWire daemon and pw-dump; run with --test-threads=1"]
fn clip_player_node_carries_both_ownership_carriers() {
use crate::audio::ownership::{self, live_test};
// SAFETY: `--test-threads=1` is documented above and in the ignore
// reason; this is the same call `main` makes, exercised for real
// rather than reimplemented, so the gate cannot pass against a
// formatter that production never uses.
unsafe { ownership::tag_this_process_alsa_audio() };
let (player, _status) = ClipPlayer::new(1.0);
// Six seconds of silence: long enough for the poll, inaudible.
player.play([0u8; 32], live_test::silent_wav(6));
let prefix = live_test::expected_prefix(ownership::CLIP_ROLE);
let found = live_test::poll_for_owned_node(&prefix, Duration::from_secs(5));
player.stop();
let (name, owned) = found.unwrap_or_else(|| {
panic!("no live clip-player node named {prefix:?} appeared within 5s")
});
assert!(
name.starts_with(ownership::OWNED_NODE_NAME_PREFIX),
"{name}"
);
assert_eq!(
owned.as_deref(),
Some(ownership::OWNED_PROP_VALUE),
"carrier 1 must be on the live node, not just carrier 2"
);
}
}
+4
View File
@@ -65,6 +65,10 @@ pub mod eq;
pub mod gate;
pub mod limiter;
pub mod multitrack;
// The cross-repo ownership tag (plan §5.1). Platform-neutral on purpose: the
// carriers only matter on PipeWire, but the literals are a wire contract and
// their test must run on every platform so a rename can't pass CI elsewhere.
pub mod ownership;
pub mod pan;
// Linear resamplers used by the Windows/cpal backend (W4). Platform-neutral and
// pure, so it builds (and its tests run) everywhere even though only the cpal
File diff suppressed because it is too large Load Diff
+75
View File
@@ -1,3 +1,4 @@
use crate::audio::ownership;
use crate::audio::{AudioBackend, AudioError};
use pipewire as pw;
use pw::{properties::properties, spa};
@@ -371,6 +372,11 @@ fn run_playback(
mainloop_clone.quit();
});
// Ownership tag, both carriers (`crate::audio::ownership`, plan §5.1).
// This is the node that carries the far end's voice, so it is the single
// most important thing for pixelpass to refuse to fan out: sharing it
// would send the call back to the person already speaking on it.
let owned_node_name = ownership::owned_node_name(ownership::NATIVE_PLAYBACK_ROLE);
let mut props = properties! {
*pw::keys::MEDIA_TYPE => "Audio",
*pw::keys::MEDIA_CATEGORY => "Playback",
@@ -379,6 +385,19 @@ fn run_playback(
// buffer — the real fix is the explicit Buffers param below — but it
// expresses the intended quantum for any node that honours it.
*pw::keys::NODE_LATENCY => "1024/48000",
ownership::OWNED_PROP_KEY => ownership::OWNED_PROP_VALUE,
// Set explicitly rather than relying on the stream name passed to
// `StreamBox::new` below: props win over that name, and this one has
// to be exact.
*pw::keys::NODE_NAME => owned_node_name.as_str(),
// Measured: this stream sets neither `application.name` nor a
// description, so a mixer falls back to `node.name` — which the line
// above just turned into an internal identifier. The plan's rule is
// that the ownership prefix must not reach `node.description`; a
// human label there is what keeps that rule's *intent* (mixers stay
// readable) true for our own stream, exactly as mpv's own
// description does for the spawned players.
*pw::keys::NODE_DESCRIPTION => "PeerSpeak",
};
if let Some(target) = target_node {
props.insert("node.target", target);
@@ -637,6 +656,62 @@ mod tests {
use std::time::Duration;
use std::{sync::mpsc, thread};
/// Phase-1 exit gate, native-playback half (impl plan §3): the stream
/// that carries the far end's voice appears on the graph with **both**
/// ownership carriers, and still with the `Communication` media role.
///
/// The third and most important of the three tagged paths — this is the
/// node whose audio, if fanned out, would send the call back to whoever
/// is speaking on it.
///
/// Feeds silence, so the gate is inaudible. Live: needs PipeWire and
/// `pw-dump`. `cargo test --lib -- --ignored native_playback`
#[test]
#[ignore = "live: requires a running PipeWire daemon and pw-dump"]
fn native_playback_node_carries_both_ownership_carriers() {
use crate::audio::ownership::{self, live_test};
use crate::audio::{AudioBackend, PLAYBACK_TARGET_SAMPLES};
let backend = super::PipeWireBackend::new();
let (tx, rx) = mpsc::channel::<Vec<i16>>();
let ring_fill = Arc::new(AtomicUsize::new(0));
backend
.start_playback(rx, None, ring_fill.clone())
.expect("playback starts");
// Keep the ring fed so the node stays live for the whole poll; the
// stream is created on connect, but a starved one is not a fair test
// of what a real call looks like on the graph.
let feeder = thread::spawn(move || {
let silence = vec![0i16; 960 * 2];
for _ in 0..300 {
if ring_fill.load(Ordering::Relaxed) < PLAYBACK_TARGET_SAMPLES
&& tx.send(silence.clone()).is_err()
{
return;
}
thread::sleep(Duration::from_millis(20));
}
});
let prefix = live_test::expected_prefix(ownership::NATIVE_PLAYBACK_ROLE);
let found = live_test::poll_for_owned_node(&prefix, Duration::from_secs(5));
let _ = backend.stop();
let _ = feeder.join();
let (name, owned) =
found.unwrap_or_else(|| panic!("no live node named {prefix:?} appeared within 5s"));
assert!(
name.starts_with(ownership::OWNED_NODE_NAME_PREFIX),
"{name}"
);
assert_eq!(
owned.as_deref(),
Some(ownership::OWNED_PROP_VALUE),
"carrier 1 must be on the live node, not just carrier 2"
);
}
#[test]
fn requested_in_range_is_honored() {
// The graph's requested quantum is produced verbatim when it fits.
+181
View File
@@ -168,6 +168,139 @@ impl std::fmt::Display for NetworkMode {
}
}
/// Pixelpass host quality preset for screen shares. `Auto` leaves pixelpass free
/// to choose from its bandwidth pre-flight; fixed presets are passed as CLI flags.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, Default)]
#[serde(rename_all = "snake_case")]
pub enum ShareQuality {
#[default]
Auto,
Low,
Medium,
High,
Source,
}
impl ShareQuality {
pub const ALL: [ShareQuality; 5] = [
ShareQuality::Auto,
ShareQuality::Low,
ShareQuality::Medium,
ShareQuality::High,
ShareQuality::Source,
];
}
impl std::fmt::Display for ShareQuality {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
ShareQuality::Auto => "Auto",
ShareQuality::Low => "Low",
ShareQuality::Medium => "Medium",
ShareQuality::High => "High",
ShareQuality::Source => "Source",
})
}
}
/// Preferred local player for watching a peer's screen share.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, Default)]
#[serde(rename_all = "snake_case")]
pub enum SharePlayer {
#[default]
Mpv,
Vlc,
}
impl SharePlayer {
pub const ALL: [SharePlayer; 2] = [SharePlayer::Mpv, SharePlayer::Vlc];
}
impl std::fmt::Display for SharePlayer {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
SharePlayer::Mpv => "mpv",
SharePlayer::Vlc => "VLC",
})
}
}
/// Local player buffering posture for screen-share playback.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, Default)]
#[serde(rename_all = "snake_case")]
pub enum ShareBuffering {
#[default]
LowLatency,
Smooth,
}
impl ShareBuffering {
pub const ALL: [ShareBuffering; 2] = [ShareBuffering::LowLatency, ShareBuffering::Smooth];
}
impl std::fmt::Display for ShareBuffering {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
ShareBuffering::LowLatency => "Low latency",
ShareBuffering::Smooth => "Smooth",
})
}
}
fn default_screen_share_cache_mb() -> u32 {
2
}
/// Local-only screen-share preferences. Host fields become pixelpass host CLI
/// flags; viewer fields shape local mpv/VLC launch. None/empty/default values
/// deliberately let pixelpass/player defaults stand.
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)]
pub struct ScreenShareSettings {
#[serde(default)]
pub quality: ShareQuality,
#[serde(default)]
pub bitrate_mbps: Option<u32>,
#[serde(default)]
pub framerate: Option<u32>,
#[serde(default)]
pub max_height: Option<u32>,
#[serde(default)]
pub max_viewers: Option<u32>,
#[serde(default)]
pub force_software_encode: bool,
#[serde(default)]
pub extra_host_args: String,
#[serde(default)]
pub player: SharePlayer,
#[serde(default)]
pub hardware_decode: bool,
#[serde(default)]
pub buffering: ShareBuffering,
#[serde(default = "default_screen_share_cache_mb")]
pub cache_mb: u32,
#[serde(default)]
pub extra_mpv_args: String,
}
impl Default for ScreenShareSettings {
fn default() -> Self {
Self {
quality: ShareQuality::default(),
bitrate_mbps: None,
framerate: None,
max_height: None,
max_viewers: None,
force_software_encode: false,
extra_host_args: String::new(),
player: SharePlayer::default(),
hardware_decode: false,
buffering: ShareBuffering::default(),
cache_mb: default_screen_share_cache_mb(),
extra_mpv_args: String::new(),
}
}
}
fn default_true() -> bool {
true
}
@@ -345,6 +478,14 @@ pub struct AppConfig {
pub custom_sound_mic_toggle: Option<String>,
#[serde(default)]
pub custom_sound_reconnect_failed: Option<String>,
#[serde(default)]
pub custom_sound_chat_sent: Option<String>,
#[serde(default)]
pub custom_sound_chat_received: Option<String>,
#[serde(default)]
pub custom_sound_contact_online: Option<String>,
#[serde(default)]
pub custom_sound_contact_offline: Option<String>,
/// Per-sound enable flags (W6). The master `notifications_enabled` toggle
/// gates ALL chimes; these let the user silence individual events while the
/// master stays on. A chime plays only if the master AND its flag are true.
@@ -365,10 +506,21 @@ pub struct AppConfig {
pub sound_mic_toggle_enabled: bool,
#[serde(default = "default_true")]
pub sound_reconnect_failed_enabled: bool,
#[serde(default = "default_true")]
pub sound_chat_sent_enabled: bool,
#[serde(default = "default_true")]
pub sound_chat_received_enabled: bool,
#[serde(default = "default_true")]
pub sound_contact_online_enabled: bool,
#[serde(default = "default_true")]
pub sound_contact_offline_enabled: bool,
/// Optional override for the `pixelpass` binary location (screen share).
/// Empty / unset = look it up on `$PATH`. Hand-editable; no Settings UI yet.
#[serde(default)]
pub pixelpass_path: Option<String>,
/// Local-only host/player controls for screen sharing.
#[serde(default)]
pub screen_share: ScreenShareSettings,
/// Recently-joined rooms (W7), most-recent-first. Purely local UI state for a
/// one-click rejoin; never sent over the wire. De-duped by room topic and
/// capped (see `recents`). Defaulted empty so older configs upgrade cleanly.
@@ -465,6 +617,10 @@ impl Default for AppConfig {
custom_sound_self_leave: None,
custom_sound_mic_toggle: None,
custom_sound_reconnect_failed: None,
custom_sound_chat_sent: None,
custom_sound_chat_received: None,
custom_sound_contact_online: None,
custom_sound_contact_offline: None,
sound_self_join_enabled: true,
sound_peer_join_enabled: true,
sound_peer_leave_enabled: true,
@@ -473,7 +629,12 @@ impl Default for AppConfig {
sound_self_leave_enabled: true,
sound_mic_toggle_enabled: true,
sound_reconnect_failed_enabled: true,
sound_chat_sent_enabled: true,
sound_chat_received_enabled: true,
sound_contact_online_enabled: true,
sound_contact_offline_enabled: true,
pixelpass_path: None,
screen_share: ScreenShareSettings::default(),
recents: Vec::new(),
peer_eq: HashMap::new(),
peer_pan: HashMap::new(),
@@ -502,6 +663,10 @@ impl AppConfig {
Sound::SelfLeave => self.sound_self_leave_enabled,
Sound::MicToggle => self.sound_mic_toggle_enabled,
Sound::ReconnectFailed => self.sound_reconnect_failed_enabled,
Sound::ChatSent => self.sound_chat_sent_enabled,
Sound::ChatReceived => self.sound_chat_received_enabled,
Sound::ContactOnline => self.sound_contact_online_enabled,
Sound::ContactOffline => self.sound_contact_offline_enabled,
}
}
@@ -516,6 +681,10 @@ impl AppConfig {
Sound::SelfLeave => self.sound_self_leave_enabled = enabled,
Sound::MicToggle => self.sound_mic_toggle_enabled = enabled,
Sound::ReconnectFailed => self.sound_reconnect_failed_enabled = enabled,
Sound::ChatSent => self.sound_chat_sent_enabled = enabled,
Sound::ChatReceived => self.sound_chat_received_enabled = enabled,
Sound::ContactOnline => self.sound_contact_online_enabled = enabled,
Sound::ContactOffline => self.sound_contact_offline_enabled = enabled,
}
}
@@ -803,6 +972,18 @@ mod tests {
assert!(deserialized.custom_sound_self_leave.is_none());
assert!(deserialized.custom_sound_mic_toggle.is_none());
assert!(deserialized.custom_sound_reconnect_failed.is_none());
assert!(deserialized.custom_sound_chat_sent.is_none());
assert!(deserialized.custom_sound_chat_received.is_none());
assert!(deserialized.custom_sound_contact_online.is_none());
assert!(deserialized.custom_sound_contact_offline.is_none());
assert_eq!(deserialized.screen_share, ScreenShareSettings::default());
assert_eq!(deserialized.screen_share.quality, ShareQuality::Auto);
assert_eq!(deserialized.screen_share.player, SharePlayer::Mpv);
assert_eq!(
deserialized.screen_share.buffering,
ShareBuffering::LowLatency
);
assert_eq!(deserialized.screen_share.cache_mb, 2);
// Configs predating the per-sound flags (W6) enable every chime, so an
// upgrade is silent-change-free.
for sound in Sound::ALL {
+132
View File
@@ -0,0 +1,132 @@
//! Roster-bound chat identity (chat-hardening plan, Phase 2).
//!
//! The wire `GossipMessage::Chat` carries a sender-CLAIMED display name, which
//! any insider could set to another member's name. This map is the antidote:
//! the core event task records each authenticated member's latest sanitized
//! presence name here (from `PeerJoined`/`PeerUpdated`, the events that only
//! fire for a verified signed `Announce`), and chat renders under THAT name —
//! the embedded wire name is never displayed.
//!
//! Shared (`Arc<Mutex<…>>`) because eviction happens in two places: the event
//! task itself (graceful `PeerLeft`) and the detached reconnect-grace timer
//! (terminal eviction). A peer mid-reconnect-grace keeps its entry, so its
//! chat stays admitted until the grace actually expires.
use iroh::EndpointId;
use std::collections::HashMap;
use std::sync::{Arc, Mutex};
/// Bound on tracked names. Mirrors the gossip roster cap (`MAX_ACTIVE_PEERS`):
/// insertions only follow cap-gated roster admissions, so this is pure defense
/// in depth against that invariant breaking.
const CHAT_ROSTER_CAP: usize = 32;
/// The authoritative id → display-name map for the current room. Cheap to
/// clone; all clones share one map.
#[derive(Debug, Clone, Default)]
pub struct ChatRoster {
names: Arc<Mutex<HashMap<EndpointId, String>>>,
}
impl ChatRoster {
/// Record (or refresh) a member's display name. The name is re-sanitized
/// here (idempotent — gossip ingress already did) and an empty result falls
/// back to the short node id so a chat line is never label-less. A NEW id
/// is refused past the cap; updates to a present id always land.
pub fn upsert(&self, id: EndpointId, name: &str) {
let clean = crate::sanitize::sanitize_name(name);
let label = if clean.is_empty() {
crate::short_id(&id.to_string())
} else {
clean
};
let mut names = self.names.lock().unwrap();
if names.contains_key(&id) || names.len() < CHAT_ROSTER_CAP {
names.insert(id, label);
}
}
/// Drop a member on graceful leave or terminal (grace-expired) eviction.
pub fn remove(&self, id: &EndpointId) {
self.names.lock().unwrap().remove(id);
}
/// The roster-bound name for an id, or `None` if the author is not a
/// current member — the caller must then drop the chat entirely.
pub fn name_of(&self, id: &EndpointId) -> Option<String> {
self.names.lock().unwrap().get(id).cloned()
}
}
#[cfg(test)]
mod tests {
use super::*;
use iroh::SecretKey;
fn fresh_id() -> EndpointId {
SecretKey::generate().public()
}
#[test]
fn upsert_then_lookup_returns_sanitized_name() {
let roster = ChatRoster::default();
let a = fresh_id();
roster.upsert(a, "Alice");
assert_eq!(roster.name_of(&a), Some("Alice".to_string()));
// Bidi override / zero-width spoofing characters are stripped.
roster.upsert(a, "Al\u{202E}ice\u{200B}");
assert_eq!(roster.name_of(&a), Some("Alice".to_string()));
}
#[test]
fn name_update_affects_future_lookups() {
let roster = ChatRoster::default();
let a = fresh_id();
roster.upsert(a, "Alice");
roster.upsert(a, "Alice2");
assert_eq!(roster.name_of(&a), Some("Alice2".to_string()));
}
#[test]
fn unknown_author_has_no_name() {
let roster = ChatRoster::default();
roster.upsert(fresh_id(), "Alice");
assert_eq!(roster.name_of(&fresh_id()), None);
}
#[test]
fn removed_author_is_no_longer_a_member() {
let roster = ChatRoster::default();
let a = fresh_id();
roster.upsert(a, "Alice");
roster.remove(&a);
assert_eq!(roster.name_of(&a), None);
}
#[test]
fn empty_sanitized_name_falls_back_to_short_id() {
let roster = ChatRoster::default();
let a = fresh_id();
roster.upsert(a, "\u{0}\r\n\t ");
let label = roster.name_of(&a).unwrap();
assert!(!label.is_empty());
assert_eq!(label, crate::short_id(&a.to_string()));
}
#[test]
fn new_ids_are_refused_past_the_cap_but_updates_land() {
let roster = ChatRoster::default();
let first = fresh_id();
roster.upsert(first, "member");
for _ in 1..CHAT_ROSTER_CAP {
roster.upsert(fresh_id(), "member");
}
// A brand-new 33rd id is refused...
let overflow = fresh_id();
roster.upsert(overflow, "overflow");
assert_eq!(roster.name_of(&overflow), None);
// ...but an update to a present id still lands at the cap.
roster.upsert(first, "renamed");
assert_eq!(roster.name_of(&first), Some("renamed".to_string()));
}
}
+187
View File
@@ -0,0 +1,187 @@
//! Per-peer connection-transparency derivation.
//!
//! The transport hands us cumulative counters for each peer's selected QUIC
//! path ([`PathSnapshot`]); this module turns two consecutive snapshots into
//! the human-facing [`PeerConnInfo`] the UI renders (badge + tooltip): path
//! type, RTT, and loss/bitrate over the poll window. Pure functions only —
//! the polling task in `core::mod` owns the clock and the previous-snapshot
//! map.
use crate::network::PathSnapshot;
use std::time::Duration;
/// How often the core polls the transport for path snapshots.
pub const POLL_INTERVAL: Duration = Duration::from_secs(1);
/// Derived, display-ready connection info for one peer, sent to the UI via
/// `UiEvent::ConnectionStats`. Window-relative fields are `None` when they
/// can't be derived yet (first poll, path switch, or an idle window).
#[derive(Debug, Clone, PartialEq)]
pub struct PeerConnInfo {
/// True = relayed path, false = direct IP path.
pub relay: bool,
/// `ip:port` for a direct path, the relay URL for a relayed one.
pub remote_addr: String,
/// Path round-trip time, rounded to whole milliseconds.
pub rtt_ms: u32,
/// Percentage of packets sent in the window that were detected lost.
pub loss_pct: Option<f32>,
/// Outbound bitrate over the window, kilobits per second.
pub up_kbps: Option<f32>,
/// Inbound bitrate over the window, kilobits per second.
pub down_kbps: Option<f32>,
}
/// Derive display info from the current snapshot and (when comparable) the
/// previous one. `prev` is comparable only if it's the same path — a relay→
/// direct migration or a reconnect resets the counters, so those windows
/// yield `None` rates rather than garbage (negative deltas show up as
/// `cur < prev` and are treated the same way).
pub fn derive(prev: Option<&PathSnapshot>, cur: &PathSnapshot, elapsed: Duration) -> PeerConnInfo {
let rates = prev
.filter(|p| comparable(p, cur))
.and_then(|p| window_rates(p, cur, elapsed));
PeerConnInfo {
relay: cur.is_relay,
remote_addr: cur.remote_addr.clone(),
rtt_ms: cur.rtt.as_millis().min(u128::from(u32::MAX)) as u32,
loss_pct: rates.and_then(|r| r.loss_pct),
up_kbps: rates.map(|r| r.up_kbps),
down_kbps: rates.map(|r| r.down_kbps),
}
}
/// True when `cur`'s counters continue `prev`'s: same path (address) and
/// monotonically non-decreasing counters (a reconnect on the same address
/// restarts them from zero).
fn comparable(prev: &PathSnapshot, cur: &PathSnapshot) -> bool {
prev.remote_addr == cur.remote_addr
&& cur.tx_bytes >= prev.tx_bytes
&& cur.rx_bytes >= prev.rx_bytes
&& cur.tx_datagrams >= prev.tx_datagrams
&& cur.lost_packets >= prev.lost_packets
}
#[derive(Debug, Clone, Copy)]
struct WindowRates {
loss_pct: Option<f32>,
up_kbps: f32,
down_kbps: f32,
}
fn window_rates(prev: &PathSnapshot, cur: &PathSnapshot, elapsed: Duration) -> Option<WindowRates> {
let secs = elapsed.as_secs_f64();
if secs <= 0.0 {
return None;
}
let sent = cur.tx_datagrams - prev.tx_datagrams;
let lost = cur.lost_packets - prev.lost_packets;
// Loss detection lags sending (it needs ACK timeouts), so a window can see
// more losses than sends; clamp to 100% rather than exceeding it. An idle
// window (nothing sent or lost) has no loss story to tell.
let loss_pct = if sent == 0 && lost == 0 {
None
} else {
Some(((lost as f64 / (sent.max(lost)) as f64) * 100.0) as f32)
};
let kbps = |bytes: u64| ((bytes as f64 * 8.0 / 1000.0) / secs) as f32;
Some(WindowRates {
loss_pct,
up_kbps: kbps(cur.tx_bytes - prev.tx_bytes),
down_kbps: kbps(cur.rx_bytes - prev.rx_bytes),
})
}
#[cfg(test)]
mod tests {
use super::*;
fn snap(addr: &str, tx_b: u64, rx_b: u64, tx_d: u64, lost: u64) -> PathSnapshot {
PathSnapshot {
is_relay: false,
remote_addr: addr.to_string(),
rtt: Duration::from_millis(12),
tx_bytes: tx_b,
rx_bytes: rx_b,
tx_datagrams: tx_d,
lost_packets: lost,
}
}
#[test]
fn first_poll_has_type_and_rtt_but_no_rates() {
let cur = snap("1.2.3.4:5", 1000, 2000, 50, 0);
let info = derive(None, &cur, POLL_INTERVAL);
assert_eq!(info.rtt_ms, 12);
assert!(!info.relay);
assert_eq!(info.remote_addr, "1.2.3.4:5");
assert_eq!(info.loss_pct, None);
assert_eq!(info.up_kbps, None);
assert_eq!(info.down_kbps, None);
}
#[test]
fn steady_window_yields_rates_and_loss() {
let prev = snap("1.2.3.4:5", 0, 0, 0, 0);
// 1s window: 4000 bytes up (32 kbps), 2000 down (16 kbps), 2 of 100 lost.
let cur = snap("1.2.3.4:5", 4000, 2000, 100, 2);
let info = derive(Some(&prev), &cur, Duration::from_secs(1));
assert_eq!(info.up_kbps, Some(32.0));
assert_eq!(info.down_kbps, Some(16.0));
assert_eq!(info.loss_pct, Some(2.0));
}
#[test]
fn idle_window_has_no_loss_story() {
let prev = snap("1.2.3.4:5", 4000, 2000, 100, 2);
let cur = prev.clone();
let info = derive(Some(&prev), &cur, Duration::from_secs(1));
assert_eq!(info.loss_pct, None);
assert_eq!(info.up_kbps, Some(0.0));
}
#[test]
fn loss_detected_in_an_idle_window_clamps_to_full() {
// Losses can be *detected* after sending stops (ACK timeouts fire late).
let prev = snap("1.2.3.4:5", 4000, 2000, 100, 0);
let cur = snap("1.2.3.4:5", 4000, 2000, 100, 3);
let info = derive(Some(&prev), &cur, Duration::from_secs(1));
assert_eq!(info.loss_pct, Some(100.0));
}
#[test]
fn path_switch_resets_the_window() {
let prev = snap("relay.example:443", 9000, 9000, 900, 5);
let cur = snap("1.2.3.4:5", 100, 100, 10, 0);
let info = derive(Some(&prev), &cur, Duration::from_secs(1));
assert_eq!(info.up_kbps, None);
assert_eq!(info.loss_pct, None);
}
#[test]
fn counter_reset_on_same_address_resets_the_window() {
// Same address but the connection was rebuilt → counters restarted.
let prev = snap("1.2.3.4:5", 9000, 9000, 900, 5);
let cur = snap("1.2.3.4:5", 100, 100, 10, 0);
let info = derive(Some(&prev), &cur, Duration::from_secs(1));
assert_eq!(info.up_kbps, None);
assert_eq!(info.loss_pct, None);
}
#[test]
fn zero_elapsed_yields_no_rates() {
let prev = snap("1.2.3.4:5", 0, 0, 0, 0);
let cur = snap("1.2.3.4:5", 4000, 2000, 100, 2);
let info = derive(Some(&prev), &cur, Duration::ZERO);
assert_eq!(info.up_kbps, None);
assert_eq!(info.loss_pct, None);
}
#[test]
fn oversized_rtt_saturates_instead_of_wrapping() {
let mut cur = snap("1.2.3.4:5", 0, 0, 0, 0);
cur.rtt = Duration::from_secs(u64::MAX);
let info = derive(None, &cur, POLL_INTERVAL);
assert_eq!(info.rtt_ms, u32::MAX);
}
}
+252
View File
@@ -0,0 +1,252 @@
//! Byte/request budgets for AUTOMATIC chat-attachment fetches (Phase 3B).
//!
//! The four-permit semaphore bounds how many auto-fetch tasks run at once, but
//! not how much a peer can make us download over time: with permits released
//! after each transfer, an insider could stream distinct ≤4 MiB images
//! sequentially forever. This budget adds per-author and session (room-wide)
//! token buckets over both request COUNT and declared BYTES. Like the Phase 2
//! chat gate, time is passed in — never read from a clock — so every refill
//! boundary is unit-testable.
//!
//! Only the automatic path consults this; a user's explicit click (Save /
//! Download / Load image) is human-rate-limited and always allowed through to
//! the fetch (still subject to the transfer cap and cache/decoder budgets).
use iroh::EndpointId;
use std::collections::HashMap;
/// Per-author request burst: how many auto-fetches one author can trigger
/// back-to-back before refill pacing binds.
pub const AUTHOR_REQ_BURST: f64 = 8.0;
/// Per-author request refill: one recovered every 10 s.
pub const AUTHOR_REQ_REFILL_PER_MS: f64 = 1.0 / 10_000.0;
/// Per-author byte burst (declared sizes): a couple of full-size auto images
/// plus a normal working set.
pub const AUTHOR_BYTES_BURST: f64 = (16 * 1024 * 1024) as f64;
/// Per-author byte refill: 64 KiB/s (~one 4 MiB auto image per minute).
pub const AUTHOR_BYTES_REFILL_PER_MS: f64 = (64 * 1024) as f64 / 1000.0;
/// Session-wide request burst across all authors.
pub const SESSION_REQ_BURST: f64 = 16.0;
/// Session-wide request refill: one recovered every 5 s.
pub const SESSION_REQ_REFILL_PER_MS: f64 = 1.0 / 5_000.0;
/// Session-wide byte burst across all authors.
pub const SESSION_BYTES_BURST: f64 = (48 * 1024 * 1024) as f64;
/// Session-wide byte refill: 128 KiB/s.
pub const SESSION_BYTES_REFILL_PER_MS: f64 = (128 * 1024) as f64 / 1000.0;
/// Bound on the per-author bucket map. Authors are roster members (≤32 live),
/// so this tracks the roster plus recently departed; the least-recently-active
/// entry is pruned past the cap.
pub const AUTHOR_MAP_CAP: usize = 64;
/// A deterministic token bucket that can take a WEIGHTED cost (bytes), unlike
/// the unit-cost bucket in the gossip chat gate.
#[derive(Debug, Clone, Copy)]
struct WeightedBucket {
tokens: f64,
last_ms: u64,
}
impl WeightedBucket {
fn full(burst: f64, now_ms: u64) -> Self {
Self {
tokens: burst,
last_ms: now_ms,
}
}
/// Refill for elapsed time (capped at `burst`) without consuming.
fn refill(&mut self, burst: f64, refill_per_ms: f64, now_ms: u64) {
let elapsed = now_ms.saturating_sub(self.last_ms) as f64;
self.tokens = (self.tokens + elapsed * refill_per_ms).min(burst);
self.last_ms = now_ms;
}
fn has(&self, cost: f64) -> bool {
self.tokens >= cost
}
fn take(&mut self, cost: f64) {
self.tokens -= cost;
}
}
/// One author's pair of buckets plus last activity (for idle pruning).
#[derive(Debug)]
struct AuthorBudget {
reqs: WeightedBucket,
bytes: WeightedBucket,
last_seen_ms: u64,
}
/// Admission budget for automatic attachment fetches. All four buckets are
/// checked BEFORE any is consumed, so a rejection never burns tokens (no
/// refund bookkeeping — the check-then-take is atomic within `admit`).
#[derive(Debug)]
pub struct AutoFetchBudget {
session_reqs: WeightedBucket,
session_bytes: WeightedBucket,
authors: HashMap<EndpointId, AuthorBudget>,
}
impl AutoFetchBudget {
pub fn new(now_ms: u64) -> Self {
Self {
session_reqs: WeightedBucket::full(SESSION_REQ_BURST, now_ms),
session_bytes: WeightedBucket::full(SESSION_BYTES_BURST, now_ms),
authors: HashMap::new(),
}
}
/// Whether an auto-fetch of `size` declared bytes for `author` may start
/// now. Consumes one request token and `size` byte tokens from BOTH the
/// author's and the session's buckets — or nothing at all on rejection.
pub fn admit(&mut self, author: EndpointId, size: u64, now_ms: u64) -> bool {
self.prune(author, now_ms);
let entry = self.authors.entry(author).or_insert_with(|| AuthorBudget {
reqs: WeightedBucket::full(AUTHOR_REQ_BURST, now_ms),
bytes: WeightedBucket::full(AUTHOR_BYTES_BURST, now_ms),
last_seen_ms: now_ms,
});
entry.last_seen_ms = now_ms;
entry
.reqs
.refill(AUTHOR_REQ_BURST, AUTHOR_REQ_REFILL_PER_MS, now_ms);
entry
.bytes
.refill(AUTHOR_BYTES_BURST, AUTHOR_BYTES_REFILL_PER_MS, now_ms);
self.session_reqs
.refill(SESSION_REQ_BURST, SESSION_REQ_REFILL_PER_MS, now_ms);
self.session_bytes
.refill(SESSION_BYTES_BURST, SESSION_BYTES_REFILL_PER_MS, now_ms);
let cost = size as f64;
let ok = entry.reqs.has(1.0)
&& entry.bytes.has(cost)
&& self.session_reqs.has(1.0)
&& self.session_bytes.has(cost);
if ok {
let entry = self.authors.get_mut(&author).expect("just inserted");
entry.reqs.take(1.0);
entry.bytes.take(cost);
self.session_reqs.take(1.0);
self.session_bytes.take(cost);
}
ok
}
/// Keep the author map bounded: past the cap, drop the least-recently
/// active entry that isn't the author being admitted. A pruned author
/// returns with full buckets, but authors are roster-gated upstream, so
/// the map can't be churned by strangers.
fn prune(&mut self, keep: EndpointId, _now_ms: u64) {
while self.authors.len() >= AUTHOR_MAP_CAP {
let Some(victim) = self
.authors
.iter()
.filter(|(id, _)| **id != keep)
.min_by_key(|(_, b)| b.last_seen_ms)
.map(|(id, _)| *id)
else {
break;
};
self.authors.remove(&victim);
}
}
#[cfg(test)]
fn author_count(&self) -> usize {
self.authors.len()
}
}
#[cfg(test)]
mod tests {
use super::*;
use iroh::SecretKey;
const T0: u64 = 1_000_000;
const MIB: u64 = 1024 * 1024;
fn author() -> EndpointId {
SecretKey::generate().public()
}
#[test]
fn author_request_burst_then_refill_recovers() {
let mut b = AutoFetchBudget::new(T0);
let a = author();
// Tiny sizes so only the REQUEST buckets can bind.
for _ in 0..AUTHOR_REQ_BURST as usize {
assert!(b.admit(a, 1, T0));
}
assert!(!b.admit(a, 1, T0), "author request burst exhausted");
// One request refills after 10 s.
assert!(b.admit(a, 1, T0 + 10_000));
assert!(!b.admit(a, 1, T0 + 10_000));
}
#[test]
fn author_byte_budget_binds_and_recovers() {
let mut b = AutoFetchBudget::new(T0);
let a = author();
// 4 × 4 MiB = the full 16 MiB author byte burst (well under the
// 8-request burst, so bytes are the binding constraint).
for _ in 0..4 {
assert!(b.admit(a, 4 * MIB, T0));
}
assert!(!b.admit(a, 4 * MIB, T0), "author byte burst exhausted");
// 64 KiB/s → a 4 MiB image is affordable again after 64 s (which also
// refills 6 request tokens, so bytes stay the binding constraint).
assert!(!b.admit(a, 4 * MIB, T0 + 32_000));
assert!(b.admit(a, 4 * MIB, T0 + 64_000));
}
#[test]
fn session_budget_binds_across_authors_without_burning_author_tokens() {
let mut b = AutoFetchBudget::new(T0);
// Three authors × 16 MiB exhausts the 48 MiB session byte burst even
// though each author is within their own budget.
for _ in 0..3 {
let a = author();
for _ in 0..4 {
assert!(b.admit(a, 4 * MIB, T0));
}
}
let fresh = author();
assert!(!b.admit(fresh, 4 * MIB, T0), "session bytes exhausted");
// The rejection consumed NOTHING: once the session refills enough for
// one image (4 MiB / 128 KiB/s = 32 s), the fresh author's own full
// burst is intact and admits immediately.
assert!(b.admit(fresh, 4 * MIB, T0 + 32_000));
}
#[test]
fn session_request_bucket_binds_across_authors() {
let mut b = AutoFetchBudget::new(T0);
// 16 tiny requests from distinct authors exhaust the session request
// burst while every author bucket stays nearly full.
for _ in 0..SESSION_REQ_BURST as usize {
assert!(b.admit(author(), 1, T0));
}
assert!(!b.admit(author(), 1, T0), "session requests exhausted");
assert!(b.admit(author(), 1, T0 + 5_000), "one recovers after 5 s");
}
#[test]
fn author_map_stays_bounded_pruning_least_recent() {
let mut b = AutoFetchBudget::new(T0);
// Session request refill would bind over a naive loop; space the
// admissions out so only the map bound is under test.
let mut t = T0;
let first = author();
assert!(b.admit(first, 1, t));
for _ in 0..(AUTHOR_MAP_CAP + 10) {
t += 10_000;
assert!(b.admit(author(), 1, t));
assert!(b.author_count() <= AUTHOR_MAP_CAP);
}
assert!(b.author_count() <= AUTHOR_MAP_CAP);
}
}
+92 -7
View File
@@ -202,15 +202,20 @@ impl JitterBuffer {
None
} else {
// Gap with later packets already buffered: a packet was lost
// or reordered out of window. First try Opus in-band FEC from
// the next packet; if unavailable, fall back to plain PLC.
// or reordered out of window. Try Opus in-band FEC from the
// packet right after the gap; if that packet isn't buffered
// (burst loss) or FEC fails, fall back to plain PLC.
self.next_seq = Some(next.wrapping_add(1));
self.note_disruption();
let next_payload = self.packets.values().next().expect("non-empty");
self.decoder
.decode_fec(next_payload)
.or_else(|_| self.decoder.decode(None))
.ok()
let (&smallest, next_payload) = self.packets.iter().next().expect("non-empty");
if fec_covers_gap(next, smallest) {
self.decoder
.decode_fec(next_payload)
.or_else(|_| self.decoder.decode(None))
.ok()
} else {
self.decoder.decode(None).ok()
}
}
}
}
@@ -222,6 +227,15 @@ impl JitterBuffer {
}
}
/// Opus in-band FEC in packet N carries a low-fidelity copy of frame N-1 and
/// nothing else — a lost frame `next` is FEC-recoverable solely from packet
/// `next+1`. Any later successor's FEC data is a different frame's audio, and
/// splicing it into this gap plays sound from the wrong position; the caller
/// must conceal with plain PLC instead.
fn fec_covers_gap(next: u32, smallest_buffered: u32) -> bool {
smallest_buffered == next.wrapping_add(1)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -379,6 +393,77 @@ mod tests {
);
}
#[test]
fn fec_covers_gap_only_for_the_immediate_successor() {
// Packet next+1 is the only one whose in-band FEC describes frame `next`.
assert!(fec_covers_gap(4, 5));
// A burst gap: the smallest survivor's FEC is some other frame's audio.
assert!(!fec_covers_gap(3, 5));
assert!(!fec_covers_gap(3, 3_000));
// Sequence wraparound still counts as adjacent.
assert!(fec_covers_gap(u32::MAX, 0));
}
#[test]
fn burst_gap_falls_back_to_plc_not_wrong_position_fec() {
let mut enc = OpusEncoder::new(48000, Channels::Mono, Application::Voip).unwrap();
enc.apply_params(&OpusParams {
bitrate: 20_000,
inband_fec: true,
packet_loss_perc: 60,
dtx: false,
})
.unwrap();
// Frames 0..=6; 3 and 4 are lost as a burst, so when playout reaches
// seq 3 the smallest buffered packet is 5 — whose FEC data is frame 4,
// NOT frame 3. The buffer must conceal 3 with plain PLC rather than
// splice frame 4's audio into the wrong position.
let packets: Vec<Vec<u8>> = (0..7).map(|seq| tone_frame(&mut enc, 8_000, seq)).collect();
// Twin decoder replaying the exact call sequence the jitter buffer
// should make for seq 3: decode 0,1,2 then a plain PLC conceal.
let mut twin = OpusDecoder::new(48000, Channels::Mono, FRAME_SAMPLES).unwrap();
for packet in packets.iter().take(3) {
twin.decode(Some(packet)).unwrap();
}
let expected_plc = twin.decode(None).unwrap();
let mut jb = JitterBuffer::new().unwrap();
for (seq, packet) in packets.iter().enumerate() {
if seq != 3 && seq != 4 {
jb.insert(seq as u32, packet.clone());
}
}
for _ in 0..3 {
assert_eq!(jb.pop_frame().map(|frame| frame.len()), Some(FRAME_SAMPLES));
}
// Seq 3: burst gap — bit-exact PLC (same decoder state, same inputs),
// which decode_fec(packet 5) could never produce.
let concealed = jb.pop_frame().expect("gap should be concealed");
assert_eq!(concealed, expected_plc, "burst gap must use plain PLC");
// Seq 4: packet 5 IS the immediate successor, so its FEC data is
// frame 4's audio — the correctly-positioned recovery still applies.
let recovered = jb
.pop_frame()
.expect("adjacent gap should be reconstructed");
let mut fec_twin = OpusDecoder::new(48000, Channels::Mono, FRAME_SAMPLES).unwrap();
for packet in packets.iter().take(3) {
fec_twin.decode(Some(packet)).unwrap();
}
fec_twin.decode(None).unwrap();
let expected_fec = fec_twin.decode_fec(&packets[5]).unwrap();
assert_eq!(recovered, expected_fec, "adjacent gap should still use FEC");
// Then 5 and 6 play normally.
assert_eq!(jb.pop_frame().map(|frame| frame.len()), Some(FRAME_SAMPLES));
assert_eq!(jb.pop_frame().map(|frame| frame.len()), Some(FRAME_SAMPLES));
assert!(jb.pop_frame().is_none());
}
#[test]
fn drops_packets_already_played() {
let mut enc = OpusEncoder::new(48000, Channels::Mono, Application::Voip).unwrap();
+78 -14
View File
@@ -1,4 +1,4 @@
use crate::config::{AudioProfile, NetworkMode, RecordingMode};
use crate::config::{AudioProfile, NetworkMode, RecordingMode, ScreenShareSettings, ShareQuality};
use crate::friends::Friend;
use crate::network::PeerState;
use crate::presence::{FriendPresence, PresenceMode};
@@ -66,15 +66,27 @@ pub enum CoreCommand {
/// Set what a recording captures (mixed / per-peer stems / both). Takes
/// effect on the next recording start. Sent at startup from config.
SetRecordingMode(RecordingMode),
/// Broadcast a room text-chat message. No-op when not in a call.
SendChat(String),
/// Broadcast a room text-chat message. `local_id` is the app's local-only
/// handle for this send — it never goes on the wire; the core echoes it back
/// in [`UiEvent::ChatSendResult`] so the UI can mark the matching local echo
/// honestly (chat-hardening Phase 5). Not being in a call is a FAILURE
/// result, not a silent no-op.
SendChat {
local_id: u64,
text: String,
},
/// Send a chat message carrying a file attachment. The app has already read +
/// capped the file and built the descriptor; core makes the bytes available
/// on the file plane and broadcasts the descriptor.
/// on the file plane and broadcasts the descriptor. `local_id` as in
/// [`CoreCommand::SendChat`].
SendChatFile {
local_id: u64,
text: String,
attachment: crate::files::ChatAttachment,
data: Vec<u8>,
/// Shared, not owned: the same allocation is retained by the UI cache
/// and handed to the serve store, so a 25 MiB attachment is held once,
/// not copied across UI / command queue / serve store (Phase 3C).
data: std::sync::Arc<Vec<u8>>,
},
/// Fetch a received attachment's bytes from its sender over the file plane
/// (used for on-demand file/chip downloads; images are auto-fetched on
@@ -119,13 +131,18 @@ pub enum CoreCommand {
/// whole desktop audio (the legacy behavior).
StartScreenShare {
audio_app: Option<String>,
settings: ScreenShareSettings,
quality: ShareQuality,
},
/// Stop sharing our screen: kill the pixelpass host and clear the presence
/// ticket. No-op when not sharing.
StopScreenShare,
/// Watch a peer's screen share: spawn a pixelpass viewer for `ticket` and
/// open it in a local player.
ViewShare(String),
ViewShare {
ticket: String,
settings: ScreenShareSettings,
},
/// Mint a fresh persistent identity (W7), discarding the old one. Takes effect
/// on the next room join (the endpoint is rebuilt then). The core replies with
/// an updated [`UiEvent::IdentityStatus`].
@@ -219,8 +236,12 @@ pub fn delivery_class(cmd: &CoreCommand) -> DeliveryClass {
| CoreCommand::SetAudioProfile(_)
| CoreCommand::SetRecording(_)
| CoreCommand::SetRecordingMode(_)
| CoreCommand::SendChat(_)
| CoreCommand::SendChat {
local_id: _,
text: _,
}
| CoreCommand::SendChatFile {
local_id: _,
text: _,
attachment: _,
data: _,
@@ -244,9 +265,16 @@ pub fn delivery_class(cmd: &CoreCommand) -> DeliveryClass {
}
| CoreCommand::SetPixelpassPath(_)
| CoreCommand::ListAudioApps
| CoreCommand::StartScreenShare { audio_app: _ }
| CoreCommand::StartScreenShare {
audio_app: _,
settings: _,
quality: _,
}
| CoreCommand::StopScreenShare
| CoreCommand::ViewShare(_)
| CoreCommand::ViewShare {
ticket: _,
settings: _,
}
| CoreCommand::RegenerateIdentity
| CoreCommand::AddFriend {
id: _,
@@ -300,8 +328,12 @@ pub fn coalesce_key(cmd: &CoreCommand) -> Option<CoalesceKey> {
| CoreCommand::SetAudioProfile(_)
| CoreCommand::SetRecording(_)
| CoreCommand::SetRecordingMode(_)
| CoreCommand::SendChat(_)
| CoreCommand::SendChat {
local_id: _,
text: _,
}
| CoreCommand::SendChatFile {
local_id: _,
text: _,
attachment: _,
data: _,
@@ -325,9 +357,16 @@ pub fn coalesce_key(cmd: &CoreCommand) -> Option<CoalesceKey> {
}
| CoreCommand::SetPixelpassPath(_)
| CoreCommand::ListAudioApps
| CoreCommand::StartScreenShare { audio_app: _ }
| CoreCommand::StartScreenShare {
audio_app: _,
settings: _,
quality: _,
}
| CoreCommand::StopScreenShare
| CoreCommand::ViewShare(_)
| CoreCommand::ViewShare {
ticket: _,
settings: _,
}
| CoreCommand::RegenerateIdentity
| CoreCommand::AddFriend {
id: _,
@@ -382,6 +421,11 @@ pub enum UiEvent {
id: EndpointId,
},
AudioLevels(Vec<(EndpointId, f32)>),
/// Periodic per-peer connection transparency snapshot (~1/sec): path type
/// (direct/relay), RTT, and window loss/bitrate for every peer with a live
/// audio link. A FULL replacement each time — a peer absent from the list
/// has no live link right now, so its badge should disappear.
ConnectionStats(Vec<(EndpointId, crate::core::connstats::PeerConnInfo)>),
/// Raw (pre-gate, pre-mute) normalized RMS of the local mic, `0.0..=1.0`,
/// for the settings level meter. Throttled to ~10/sec.
MicLevel(f32),
@@ -393,6 +437,15 @@ pub enum UiEvent {
RecordingStopped {
path: String,
},
/// The outcome of one locally initiated chat send (chat-hardening Phase 5).
/// `error = None` means our signed broadcast was handed to the gossip swarm
/// — deliberately NOT a delivery/read receipt; PeerSpeak has no peer
/// acknowledgements. `local_id` is the app's own handle from the
/// `SendChat`/`SendChatFile` command and never appears on the wire.
ChatSendResult {
local_id: u64,
error: Option<String>,
},
/// A room text-chat message arrived from a peer (never our own — local
/// messages are echoed by the UI on send). `from` is the sender's node id
/// string, used to key their avatar (W4).
@@ -410,7 +463,15 @@ pub enum UiEvent {
AttachmentReady {
from: EndpointId,
id: crate::files::AttachmentId,
data: Vec<u8>,
data: std::sync::Arc<Vec<u8>>,
},
/// An attachment fetch task was spawned (auto or on demand). Lets the UI
/// show a real "loading" state instead of inferring it from cache absence —
/// absence now means NOT fetched (e.g. auto-fetch was skipped), which
/// renders a Load button rather than an indefinite "loading…" (Phase 3B).
AttachmentFetchStarted {
from: EndpointId,
id: crate::files::AttachmentId,
},
/// An attachment fetch failed (sender gone, too large, decode error, etc.).
AttachmentFailed {
@@ -581,7 +642,10 @@ mod tests {
CoreCommand::SetPeerMuted(peer, true),
CoreCommand::SetPresenceMode(PresenceMode::Normal),
CoreCommand::SetAudioProfile(crate::config::AudioProfile::BadNetwork),
CoreCommand::SendChat("hello".to_string()),
CoreCommand::SendChat {
local_id: 1,
text: "hello".to_string(),
},
];
for cmd in commands {
+566 -93
View File
@@ -1,6 +1,10 @@
pub mod chatroster;
pub mod connstats;
pub mod fetchbudget;
pub mod jitter;
pub mod messages;
mod recovery;
mod teardown;
use crate::audio::eq::{Eq, EqSettings};
use crate::audio::{AudioBackend, PlatformAudioBackend};
@@ -299,6 +303,9 @@ struct GraceExpiry<'a> {
jitter: &'a Arc<Mutex<HashMap<EndpointId, JitterBuffer>>>,
ui_tx: &'a mpsc::Sender<UiEvent>,
recovery: Option<&'a RecoveryContext>,
/// Terminal eviction also revokes the peer's chat authority (Phase 2):
/// the roster-bound name map entry goes with the peer.
chat_roster: &'a chatroster::ChatRoster,
}
fn arm_grace_timer(
@@ -318,6 +325,7 @@ fn arm_grace_timer(
let timers_evict = timers.clone();
let seen_evict = seen_connected.clone();
let recovery_evict = expiry.recovery.cloned();
let chat_roster_evict = expiry.chat_roster.clone();
let handle = tokio::spawn(async move {
tokio::time::sleep(grace).await;
crate::log_msg(&format!("Reconnect grace expired for peer {:?}", peer_id));
@@ -329,6 +337,9 @@ fn arm_grace_timer(
}
transport_evict.remove_audio_sender(peer_id);
// Terminal eviction revokes chat authority too (Phase 2): a readmission
// via fresh authenticated Announce re-registers the name on PeerJoined.
chat_roster_evict.remove(&peer_id);
if let Some(recovery) = &recovery_evict {
// Revoke roster authority before the first await in teardown. A
// verified Announce racing after this point is then a PeerJoined and
@@ -554,6 +565,7 @@ pub struct ConnEventHandler {
jitter: Arc<Mutex<HashMap<EndpointId, JitterBuffer>>>,
recovery: Option<RecoveryContext>,
grace: Duration,
chat_roster: chatroster::ChatRoster,
}
impl ConnEventHandler {
@@ -572,6 +584,7 @@ impl ConnEventHandler {
jitter,
recovery: None,
grace: RECONNECT_GRACE,
chat_roster: chatroster::ChatRoster::default(),
}
}
@@ -581,6 +594,13 @@ impl ConnEventHandler {
self
}
/// Share the room's chat roster so a grace-expiry eviction fired from the
/// transport's link-state path also revokes chat authority (Phase 2).
pub fn with_chat_roster(mut self, chat_roster: chatroster::ChatRoster) -> Self {
self.chat_roster = chat_roster;
self
}
fn with_recovery(mut self, recovery: RecoveryContext) -> Self {
self.recovery = Some(recovery);
self
@@ -604,6 +624,7 @@ impl ConnEventHandler {
jitter: &self.jitter,
ui_tx: &self.ui_tx,
recovery: self.recovery.as_ref(),
chat_roster: &self.chat_roster,
},
self.grace,
id,
@@ -652,37 +673,43 @@ struct ActiveSession {
mixer_task: tokio::task::JoinHandle<()>,
event_task: tokio::task::JoinHandle<()>,
conn_event_task: tokio::task::JoinHandle<()>,
conn_stats_task: tokio::task::JoinHandle<()>,
recovery_task: tokio::task::JoinHandle<()>,
recovery_terminal_task: tokio::task::JoinHandle<()>,
grace_timers: GraceTimers,
transport: Arc<IrohTransport>,
/// Loaded PipeWire echo-cancel module (if enabled); unloads on drop.
#[cfg(target_os = "linux")]
echo_cancel: Option<crate::audio::echo_cancel::EchoCancelGuard>,
/// Our pixelpass screen-share host child while sharing (`kill_on_drop`, so it
/// also dies if the session is dropped without an explicit stop).
screenshare_host: Option<tokio::process::Child>,
/// pixelpass viewer children we spawned to watch peers' shares; killed on
/// session teardown (each also self-exits when its player window closes).
screenshare_viewers: Vec<tokio::process::Child>,
/// The screen-share children and the echo-cancel module, held together
/// because their **destruction order** is load-bearing: the AEC module must
/// not unload while a pixelpass host is alive and fanning out (design v3.4
/// §7.1). `teardown` owns that ordering; see `core::teardown`.
teardown: SessionTeardown,
}
/// The session's teardown set, with the echo-cancel guard the platform actually
/// has. On non-Linux there is no AEC module, and `Infallible` makes that
/// structural — the `Option` cannot be `Some`.
#[cfg(target_os = "linux")]
type SessionTeardown = teardown::ScreenshareTeardown<
tokio::process::Child,
crate::audio::echo_cancel::EchoCancelGuard,
>;
#[cfg(not(target_os = "linux"))]
type SessionTeardown =
teardown::ScreenshareTeardown<tokio::process::Child, std::convert::Infallible>;
impl ActiveSession {
async fn shutdown(mut self, audio_backend: Arc<PlatformAudioBackend>) {
crate::log_msg("ActiveSession::shutdown started");
// Tear down any screen-share children first so the host stops streaming
// promptly (kill_on_drop is the backstop, but kill explicitly so viewers
// see the stream end without waiting on drop ordering).
if let Some(mut host) = self.screenshare_host.take() {
let _ = host.kill().await;
}
for mut viewer in self.screenshare_viewers.drain(..) {
let _ = viewer.kill().await;
}
// promptly, and so they are dead *and reaped* well before the AEC guard
// unloads at the end of this function (design v3.4 §7.1). Drop ordering
// is the backstop for the unwind path; this is the path we control.
self.teardown.shutdown_children().await;
self.datagram_task.abort();
self.mixer_task.abort();
self.event_task.abort();
self.conn_event_task.abort();
self.conn_stats_task.abort();
// Abort any pending reconnect grace timers so they can't fire a stray
// eviction (or touch a torn-down transport) after the session is gone.
for (_, handle) in self.grace_timers.lock().unwrap().drain() {
@@ -702,8 +729,9 @@ impl ActiveSession {
// Unload the echo-cancel module now that the audio streams releasing its
// virtual nodes have stopped. (Dropping the guard runs `pactl unload`.)
#[cfg(target_os = "linux")]
drop(self.echo_cancel);
// The screen-share children were killed *and reaped* at the top of this
// function, so nothing pixelpass-side is alive to see the module vanish.
drop(self.teardown);
crate::log_msg("Leaving room...");
let _ = self.room_state.leave().await;
@@ -883,6 +911,103 @@ async fn build_net_stack(
})
}
/// Retry policy for a live net-stack replacement, generic over the builder so
/// it is unit-testable without binding sockets: build for `requested`; if that
/// fails, build for `live` (the posture the old stack was actually running) so
/// a bad posture change degrades to the previous posture instead of leaving no
/// stack at all. When `requested == live` the second attempt is a plain retry.
///
/// `Ok((stack, mode, primary_err))` — a stack is up on `mode`; `primary_err`
/// is `Some` when the first attempt failed. `Err((primary, fallback))` — both
/// attempts failed and networking is gone.
async fn rebuild_with_fallback<T, E, F, Fut>(
mut build: F,
requested: NetworkMode,
live: NetworkMode,
) -> Result<(T, NetworkMode, Option<E>), (E, E)>
where
F: FnMut(NetworkMode) -> Fut,
Fut: std::future::Future<Output = Result<T, E>>,
{
match build(requested).await {
Ok(stack) => Ok((stack, requested, None)),
Err(primary) => match build(live).await {
Ok(stack) => Ok((stack, live, Some(primary))),
Err(fallback) => Err((primary, fallback)),
},
}
}
/// Tear down `old` and stand up a replacement stack for `requested_mode`.
///
/// A build failure here is rare (only the local socket bind can fail; the
/// relay handshake is backgrounded), but it used to propagate straight out of
/// `run_core_loop` with no `UiEvent`, silently killing every future command —
/// the app looked alive and did nothing. Instead, fall back to `live_mode`
/// via `rebuild_with_fallback`, tell the UI when the requested change did not
/// stick, and return the mode the new stack actually runs so the caller can
/// keep its state honest. `Err` only when both builds fail: networking is
/// gone (already reported to the UI as fatal) and the caller should exit.
#[allow(clippy::too_many_arguments)]
async fn replace_net_stack(
old: NetStack,
what: &str,
secret_key: &SecretKey,
requested_mode: NetworkMode,
live_mode: NetworkMode,
friends_handler: &crate::presence_net::Handler,
publish: bool,
ui_tx: &mpsc::Sender<UiEvent>,
) -> Result<(NetStack, NetworkMode), anyhow::Error> {
let lookup = old.memory_lookup.clone();
old.shutdown().await;
let outcome = rebuild_with_fallback(
|mode| {
build_net_stack(
secret_key.clone(),
mode,
lookup.clone(),
friends_handler.clone(),
publish,
)
},
requested_mode,
live_mode,
)
.await;
match outcome {
Ok((stack, mode, None)) => Ok((stack, mode)),
Ok((stack, mode, Some(primary))) => {
if mode == requested_mode {
// Same-posture retry succeeded — everything the user asked for
// is in effect, so log it rather than raising a UI error.
crate::log_msg(&format!(
"{what}: net stack build failed once ({primary:#}); retry succeeded"
));
} else {
let _ = ui_tx
.send(UiEvent::Error(format!(
"{what} failed ({primary:#}); staying on the previous \
network mode for this session"
)))
.await;
}
Ok((stack, mode))
}
Err((primary, fallback)) => {
let _ = ui_tx
.send(UiEvent::Error(format!(
"Networking lost: {primary:#} (recovery attempt also failed: \
{fallback:#}). Restart PeerSpeak to reconnect."
)))
.await;
Err(anyhow::anyhow!(
"net stack rebuild failed: {primary:#}; fallback: {fallback:#}"
))
}
}
}
/// Maximum number of *automatic* chat-attachment fetches in flight at once.
///
/// Auto-fetch (inline image preview) is triggered by an untrusted peer's chat
@@ -897,6 +1022,15 @@ const MAX_INFLIGHT_ATTACHMENT_FETCHES: usize = 4;
/// fetches (Tier C F-02). Bounded by [`MAX_INFLIGHT_ATTACHMENT_FETCHES`].
type InflightAttachments = Arc<std::sync::Mutex<HashSet<(EndpointId, crate::files::AttachmentId)>>>;
/// Milliseconds since the Unix epoch — the time source handed to the
/// deterministic auto-fetch budget (mirrors the gossip plane's timestamps).
fn unix_now_ms() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0)
}
/// RAII bookkeeping for one bounded auto-fetch: holds the concurrency permit for
/// the task's lifetime and clears the in-flight `(author, id)` marker when the
/// fetch finishes (success OR failure), so the same image can be retried later.
@@ -913,11 +1047,22 @@ impl Drop for AutoFetchGuard {
}
/// Whether to AUTO-fetch a chat image attachment. Only authenticated roster
/// authors qualify (closing the non-roster injection vector), and a `(author,
/// id)` already being fetched is skipped (dedup). The concurrency bound itself is
/// enforced separately by the permit. Pure → unit-testable (Tier C F-02).
fn should_auto_fetch(is_image: bool, author_in_roster: bool, already_inflight: bool) -> bool {
is_image && author_in_roster && !already_inflight
/// authors qualify (closing the non-roster injection vector), only declared
/// sizes at or under [`crate::files::MAX_AUTO_IMAGE_BYTES`] (larger images get
/// a Load button instead, Phase 3B), and a `(author, id)` already being fetched
/// is skipped (dedup). The concurrency bound (permits) and the byte/request
/// budgets ([`fetchbudget::AutoFetchBudget`]) are enforced separately. Pure →
/// unit-testable (Tier C F-02).
fn should_auto_fetch(
is_image: bool,
author_in_roster: bool,
already_inflight: bool,
declared_size: u64,
) -> bool {
is_image
&& author_in_roster
&& !already_inflight
&& declared_size <= crate::files::MAX_AUTO_IMAGE_BYTES
}
/// Fetch a chat attachment's bytes from `from` over the file plane in a detached
@@ -940,6 +1085,12 @@ fn spawn_attachment_fetch(
tokio::spawn(async move {
// Held for the whole fetch; dropped here on completion (Tier C F-02).
let _guard = guard;
// Tell the UI a real fetch task exists for this key, so it can show a
// genuine loading state and dedup further clicks (Phase 3B). Sent from
// this task's channel handle, so it always precedes Ready/Failed.
let _ = ui_tx
.send(UiEvent::AttachmentFetchStarted { from, id: att.id })
.await;
match transport.fetch_attachment(from, &att).await {
Ok(data) => {
if is_image && crate::files::validate_image_bytes(&data).is_none() {
@@ -956,7 +1107,7 @@ fn spawn_attachment_fetch(
.send(UiEvent::AttachmentReady {
from,
id: att.id,
data,
data: Arc::new(data),
})
.await;
}
@@ -1330,6 +1481,10 @@ async fn run_core_loop(
// rebuilt on the next Leave (or before the next Join), preserving the old
// "applies on next join" semantics while keeping the endpoint up while idle.
let mut net_rebuild_pending = false;
// The posture the live stack was actually built with. Trails `network_mode`
// while a rebuild is pending, and is the fallback posture when a rebuild
// fails (see `replace_net_stack`).
let mut net_mode = network_mode;
// When Discoverable is on, the instant it auto-reverts to Normal (W7 P6 time-box).
// `None` = not Discoverable, no pending revert. Set on SetPresenceMode(Discoverable),
@@ -1360,6 +1515,8 @@ async fn run_core_loop(
biased;
maybe_cmd = reliable_rx.recv() => match maybe_cmd {
Some(cmd) => cmd,
// Every `CoreController`/`CoreCommandSender` is gone — the UI has
// dropped the core. Teardown happens once, after the loop.
None => break,
},
maybe_wake = besteffort_wake_rx.recv() => match maybe_wake {
@@ -1378,6 +1535,18 @@ async fn run_core_loop(
None => continue,
}
}
// ⚠️ UNREACHABLE BY CONSTRUCTION, twice over — do not mistake this
// for a live teardown path (phase 0b finding, 2026-07-26):
// 1. this function owns `besteffort_wake_tx` (cloned at the
// `CoreController::new` spawn site, used just above for the
// `has_more` re-arm), so the channel can never close while
// this loop is running;
// 2. even without that, every holder of a wake sender —
// `CoreController` and `CoreCommandSender` — holds
// `reliable_tx` too, and the `biased` select polls that one
// first, so the reliable arm always wins the race to exit.
// Teardown is hoisted after the loop, so if this arm is ever made
// reachable it is already covered — nothing to add here.
None => break,
},
game_change = next_game_change(&mut game_rx) => {
@@ -1547,17 +1716,24 @@ async fn run_core_loop(
// active, rebuild the persistent stack now — after the old session is
// gone, before the new one binds — so this join uses the new posture.
if net_rebuild_pending {
let lookup = net.memory_lookup.clone();
net.shutdown().await;
let publish = presence_mode.lock().unwrap().publishes_to_discovery();
net = build_net_stack(
secret_key.clone(),
let (stack, live) = replace_net_stack(
net,
"Applying deferred network settings",
&secret_key,
network_mode,
lookup,
friends_handler.clone(),
net_mode,
&friends_handler,
publish,
&ui_tx,
)
.await?;
net = stack;
net_mode = live;
// If the new posture failed and we fell back, keep the mode
// state honest (and re-attemptable) rather than pretending
// the change applied. The join proceeds on the live stack.
network_mode = live;
net_rebuild_pending = false;
}
@@ -2220,17 +2396,27 @@ async fn run_core_loop(
Arc::new(tokio::sync::Semaphore::new(MAX_INFLIGHT_ATTACHMENT_FETCHES));
let inflight_attachments: InflightAttachments =
Arc::new(std::sync::Mutex::new(HashSet::new()));
// The authenticated chat roster for this room: id → roster-bound
// display name (chat-hardening Phase 2). Maintained from the same
// sequential event stream; shared because the detached grace-expiry
// timers (here and in the conn-event handler) must also revoke a
// terminally evicted peer's entry. Gates BOTH the chat text (only
// members render, under their roster name — never the wire name)
// and the automatic attachment fetch (Tier C F-02).
let chat_roster = chatroster::ChatRoster::default();
let chat_roster_events = chat_roster.clone();
let event_task = tokio::spawn(async move {
// The authenticated roster for this room, maintained from the
// same sequential event stream. Only its members may trigger an
// automatic attachment fetch (Tier C F-02).
let mut roster: HashSet<EndpointId> = HashSet::new();
let roster = chat_roster_events;
// Byte/request budgets for automatic attachment fetches
// (Phase 3B). Only this sequential task consults it, so it
// needs no lock; time is passed in for testability.
let mut auto_fetch_budget = fetchbudget::AutoFetchBudget::new(unix_now_ms());
while let Some(event) = room_events.recv().await {
match event {
RoomEvent::PeerJoined(peer_id, state) => {
// A (re)join means the peer is back — cancel any
// pending reconnect grace timer before re-adding it.
roster.insert(peer_id);
roster.upsert(peer_id, &state.name);
cancel_grace_timer(&grace_timers_events, &peer_id);
recovery_events.cancel(peer_id);
transport_events.admit_audio_sender(peer_id);
@@ -2286,7 +2472,8 @@ async fn run_core_loop(
.await;
}
RoomEvent::PeerLeft(peer_id) => {
// Graceful leave — evict immediately.
// Graceful leave — evict immediately (chat authority
// and roster-bound name included).
roster.remove(&peer_id);
cancel_grace_timer(&grace_timers_events, &peer_id);
seen_connected_events.lock().unwrap().remove(&peer_id);
@@ -2299,6 +2486,11 @@ async fn run_core_loop(
let _ = ui_tx_events.send(UiEvent::PeerLeft { id: peer_id }).await;
}
RoomEvent::PeerUpdated(peer_id, state) => {
// Keep the roster-bound chat name current: a rename
// lands here as a state update (Phase 2). Future
// messages render under the new name; history keeps
// its stored snapshots.
roster.upsert(peer_id, &state.name);
// A re-announce means the peer is alive — cancel any
// pending grace timer. It may also carry a fresh
// address (peer back on a new network); refresh the
@@ -2346,11 +2538,26 @@ async fn run_core_loop(
}
RoomEvent::ChatMessage {
from,
name,
// The wire name is sender-claimed and NEVER rendered:
// the roster-bound name below is the author label
// (chat-hardening Phase 2, the impersonation fix).
name: _,
text,
ts: _,
attachment,
} => {
// Final-authority roster gate: only a current
// authenticated member (including one inside its
// reconnect grace) may create chat UI work. The
// gossip loop's early known-author gate is defense
// in depth; this map is what actually decides.
let Some(name) = roster.name_of(&from) else {
crate::log_msg(&format!(
"Dropped chat from non-roster author {}",
crate::short_id(&from.to_string())
));
continue;
};
// Auto-fetch image attachments so they render inline
// without a click; non-image files wait for an explicit
// FetchAttachment (the "Save" chip). The descriptor was
@@ -2369,9 +2576,14 @@ async fn run_core_loop(
inflight_attachments.lock().unwrap().contains(&key);
if should_auto_fetch(
is_image,
roster.contains(&from),
// Membership was proven by the roster name
// gate above, which drops non-members before
// any attachment handling.
true,
already_inflight,
) {
att.size,
) && auto_fetch_budget.admit(from, att.size, unix_now_ms())
{
// Reserve the dedup slot, then a permit. If the
// pool is exhausted, drop the auto-fetch (and the
// dedup marker) — the descriptor still shows and
@@ -2443,6 +2655,7 @@ async fn run_core_loop(
jitter: &jitter_events,
ui_tx: &ui_tx_events,
recovery: Some(&recovery_events),
chat_roster: &roster,
},
RECONNECT_GRACE,
peer_id,
@@ -2476,13 +2689,48 @@ async fn run_core_loop(
transport.clone(),
jitter.clone(),
)
.with_recovery(recovery_context);
.with_recovery(recovery_context)
.with_chat_roster(chat_roster.clone());
let conn_event_task = tokio::spawn(async move {
while let Some(event) = conn_events.recv().await {
conn_handler.handle(event).await;
}
});
// Connection-transparency poll: ~1/sec, snapshot every live audio
// link's selected path and hand the UI derived badge info (path
// type, RTT, window loss/bitrate). Read-only against the
// transport; owns the previous-snapshot map the derivation diffs
// against.
let transport_stats = transport.clone();
let ui_tx_stats = ui_tx.clone();
let conn_stats_task = tokio::spawn(async move {
let mut prev: HashMap<EndpointId, crate::network::PathSnapshot> =
HashMap::new();
let mut last = tokio::time::Instant::now();
let mut ticker = tokio::time::interval(connstats::POLL_INTERVAL);
ticker.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
loop {
ticker.tick().await;
let now = tokio::time::Instant::now();
let elapsed = now - last;
last = now;
let snaps = transport_stats.connection_stats();
let infos = snaps
.iter()
.map(|(id, cur)| (*id, connstats::derive(prev.get(id), cur, elapsed)))
.collect();
prev = snaps.into_iter().collect();
if ui_tx_stats
.send(UiEvent::ConnectionStats(infos))
.await
.is_err()
{
break;
}
}
});
let session = ActiveSession {
room_state: room_state.clone(),
capture_thread,
@@ -2490,14 +2738,15 @@ async fn run_core_loop(
mixer_task,
event_task,
conn_event_task,
conn_stats_task,
recovery_task,
recovery_terminal_task,
grace_timers,
transport: transport.clone(),
#[cfg(target_os = "linux")]
echo_cancel: echo_cancel_guard,
screenshare_host: None,
screenshare_viewers: Vec::new(),
teardown: SessionTeardown::new(echo_cancel_guard),
#[cfg(not(target_os = "linux"))]
teardown: SessionTeardown::new(None),
};
let self_id = endpoint.id().to_string();
@@ -2554,17 +2803,21 @@ async fn run_core_loop(
// Apply any network-mode / identity change that was deferred while we
// were in the call (rebuild while idle keeps the endpoint reachable).
if net_rebuild_pending {
let lookup = net.memory_lookup.clone();
net.shutdown().await;
let publish = presence_mode.lock().unwrap().publishes_to_discovery();
net = build_net_stack(
secret_key.clone(),
let (stack, live) = replace_net_stack(
net,
"Applying deferred network settings",
&secret_key,
network_mode,
lookup,
friends_handler.clone(),
net_mode,
&friends_handler,
publish,
&ui_tx,
)
.await?;
net = stack;
net_mode = live;
network_mode = live;
net_rebuild_pending = false;
}
}
@@ -2710,17 +2963,21 @@ async fn run_core_loop(
// idle; if a call is active, defer to the next Leave/Join so the
// live call isn't disrupted (preserves "applies on next join").
if active_session.is_none() {
let lookup = net.memory_lookup.clone();
net.shutdown().await;
let publish = presence_mode.lock().unwrap().publishes_to_discovery();
net = build_net_stack(
secret_key.clone(),
let (stack, live) = replace_net_stack(
net,
"Network mode change",
&secret_key,
network_mode,
lookup,
friends_handler.clone(),
net_mode,
&friends_handler,
publish,
&ui_tx,
)
.await?;
net = stack;
net_mode = live;
network_mode = live;
} else {
net_rebuild_pending = true;
}
@@ -2758,17 +3015,22 @@ async fn run_core_loop(
// key unchanged, so a rebuild would be pointless churn).
if regenerated {
if active_session.is_none() {
let lookup = net.memory_lookup.clone();
net.shutdown().await;
let publish = presence_mode.lock().unwrap().publishes_to_discovery();
net = build_net_stack(
secret_key.clone(),
// Same mode both attempts — the fallback is a plain
// retry under the (already persisted) new key.
let (stack, live) = replace_net_stack(
net,
"Endpoint restart after identity change",
&secret_key,
network_mode,
lookup,
friends_handler.clone(),
net_mode,
&friends_handler,
publish,
&ui_tx,
)
.await?;
net = stack;
net_mode = live;
} else {
net_rebuild_pending = true;
}
@@ -3015,29 +3277,54 @@ async fn run_core_loop(
}
}
CoreCommand::SendChat(text) => {
if let Some(session) = &active_session
&& let Err(e) = session.room_state.send_chat(text, None).await
{
CoreCommand::SendChat { local_id, text } => {
// Honest result either way (Phase 5): no active session is a
// FAILURE the sender must see, not a silent drop.
let error = match &active_session {
Some(session) => session
.room_state
.send_chat(text, None)
.await
.err()
.map(|e| e.to_string()),
None => Some("not in a room".to_string()),
};
if let Some(e) = &error {
crate::log_msg(&format!("Failed to send chat: {e}"));
}
let _ = ui_tx
.send(UiEvent::ChatSendResult { local_id, error })
.await;
}
CoreCommand::SendChatFile {
local_id,
text,
attachment,
data,
} => {
if let Some(session) = &active_session {
// Make the bytes fetchable by room members, then broadcast the
// descriptor alongside the (possibly empty) caption text.
session
.transport
.serve_attachment(attachment.id, Arc::new(data));
if let Err(e) = session.room_state.send_chat(text, Some(attachment)).await {
crate::log_msg(&format!("Failed to send chat file: {e}"));
let error = match &active_session {
Some(session) => {
// Make the bytes fetchable by room members, then broadcast
// the descriptor alongside the (possibly empty) caption
// text. Re-serving the same id on a retry REPLACES the
// store entry (same Arc), never double-counts it.
session.transport.serve_attachment(attachment.id, data);
session
.room_state
.send_chat(text, Some(attachment))
.await
.err()
.map(|e| e.to_string())
}
None => Some("not in a room".to_string()),
};
if let Some(e) = &error {
crate::log_msg(&format!("Failed to send chat file: {e}"));
}
let _ = ui_tx
.send(UiEvent::ChatSendResult { local_id, error })
.await;
}
CoreCommand::FetchAttachment { from, attachment } => {
@@ -3121,7 +3408,11 @@ async fn run_core_loop(
.await;
}
CoreCommand::StartScreenShare { audio_app } => {
CoreCommand::StartScreenShare {
audio_app,
settings,
quality,
} => {
let Some(session) = &mut active_session else {
let _ = ui_tx
.send(UiEvent::Error(
@@ -3130,7 +3421,7 @@ async fn run_core_loop(
.await;
continue;
};
if session.screenshare_host.is_some() {
if session.teardown.is_sharing() {
continue; // already sharing
}
let bin = match crate::screenshare::pixelpass_path(pixelpass_override.as_deref()) {
@@ -3171,10 +3462,18 @@ async fn run_core_loop(
});
tx
});
match crate::screenshare::spawn_host(&bin, audio_app.as_deref(), notices).await {
match crate::screenshare::spawn_host(
&bin,
audio_app.as_deref(),
&settings,
quality,
notices,
)
.await
{
Ok((child, ticket)) => {
crate::log_msg("Screen share host started");
session.screenshare_host = Some(child);
session.teardown.set_host(child);
current_sharing = Some(ticket.clone());
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
@@ -3195,9 +3494,24 @@ async fn run_core_loop(
CoreCommand::StopScreenShare => {
current_sharing = None;
if let Some(session) = &mut active_session {
if let Some(mut child) = session.screenshare_host.take() {
let _ = child.kill().await;
crate::log_msg("Screen share host stopped");
match session.teardown.stop_host().await {
None => {}
Some(teardown::StopOutcome::Reaped) => {
crate::log_msg("Screen share host stopped");
}
// We gave up waiting rather than freeze the client, so
// pixelpass may still be alive and serving. Saying
// "stopped" and nothing else would be a lie the user
// cannot see through (round-16 review, P3-2).
Some(teardown::StopOutcome::Unconfirmed) => {
let _ = ui_tx
.send(UiEvent::Error(
"Couldn't confirm the screen-share process exited — \
it may still be sharing. Check for a stray pixelpass."
.into(),
))
.await;
}
}
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
@@ -3209,7 +3523,7 @@ async fn run_core_loop(
let _ = ui_tx.send(UiEvent::ScreenShareStopped).await;
}
CoreCommand::ViewShare(ticket) => {
CoreCommand::ViewShare { ticket, settings } => {
let bin = match crate::screenshare::pixelpass_path(pixelpass_override.as_deref()) {
Some(b) => b,
None => {
@@ -3221,11 +3535,23 @@ async fn run_core_loop(
continue;
}
};
match crate::screenshare::spawn_viewer(&bin, &ticket).await {
if let Some(session) = &mut active_session {
// Drop viewers whose player window has already closed so the
// list only tracks live players.
session.teardown.sweep_exited_viewers();
// One player per share: a second Watch click on a share we're
// already viewing is a retry (usually because the first window
// froze), so replace the existing player rather than stacking a
// second mpv — two players would double the shared audio.
if session.teardown.replace_viewer(&ticket).await {
crate::log_msg("Screen share viewer replaced (re-watch)");
}
}
match crate::screenshare::spawn_viewer(&bin, &ticket, &settings).await {
Ok(child) => {
crate::log_msg("Screen share viewer started");
if let Some(session) = &mut active_session {
session.screenshare_viewers.push(child);
session.teardown.push_viewer(ticket, child);
}
}
Err(e) => {
@@ -3238,17 +3564,41 @@ async fn run_core_loop(
}
}
// The command loop has exited, by any route. Tear the session down
// explicitly rather than letting it drop on the way out of this function:
// an implicit drop unloads the echo-cancel module without first reaping the
// pixelpass host (design v3.4 §7.2, decision D4).
//
// This sits *after* the loop rather than in the close arm on purpose. The
// impl plan pinned one teardown per channel-close arm, but the best-effort
// wake arm is unreachable by construction (see the comment at that arm), so
// that shape would have duplicated teardown to cover one live path and one
// dead one. Here every `break` is covered structurally, including any added
// later. Adjudication: impl plan §10, 2026-07-26.
if let Some(session) = active_session.take() {
session.shutdown(audio_backend.clone()).await;
}
Ok(())
}
/// Index of an existing viewer for `ticket` in the live-viewers list, if any.
/// A re-watch of the same share replaces that player instead of stacking a
/// second one — two players decoding the same stream would double the shared
/// audio. Generic over the child value so the dedup rule is unit-testable
/// without spawning real player processes.
fn replace_viewer_index<T>(viewers: &[(String, T)], ticket: &str) -> Option<usize> {
viewers.iter().position(|(t, _)| t == ticket)
}
#[cfg(test)]
mod tests {
use super::{
KnownPeers, MAX_OPUS_PAYLOAD, MAX_RETAINED_PEERS, MIC_LEVEL_REPORT_SAMPLES, MicLevelMeter,
PLAYBACK_HANDOFF_QUEUE_FRAMES, PeerSpeakTicket, admit_retained, apply_peer_volume,
apply_volume, audio_datagram_len_ok, coalesce_insert, coalesce_pop, frame_level,
mix_frames, mix_stereo_frames, next_game_change, send_playback_frame, should_auto_fetch,
stereo_to_mono,
NetworkMode, PLAYBACK_HANDOFF_QUEUE_FRAMES, PeerSpeakTicket, admit_retained,
apply_peer_volume, apply_volume, audio_datagram_len_ok, coalesce_insert, coalesce_pop,
frame_level, mix_frames, mix_stereo_frames, next_game_change, rebuild_with_fallback,
replace_viewer_index, send_playback_frame, should_auto_fetch, stereo_to_mono,
};
use crate::core::messages::{CoalesceKey, CoreCommand, coalesce_key};
use std::collections::{HashMap, HashSet};
@@ -3259,6 +3609,114 @@ mod tests {
iroh::SecretKey::generate().public()
}
#[test]
fn re_watch_replaces_existing_viewer_for_same_ticket() {
// The value type stands in for a viewer Child; only the ticket matters.
let viewers = vec![("ticket-A".to_string(), 0u8), ("ticket-B".to_string(), 1u8)];
// Re-watching an already-open share finds the existing player to replace.
assert_eq!(replace_viewer_index(&viewers, "ticket-A"), Some(0));
assert_eq!(replace_viewer_index(&viewers, "ticket-B"), Some(1));
// A different (new) share has nothing to replace — it opens fresh.
assert_eq!(replace_viewer_index(&viewers, "ticket-C"), None);
// Empty list: first watch of anything opens fresh.
assert_eq!(replace_viewer_index::<u8>(&[], "ticket-A"), None);
}
// --- rebuild_with_fallback: the retry policy behind replace_net_stack ---
// The builder is injected, so these cover the policy without sockets. The
// closure does its bookkeeping synchronously and returns a ready future.
#[tokio::test]
async fn rebuild_keeps_requested_posture_on_first_success() {
let calls = std::cell::RefCell::new(Vec::new());
let out = rebuild_with_fallback(
|mode| {
calls.borrow_mut().push(mode);
std::future::ready(Ok::<u8, String>(7))
},
NetworkMode::DirectOnly,
NetworkMode::N0Full,
)
.await;
assert_eq!(out, Ok((7, NetworkMode::DirectOnly, None)));
// No second build: the live posture is only a fallback.
assert_eq!(*calls.borrow(), vec![NetworkMode::DirectOnly]);
}
#[tokio::test]
async fn rebuild_falls_back_to_the_live_posture_when_the_requested_one_fails() {
let calls = std::cell::RefCell::new(Vec::new());
let out = rebuild_with_fallback(
|mode| {
calls.borrow_mut().push(mode);
std::future::ready(if mode == NetworkMode::DirectOnly {
Err("bind failed".to_string())
} else {
Ok(7u8)
})
},
NetworkMode::DirectOnly,
NetworkMode::N0Full,
)
.await;
// A stack is up on the OLD posture and the caller learns both that it
// fell back (mode) and why (the primary error) — no silent zombie.
assert_eq!(
out,
Ok((7, NetworkMode::N0Full, Some("bind failed".to_string())))
);
assert_eq!(
*calls.borrow(),
vec![NetworkMode::DirectOnly, NetworkMode::N0Full]
);
}
#[tokio::test]
async fn rebuild_reports_both_errors_when_networking_is_gone() {
let out = rebuild_with_fallback(
|_| std::future::ready(Err::<u8, String>("bind failed".to_string())),
NetworkMode::DirectOnly,
NetworkMode::N0Full,
)
.await;
assert_eq!(
out,
Err(("bind failed".to_string(), "bind failed".to_string()))
);
}
#[tokio::test]
async fn rebuild_with_equal_postures_is_a_plain_retry() {
// RegenerateIdentity rebuilds under the same mode: the fallback is a
// second attempt with identical parameters, not a posture change.
let calls = std::cell::Cell::new(0u8);
let out = rebuild_with_fallback(
|mode| {
calls.set(calls.get() + 1);
assert_eq!(mode, NetworkMode::RelayNoDiscovery);
std::future::ready(if calls.get() == 1 {
Err("transient".to_string())
} else {
Ok(7u8)
})
},
NetworkMode::RelayNoDiscovery,
NetworkMode::RelayNoDiscovery,
)
.await;
// Succeeded on the requested posture, so the caller treats the change
// as applied (the Some(err) is logged, not surfaced as a UI error).
assert_eq!(
out,
Ok((
7,
NetworkMode::RelayNoDiscovery,
Some("transient".to_string())
))
);
assert_eq!(calls.get(), 2);
}
#[test]
fn admit_retained_rejects_only_new_ids_at_the_cap() {
// Below the cap, a brand-new identity is retained.
@@ -3397,15 +3855,30 @@ mod tests {
#[test]
fn auto_fetch_only_for_roster_images_not_already_inflight() {
// The happy path: a roster author's brand-new image attachment.
assert!(should_auto_fetch(true, true, false));
const OK_SIZE: u64 = 1024;
// The happy path: a roster author's brand-new, small-enough image.
assert!(should_auto_fetch(true, true, false, OK_SIZE));
// A non-image (generic file) never auto-fetches — it waits for "Save".
assert!(!should_auto_fetch(false, true, false));
assert!(!should_auto_fetch(false, true, false, OK_SIZE));
// A non-roster author (e.g. a sock puppet that never announced) is rejected,
// closing the F-02 unbounded-task vector.
assert!(!should_auto_fetch(true, false, false));
assert!(!should_auto_fetch(true, false, false, OK_SIZE));
// An identical (author,id) already being fetched is deduped.
assert!(!should_auto_fetch(true, true, true));
assert!(!should_auto_fetch(true, true, true, OK_SIZE));
// The declared-size gate (Phase 3B): at the cap auto-fetches, the first
// byte over requires a click.
assert!(should_auto_fetch(
true,
true,
false,
crate::files::MAX_AUTO_IMAGE_BYTES
));
assert!(!should_auto_fetch(
true,
true,
false,
crate::files::MAX_AUTO_IMAGE_BYTES + 1
));
}
#[test]
+889
View File
@@ -0,0 +1,889 @@
//! Destruction-order guarantees for the screen-share children and the
//! echo-cancel module (phase 0b of the screenshare audio-exclusion plan;
//! design v3.4 §7.1–§7.2, decision D4).
//!
//! # The invariant
//!
//! > **The echo-cancel module must not unload while a pixelpass host is alive
//! > and fanning out.**
//!
//! If it does, the AEC's virtual nodes vanish from under a live pixelpass that
//! still holds link proxies and a stale module index. Phase 6 makes this sharp
//! — it is the first phase whose objects live only as long as pixelpass does —
//! so the ordering guarantee has to exist *before* it.
//!
//! Two paths have to honour it, and only one of them is code we get to run:
//!
//! 1. **The explicit path** — [`ScreenshareTeardown::shutdown_children`], awaited
//! by `ActiveSession::shutdown` before the guard is dropped.
//! 2. **The drop/unwind path** — nobody calls anything. The core has numerous
//! `unwrap()` sites and no `panic=abort` profile, so unwind is reachable, and
//! on that path the only thing standing between us and a violated invariant
//! is *field declaration order* plus [`ReapOnDrop`].
//!
//! Hence the two structural rules enforced here:
//!
//! - `echo_cancel` is the **last declared field** of [`ScreenshareTeardown`].
//! Rust drops fields in declaration order, so last-declared is last-dropped.
//! This is not a style choice; reversing it reintroduces the bug.
//! - Killing is not enough — a child must be **reaped**. `kill_on_drop(true)`
//! only *signals*; it hands the child to the runtime's orphan queue and
//! returns, which on an unwinding runtime may never be drained. [`ReapOnDrop`]
//! therefore blocks, briefly and boundedly, until the child is actually gone.
//!
//! Everything here is generic over [`ChildProcess`] and over the guard type so
//! the ordering is unit-testable without spawning processes or loading PipeWire
//! modules — the same seam idiom as `replace_viewer_index` and
//! `rebuild_with_fallback` in the parent module.
use std::future::Future;
use std::time::{Duration, Instant};
/// How long [`ReapOnDrop::drop`] will block waiting for a killed child to be
/// reaped before giving up and logging. This runs on the unwind path, so it is
/// a deliberate trade: a bounded stall is preferable to unloading the AEC out
/// from under a live pixelpass, and unbounded blocking in a `Drop` is not.
const REAP_BUDGET: Duration = Duration::from_millis(250);
/// Poll interval while waiting out [`REAP_BUDGET`].
const REAP_POLL: Duration = Duration::from_millis(5);
/// How long a child gets to honour the graceful stop before it is killed.
///
/// A healthy pixelpass exits in well under this, so the normal path never
/// spends it; only a wedged child does. It is awaited inline in the core
/// command loop, so it is also how long a wedged child can delay other
/// commands — hence seconds, not tens of seconds.
const STOP_GRACE: Duration = Duration::from_secs(2);
/// The child-process operations the teardown ordering actually depends on.
///
/// Deliberately narrow, and deliberately not `ExitStatus`-shaped: the ordering
/// rules care only about *whether* a child has been signalled and *whether* it
/// has been reaped, so the test double is a few lines instead of a fabricated
/// exit status.
pub(super) trait ChildProcess {
/// Ask the child to exit **gracefully**, so it can run its own cleanup.
/// Does **not** wait, and is not guaranteed to be honoured.
fn request_stop(&mut self) -> std::io::Result<()>;
/// Signal the child to die. Does **not** wait.
fn start_kill(&mut self) -> std::io::Result<()>;
/// Poll once. `true` once the child has exited **and been reaped**.
fn try_reap(&mut self) -> bool;
/// Wait until the child has exited and been reaped.
///
/// The `io::Result` is load-bearing and must not be discarded by callers:
/// a failed wait is *not* a confirmed reap, and treating it as one is how
/// the AEC ends up unloading over a live child.
fn wait_reaped(&mut self) -> impl Future<Output = std::io::Result<()>> + Send;
}
impl ChildProcess for tokio::process::Child {
/// **SIGINT, not SIGTERM.** pixelpass installs only a `tokio::signal::ctrl_c()`
/// handler (`pixelpass/src/common/signal.rs`), so SIGTERM would be the default
/// disposition — instant death, no cleanup — which is indistinguishable from
/// SIGKILL for our purposes.
///
/// Signalling by pid is safe against pid reuse here because we have not
/// reaped this child: an exited-but-unreaped child is a zombie whose pid the
/// kernel reserves until we `wait` it, so the pid cannot name a stranger.
#[cfg(unix)]
fn request_stop(&mut self) -> std::io::Result<()> {
let Some(pid) = self.id() else {
// Already reaped — nothing to signal.
return Ok(());
};
// SAFETY: `kill` is async-signal-safe and takes no pointers; the pid is
// this process's own unreaped child (see above).
if unsafe { libc::kill(pid as libc::pid_t, libc::SIGINT) } == 0 {
Ok(())
} else {
Err(std::io::Error::last_os_error())
}
}
/// Windows has no SIGINT to send to another process without attaching to its
/// console, so the graceful request degrades to the hard kill and the
/// bounded wait below simply returns early.
#[cfg(not(unix))]
fn request_stop(&mut self) -> std::io::Result<()> {
tokio::process::Child::start_kill(self)
}
fn start_kill(&mut self) -> std::io::Result<()> {
tokio::process::Child::start_kill(self)
}
fn try_reap(&mut self) -> bool {
matches!(self.try_wait(), Ok(Some(_)))
}
async fn wait_reaped(&mut self) -> std::io::Result<()> {
self.wait().await.map(|_| ())
}
}
/// Did the explicit stop path actually confirm the child was reaped?
///
/// The distinction is not cosmetic: on [`Unconfirmed`](Self::Unconfirmed) we
/// deliberately stopped waiting (see [`ReapOnDrop::shutdown`]), so pixelpass may
/// still be alive and fanning out. A user-initiated Stop Share must not report
/// that as a clean stop.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[must_use = "an unconfirmed stop means the child may still be sharing"]
pub(super) enum StopOutcome {
/// The child is gone and has been reaped.
Reaped,
/// We could not confirm the reap within the bound and gave up waiting.
Unconfirmed,
}
/// A child that is killed **and reaped** when it is dropped.
///
/// The explicit path calls [`shutdown`](Self::shutdown), which releases the
/// child only once its reap is *confirmed*, so the `Drop` below is a no-op
/// afterwards but stays armed through every await until then. `Drop` is the
/// last-ditch protection for the panic/unwind/cancellation paths.
pub(super) struct ReapOnDrop<C: ChildProcess> {
/// `None` once the child has been reaped through the explicit path.
child: Option<C>,
/// Names the child in the reap-timeout log line.
label: &'static str,
}
impl<C: ChildProcess> ReapOnDrop<C> {
pub(super) fn new(child: C, label: &'static str) -> Self {
Self {
child: Some(child),
label,
}
}
/// Poll once, without killing. `true` if the child has exited on its own —
/// used to sweep player windows the user has already closed.
pub(super) fn has_exited(&mut self) -> bool {
match &mut self.child {
Some(child) => {
if child.try_reap() {
self.child = None;
true
} else {
false
}
}
// Already reaped through the explicit path.
None => true,
}
}
/// Stop the child gracefully if it will go, and by force if it will not.
/// Waits for it to be reaped either way. Idempotent.
///
/// Ask, then insist (design v3.4 §7.4): a pixelpass host that gets SIGINT
/// unloads its capture sink on the way out, whereas SIGKILL skips that and
/// leaks a null-sink module on every Stop Share.
///
/// The wait is the point: returning after signalling would let the caller
/// proceed to unload the AEC while the child is still running.
///
/// ⚠️ The child stays owned by `self` across every `.await`, and is released
/// **only after a confirmed reap**. Taking it out first would disarm the
/// `Drop` fallback for exactly as long as the wait lasts: cancel or unwind
/// this future at that moment and the raw child would drop with nothing but
/// `kill_on_drop` (which signals without reaping) while `Drop` below found
/// `None` and did nothing — the precise hole this type exists to close.
pub(super) async fn shutdown(&mut self) -> StopOutcome {
let Some(child) = self.child.as_mut() else {
return StopOutcome::Reaped;
};
// Three different things can go wrong here and they want three
// different operator diagnoses: the signal never left (a runtime or
// permission fault), the child ignored it (a wedged pixelpass), or the
// wait itself broke (we no longer know anything about the child).
// Collapsing them into one line was P3-1 of the round-16 review.
if let Err(e) = child.request_stop() {
crate::log_msg(&format!(
"teardown: could not ask {} to stop: {e}",
self.label
));
}
match tokio::time::timeout(STOP_GRACE, child.wait_reaped()).await {
Ok(Ok(())) => {
self.child = None;
return StopOutcome::Reaped;
}
Ok(Err(e)) => crate::log_msg(&format!(
"teardown: waiting for {} failed ({e}); killing it",
self.label
)),
Err(_) => crate::log_msg(&format!(
"teardown: {} ignored the graceful stop within {STOP_GRACE:?}; killing it",
self.label
)),
}
if let Err(e) = child.start_kill() {
crate::log_msg(&format!(
"teardown: {} could not be killed: {e}",
self.label
));
}
// The second wait is bounded too. An unbounded one lets a process stuck
// in uninterruptible sleep wedge the core command loop forever, and a
// permanently frozen app is a worse failure than the risk below.
if let Ok(Ok(())) = tokio::time::timeout(STOP_GRACE, child.wait_reaped()).await {
self.child = None;
return StopOutcome::Reaped;
}
// Explicit policy for the one case where the two guarantees conflict:
// we could not confirm the reap and will NOT block indefinitely, so we
// give up availability-first and leave the child owned — `Drop`'s
// bounded retry stays armed, and the AEC may unload over a child that
// is still somehow alive. That residual risk is logged, not silent —
// and, for a user-initiated stop, reported to the caller rather than
// dressed up as success.
crate::log_msg(&format!(
"teardown: {} could not be confirmed dead; the echo-cancel module \
may unload while it lives",
self.label
));
StopOutcome::Unconfirmed
}
/// Is the `Drop` fallback still armed? Test-only: the arming rule is the
/// whole point of holding the child across the waits.
#[cfg(test)]
fn is_armed(&self) -> bool {
self.child.is_some()
}
}
impl<C: ChildProcess> Drop for ReapOnDrop<C> {
fn drop(&mut self) {
let Some(child) = self.child.as_mut() else {
return;
};
let _ = child.start_kill();
// `Drop` cannot await, so poll on a bounded budget. See `REAP_BUDGET`.
let deadline = Instant::now() + REAP_BUDGET;
loop {
if child.try_reap() {
return;
}
if Instant::now() >= deadline {
crate::log_msg(&format!(
"teardown: {} did not exit within the reap budget; \
continuing (the echo-cancel module may unload while it lives)",
self.label
));
return;
}
std::thread::sleep(REAP_POLL);
}
}
}
/// Everything in an `ActiveSession` whose **destruction order** is load-bearing.
///
/// ⚠️ Field order below **is** the invariant. `echo_cancel` is declared last so
/// it is dropped last, after every screen-share child has been killed and
/// reaped. Do not reorder these fields.
pub(super) struct ScreenshareTeardown<C: ChildProcess, G> {
/// Our pixelpass screen-share host child while sharing.
host: Option<ReapOnDrop<C>>,
/// pixelpass viewer children we spawned to watch peers' shares, each paired
/// with the share ticket it is viewing so a re-watch of the same share can
/// replace (not stack) its player.
viewers: Vec<(String, ReapOnDrop<C>)>,
/// Loaded PipeWire echo-cancel module (if enabled); unloads on drop.
///
/// ⚠️ **LAST FIELD ON PURPOSE** — see the module docs and the struct note.
///
/// Never read, and that is the design: the guard is held only so that its
/// `Drop` runs, and only so that it runs *here*, last. `dead_code` is right
/// that nothing reads it and wrong that it does nothing.
#[allow(dead_code)]
echo_cancel: Option<G>,
}
impl<C: ChildProcess, G> ScreenshareTeardown<C, G> {
pub(super) fn new(echo_cancel: Option<G>) -> Self {
Self {
host: None,
viewers: Vec::new(),
echo_cancel,
}
}
pub(super) fn is_sharing(&self) -> bool {
self.host.is_some()
}
pub(super) fn set_host(&mut self, child: C) {
self.host = Some(ReapOnDrop::new(child, "screen-share host"));
}
/// Stop sharing: kill the host and wait for it to be reaped. `None` if we
/// were not sharing; otherwise whether the reap was actually confirmed —
/// the caller owns telling the user, since an unconfirmed stop may leave
/// pixelpass fanning out after the UI says sharing ended.
pub(super) async fn stop_host(&mut self) -> Option<StopOutcome> {
let mut host = self.host.take()?;
Some(host.shutdown().await)
}
/// Drop viewers whose player window has already closed, so the list only
/// tracks live players.
pub(super) fn sweep_exited_viewers(&mut self) {
self.viewers.retain_mut(|(_, child)| !child.has_exited());
}
/// Kill and reap the viewer already showing `ticket`, if any, so a re-watch
/// replaces its player instead of stacking a second one.
pub(super) async fn replace_viewer(&mut self, ticket: &str) -> bool {
let Some(pos) = super::replace_viewer_index(&self.viewers, ticket) else {
return false;
};
let (_, mut old) = self.viewers.remove(pos);
// A viewer is our own player window, not the thing peers are watching:
// an unconfirmed reap is already logged, and there is no user decision
// riding on it the way there is for Stop Share.
let _ = old.shutdown().await;
true
}
pub(super) fn push_viewer(&mut self, ticket: String, child: C) {
self.viewers
.push((ticket, ReapOnDrop::new(child, "screen-share viewer")));
}
/// Kill and reap **every** screen-share child, host first so viewers see the
/// stream end promptly.
///
/// The caller must await this before the echo-cancel guard is dropped. On
/// the drop/unwind path nothing calls it and field order carries the
/// invariant instead.
pub(super) async fn shutdown_children(&mut self) {
// Outcomes are discarded on purpose: this runs on the session/teardown
// path, where the policy is already availability-first and the residual
// risk is logged by `shutdown` itself. There is no user still waiting
// on an answer here, unlike `stop_host`.
if let Some(host) = &mut self.host {
let _ = host.shutdown().await;
}
self.host = None;
for (_, viewer) in self.viewers.iter_mut() {
let _ = viewer.shutdown().await;
}
self.viewers.clear();
}
}
#[cfg(test)]
mod tests {
use super::{ChildProcess, ReapOnDrop, STOP_GRACE, ScreenshareTeardown, StopOutcome};
use std::future::Future;
use std::sync::{Arc, Mutex};
use std::time::Duration;
type Log = Arc<Mutex<Vec<String>>>;
fn log() -> Log {
Arc::new(Mutex::new(Vec::new()))
}
fn entries(log: &Log) -> Vec<String> {
log.lock().unwrap().clone()
}
fn position(log: &Log, entry: &str) -> Option<usize> {
entries(log).iter().position(|e| e == entry)
}
/// Records the events the ordering rules turn on. Death is gated on an
/// actual signal, so the double cannot report a reap that nothing caused.
struct FakeChild {
log: Log,
label: &'static str,
interrupted: bool,
killed: bool,
reaped: bool,
/// A well-behaved child exits on SIGINT. A wedged one ignores it and
/// dies only to SIGKILL.
honours_interrupt: bool,
/// When true the child is already dead before anyone signals it — the
/// closed-player-window case that `sweep_exited_viewers` looks for.
exited_on_its_own: bool,
/// Death is not instantaneous: `try_reap` reports the child alive this
/// many more times before it goes.
polls_before_death: u32,
/// `wait` reports an error instead of a reap.
wait_fails: bool,
}
impl FakeChild {
/// A well-behaved child: exits when asked.
fn new(log: &Log, label: &'static str) -> Self {
Self {
log: log.clone(),
label,
interrupted: false,
killed: false,
reaped: false,
honours_interrupt: true,
exited_on_its_own: false,
polls_before_death: 0,
wait_fails: false,
}
}
/// A child that ignores the graceful stop entirely.
fn wedged(log: &Log, label: &'static str) -> Self {
Self {
honours_interrupt: false,
..Self::new(log, label)
}
}
/// A child that does not die the instant it is signalled: `try_reap`
/// reports it alive for `polls` calls first. Without this the `Drop`
/// polling loop could be replaced by a single `try_reap` and no test
/// would notice.
fn reaps_after_polls(log: &Log, label: &'static str, polls: u32) -> Self {
Self {
polls_before_death: polls,
..Self::new(log, label)
}
}
/// A child that ignores SIGINT *and* does not die the instant it is
/// killed — the only shape that lets a test reach the post-SIGKILL
/// wait and still be reaped by the `Drop` poll loop afterwards.
fn wedged_then_dies_after_polls(log: &Log, label: &'static str, polls: u32) -> Self {
Self {
honours_interrupt: false,
polls_before_death: polls,
..Self::new(log, label)
}
}
/// A child whose `wait` fails. A failed wait is not a confirmed reap,
/// so it must not be reported as one.
fn wait_fails(log: &Log, label: &'static str) -> Self {
Self {
wait_fails: true,
..Self::new(log, label)
}
}
fn already_exited(log: &Log, label: &'static str) -> Self {
Self {
exited_on_its_own: true,
..Self::new(log, label)
}
}
/// Has anything actually made this child exit yet? A signalled child
/// still has to burn through `polls_before_death` first.
fn is_dead(&self) -> bool {
let signalled = self.killed
|| self.exited_on_its_own
|| (self.interrupted && self.honours_interrupt);
signalled && self.polls_before_death == 0
}
/// One observation of a dying-but-not-yet-dead child.
fn tick(&mut self) {
self.polls_before_death = self.polls_before_death.saturating_sub(1);
}
fn record(&self, event: &str) {
self.log
.lock()
.unwrap()
.push(format!("{}:{event}", self.label));
}
fn mark_reaped(&mut self) {
if !self.reaped {
self.reaped = true;
self.record("reap");
}
}
}
impl ChildProcess for FakeChild {
fn request_stop(&mut self) -> std::io::Result<()> {
if !self.interrupted {
self.interrupted = true;
self.record("sigint");
}
Ok(())
}
fn start_kill(&mut self) -> std::io::Result<()> {
if !self.killed {
self.killed = true;
self.record("kill");
}
Ok(())
}
fn try_reap(&mut self) -> bool {
if self.is_dead() {
self.mark_reaped();
return true;
}
self.tick();
false
}
/// Pending until something actually kills the child, so a wedged child
/// really does make the caller wait out `STOP_GRACE`. No waker is
/// registered: under `start_paused` the runtime auto-advances its clock
/// when every task is idle, which is exactly what fires the timeout.
fn wait_reaped(&mut self) -> impl Future<Output = std::io::Result<()>> + Send {
std::future::poll_fn(move |_cx| {
if self.wait_fails {
return std::task::Poll::Ready(Err(std::io::Error::other("wait failed")));
}
if self.is_dead() {
self.mark_reaped();
std::task::Poll::Ready(Ok(()))
} else {
std::task::Poll::Pending
}
})
}
}
/// Stands in for `EchoCancelGuard`, whose real `Drop` runs `pactl unload`.
struct FakeAec(Log);
impl Drop for FakeAec {
fn drop(&mut self) {
self.0.lock().unwrap().push("aec:unload".to_string());
}
}
fn teardown(log: &Log) -> ScreenshareTeardown<FakeChild, FakeAec> {
ScreenshareTeardown::new(Some(FakeAec(log.clone())))
}
// --- The drop/unwind path: field order + ReapOnDrop carry the invariant ---
/// Mutation gate #5 (remove the reap loop from `ReapOnDrop::drop`).
///
/// Asserts only that dropping a guard reaps, and reaps *after* killing —
/// deliberately says nothing about the AEC, so reversing the struct's field
/// order leaves this test green and only the ordering test below fails.
#[test]
fn dropping_a_guard_kills_and_then_reaps_the_child() {
let log = log();
drop(ReapOnDrop::new(FakeChild::new(&log, "host"), "host"));
assert_eq!(entries(&log), vec!["host:kill", "host:reap"]);
}
/// Mutation gate #4 (reverse the field order of `ScreenshareTeardown`).
///
/// Asserts only kill-before-unload, so removing the reap loop leaves this
/// test green and only the reap test above fails.
#[test]
fn the_aec_unloads_after_the_children_on_the_drop_path() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
t.push_viewer("ticket-A".to_string(), FakeChild::new(&log, "viewer"));
drop(t);
let unload = position(&log, "aec:unload").expect("the AEC guard must be dropped");
let host_kill = position(&log, "host:kill").expect("the host must be killed");
let viewer_kill = position(&log, "viewer:kill").expect("the viewer must be killed");
assert!(
host_kill < unload,
"the AEC unloaded while the host was alive: {:?}",
entries(&log)
);
assert!(
viewer_kill < unload,
"the AEC unloaded while a viewer was alive: {:?}",
entries(&log)
);
}
/// The whole invariant in one sequence, as documentation.
#[test]
fn the_drop_path_reaps_every_child_before_unloading_the_aec() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
drop(t);
assert_eq!(entries(&log), vec!["host:kill", "host:reap", "aec:unload"]);
}
// --- The explicit path: ask, then insist ---
/// A healthy child must be *asked*, never killed. If Stop Share went
/// straight to SIGKILL, pixelpass would skip its own cleanup and leak a
/// null-sink module every time (design v3.4 §7.4).
#[tokio::test]
async fn a_healthy_child_is_asked_to_stop_and_never_killed() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
assert_eq!(t.stop_host().await, Some(StopOutcome::Reaped));
assert_eq!(entries(&log), vec!["host:sigint", "host:reap"]);
assert!(
!entries(&log).contains(&"host:kill".to_string()),
"a child that honoured the graceful stop must not be killed: {:?}",
entries(&log)
);
}
/// ...but a child that ignores the request must not be able to hold the
/// session open forever: the grace is bounded and SIGKILL follows.
#[tokio::test(start_paused = true)]
async fn a_wedged_child_is_killed_once_the_grace_expires() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::wedged(&log, "host"));
// The outer bound turns "the fallback was removed" into a failure
// rather than a hung test. Under `start_paused` no real time passes.
let start = tokio::time::Instant::now();
tokio::time::timeout(Duration::from_secs(60), t.stop_host())
.await
.expect("a wedged child must not block teardown indefinitely");
assert_eq!(entries(&log), vec!["host:sigint", "host:kill", "host:reap"]);
assert!(
start.elapsed() >= STOP_GRACE,
"the child must actually be given the grace period, waited {:?}",
start.elapsed()
);
}
/// The assertion above compares elapsed time against `STOP_GRACE` itself,
/// so it stays vacuously true if the constant is set to zero — both sides
/// move together. Pin the constant independently: the whole point of the
/// graceful stop is that pixelpass gets a real interval in which to unload
/// its capture sink, and zero is not one.
#[test]
fn the_grace_is_a_real_interval() {
assert!(
STOP_GRACE >= Duration::from_millis(500),
"too short to let pixelpass tear its pipeline down: {STOP_GRACE:?}"
);
// ...and short enough that a wedged child cannot visibly stall the core
// command loop, which awaits this inline.
assert!(
STOP_GRACE <= Duration::from_secs(5),
"long enough to freeze the UI's command handling: {STOP_GRACE:?}"
);
}
/// The hole the whole type exists to close, and the one place the old
/// implementation left open: if `shutdown` is cancelled while waiting, the
/// child must still be owned, so dropping the guard still kills and reaps.
#[tokio::test(start_paused = true)]
async fn cancelling_shutdown_mid_wait_leaves_the_fallback_armed() {
let log = log();
let mut guard = ReapOnDrop::new(FakeChild::wedged(&log, "host"), "host");
// Cancel well inside the grace, while it is still waiting.
assert!(
tokio::time::timeout(STOP_GRACE / 4, guard.shutdown())
.await
.is_err(),
"the wedged child should still have been waiting when we cancelled"
);
assert!(
guard.is_armed(),
"a cancelled shutdown must not disarm the drop fallback"
);
drop(guard);
assert_eq!(entries(&log), vec!["host:sigint", "host:kill", "host:reap"]);
}
/// The test above only ever cancels during the *graceful* wait, so a
/// mutation that disarmed the wrapper between the two waits would survive
/// it (round-16 review, P3-3). This one cancels during the post-SIGKILL
/// wait — the window where we have already given up on cooperation and the
/// `Drop` fallback is the only thing left.
#[tokio::test(start_paused = true)]
async fn cancelling_shutdown_after_the_kill_leaves_the_fallback_armed() {
let log = log();
// Ignores SIGINT, so the grace expires and we reach the kill; then
// survives three polls, so the second wait is still pending when we
// cancel, and the drop loop still gets to reap it.
let mut guard = ReapOnDrop::new(
FakeChild::wedged_then_dies_after_polls(&log, "host", 3),
"host",
);
assert!(
tokio::time::timeout(STOP_GRACE + STOP_GRACE / 4, guard.shutdown())
.await
.is_err(),
"we should have been cancelled inside the post-kill wait"
);
assert_eq!(
entries(&log),
vec!["host:sigint", "host:kill"],
"the graceful stop must have expired and escalated before we cancelled"
);
assert!(
guard.is_armed(),
"cancelling after the kill must not disarm the drop fallback either"
);
drop(guard);
// The fake's `start_kill` is idempotent, so `Drop` re-signalling an
// already-killed child adds no entry; the *reap* is what proves the
// fallback ran to completion after we abandoned the wait.
assert_eq!(
entries(&log),
vec!["host:sigint", "host:kill", "host:reap"],
"Drop must poll until the child is actually gone"
);
}
/// A failed wait is not a reap. Reporting it as one is how the AEC ends up
/// unloading over a child that is still alive.
#[tokio::test(start_paused = true)]
async fn a_failed_wait_is_not_treated_as_a_confirmed_reap() {
let log = log();
let mut guard = ReapOnDrop::new(FakeChild::wait_fails(&log, "host"), "host");
assert_eq!(
guard.shutdown().await,
StopOutcome::Unconfirmed,
"a stop we could not confirm must not be reported as a clean one"
);
assert!(
!entries(&log).contains(&"host:reap".to_string()),
"nothing confirmed the reap: {:?}",
entries(&log)
);
assert!(
entries(&log).contains(&"host:kill".to_string()),
"a child that would not stop must still be escalated: {:?}",
entries(&log)
);
assert!(
guard.is_armed(),
"an unconfirmed reap must leave the drop fallback armed"
);
}
/// Death is not instantaneous, so the drop path has to keep polling. A
/// single `try_reap` in place of the loop must not pass.
#[test]
fn the_drop_path_polls_until_the_child_is_actually_gone() {
let log = log();
drop(ReapOnDrop::new(
FakeChild::reaps_after_polls(&log, "host", 3),
"host",
));
assert_eq!(entries(&log), vec!["host:kill", "host:reap"]);
}
/// Mutation gate #3 (remove the wait after the host kill).
#[tokio::test]
async fn explicit_shutdown_reaps_the_host_before_the_aec_can_unload() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
t.push_viewer("ticket-A".to_string(), FakeChild::new(&log, "viewer"));
t.shutdown_children().await;
// Reaped by the explicit path — before the guard is anywhere near dropped.
assert_eq!(
entries(&log),
vec!["host:sigint", "host:reap", "viewer:sigint", "viewer:reap"],
"children must be stopped and reaped by the explicit path"
);
drop(t);
let unload = position(&log, "aec:unload").expect("the AEC guard must be dropped");
let host_reap = position(&log, "host:reap").expect("the host must be reaped");
assert!(host_reap < unload);
}
#[tokio::test]
async fn explicit_shutdown_is_idempotent_with_the_drop_path() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
t.shutdown_children().await;
drop(t);
// Exactly one stop and one reap: the drop path must not re-signal a
// child the explicit path already took.
assert_eq!(
entries(&log),
vec!["host:sigint", "host:reap", "aec:unload"]
);
}
// --- Host/viewer bookkeeping ---
#[tokio::test]
async fn stop_host_reports_whether_it_was_sharing() {
let log = log();
let mut t = teardown(&log);
assert!(!t.is_sharing());
assert_eq!(t.stop_host().await, None, "not sharing: nothing to stop");
t.set_host(FakeChild::new(&log, "host"));
assert!(t.is_sharing());
assert_eq!(t.stop_host().await, Some(StopOutcome::Reaped));
assert!(!t.is_sharing());
assert_eq!(entries(&log), vec!["host:sigint", "host:reap"]);
}
#[test]
fn sweeping_drops_only_the_players_that_already_closed() {
let log = log();
let mut t = teardown(&log);
t.push_viewer(
"closed".to_string(),
FakeChild::already_exited(&log, "closed"),
);
t.push_viewer("live".to_string(), FakeChild::new(&log, "live"));
t.sweep_exited_viewers();
// The live player survives the sweep; only the closed one is dropped,
// and dropping it must not kill anything (it was already gone).
assert_eq!(t.viewers.len(), 1);
assert_eq!(t.viewers[0].0, "live");
assert_eq!(entries(&log), vec!["closed:reap"]);
}
#[tokio::test]
async fn re_watching_a_share_replaces_that_player_only() {
let log = log();
let mut t = teardown(&log);
t.push_viewer("ticket-A".to_string(), FakeChild::new(&log, "a"));
t.push_viewer("ticket-B".to_string(), FakeChild::new(&log, "b"));
assert!(t.replace_viewer("ticket-A").await);
assert_eq!(entries(&log), vec!["a:sigint", "a:reap"]);
assert_eq!(t.viewers.len(), 1);
assert_eq!(t.viewers[0].0, "ticket-B");
// A share we are not watching has nothing to replace.
assert!(!t.replace_viewer("ticket-C").await);
}
}
+341 -10
View File
@@ -22,6 +22,22 @@ pub const MAX_ATTACHMENT_BYTES: u64 = 25 * 1024 * 1024;
/// the byte cap. Applied via `image::Limits` when validating/decoding.
pub const MAX_IMAGE_PX: u32 = 4096;
/// Max total decoded pixels, applied on top of the per-side [`MAX_IMAGE_PX`]
/// limit. The per-side cap alone still admits a 4096×4096 ≈ 16.8 MP bitmap
/// (~64 MiB transient RGBA); this bounds the worst-case decode allocation while
/// still clearing common 12 MP phone photos (4032×3024 ≈ 12.2 MP).
pub const MAX_IMAGE_TOTAL_PIXELS: u64 = 14_000_000;
/// Max pixels per side of the downscaled inline preview handed to the renderer.
/// Original bytes are kept only for Save; the chat column never needs more than
/// this (it displays at ~260 px, and the lightbox at window size).
pub const IMAGE_PREVIEW_MAX_SIDE: u32 = 1600;
/// Largest declared size an image attachment may auto-fetch at. Anything larger
/// (or any skipped/evicted image) renders a "Load image" button instead; a
/// manual click may use the full [`MAX_ATTACHMENT_BYTES`] cap.
pub const MAX_AUTO_IMAGE_BYTES: u64 = 4 * 1024 * 1024;
/// Longest filename we keep and display. Keeps the gossip descriptor compact and
/// the UI tidy; the real bytes are unaffected.
pub const MAX_FILENAME_LEN: usize = 96;
@@ -75,8 +91,14 @@ pub fn sanitize_filename(raw: &str) -> String {
.unwrap_or("")
.trim();
// Drop control chars; turn other whitespace into single spaces later.
let cleaned: String = base.chars().filter(|c| !c.is_control()).collect();
// Drop control chars and the same bidi/zero-width spoofing format chars
// stripped from display names (a U+202E override can visually reverse an
// extension, e.g. "photo\u{202E}gnp.exe" renders as "photoexe.png").
// Ordinary non-ASCII filenames pass through untouched.
let cleaned: String = base
.chars()
.filter(|c| !c.is_control() && !crate::sanitize::is_spoofing_format_char(*c))
.collect();
let collapsed = cleaned.split_whitespace().collect::<Vec<_>>().join(" ");
let collapsed = collapsed.trim_matches('.').trim();
@@ -164,20 +186,185 @@ pub fn classify(bytes: &[u8]) -> AttachmentKind {
/// regardless of [`MAX_ATTACHMENT_BYTES`]. Only PNG/JPEG are buildable in our
/// `image` feature set; anything else returns `None` and the caller shows a chip.
pub fn validate_image_bytes(bytes: &[u8]) -> Option<(u32, u32)> {
let mut limits = image::Limits::default();
limits.max_image_width = Some(MAX_IMAGE_PX);
limits.max_image_height = Some(MAX_IMAGE_PX);
let img = decode_image_bounded(bytes)?;
Some((img.width(), img.height()))
}
/// Shared bounded decode: header-check the dimensions (per-side AND total-pixel
/// limits) BEFORE decoding, then decode under `image::Limits` as defense in
/// depth. The precheck reads only the container header, so an over-limit bomb is
/// rejected without paying its decode cost.
fn decode_image_bounded(bytes: &[u8]) -> Option<image::DynamicImage> {
let reader = image::ImageReader::new(std::io::Cursor::new(bytes))
.with_guessed_format()
.ok()?;
let mut reader = reader;
reader.limits(limits);
let img = reader.decode().ok()?;
let (w, h) = (img.width(), img.height());
let (w, h) = reader.into_dimensions().ok()?;
if w == 0 || h == 0 || w > MAX_IMAGE_PX || h > MAX_IMAGE_PX {
return None;
}
Some((w, h))
if u64::from(w) * u64::from(h) > MAX_IMAGE_TOTAL_PIXELS {
return None;
}
let mut limits = image::Limits::default();
limits.max_image_width = Some(MAX_IMAGE_PX);
limits.max_image_height = Some(MAX_IMAGE_PX);
let mut reader = image::ImageReader::new(std::io::Cursor::new(bytes))
.with_guessed_format()
.ok()?;
reader.limits(limits);
let img = reader.decode().ok()?;
// Decoded size must match the header the precheck approved.
if img.width() != w || img.height() != h {
return None;
}
Some(img)
}
/// A decoded, display-ready inline preview: RGBA pixels downscaled so neither
/// side exceeds [`IMAGE_PREVIEW_MAX_SIDE`]. `rgba.len() == width * height * 4`,
/// which is also the preview's decoded-budget weight in the attachment cache.
pub struct ImagePreview {
pub width: u32,
pub height: u32,
pub rgba: Vec<u8>,
}
/// Decode image bytes under the same limits as [`validate_image_bytes`] and
/// build the downscaled inline preview. The full-resolution bitmap exists only
/// transiently here; the renderer is never handed more than
/// [`IMAGE_PREVIEW_MAX_SIDE`]² pixels. Returns `None` for anything that fails
/// validation (caller falls back to a chip / failure row).
pub fn decode_preview(bytes: &[u8]) -> Option<ImagePreview> {
let img = decode_image_bounded(bytes)?;
let img = if img.width() > IMAGE_PREVIEW_MAX_SIDE || img.height() > IMAGE_PREVIEW_MAX_SIDE {
// `thumbnail` preserves aspect ratio within the bounding box.
img.thumbnail(IMAGE_PREVIEW_MAX_SIDE, IMAGE_PREVIEW_MAX_SIDE)
} else {
img
};
let rgba = img.into_rgba8();
let (width, height) = (rgba.width(), rgba.height());
Some(ImagePreview {
width,
height,
rgba: rgba.into_raw(),
})
}
/// Estimated decoded RGBA cost of a preview, the weight counted against the
/// attachment cache's decoded-byte budget (`width * height * 4`).
pub fn preview_rgba_cost(width: u32, height: u32) -> usize {
(width as usize)
.saturating_mul(height as usize)
.saturating_mul(4)
}
/// Read at most [`MAX_ATTACHMENT_BYTES`] bytes from `r`. Returns `Ok(None)` if
/// the source holds even one byte more (detected by reading cap + 1), so a huge
/// or unbounded source is never fully buffered. Pure over `Read` for tests; the
/// picker wraps it via [`read_file_capped`].
pub fn read_capped<R: std::io::Read>(r: R) -> std::io::Result<Option<Vec<u8>>> {
use std::io::Read as _;
let mut buf = Vec::new();
let mut limited = r.take(MAX_ATTACHMENT_BYTES + 1);
limited.read_to_end(&mut buf)?;
if buf.len() as u64 > MAX_ATTACHMENT_BYTES {
return Ok(None);
}
Ok(Some(buf))
}
/// Read a picked file, bounded by [`MAX_ATTACHMENT_BYTES`]. Checks metadata
/// first to reject an obviously-oversized file without opening it, but keeps the
/// bounded read regardless — metadata can race (the file can grow after the
/// check) or be unavailable through a portal. `Ok(None)` = over the cap.
pub fn read_file_capped(path: &std::path::Path) -> std::io::Result<Option<Vec<u8>>> {
if let Ok(meta) = std::fs::metadata(path)
&& meta.len() > MAX_ATTACHMENT_BYTES
{
return Ok(None);
}
read_capped(std::fs::File::open(path)?)
}
/// Cap on how many blobs the session serve store retains at once (sent chat
/// attachments plus the current/next broadcast music tracks).
pub const SERVED_FILES_MAX_ENTRIES: usize = 16;
/// Byte budget for the serve store. Without it, a sender's own session could
/// grow unbounded at up to [`MAX_ATTACHMENT_BYTES`] per send (Phase 3C).
pub const SERVED_FILES_MAX_BYTES: usize = 128 * 1024 * 1024;
/// Count- and byte-budgeted FIFO store of blobs we serve to room members over
/// the file plane. Evicting an id makes a later request for it read as an empty
/// body — the existing "sender no longer has the file" response — never stale
/// or aliased bytes. Pure (no locks/IO) so budgets are unit-testable; the
/// transport wraps it in its own mutex.
#[derive(Debug, Default)]
pub struct ServeStore {
entries: std::collections::HashMap<AttachmentId, std::sync::Arc<Vec<u8>>>,
/// Present ids in insertion order; the front is the eviction candidate.
order: std::collections::VecDeque<AttachmentId>,
total_bytes: usize,
}
impl ServeStore {
/// Insert or replace a blob, evicting oldest entries until the count and
/// byte budgets fit. Replacement keeps the id's age and subtracts the old
/// bytes before the new ones are counted. Returns `false` for a blob that
/// alone exceeds the byte budget (not stored; an existing entry under the
/// id is dropped rather than left stale).
pub fn insert(&mut self, id: AttachmentId, bytes: std::sync::Arc<Vec<u8>>) -> bool {
if let Some(old) = self.entries.get(&id) {
self.total_bytes -= old.len();
}
if bytes.len() > SERVED_FILES_MAX_BYTES {
if self.entries.remove(&id).is_some() {
self.order.retain(|k| k != &id);
}
return false;
}
let replacing = self.entries.contains_key(&id);
loop {
let count_full = !replacing && self.entries.len() >= SERVED_FILES_MAX_ENTRIES;
let bytes_full = self.total_bytes + bytes.len() > SERVED_FILES_MAX_BYTES;
if !count_full && !bytes_full {
break;
}
let Some(victim) = self.order.iter().find(|k| **k != id).copied() else {
break;
};
self.remove(&victim);
}
if !replacing {
self.order.push_back(id);
}
self.total_bytes += bytes.len();
self.entries.insert(id, bytes);
true
}
pub fn get(&self, id: &AttachmentId) -> Option<std::sync::Arc<Vec<u8>>> {
self.entries.get(id).cloned()
}
pub fn remove(&mut self, id: &AttachmentId) {
if let Some(old) = self.entries.remove(id) {
self.total_bytes -= old.len();
self.order.retain(|k| k != id);
}
}
pub fn clear(&mut self) {
self.entries.clear();
self.order.clear();
self.total_bytes = 0;
}
#[cfg(test)]
fn len(&self) -> usize {
self.entries.len()
}
}
/// Parse a file-plane request: it must be exactly one [`AttachmentId`] (32
@@ -352,6 +539,87 @@ mod tests {
assert_eq!(validate_image_bytes(&buf.into_inner()), Some((4, 3)));
}
#[test]
fn sanitize_strips_bidi_and_zero_width_spoofing_chars() {
// U+202E would visually reverse the tail, disguising the extension.
assert_eq!(sanitize_filename("photo\u{202E}gnp.exe"), "photognp.exe");
assert_eq!(sanitize_filename("a\u{200B}b\u{FEFF}.txt"), "ab.txt");
// Ordinary Unicode filenames pass through.
assert_eq!(sanitize_filename("família_fotos.png"), "família_fotos.png");
assert_eq!(sanitize_filename("日本語.pdf"), "日本語.pdf");
}
/// Encode a solid PNG of the given dimensions for limit tests.
fn png_bytes(w: u32, h: u32) -> Vec<u8> {
let img = image::RgbImage::from_pixel(w, h, image::Rgb([10, 20, 30]));
let mut buf = std::io::Cursor::new(Vec::new());
image::DynamicImage::ImageRgb8(img)
.write_to(&mut buf, image::ImageFormat::Png)
.unwrap();
buf.into_inner()
}
#[test]
fn validate_image_rejects_excessive_total_pixels() {
// Both sides within MAX_IMAGE_PX, but 4096 * 4096 > MAX_IMAGE_TOTAL_PIXELS.
assert!(u64::from(MAX_IMAGE_PX) * u64::from(MAX_IMAGE_PX) > MAX_IMAGE_TOTAL_PIXELS);
assert_eq!(validate_image_bytes(&png_bytes(4096, 4096)), None);
// A 12 MP phone-photo shape passes both limits.
assert_eq!(
validate_image_bytes(&png_bytes(4032, 3024)),
Some((4032, 3024))
);
}
#[test]
fn preview_downscales_to_max_side_preserving_aspect() {
// Wide: 3200x400 → 1600x200.
let p = decode_preview(&png_bytes(3200, 400)).unwrap();
assert_eq!((p.width, p.height), (1600, 200));
assert_eq!(p.rgba.len(), preview_rgba_cost(1600, 200));
// Tall: 400x3200 → 200x1600.
let p = decode_preview(&png_bytes(400, 3200)).unwrap();
assert_eq!((p.width, p.height), (200, 1600));
// Square over the side cap: 2000x2000 → 1600x1600.
let p = decode_preview(&png_bytes(2000, 2000)).unwrap();
assert_eq!((p.width, p.height), (1600, 1600));
// At/under the cap is untouched.
let p = decode_preview(&png_bytes(1600, 900)).unwrap();
assert_eq!((p.width, p.height), (1600, 900));
let p = decode_preview(&png_bytes(4, 3)).unwrap();
assert_eq!((p.width, p.height), (4, 3));
assert_eq!(p.rgba.len(), preview_rgba_cost(4, 3));
}
#[test]
fn preview_rejects_what_validation_rejects() {
assert!(decode_preview(b"not an image").is_none());
assert!(decode_preview(&png_bytes(4096, 4096)).is_none());
}
#[test]
fn read_capped_stops_at_cap_plus_one() {
// Under the cap: full read.
let small = vec![7u8; 1024];
assert_eq!(
read_capped(std::io::Cursor::new(&small))
.unwrap()
.as_deref(),
Some(&small[..])
);
// Exactly at the cap: accepted. `repeat` is endless, `take` proves the
// reader is bounded rather than draining the source.
let at_cap = std::io::Read::take(std::io::repeat(1), MAX_ATTACHMENT_BYTES);
let got = read_capped(at_cap).unwrap().unwrap();
assert_eq!(got.len() as u64, MAX_ATTACHMENT_BYTES);
// One byte over: rejected, and only cap + 1 bytes were ever buffered
// (an unbounded source returns instead of allocating forever).
let over = std::io::Read::take(std::io::repeat(1), MAX_ATTACHMENT_BYTES + 1);
assert_eq!(read_capped(over).unwrap(), None);
let endless = std::io::repeat(1);
assert_eq!(read_capped(endless).unwrap(), None);
}
#[test]
fn human_size_units() {
assert_eq!(human_size(40), "40 B");
@@ -359,6 +627,69 @@ mod tests {
assert_eq!(human_size(3 * 1024 * 1024 + 300 * 1024), "3.3 MB");
}
#[test]
fn serve_store_count_and_byte_eviction_fifo() {
use std::sync::Arc;
let mut s = ServeStore::default();
let blob = |n: u8, len: usize| ([n; 32], Arc::new(vec![n; len]));
// Count cap: entry 0 is evicted when the 17th arrives.
for n in 0..=SERVED_FILES_MAX_ENTRIES as u8 {
let (id, b) = blob(n, 8);
assert!(s.insert(id, b));
}
assert_eq!(s.len(), SERVED_FILES_MAX_ENTRIES);
assert!(s.get(&[0u8; 32]).is_none(), "oldest evicted by count");
assert!(s.get(&[1u8; 32]).is_some());
// Byte budget: two ~half-budget blobs evict everything older.
let half = SERVED_FILES_MAX_BYTES / 2;
let (a, ab) = blob(100, half);
let (b, bb) = blob(101, half);
assert!(s.insert(a, ab));
assert!(s.insert(b, bb));
assert!(s.get(&a).is_some());
assert!(s.get(&b).is_some());
assert!(s.get(&[1u8; 32]).is_none(), "evicted for byte budget");
// A third half-budget blob evicts `a` (oldest), keeps `b`.
let (c, cb) = blob(102, half);
assert!(s.insert(c, cb));
assert!(s.get(&a).is_none());
assert!(s.get(&b).is_some());
assert!(s.get(&c).is_some());
}
#[test]
fn serve_store_replacement_accounting_and_remove_clear() {
use std::sync::Arc;
let mut s = ServeStore::default();
let id = [9u8; 32];
assert!(s.insert(id, Arc::new(vec![1; SERVED_FILES_MAX_BYTES - 10])));
// Replacing the near-budget blob must subtract its old bytes first —
// otherwise this same-id replacement would evict itself.
assert!(s.insert(id, Arc::new(vec![2; SERVED_FILES_MAX_BYTES - 5])));
assert_eq!(s.get(&id).unwrap()[0], 2);
assert_eq!(s.len(), 1);
s.remove(&id);
assert!(s.get(&id).is_none());
// Removed bytes were released: the budget admits a full-size blob again.
assert!(s.insert(id, Arc::new(vec![3; SERVED_FILES_MAX_BYTES])));
s.clear();
assert_eq!(s.len(), 0);
assert!(s.insert(id, Arc::new(vec![4; SERVED_FILES_MAX_BYTES])));
}
#[test]
fn serve_store_rejects_individually_overweight_blob() {
use std::sync::Arc;
let mut s = ServeStore::default();
let id = [7u8; 32];
assert!(s.insert(id, Arc::new(vec![1; 8])));
assert!(!s.insert(id, Arc::new(vec![2; SERVED_FILES_MAX_BYTES + 1])));
// The stale small blob is gone too — a fetch reads "no longer has it",
// never old bytes under a replaced id.
assert!(s.get(&id).is_none());
assert_eq!(s.len(), 0);
}
#[test]
fn attachment_descriptor_round_trips_json() {
let a = ChatAttachment {
+16
View File
@@ -3,6 +3,22 @@
#![cfg_attr(not(debug_assertions), windows_subsystem = "windows")]
fn main() {
// Tag the audio we play through ALSA (rodio's `ClipPlayer`: chat clips,
// peer music, local playlist tracks) so the screen-share exclusion engine
// can recognise it as ours and refuse to fan it back to the far end.
//
// First statement in the program, and that is load-bearing: this sets an
// environment variable, which is only sound while the process is still
// single-threaded, and PipeWire's ALSA plugin reads it when a stream is
// opened. See `audio::ownership::tag_this_process_alsa_audio`.
//
// SAFETY: nothing has been spawned yet, so no thread can be reading the
// environment concurrently.
#[cfg(target_os = "linux")]
unsafe {
peerspeak::audio::ownership::tag_this_process_alsa_audio()
};
if let Err(e) = peerspeak::app::run_gui() {
eprintln!("Error running GUI: {:?}", e);
}
+566 -22
View File
@@ -36,7 +36,8 @@ const MAX_GOSSIP_FRAME_BYTES: usize = 128 * 1024;
/// key); `sig` is that key's signature over [`signable_bytes`], so a forged
/// `author` can't validate (the attacker lacks the victim's secret key). `ts`
/// (sender-stamped unix-millis) is covered by the signature and gates replay
/// freshness — distinct from `GossipMessage::Chat.ts`, which is only for display.
/// freshness — distinct from `GossipMessage::Chat.ts`, an unauthenticated
/// duplicate kept only for wire compatibility and ignored on receive.
#[derive(Serialize, Deserialize, Clone)]
pub struct GossipPayload {
pub author: EndpointId,
@@ -254,6 +255,242 @@ impl ClockSkewMonitor {
}
}
/// Chat-hardening Phase 2 policy (docs/chat-hardening-plan.md): exact-replay
/// suppression bounds and the token-bucket rates that stop one admitted member
/// from monopolizing the event channel / UI with chat.
///
/// The replay cache is keyed on the payload's Ed25519 SIGNATURE bytes rather
/// than a separate BLAKE3 digest: ed25519 signing is deterministic (RFC 8032),
/// so the 64-byte signature is itself a collision-resistant fingerprint of the
/// exact signed bytes (topic + author + ts + msg) — same dedup power, zero new
/// dependencies. Entries are stamped with the payload's SIGNED timestamp and
/// pruned once that falls out of the freshness window, because `verify_gossip`
/// already rejects any replay whose signed `ts` is out-of-window — an expired
/// cache entry can no longer correspond to an admissible frame.
const CHAT_REPLAY_CACHE_CAP: usize = 1024;
/// Per-author chat budget: a burst of 8 absorbs a fast typist; 1 msg/s sustained
/// is well above real human chat rate while bounding a flooder to a trickle.
/// `pub(crate)` because the sender-side pacer (chat-hardening Phase 5) mirrors
/// this exact policy — one definition, so the two sides can never drift apart.
pub(crate) const CHAT_AUTHOR_BURST: f64 = 8.0;
pub(crate) const CHAT_AUTHOR_REFILL_PER_MS: f64 = 1.0 / 1000.0;
/// Room-wide chat budget across ALL authors, so a set of sock-puppet identities
/// can't multiply the per-author budget into unbounded event-channel pressure.
const CHAT_ROOM_BURST: f64 = 32.0;
const CHAT_ROOM_REFILL_PER_MS: f64 = 8.0 / 1000.0;
/// Bound on the per-author bucket map. Authors only enter it after the
/// known-author gate, so it tracks roughly the live roster plus recently
/// disconnected members; idle entries are pruned past this cap.
const CHAT_AUTHOR_BUCKETS_CAP: usize = 64;
/// Cooldown between logged chat rejections for one author (and one shared slot
/// for unknown authors), so a flood of rejected frames can't turn the log into
/// the new unbounded cost.
const CHAT_REJECT_LOG_COOLDOWN_MS: u64 = 10_000;
/// A minimal deterministic token bucket: time is passed in, never read from a
/// clock, so every boundary is unit-testable. Shared with the sender-side chat
/// pacer (`app::sendqueue`) so both sides of the rate policy use one mechanism.
#[derive(Debug, Clone, Copy)]
pub(crate) struct TokenBucket {
tokens: f64,
last_ms: u64,
}
impl TokenBucket {
pub(crate) fn full(burst: f64, now_ms: u64) -> Self {
Self {
tokens: burst,
last_ms: now_ms,
}
}
/// Refill for elapsed time (capped at `burst`), then take one token if
/// available. Returns whether a token was consumed.
pub(crate) fn try_take(&mut self, burst: f64, refill_per_ms: f64, now_ms: u64) -> bool {
let elapsed = now_ms.saturating_sub(self.last_ms) as f64;
self.tokens = (self.tokens + elapsed * refill_per_ms).min(burst);
self.last_ms = now_ms;
if self.tokens >= 1.0 {
self.tokens -= 1.0;
true
} else {
false
}
}
}
/// Why an authenticated chat payload was still refused admission (Phase 2).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum ChatReject {
/// Author is neither a live gossip peer nor one mid-reconnect. The core
/// roster gate re-checks this as the final authority; this early copy just
/// refuses the work before any sanitize/attachment handling.
UnknownAuthor,
/// Exact byte-for-byte replay of an already-admitted signed chat.
Replay,
/// Per-author or room-wide token bucket empty.
RateLimited,
}
/// Per-author rate-limit + log-squelch state (see [`ChatIngressGate`]).
#[derive(Debug)]
struct AuthorGateState {
bucket: TokenBucket,
last_seen_ms: u64,
last_reject_log_ms: Option<u64>,
}
/// Chat admission gate run after `verify_gossip`, before any sanitize work or
/// event-channel send (chat-hardening plan Phase 2): known author → exact-replay
/// dedup → per-author + room token buckets, in that order. Dedup runs BEFORE the
/// buckets so a replayed frame can never consume tokens and starve the author's
/// own legitimate next message. Pure — callers pass `now_ms` — so every branch
/// is unit-testable.
#[derive(Debug)]
struct ChatIngressGate {
seen: HashSet<[u8; 64]>,
/// FIFO of (signature, signed ts) mirroring `seen`, for TTL + cap pruning.
seen_order: std::collections::VecDeque<([u8; 64], u64)>,
room: TokenBucket,
authors: HashMap<EndpointId, AuthorGateState>,
/// Shared squelch slot for unknown-author rejects (they have no map entry).
last_unknown_log_ms: Option<u64>,
}
impl ChatIngressGate {
fn new(now_ms: u64) -> Self {
Self {
seen: HashSet::new(),
seen_order: std::collections::VecDeque::new(),
room: TokenBucket::full(CHAT_ROOM_BURST, now_ms),
authors: HashMap::new(),
last_unknown_log_ms: None,
}
}
/// Admit or reject one verified chat payload. `known_author` is the caller's
/// live-or-reconnecting membership check; `payload_ts` is the SIGNED envelope
/// timestamp (already freshness-checked by `verify_gossip`).
fn admit(
&mut self,
known_author: bool,
author: EndpointId,
sig: &[u8; 64],
payload_ts: u64,
now_ms: u64,
) -> Result<(), ChatReject> {
if !known_author {
return Err(ChatReject::UnknownAuthor);
}
self.prune_replay_cache(now_ms);
if self.seen.contains(sig) {
return Err(ChatReject::Replay);
}
// Room bucket first: it is the cheaper aggregate bound, and consuming
// from it only when the author bucket also admits keeps the two in
// lockstep — so check both, then commit both.
let author_state = self.author_entry(author, now_ms);
let author_ok =
author_state
.bucket
.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, now_ms);
if !author_ok {
return Err(ChatReject::RateLimited);
}
if !self
.room
.try_take(CHAT_ROOM_BURST, CHAT_ROOM_REFILL_PER_MS, now_ms)
{
// Refund the author token so a room-wide squeeze doesn't also debit
// every individual author's future budget.
if let Some(state) = self.authors.get_mut(&author) {
state.bucket.tokens = (state.bucket.tokens + 1.0).min(CHAT_AUTHOR_BURST);
}
return Err(ChatReject::RateLimited);
}
// Fully admitted — only now does the frame enter the replay cache, so a
// rate-limited legitimate message redelivered later is not misread as a
// replay of something that was never displayed.
self.seen.insert(*sig);
self.seen_order.push_back((*sig, payload_ts));
Ok(())
}
/// Whether this rejection should be logged: at most one log line per author
/// (or one shared line for unknown authors) per cooldown window.
fn should_log_reject(&mut self, known_author: bool, author: EndpointId, now_ms: u64) -> bool {
let slot = if known_author {
self.authors
.get_mut(&author)
.map(|state| &mut state.last_reject_log_ms)
} else {
Some(&mut self.last_unknown_log_ms)
};
let Some(slot) = slot else {
return true;
};
let due =
slot.is_none_or(|last| now_ms.saturating_sub(last) >= CHAT_REJECT_LOG_COOLDOWN_MS);
if due {
*slot = Some(now_ms);
}
due
}
/// Drop cache entries whose signed timestamp fell out of the freshness
/// window (they can no longer pass `verify_gossip`), then enforce the hard
/// cap FIFO-oldest-first.
fn prune_replay_cache(&mut self, now_ms: u64) {
let floor = now_ms.saturating_sub(GOSSIP_FRESHNESS_MS);
while let Some((sig, ts)) = self.seen_order.front() {
if *ts >= floor && self.seen_order.len() < CHAT_REPLAY_CACHE_CAP {
break;
}
self.seen.remove(sig);
self.seen_order.pop_front();
}
}
/// Get-or-create the author's bucket state, pruning the map if a flood of
/// short-lived identities has grown it past its cap: idle authors (nothing
/// admitted within the freshness window) go first, then oldest-seen.
fn author_entry(&mut self, author: EndpointId, now_ms: u64) -> &mut AuthorGateState {
if !self.authors.contains_key(&author) && self.authors.len() >= CHAT_AUTHOR_BUCKETS_CAP {
let floor = now_ms.saturating_sub(GOSSIP_FRESHNESS_MS);
self.authors.retain(|_, state| state.last_seen_ms >= floor);
while self.authors.len() >= CHAT_AUTHOR_BUCKETS_CAP {
if let Some(oldest) = self
.authors
.iter()
.min_by_key(|(_, state)| state.last_seen_ms)
.map(|(id, _)| *id)
{
self.authors.remove(&oldest);
} else {
break;
}
}
}
let state = self.authors.entry(author).or_insert(AuthorGateState {
bucket: TokenBucket::full(CHAT_AUTHOR_BURST, now_ms),
last_seen_ms: now_ms,
last_reject_log_ms: None,
});
state.last_seen_ms = now_ms;
state
}
/// Drop an author's limiter state alongside its roster eviction (a signed
/// `Leave`), so the map stays bounded by the roster's own churn.
fn evict_author(&mut self, author: &EndpointId) {
self.authors.remove(author);
}
}
/// Maximum number of distinct peers we hold in a room roster at once.
///
/// Everyone with the room ticket is an authenticated *insider*: a signature only
@@ -524,6 +761,7 @@ impl RoomState for IrohGossipState {
));
let mut state_mutations_seen = HashMap::new();
let mut clock_skew_monitor = ClockSkewMonitor::default();
let mut chat_gate = ChatIngressGate::new(now_millis());
// Broadcast initial state
let initial_payload = {
@@ -793,6 +1031,10 @@ impl RoomState for IrohGossipState {
.lock()
.unwrap()
.remove(&payload.author);
// Roster eviction also drops the author's
// chat-limiter state, keeping that map
// bounded by roster churn (Phase 2).
chat_gate.evict_author(&payload.author);
if removed || was_disconnected {
let _ = event_tx
.send(RoomEvent::PeerLeft(payload.author))
@@ -802,9 +1044,43 @@ impl RoomState for IrohGossipState {
GossipMessage::Chat {
name,
text,
ts,
// The inner ts is an unauthenticated duplicate of the
// signed envelope ts — ignored entirely; the envelope
// value is what RoomEvent carries (Phase 2).
ts: _,
attachment,
} => {
// Phase 2 ingress admission, BEFORE any sanitize or
// attachment work: known author (live or
// mid-reconnect — the core roster gate is the final
// authority) → exact-replay dedup keyed on the
// signature → per-author + room token buckets.
let known_author =
peers.lock().unwrap().contains_key(&payload.author)
|| disconnected_peers
.lock()
.unwrap()
.contains(&payload.author);
if let Err(reject) = chat_gate.admit(
known_author,
payload.author,
&payload.sig.to_bytes(),
payload.ts,
received_now_ms,
) {
if chat_gate.should_log_reject(
known_author,
payload.author,
received_now_ms,
) {
crate::log_msg(&format!(
"Dropped chat from author={}: {:?}",
crate::short_id(&payload.author.to_string()),
reject
));
}
continue;
}
crate::log_msg(&format!(
"Gossip chat from author={:?}",
payload.author
@@ -819,12 +1095,30 @@ impl RoomState for IrohGossipState {
a.name = crate::files::sanitize_filename(&a.name);
Some(a)
});
// Chat text policy at INGRESS: reject raw text
// over the byte ceiling before spending any
// sanitize work on it (a compliant sender
// sanitizes before signing), and drop a message
// with neither visible text nor an attachment.
let Some(text) = crate::sanitize::admit_chat_text(
&text,
attachment.is_some(),
) else {
crate::log_msg(&format!(
"Dropped out-of-policy chat from author={:?} (oversized or empty)",
payload.author
));
continue;
};
let _ = event_tx
.send(RoomEvent::ChatMessage {
from: payload.author,
name,
text,
ts,
// Only the SIGNED envelope timestamp travels
// downstream (never used for replay/ordering —
// the gate above already handled replay).
ts: payload.ts,
attachment,
})
.await;
@@ -966,6 +1260,14 @@ impl RoomState for IrohGossipState {
text: String,
attachment: Option<crate::files::ChatAttachment>,
) -> Result<(), NetError> {
// Enforce the chat text policy at the SIGN point, not only in the UI, so
// a future non-UI caller can't sign an out-of-policy body (chat-hardening
// plan Phase 1). Idempotent over the UI's own sanitize pass.
let text = crate::sanitize::sanitize_chat(&text);
if text.is_empty() && attachment.is_none() {
// Nothing visible to send — not an error, just nothing to do.
return Ok(());
}
let name = {
let guard = self.self_state.lock().unwrap();
match guard.as_ref() {
@@ -977,26 +1279,29 @@ impl RoomState for IrohGossipState {
let sender_opt = self.active_sender.lock().unwrap().clone();
let topic_opt = *self.active_topic_bytes.lock().unwrap();
if let (Some(sender), Some(topic)) = (sender_opt, topic_opt) {
let payload = sign_gossip(
&self.secret_key,
&topic,
// A missing sender/topic or an encode failure is a real send failure the
// caller must see (chat-hardening Phase 5) — silently returning Ok here
// would let the UI present an unsent message as broadcast.
let (Some(sender), Some(topic)) = (sender_opt, topic_opt) else {
return Err(NetError::Other("Not in a room".to_string()));
};
let payload = sign_gossip(
&self.secret_key,
&topic,
ts,
GossipMessage::Chat {
name,
text,
ts,
GossipMessage::Chat {
name,
text,
ts,
attachment,
},
);
if let Ok(bytes) = serde_json::to_vec(&payload) {
sender
.broadcast(bytes.into())
.await
.map_err(|e| NetError::Gossip(e.to_string()))?;
}
}
Ok(())
attachment,
},
);
let bytes = serde_json::to_vec(&payload)
.map_err(|e| NetError::Other(format!("Failed to encode chat: {e}")))?;
sender
.broadcast(bytes.into())
.await
.map_err(|e| NetError::Gossip(e.to_string()))
}
async fn leave(&self) -> Result<(), NetError> {
@@ -1664,4 +1969,243 @@ mod tests {
5
));
}
// ---- Chat-hardening Phase 2: ingress gate (replay dedup + token buckets) ----
/// Distinct opaque "signature" bytes; the gate never inspects them beyond
/// equality, so a counter-stamped array stands in for a real signature.
fn sig(n: u64) -> [u8; 64] {
let mut bytes = [0u8; 64];
bytes[..8].copy_from_slice(&n.to_le_bytes());
bytes
}
const T0: u64 = 1_000_000_000_000;
#[test]
fn token_bucket_burst_and_refill_boundaries() {
let mut bucket = TokenBucket::full(CHAT_AUTHOR_BURST, T0);
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert!(bucket.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, T0));
}
// Burst exhausted at the same instant.
assert!(!bucket.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, T0));
// 999 ms refills just under one token at 1/s…
assert!(!bucket.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, T0 + 999));
// …a full second refills exactly one (999 ms already banked 0.999 of it,
// so take at the accumulated boundary).
assert!(bucket.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, T0 + 1_001));
assert!(!bucket.try_take(CHAT_AUTHOR_BURST, CHAT_AUTHOR_REFILL_PER_MS, T0 + 1_001));
}
#[test]
fn chat_gate_rejects_unknown_author() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
assert_eq!(
gate.admit(false, author, &sig(1), T0, T0),
Err(ChatReject::UnknownAuthor)
);
// Same frame from a known author is fine — nothing was consumed above.
assert_eq!(gate.admit(true, author, &sig(1), T0, T0), Ok(()));
}
#[test]
fn chat_gate_suppresses_exact_replay_but_admits_distinct_same_ms() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
assert_eq!(gate.admit(true, author, &sig(1), T0, T0), Ok(()));
// Two DISTINCT chats signed in the same millisecond both land…
assert_eq!(gate.admit(true, author, &sig(2), T0, T0), Ok(()));
// …but the byte-identical frame is a replay, from any deliverer.
assert_eq!(
gate.admit(true, author, &sig(1), T0, T0),
Err(ChatReject::Replay)
);
}
#[test]
fn chat_gate_author_burst_then_refill() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
for n in 0..CHAT_AUTHOR_BURST as u64 {
assert_eq!(gate.admit(true, author, &sig(n), T0, T0), Ok(()));
}
assert_eq!(
gate.admit(true, author, &sig(99), T0, T0),
Err(ChatReject::RateLimited)
);
// One second later the author has exactly one more message.
assert_eq!(gate.admit(true, author, &sig(100), T0, T0 + 1_000), Ok(()));
assert_eq!(
gate.admit(true, author, &sig(101), T0, T0 + 1_000),
Err(ChatReject::RateLimited)
);
}
#[test]
fn chat_gate_replay_never_consumes_tokens() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
for n in 0..CHAT_AUTHOR_BURST as u64 {
assert_eq!(gate.admit(true, author, &sig(n), T0, T0), Ok(()));
}
// Replays of an admitted frame while exhausted report Replay (dedup runs
// BEFORE the buckets) and burn no tokens…
for _ in 0..50 {
assert_eq!(
gate.admit(true, author, &sig(0), T0, T0 + 1_000),
Err(ChatReject::Replay)
);
}
// …so the token refilled at +1s is still there for a NEW message.
assert_eq!(gate.admit(true, author, &sig(200), T0, T0 + 1_000), Ok(()));
}
#[test]
fn chat_gate_rate_limited_frame_is_not_marked_replayed() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
for n in 0..CHAT_AUTHOR_BURST as u64 {
assert_eq!(gate.admit(true, author, &sig(n), T0, T0), Ok(()));
}
// Rejected for rate only — NOT entered into the replay cache…
assert_eq!(
gate.admit(true, author, &sig(300), T0, T0),
Err(ChatReject::RateLimited)
);
// …so the same signed frame redelivered after refill is admitted once.
assert_eq!(gate.admit(true, author, &sig(300), T0, T0 + 1_000), Ok(()));
}
#[test]
fn chat_gate_room_bucket_bounds_sock_puppet_authors() {
let mut gate = ChatIngressGate::new(T0);
// 40 distinct authors, one message each, same instant: per-author buckets
// are all full, so only the room-wide burst bounds admission.
let mut admitted = 0;
for n in 0..40u64 {
if gate.admit(true, fresh_id(), &sig(n), T0, T0).is_ok() {
admitted += 1;
}
}
assert_eq!(admitted, CHAT_ROOM_BURST as usize);
// The room refills at 8/s: exactly 8 more land a second later, even from
// fresh authors whose own buckets are full — the room bound decides.
let mut late_admitted = 0;
for n in 100..120u64 {
if gate
.admit(true, fresh_id(), &sig(n), T0, T0 + 1_000)
.is_ok()
{
late_admitted += 1;
}
}
assert_eq!(late_admitted, 8);
// Chat admission being room-bounded is what keeps the event channel
// available for control messages: an Announce is gated independently.
let mut seen = HashMap::new();
assert!(admit_state_mutation(
&mut seen,
fresh_id(),
&GossipMessage::Leave,
T0 + 1_000
));
}
#[test]
fn chat_gate_room_reject_refunds_the_author_token() {
let mut gate = ChatIngressGate::new(T0);
// Author A drains the whole room burst alone? No — its own burst is 8.
// Use 4 authors × 8 to empty the room exactly.
let mut n = 0u64;
for _ in 0..4 {
let author = fresh_id();
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert_eq!(gate.admit(true, author, &sig(n), T0, T0), Ok(()));
n += 1;
}
}
// A 5th author is room-rejected 8 times, but its own bucket is refunded
// each time…
let victim = fresh_id();
for _ in 0..CHAT_AUTHOR_BURST as usize {
assert_eq!(
gate.admit(true, victim, &sig(n), T0, T0),
Err(ChatReject::RateLimited)
);
n += 1;
}
// …so when the room refills, the victim still has its FULL burst.
let mut admitted = 0;
for _ in 0..CHAT_AUTHOR_BURST as usize {
if gate.admit(true, victim, &sig(n), T0, T0 + 1_000).is_ok() {
admitted += 1;
}
n += 1;
}
assert_eq!(admitted, CHAT_AUTHOR_BURST as usize);
}
#[test]
fn chat_replay_cache_prunes_by_ttl_and_cap() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
// TTL: an admitted frame's entry is dropped once its SIGNED ts falls out
// of the freshness window (it could no longer pass verify_gossip anyway).
assert_eq!(gate.admit(true, author, &sig(1), T0, T0), Ok(()));
assert_eq!(gate.seen.len(), 1);
gate.prune_replay_cache(T0 + GOSSIP_FRESHNESS_MS + 1);
assert!(gate.seen.is_empty() && gate.seen_order.is_empty());
// Hard cap: stuff the cache directly (admission itself is rate-limited
// far below the cap) and verify FIFO-oldest eviction bounds it.
for n in 0..(CHAT_REPLAY_CACHE_CAP as u64 + 100) {
gate.seen.insert(sig(n));
gate.seen_order.push_back((sig(n), T0));
}
gate.prune_replay_cache(T0);
assert!(gate.seen_order.len() < CHAT_REPLAY_CACHE_CAP);
assert_eq!(gate.seen.len(), gate.seen_order.len());
// The oldest entries went first.
assert!(!gate.seen.contains(&sig(0)));
assert!(gate.seen.contains(&sig(CHAT_REPLAY_CACHE_CAP as u64 + 99)));
}
#[test]
fn chat_gate_author_bucket_map_stays_bounded() {
let mut gate = ChatIngressGate::new(T0);
for n in 0..(CHAT_AUTHOR_BUCKETS_CAP as u64 * 2) {
let _ = gate.admit(true, fresh_id(), &sig(n), T0, T0 + n);
}
assert!(gate.authors.len() <= CHAT_AUTHOR_BUCKETS_CAP);
}
#[test]
fn chat_gate_evict_author_drops_limiter_state() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
assert_eq!(gate.admit(true, author, &sig(1), T0, T0), Ok(()));
assert!(gate.authors.contains_key(&author));
gate.evict_author(&author);
assert!(!gate.authors.contains_key(&author));
}
#[test]
fn chat_gate_reject_logging_is_squelched_per_author() {
let mut gate = ChatIngressGate::new(T0);
let author = fresh_id();
// Establish limiter state, then exhaust it.
for n in 0..=CHAT_AUTHOR_BURST as u64 {
let _ = gate.admit(true, author, &sig(n), T0, T0);
}
assert!(gate.should_log_reject(true, author, T0));
assert!(!gate.should_log_reject(true, author, T0 + 1));
assert!(gate.should_log_reject(true, author, T0 + CHAT_REJECT_LOG_COOLDOWN_MS));
// Unknown authors share one squelch slot (they have no map entry).
let stranger = fresh_id();
assert!(gate.should_log_reject(false, stranger, T0));
assert!(!gate.should_log_reject(false, fresh_id(), T0 + 1));
}
}
+69 -4
View File
@@ -69,7 +69,9 @@ struct Shared {
/// the random attachment id. Populated when we send a chat file; read by the
/// file protocol handler to answer a member's fetch. Cleared on leave. Each
/// blob is already byte-capped at send time.
served_files: StdMutex<HashMap<crate::files::AttachmentId, Arc<Vec<u8>>>>,
/// Blobs we serve to room members, bounded by count and byte budgets
/// (Phase 3C) — an evicted id reads as "sender no longer has the file".
served_files: StdMutex<crate::files::ServeStore>,
incoming_tx: mpsc::Sender<(EndpointId, Bytes)>,
/// Best-effort link-state notifications for the UI (connecting / connected).
conn_events_tx: mpsc::Sender<ConnEvent>,
@@ -518,7 +520,7 @@ impl iroh::protocol::ProtocolHandler for FileRouter {
let Some(id) = crate::files::parse_request(&req) else {
return Ok(());
};
let blob = shared.served_files.lock().unwrap().get(&id).cloned();
let blob = shared.served_files.lock().unwrap().get(&id);
if let Some(blob) = blob {
let _ = send.write_all(&blob).await;
}
@@ -563,7 +565,7 @@ impl IrohTransport {
peers: tokio::sync::Mutex::new(HashMap::new()),
live_conns: StdMutex::new(HashMap::new()),
admitted_audio: StdMutex::new(HashSet::new()),
served_files: StdMutex::new(HashMap::new()),
served_files: StdMutex::new(crate::files::ServeStore::default()),
incoming_tx,
conn_events_tx,
});
@@ -634,7 +636,9 @@ impl IrohTransport {
/// session (served by the [`FileRouter`] handler). Called by core when we
/// send a chat file. The blob is cleared on leave.
pub fn serve_attachment(&self, id: AttachmentId, bytes: Arc<Vec<u8>>) {
self.shared.served_files.lock().unwrap().insert(id, bytes);
if !self.shared.served_files.lock().unwrap().insert(id, bytes) {
crate::log_msg("Transport: refused to serve an over-budget blob");
}
}
/// Drop a previously-served blob (e.g. a music track no longer current-or-next).
@@ -676,6 +680,8 @@ impl IrohTransport {
send.finish()
.map_err(|e| NetError::Other(format!("file fetch: request finish failed: {e}")))?;
// `read_to_end(size)` errors if the stream exceeds `size`, rejecting an
// overlong transfer; the exact-length check below rejects a short one.
let read = recv.read_to_end(size as usize);
let bytes = tokio::time::timeout(FILE_FETCH_TIMEOUT, read)
.await
@@ -686,9 +692,68 @@ impl IrohTransport {
"file fetch: sender no longer has the file".to_string(),
));
}
// Exact transfer required (Phase 3C): a truncated body must not be
// cached/saved/decoded as if it were the declared attachment.
if bytes.len() as u64 != size {
return Err(NetError::Other(format!(
"file fetch: incomplete transfer ({} of {size} bytes)",
bytes.len()
)));
}
Ok(bytes)
}
/// Snapshot the selected QUIC path of every live audio connection, for the
/// UI's per-peer connection badge (direct/relay, RTT, loss, bitrate).
/// Cheap and lock-light: the `live_conns` guard is released before touching
/// any connection, and `Connection::paths()` reads shared state without I/O.
pub fn connection_stats(&self) -> Vec<(EndpointId, crate::network::PathSnapshot)> {
// Clone the connections out so the map lock isn't held while we inspect
// paths (a supervisor inserts/removes entries as links come and go).
let conns: Vec<(EndpointId, Connection)> = self
.shared
.live_conns
.lock()
.unwrap()
.iter()
.map(|(id, conn)| (*id, conn.clone()))
.collect();
conns
.into_iter()
.filter_map(|(id, conn)| {
let paths = conn.paths();
// The selected path is the one carrying application data. In the
// brief window where none is flagged (e.g. mid-migration), fall
// back to the first open path rather than dropping the badge.
let path = paths
.iter()
.find(|p| p.is_selected())
.or_else(|| paths.iter().next())?;
let stats = path.stats();
// Per-variant display: `TransportAddr`'s own `Display` prefixes
// a scheme ("ip:1.2.3.4:5") that's noise next to the badge's
// Direct/Relay label.
let remote_addr = match path.remote_addr() {
iroh::TransportAddr::Ip(sock) => sock.to_string(),
iroh::TransportAddr::Relay(url) => url.to_string(),
other => other.to_string(),
};
Some((
id,
crate::network::PathSnapshot {
is_relay: path.remote_addr().is_relay(),
remote_addr,
rtt: stats.rtt,
tx_bytes: stats.udp_tx.bytes,
rx_bytes: stats.udp_rx.bytes,
tx_datagrams: stats.udp_tx.datagrams,
lost_packets: stats.lost_packets,
},
))
})
.collect()
}
/// Fetch a chat attachment's bytes from its sender over the file plane.
pub async fn fetch_attachment(
&self,
+30 -3
View File
@@ -152,9 +152,10 @@ pub enum RoomEvent {
author: EndpointId,
skew_ms: i64,
},
/// A peer sent a room text-chat message. Carries the sender's id, their
/// display name (embedded so it shows even without a presence entry), the
/// text, and a sender-stamped millisecond timestamp.
/// A peer sent a room text-chat message. Carries the sender's id, the
/// sender-CLAIMED display name (untrusted; the core replaces it with the
/// roster-bound name before the UI sees it — chat-hardening Phase 2), the
/// text, and the signed envelope timestamp (display only, never ordering).
ChatMessage {
from: EndpointId,
name: String,
@@ -184,6 +185,32 @@ pub enum ConnEvent {
Left(EndpointId),
}
/// Owned snapshot of a peer's *selected* QUIC path (the one currently carrying
/// application data), taken from the live audio connection for the UI's
/// connection-transparency badge. Counters are cumulative for the path's
/// lifetime; rate/loss derivation over a poll window happens in
/// `core::connstats` (which also detects path switches via `remote_addr`).
#[derive(Debug, Clone, PartialEq)]
pub struct PathSnapshot {
/// True when the path runs through a relay server, false for a direct
/// (holepunched or local) IP path.
pub is_relay: bool,
/// The path's remote transport address: `ip:port` for a direct path, the
/// relay URL for a relayed one.
pub remote_addr: String,
/// Current QUIC round-trip-time estimate for the path.
pub rtt: std::time::Duration,
/// Cumulative bytes sent in UDP datagrams on the path.
pub tx_bytes: u64,
/// Cumulative bytes received in UDP datagrams on the path.
pub rx_bytes: u64,
/// Cumulative UDP datagrams sent on the path (the loss denominator: for our
/// small voice frames these map ~1:1 to QUIC packets).
pub tx_datagrams: u64,
/// Cumulative packets detected lost on the path.
pub lost_packets: u64,
}
#[derive(Serialize, Deserialize, Clone, Debug)]
pub struct PeerSpeakTicket {
pub host_addr: iroh::EndpointAddr,
+81 -4
View File
@@ -10,6 +10,8 @@
//! leaves a zombie. Any failure (no player, no audio) is silent by design — a
//! missing chime should never disrupt a call.
#[cfg(not(windows))]
use crate::audio::ownership;
use std::collections::HashMap;
use std::fs::OpenOptions;
use std::io::Write;
@@ -80,6 +82,14 @@ pub enum Sound {
MicToggle,
/// Reconnect failed / peer evicted.
ReconnectFailed,
/// One of our chat messages was broadcast to the room.
ChatSent,
/// A chat message from another participant was admitted.
ChatReceived,
/// A saved contact was detected online on the home screen.
ContactOnline,
/// A saved contact previously seen online went offline on the home screen.
ContactOffline,
}
impl Sound {
@@ -93,10 +103,14 @@ impl Sound {
Sound::SelfLeave,
Sound::MicToggle,
Sound::ReconnectFailed,
Sound::ChatSent,
Sound::ChatReceived,
Sound::ContactOnline,
Sound::ContactOffline,
];
/// Number of distinct notification events.
pub const COUNT: usize = 8;
pub const COUNT: usize = 12;
/// Stable 0-based index into the per-sound flag array. Must match `ALL`.
fn index(self) -> usize {
@@ -109,6 +123,10 @@ impl Sound {
Sound::SelfLeave => 5,
Sound::MicToggle => 6,
Sound::ReconnectFailed => 7,
Sound::ChatSent => 8,
Sound::ChatReceived => 9,
Sound::ContactOnline => 10,
Sound::ContactOffline => 11,
}
}
@@ -123,6 +141,10 @@ impl Sound {
Sound::SelfLeave => include_bytes!("../assets/sounds/self-leave.wav"),
Sound::MicToggle => include_bytes!("../assets/sounds/mic-toggle.wav"),
Sound::ReconnectFailed => include_bytes!("../assets/sounds/reconnect-failed.wav"),
Sound::ChatSent => include_bytes!("../assets/sounds/chat-sent.wav"),
Sound::ChatReceived => include_bytes!("../assets/sounds/chat-received.wav"),
Sound::ContactOnline => include_bytes!("../assets/sounds/contact-online.wav"),
Sound::ContactOffline => include_bytes!("../assets/sounds/contact-offline.wav"),
}
}
@@ -137,6 +159,10 @@ impl Sound {
Sound::SelfLeave => "self-leave",
Sound::MicToggle => "mic-toggle",
Sound::ReconnectFailed => "reconnect-failed",
Sound::ChatSent => "chat-sent",
Sound::ChatReceived => "chat-received",
Sound::ContactOnline => "contact-online",
Sound::ContactOffline => "contact-offline",
}
}
}
@@ -240,12 +266,20 @@ fn escape_powershell_single_quoted(s: &str) -> String {
#[cfg(not(windows))]
fn spawn_player(path: &Path) {
for player in ["pw-play", "paplay", "aplay"] {
let started = Command::new(player)
let mut command = Command::new(player);
command
.arg(path)
.stdin(Stdio::null())
.stdout(Stdio::null())
.stderr(Stdio::null())
.status();
.stderr(Stdio::null());
// Ownership tag (plan §5.1). A chime is short, but it is still our
// audio on the default sink, and an untagged one is an unowned root
// the exclusion engine would have to reason about from scratch.
// Measured on this host: all three fallbacks tag correctly, `aplay`
// included — it reaches the graph through PipeWire's ALSA plugin,
// which honours `PIPEWIRE_PROPS` like any other client.
ownership::tag_child(&mut command, ownership::NOTIFICATION_ROLE);
let started = command.status();
// `status()` errors only if the player binary isn't present; on a real
// playback error it still returns (non-zero), so a started player ends
// the loop either way — we don't want to double-play through fallbacks.
@@ -289,6 +323,49 @@ mod tests {
dir
}
/// Phase-1 exit gate, notification half (impl plan §3): a chime peerspeak
/// actually plays produces a live PipeWire node carrying **both**
/// ownership carriers.
///
/// ⚠️ Deliberately drives `play()`, not `tag_child()`. The unit test in
/// `audio::ownership` proves the environment is built correctly; only a
/// live run proves this module *uses* it and that the audio stack honours
/// it end to end. The chime is silent (a zero-filled WAV), so running it
/// never makes noise.
///
/// Live: needs a running PipeWire daemon, `pw-play`/`paplay` and
/// `pw-dump`. `cargo test --lib -- --ignored notification_chime`
#[test]
#[ignore = "live: requires a running PipeWire daemon and pw-dump"]
#[cfg(not(windows))]
fn notification_chime_node_carries_both_ownership_carriers() {
use crate::audio::ownership::{self, live_test};
let dir = temp_wav_dir("ownership");
let path = dir.join("silence.wav");
std::fs::write(&path, live_test::silent_wav(6)).unwrap();
set_enabled(true);
set_sound_enabled(Sound::PeerJoin, true);
play(Sound::PeerJoin, Some(path.to_str().unwrap()));
let prefix = live_test::expected_prefix(ownership::NOTIFICATION_ROLE);
let found = live_test::poll_for_owned_node(&prefix, std::time::Duration::from_secs(5));
std::fs::remove_dir_all(&dir).ok();
let (name, owned) =
found.unwrap_or_else(|| panic!("no live node named {prefix:?} appeared within 5s"));
assert!(
name.starts_with(ownership::OWNED_NODE_NAME_PREFIX),
"{name}"
);
assert_eq!(
owned.as_deref(),
Some(ownership::OWNED_PROP_VALUE),
"carrier 1 must be on the live node too, not just carrier 2"
);
}
#[test]
fn test_should_play_truth_table() {
// Plays only when BOTH the master and the per-sound flag are on.
+423 -124
View File
@@ -3,28 +3,40 @@
//! Peer display names ride the gossip presence plane (`PeerState.name`), which is
//! untrusted and spoofable, yet they're rendered directly in the roster. This
//! module cleans a name at the gossip ingest point so every downstream consumer
//! gets a safe value (security finding S4). Chat text has its own sanitizer in
//! the UI layer (`app::sanitize_chat`).
//! gets a safe value (security finding S4). Chat text policy ([`sanitize_chat`],
//! [`cap_chat_input`], [`admit_chat_text`]) also lives here so the UI, the gossip
//! sign point, and the gossip ingress all enforce the same ceilings.
/// Max characters kept for a peer's display name after sanitizing. Names are
/// short labels, so a tight cap both prevents UI/layout/memory abuse and keeps
/// the roster readable.
pub const NAME_MAX_CHARS: usize = 48;
/// Bidirectional override/isolate format characters (`General_Category=Cf`, NOT
/// caught by [`char::is_control`]) that can visually reorder surrounding text.
/// Stripped even from expressive chat bodies (security finding S14): unlike the
/// benign zero-width joiners/marks, these let a sender make rendered text read
/// differently from what was actually sent.
pub(crate) fn is_bidi_override_char(c: char) -> bool {
matches!(c,
'\u{202A}'..='\u{202E}' // LRE, RLE, PDF, LRO, RLO (bidi overrides)
| '\u{2066}'..='\u{2069}' // LRI, RLI, FSI, PDI (bidi isolates)
)
}
/// Unicode *format* characters (`General_Category=Cf`) that can spoof or garble a
/// rendered name even though they are NOT caught by [`char::is_control`]:
/// bidirectional overrides/isolates (text-direction spoofing) and
/// zero-width / BOM characters (invisible, can hide or fake content). Listed
/// explicitly so the sanitizer stays dependency-free (std exposes no category
/// query). Stripped outright rather than replaced.
fn is_spoofing_format_char(c: char) -> bool {
matches!(c,
'\u{200B}'..='\u{200F}' // zero-width space, ZWNJ, ZWJ, LRM, RLM
| '\u{202A}'..='\u{202E}' // LRE, RLE, PDF, LRO, RLO (bidi overrides)
| '\u{2060}'..='\u{2064}' // word joiner .. invisible plus
| '\u{2066}'..='\u{2069}' // LRI, RLI, FSI, PDI (bidi isolates)
| '\u{FEFF}' // BOM / zero-width no-break space
)
pub(crate) fn is_spoofing_format_char(c: char) -> bool {
is_bidi_override_char(c)
|| matches!(c,
'\u{200B}'..='\u{200F}' // zero-width space, ZWNJ, ZWJ, LRM, RLM
| '\u{2060}'..='\u{2064}' // word joiner .. invisible plus
| '\u{FEFF}' // BOM / zero-width no-break space
)
}
/// Max characters kept for a broadcast game-presence label after sanitizing
@@ -73,15 +85,98 @@ pub fn sanitize_game_label(input: &str) -> String {
out
}
/// A piece of a chat message after URL detection: literal text or a link.
#[derive(Debug, PartialEq, Eq, Clone)]
pub enum Segment {
/// Plain text to render as-is.
Text(String),
/// A detected URL to render as a clickable link (also its href).
Link(String),
/// Max characters kept for a single chat message after sanitizing.
pub const CHAT_MSG_MAX_CHARS: usize = 2000;
/// Max UTF-8 bytes kept for a single chat message, enforced alongside
/// [`CHAT_MSG_MAX_CHARS`] (2,000 four-byte scalars would otherwise reach 8,000
/// bytes). This is also the ingress bound: signed peers never produce more, so
/// raw incoming text above it is rejected outright (see [`admit_chat_text`]).
pub const CHAT_MSG_MAX_BYTES: usize = 8 * 1024;
/// Sanitize a chat message body, applied to BOTH our outgoing text (before local
/// echo, and again at the gossip sign point) and incoming peer text (untrusted —
/// a buggy/malicious sender could include control characters or an enormous
/// payload). Single pass: bidi overrides/isolates are stripped outright (S14 —
/// they can visually reorder the rendered line), control characters become
/// spaces, any whitespace run collapses to a single space, the ends are trimmed,
/// and both the character and UTF-8 byte ceilings are enforced without ever
/// splitting a scalar. Message bodies deliberately keep the OTHER format
/// characters (ZWJ/ZWNJ/LRM/RLM etc.) that the short-label sanitizers strip —
/// chat is expressive text, not a label, and those are needed for emoji
/// sequences and joining scripts. Returns `""` for input with no visible text
/// (callers drop empty messages). Idempotent, so layered application converges
/// on the same result.
pub fn sanitize_chat(input: &str) -> String {
let mut out = String::new();
let mut chars = 0usize;
let mut pending_space = false;
for c in input.chars() {
if is_bidi_override_char(c) {
continue;
}
let c = if c.is_control() { ' ' } else { c };
if c.is_whitespace() {
// Trim: only mark a separator once visible text exists; a trailing
// run is never emitted because the space lands with the NEXT char.
pending_space = !out.is_empty();
continue;
}
let sep = usize::from(pending_space);
if chars + sep + 1 > CHAT_MSG_MAX_CHARS
|| out.len() + sep + c.len_utf8() > CHAT_MSG_MAX_BYTES
{
break;
}
if pending_space {
out.push(' ');
chars += 1;
pending_space = false;
}
out.push(c);
chars += 1;
}
out
}
/// Cap the LIVE chat-input text (typing, clipboard/primary-selection paste,
/// context-menu paste) at the chat ceilings. Unlike [`sanitize_chat`] this
/// preserves the user's whitespace exactly — normalization stays a submit-time
/// operation so the visible text never jumps while editing — and only truncates,
/// always on a scalar boundary. Returns the input unchanged when within bounds.
pub fn cap_chat_input(input: String) -> String {
if input.len() <= CHAT_MSG_MAX_BYTES && input.chars().count() <= CHAT_MSG_MAX_CHARS {
return input;
}
let mut out = String::new();
for (chars, c) in input.chars().enumerate() {
if chars >= CHAT_MSG_MAX_CHARS || out.len() + c.len_utf8() > CHAT_MSG_MAX_BYTES {
break;
}
out.push(c);
}
out
}
/// Gossip-ingress admission for an untrusted incoming chat body. `None` drops
/// the message: raw text over the byte ceiling is rejected BEFORE any
/// sanitization work (a compliant sender sanitizes before signing, so oversized
/// text is a protocol violation, not something to repair), and a message with
/// neither visible text nor an attachment carries nothing to show. Otherwise
/// yields the sanitized (possibly empty, attachment-only) body to forward.
pub fn admit_chat_text(raw: &str, has_attachment: bool) -> Option<String> {
if raw.len() > CHAT_MSG_MAX_BYTES {
return None;
}
let text = sanitize_chat(raw);
(!text.is_empty() || has_attachment).then_some(text)
}
/// Max clickable links rendered per chat message. Later URL candidates stay
/// selectable plain text — bounds both the span count a message can force the
/// renderer to build and the opener targets one line can carry.
pub const CHAT_MSG_MAX_LINKS: usize = 8;
/// Trailing characters commonly adjacent to a URL in prose that should NOT be
/// part of the link (so "see http://x.com." or "(http://x.com)" linkify cleanly).
fn is_url_trailing_punct(c: char) -> bool {
@@ -91,42 +186,85 @@ fn is_url_trailing_punct(c: char) -> bool {
)
}
/// Find the byte index of the earliest `http://` or `https://` scheme in `s`,
/// Find the byte index of the earliest `http://` or `https://` scheme in `s`
/// (ASCII-case-insensitive, so a sentence-capitalized "Http://…" still counts),
/// scanning only on char boundaries so slicing is always safe.
fn find_scheme(s: &str) -> Option<usize> {
s.char_indices().find_map(|(i, _)| {
let tail = &s[i..];
(tail.starts_with("http://") || tail.starts_with("https://")).then_some(i)
let matches_prefix = |p: &str| {
tail.get(..p.len())
.is_some_and(|t| t.eq_ignore_ascii_case(p))
};
(matches_prefix("http://") || matches_prefix("https://")).then_some(i)
})
}
/// Split an (already chat-sanitized) message into plain-text and URL [`Segment`]s
/// for rendering. **Conservative on purpose:** only `http://` / `https://` runs
/// are treated as links, each ending at the first whitespace, with trailing prose
/// punctuation peeled back into the following text. Concatenating every segment's
/// inner string reproduces the input exactly (no characters added or dropped), so
/// it's purely a presentational split. Linkify AFTER sanitizing so control/format
/// chars are already gone (the URL can't smuggle them). Pure → unit-testable.
pub fn linkify(input: &str) -> Vec<Segment> {
/// The clickable-link policy, shared by link detection ([`link_ranges`]) and the
/// opener's defence-in-depth re-check (`AppMessage::OpenUrl`): the candidate must
/// parse as a URL with an `http`/`https` scheme, a non-empty host, and NO
/// username/password syntax (`http://user@host` reads as a credential but is a
/// classic destination-spoof — such text stays plain, never clickable).
pub fn is_safe_web_url(s: &str) -> bool {
let Ok(u) = url::Url::parse(s) else {
return false;
};
matches!(u.scheme(), "http" | "https")
&& u.host_str().is_some_and(|h| !h.is_empty())
&& u.username().is_empty()
&& u.password().is_none()
}
/// Detect clickable links in an (already chat-sanitized) message, returning the
/// byte range of each — computed ONCE when a message enters history and cached
/// on its entry, so redraws slice instead of rescanning. **Conservative on
/// purpose:** only `http://` / `https://` runs count, each ending at the first
/// whitespace with trailing prose punctuation peeled off, and only candidates
/// passing [`is_safe_web_url`] become links — a failing candidate's whole
/// whitespace-delimited run stays plain text (its interior is not re-scanned).
/// At most [`CHAT_MSG_MAX_LINKS`] ranges; ranges are ascending, non-overlapping,
/// and always on char boundaries. The href is exactly the displayed slice, so
/// what the user sees IS what the opener receives.
pub fn link_ranges(text: &str) -> Vec<std::ops::Range<usize>> {
let mut out = Vec::new();
let mut rest = input;
while !rest.is_empty() {
let Some(start) = find_scheme(rest) else {
out.push(Segment::Text(rest.to_string()));
let mut base = 0usize;
while out.len() < CHAT_MSG_MAX_LINKS {
let Some(start) = find_scheme(&text[base..]) else {
break;
};
if start > 0 {
out.push(Segment::Text(rest[..start].to_string()));
let run_start = base + start;
let run = &text[run_start..];
let run_end = run.find(char::is_whitespace).unwrap_or(run.len());
// Peel trailing punctuation back out of the candidate; a run is at least
// the 7-byte scheme long, so `base` always advances.
let candidate = run[..run_end].trim_end_matches(is_url_trailing_punct);
if is_safe_web_url(candidate) {
out.push(run_start..run_start + candidate.len());
base = run_start + candidate.len();
} else {
base = run_start + run_end;
}
let after = &rest[start..];
let end = after.find(char::is_whitespace).unwrap_or(after.len());
let candidate = &after[..end];
// Peel trailing punctuation back out of the link.
let url = candidate.trim_end_matches(is_url_trailing_punct);
out.push(Segment::Link(url.to_string()));
// Continue past just the URL; any peeled punctuation + the rest (incl. the
// whitespace) is reconsidered as ordinary text on the next iteration.
rest = &after[url.len()..];
}
out
}
/// Split `text` into `(slice, is_link)` pieces from cached [`link_ranges`]
/// output. Concatenating the slices reproduces `text` exactly (purely a
/// presentational split — no characters added or dropped). Borrows, so a redraw
/// allocates nothing for plain text. `ranges` must come from [`link_ranges`] on
/// this same `text` (ascending, non-overlapping, char-boundary ranges).
pub fn segments<'a>(text: &'a str, ranges: &[std::ops::Range<usize>]) -> Vec<(&'a str, bool)> {
let mut out = Vec::new();
let mut pos = 0usize;
for r in ranges {
if r.start > pos {
out.push((&text[pos..r.start], false));
}
out.push((&text[r.clone()], true));
pos = r.end;
}
if pos < text.len() {
out.push((&text[pos..], false));
}
out
}
@@ -215,93 +353,199 @@ mod tests {
assert_eq!(sanitize_game_label("\u{0}\r\n\t "), "");
}
// --- linkify -----------------------------------------------------------
// --- chat body policy ---------------------------------------------------
/// Concatenating every segment's inner text must reproduce the input exactly.
fn reassemble(segs: &[Segment]) -> String {
segs.iter()
.map(|s| match s {
Segment::Text(t) | Segment::Link(t) => t.as_str(),
})
#[test]
fn chat_keeps_ordinary_text_and_unicode() {
assert_eq!(sanitize_chat("hello world"), "hello world");
assert_eq!(sanitize_chat("héllo 🎙 世界"), "héllo 🎙 世界");
// Bodies keep format characters that label sanitizers strip: a ZWJ emoji
// family sequence survives intact.
let family = "👨\u{200D}👩\u{200D}👧";
assert_eq!(sanitize_chat(family), family);
}
#[test]
fn chat_strips_control_chars_and_collapses_whitespace() {
assert_eq!(sanitize_chat(" hi there "), "hi there");
assert_eq!(sanitize_chat("a\u{0}b\r\nc\td\u{1b}[31m"), "a b c d [31m");
assert_eq!(sanitize_chat("\u{0}\r\n\t "), "");
assert_eq!(sanitize_chat(""), "");
}
#[test]
fn chat_caps_chars_at_exact_boundary_without_trailing_space() {
let long = "x".repeat(CHAT_MSG_MAX_CHARS + 500);
assert_eq!(sanitize_chat(&long).chars().count(), CHAT_MSG_MAX_CHARS);
assert_eq!(
sanitize_chat(&"x".repeat(CHAT_MSG_MAX_CHARS))
.chars()
.count(),
CHAT_MSG_MAX_CHARS
);
// Truncation never leaves a dangling separator: with "word " units the
// cut lands mid-run, and the output still ends on visible text.
let words = "word ".repeat(1000);
let out = sanitize_chat(&words);
assert!(out.chars().count() <= CHAT_MSG_MAX_CHARS);
assert!(!out.ends_with(' '));
}
#[test]
fn chat_ceilings_never_split_a_scalar() {
// Four-byte scalars: the char cap bites first (2,000 × 4 = 8,000 bytes,
// inside the byte ceiling by design) and the last emoji is kept whole.
let emoji = "🎮".repeat(CHAT_MSG_MAX_CHARS + 100);
let out = sanitize_chat(&emoji);
assert_eq!(out.chars().count(), CHAT_MSG_MAX_CHARS);
assert!(out.len() <= CHAT_MSG_MAX_BYTES);
assert!(out.chars().all(|c| c == '🎮'));
// Three-byte scalars at the char boundary.
let cjk = "".repeat(CHAT_MSG_MAX_CHARS + 1);
let out = sanitize_chat(&cjk);
assert_eq!(out.chars().count(), CHAT_MSG_MAX_CHARS);
assert!(out.is_char_boundary(out.len()));
}
#[test]
fn chat_sanitize_is_idempotent() {
for input in [
"plain text",
" spaced \t out\r\n text ",
"unicode 🎙 世界 👨\u{200D}👩\u{200D}👧",
&"word ".repeat(1000),
&"🎮".repeat(CHAT_MSG_MAX_CHARS + 100),
] {
let once = sanitize_chat(input);
assert_eq!(sanitize_chat(&once), once, "not idempotent for {input:?}");
}
}
#[test]
fn cap_chat_input_preserves_whitespace_within_bounds() {
// In-bounds input comes back byte-identical — no normalization while
// the user is still editing.
let draft = " hello world \t ".to_string();
assert_eq!(cap_chat_input(draft.clone()), draft);
}
#[test]
fn cap_chat_input_truncates_oversized_paste_on_scalar_boundary() {
let paste = "x".repeat(CHAT_MSG_MAX_CHARS + 5000);
let out = cap_chat_input(paste);
assert_eq!(out.chars().count(), CHAT_MSG_MAX_CHARS);
let emoji_paste = "🎮".repeat(CHAT_MSG_MAX_CHARS + 100);
let out = cap_chat_input(emoji_paste);
assert_eq!(out.chars().count(), CHAT_MSG_MAX_CHARS);
assert!(out.len() <= CHAT_MSG_MAX_BYTES);
assert!(out.chars().all(|c| c == '🎮'));
}
#[test]
fn admit_rejects_oversized_raw_bytes_before_sanitizing() {
// One byte over the ceiling → rejected outright, attachment or not.
let over = "x".repeat(CHAT_MSG_MAX_BYTES + 1);
assert_eq!(admit_chat_text(&over, false), None);
assert_eq!(admit_chat_text(&over, true), None);
// Exactly at the ceiling → admitted (then sanitized/capped).
let at = "x".repeat(CHAT_MSG_MAX_BYTES);
let admitted = admit_chat_text(&at, false).expect("at-ceiling text admitted");
assert_eq!(admitted.chars().count(), CHAT_MSG_MAX_CHARS);
// Multibyte raw over the ceiling → rejected.
let cjk_over = "".repeat(CHAT_MSG_MAX_BYTES / 3 + 1);
assert!(cjk_over.len() > CHAT_MSG_MAX_BYTES);
assert_eq!(admit_chat_text(&cjk_over, false), None);
}
#[test]
fn admit_keeps_attachment_only_messages_and_drops_truly_empty_ones() {
// No visible text + no attachment → nothing to show, dropped.
assert_eq!(admit_chat_text("", false), None);
assert_eq!(admit_chat_text("\u{0}\r\n\t ", false), None);
// Same bodies WITH an attachment → kept as an empty caption.
assert_eq!(admit_chat_text("", true), Some(String::new()));
assert_eq!(admit_chat_text("\u{0}\r\n\t ", true), Some(String::new()));
// Normal text converges on the same result as direct sanitization.
assert_eq!(
admit_chat_text(" hi there ", false),
Some(sanitize_chat(" hi there "))
);
}
#[test]
fn chat_strips_bidi_overrides_but_keeps_benign_format_chars() {
// Overrides and isolates are removed outright (S14) …
assert_eq!(sanitize_chat("pay \u{202E}gpj.exe now"), "pay gpj.exe now");
assert_eq!(sanitize_chat("a\u{2066}b\u{2069}c"), "abc");
assert_eq!(
sanitize_chat("\u{202A}\u{202B}\u{202C}\u{202D}\u{202E}"),
""
);
// … while the expressive format characters chat promises to keep — ZWJ
// (emoji sequences), ZWNJ (joining scripts), LRM/RLM (bidi *marks*, which
// cannot reorder text) — survive.
for kept in ['\u{200D}', '\u{200C}', '\u{200E}', '\u{200F}'] {
let msg = format!("a{kept}b");
assert_eq!(sanitize_chat(&msg), msg, "stripped benign {kept:?}");
}
}
// --- link policy ---------------------------------------------------------
/// Concatenating every segment's slice must reproduce the input exactly.
fn reassemble(text: &str) -> String {
segments(text, &link_ranges(text))
.iter()
.map(|(s, _)| *s)
.collect()
}
/// The link slices of a message, in order.
fn links(text: &str) -> Vec<&str> {
segments(text, &link_ranges(text))
.into_iter()
.filter_map(|(s, is_link)| is_link.then_some(s))
.collect()
}
#[test]
fn linkify_plain_text_has_no_links() {
let segs = linkify("just a normal message, nothing here");
assert_eq!(
segs,
vec![Segment::Text("just a normal message, nothing here".into())]
);
fn url_policy_accepts_only_wellformed_web_urls() {
for ok in [
"http://example.com",
"https://a.test/path?q=1&w=2",
"HTTP://EXAMPLE.COM", // mixed case scheme+host
"https://x.com:8443/p", // explicit port
"https://d.com/路径?q=世界#frag", // unicode path/query/fragment
// WHATWG parsing (what browsers do) collapses the extra slash into
// host "path" — a valid, if odd, destination; not an empty host.
"http:///path",
] {
assert!(is_safe_web_url(ok), "rejected {ok:?}");
}
for bad in [
"",
"example.com", // no scheme
"http://", // empty host
"ftp://x.com", // non-web scheme
"file:///etc/passwd", // no host, wrong scheme
"javascript:alert(1)", // opener must never see this
"http://user@good.com", // userinfo → destination spoof risk
"http://user:pw@good.com", // credentials
"http://exa mple.com", // malformed host
] {
assert!(!is_safe_web_url(bad), "accepted {bad:?}");
}
}
#[test]
fn linkify_detects_http_and_https() {
fn link_ranges_detects_http_and_https_with_exact_roundtrip() {
assert_eq!(links("see http://example.com now"), ["http://example.com"]);
assert_eq!(
linkify("see http://example.com now"),
vec![
Segment::Text("see ".into()),
Segment::Link("http://example.com".into()),
Segment::Text(" now".into()),
]
links("a http://one.com b https://two.com c"),
["http://one.com", "https://two.com"]
);
assert_eq!(
linkify("https://a.test/path?q=1"),
vec![Segment::Link("https://a.test/path?q=1".into())]
);
}
#[test]
fn linkify_peels_trailing_punctuation() {
// Sentence-final period is not part of the link.
assert_eq!(
linkify("go to https://x.com."),
vec![
Segment::Text("go to ".into()),
Segment::Link("https://x.com".into()),
Segment::Text(".".into()),
]
);
// Parenthesized URL.
assert_eq!(
linkify("(https://x.com)"),
vec![
Segment::Text("(".into()),
Segment::Link("https://x.com".into()),
Segment::Text(")".into()),
]
);
}
#[test]
fn linkify_handles_multiple_urls() {
let segs = linkify("a http://one.com b https://two.com c");
assert_eq!(
segs,
vec![
Segment::Text("a ".into()),
Segment::Link("http://one.com".into()),
Segment::Text(" b ".into()),
Segment::Link("https://two.com".into()),
Segment::Text(" c".into()),
]
);
}
#[test]
fn linkify_only_matches_http_schemes() {
// Non-web schemes and bare domains are NOT linkified (conservative).
let segs = linkify("email me@x.com or ftp://x.com or visit x.com");
assert_eq!(
segs,
vec![Segment::Text(
"email me@x.com or ftp://x.com or visit x.com".into()
)]
);
}
#[test]
fn linkify_preserves_input_exactly() {
// Sentence-capitalized scheme still detected; href = the displayed slice.
assert_eq!(links("go to Http://example.com"), ["Http://example.com"]);
for msg in [
"",
"no urls at all",
@@ -309,12 +553,67 @@ mod tests {
"pre http://a.com/x?y=z&w=1 mid https://b.org/p, end!",
"weird))) http://c.com]]] tail",
"unicode 世界 http://d.com/路径 more 世界",
"bad http:// and http://user@x.com around https://ok.org here",
] {
assert_eq!(
reassemble(&linkify(msg)),
msg,
"roundtrip failed for {msg:?}"
);
assert_eq!(reassemble(msg), msg, "roundtrip failed for {msg:?}");
}
}
#[test]
fn link_ranges_peels_trailing_punctuation() {
assert_eq!(links("go to https://x.com."), ["https://x.com"]);
assert_eq!(links("(https://x.com)"), ["https://x.com"]);
}
#[test]
fn link_ranges_leaves_invalid_candidates_as_plain_text() {
// Non-web schemes and bare domains never linkify (conservative).
assert_eq!(
links("email me@x.com or ftp://x.com or visit x.com"),
[] as [&str; 0]
);
// A malformed/deceptive candidate stays text WITHOUT eating a later
// valid link.
assert_eq!(links("http:// then https://ok.org"), ["https://ok.org"]);
assert_eq!(
links("http://user:pw@evil.com vs https://good.com"),
["https://good.com"]
);
// An invalid run's interior is not re-scanned for nested schemes.
assert_eq!(links("http://a@http://b.com"), [] as [&str; 0]);
}
#[test]
fn link_ranges_caps_clickable_links_per_message() {
let many = (0..CHAT_MSG_MAX_LINKS + 4)
.map(|i| format!("https://site{i}.test"))
.collect::<Vec<_>>()
.join(" ");
let ranges = link_ranges(&many);
assert_eq!(ranges.len(), CHAT_MSG_MAX_LINKS);
// The 9th+ URLs remain, but as plain selectable text.
assert_eq!(reassemble(&many), many);
let l = links(&many);
assert_eq!(l.last(), Some(&"https://site7.test"));
// Exactly at the cap: all clickable.
let at_cap = (0..CHAT_MSG_MAX_LINKS)
.map(|i| format!("https://site{i}.test"))
.collect::<Vec<_>>()
.join(" ");
assert_eq!(link_ranges(&at_cap).len(), CHAT_MSG_MAX_LINKS);
}
#[test]
fn link_ranges_survives_adversarial_many_link_input() {
// A ceiling-length message packed with minimal URLs: bounded output,
// exact reconstruction, and every range on char boundaries.
let flood = "http://a.io ".repeat(CHAT_MSG_MAX_BYTES / 12 + 1);
let msg = sanitize_chat(&flood);
let ranges = link_ranges(&msg);
assert_eq!(ranges.len(), CHAT_MSG_MAX_LINKS);
for r in &ranges {
assert!(msg.is_char_boundary(r.start) && msg.is_char_boundary(r.end));
}
assert_eq!(reassemble(&msg), msg);
}
}
+314
View File
@@ -0,0 +1,314 @@
//! Live-edge catch-up for the screen-share viewer.
//!
//! PixelPass carries the share as MPEG-TS over a reliable, ordered transport. On
//! a lossy link (satellite handovers are the pathological case) every loss burst
//! becomes retransmission plus head-of-line blocking, and the viewer absorbs the
//! stall as buffered latency. Nothing in the chain ever trims that buffer back,
//! so the picture ends up seconds behind the host and stays there.
//!
//! Measured on a `tc netem` rig that simulates a satellite link (40 ms +/- 20 ms
//! jitter, 0.5% loss, a 250 ms/30%-loss handover burst every 15 s): a viewer with
//! ordinary timestamp pacing settles ~1.24 s behind. mpv's `--untimed` does NOT
//! help (~1.38 s, marginally worse) because it only removes pacing at
//! *presentation* while audio still drains at 1x the DAC rate, so an accumulated
//! buffer never shrinks. Returning to the live edge requires consuming the
//! backlog faster than it arrives.
//!
//! So we nudge playback slightly faster than realtime while the buffer is deep,
//! and drop back to 1x once it has drained. mpv's default pitch correction
//! (`scaletempo2`) keeps a 5% speedup inaudible, and because audio and video are
//! sped up together A/V sync is preserved — unlike `--untimed`.
//!
//! The control law and the JSON-IPC message handling are pure functions with
//! tests; the only I/O is [`drive`], which talks to mpv's `--input-ipc-server`
//! socket.
use std::path::{Path, PathBuf};
use std::time::Duration;
/// Buffer depth (seconds) above which we start draining.
pub const CACHE_HIGH_S: f64 = 1.0;
/// Buffer depth (seconds) below which we return to realtime.
pub const CACHE_LOW_S: f64 = 0.4;
/// The buffer depth we aim to sit at; the drain rate is proportional to how far
/// above this the buffer actually is.
pub const CACHE_TARGET_S: f64 = 0.5;
/// Extra playback rate per second of excess buffer.
pub const CATCHUP_GAIN: f64 = 0.05;
/// Hard ceiling on the drain rate. Beyond this the speedup stops being
/// unnoticeable, and a share that far behind is better served by the operator
/// restarting it than by a chipmunk impression.
pub const MAX_CATCHUP_SPEED: f64 = 1.15;
/// Normal realtime playback.
pub const NORMAL_SPEED: f64 = 1.0;
/// How often we sample the buffer depth.
pub const POLL_INTERVAL: Duration = Duration::from_millis(500);
/// Smallest rate change worth sending to the player.
pub const SPEED_EPSILON: f64 = 0.005;
/// The property we watch on the viewer.
const CACHE_PROPERTY: &str = "demuxer-cache-duration";
/// Decide the playback rate for the next interval.
///
/// Proportional, because a fixed small speedup cannot recover a large backlog in
/// any reasonable time: draining 6 s at 1.05x takes two minutes, which a viewer
/// experiences as "still broken". The drain rate instead scales with how deep
/// the buffer is, so a bad handover is cleared in tens of seconds while a small
/// excursion still gets only a gentle, inaudible nudge.
///
/// Deliberately hysteretic: between [`CACHE_LOW_S`] and [`CACHE_HIGH_S`] the
/// current rate is held, so a buffer hovering near a single threshold cannot
/// oscillate the speed (and with it the audio pitch) every poll. Pure.
///
/// A non-finite reading (mpv reports `null` before playback starts, and the
/// caller maps that to NaN) holds the current rate rather than guessing.
pub fn catchup_speed(cache_s: f64, current: f64) -> f64 {
if !cache_s.is_finite() {
return current;
}
if cache_s < CACHE_LOW_S {
return NORMAL_SPEED;
}
if cache_s <= CACHE_HIGH_S {
return current;
}
let excess = cache_s - CACHE_TARGET_S;
(NORMAL_SPEED + CATCHUP_GAIN * excess).clamp(NORMAL_SPEED, MAX_CATCHUP_SPEED)
}
/// Where mpv should create its IPC socket. Kept separate from the runtime
/// lookup so tests can pin a directory. Pure.
pub fn socket_path(dir: &Path, token: u64) -> PathBuf {
dir.join(format!("peerspeak-mpv-{token}.sock"))
}
/// The directory for the IPC socket: the XDG runtime dir when the session
/// provides one (tmpfs, user-private, cleaned at logout), else the temp dir.
pub fn socket_dir() -> PathBuf {
std::env::var_os("XDG_RUNTIME_DIR")
.map(PathBuf::from)
.unwrap_or_else(std::env::temp_dir)
}
/// A `get_property` request for the buffer depth. Pure.
pub fn get_cache_request(request_id: u64) -> String {
format!(r#"{{"command":["get_property","{CACHE_PROPERTY}"],"request_id":{request_id}}}"#)
}
/// A `set_property` request for the playback rate. Pure.
pub fn set_speed_request(request_id: u64, speed: f64) -> String {
format!(r#"{{"command":["set_property","speed",{speed}],"request_id":{request_id}}}"#)
}
/// Extract the buffer depth from one line of mpv's IPC output.
///
/// mpv interleaves unsolicited event lines with command replies, so a line is
/// only ours when it carries the matching `request_id`. Returns:
/// - `Some(Some(secs))` — our reply, with a usable number,
/// - `Some(None)` — our reply, but no number (mpv sends `"data":null` before
/// playback starts, and reports `error` while the demuxer has no cache yet),
/// - `None` — not our reply (an event, or another command's response).
///
/// Pure.
pub fn parse_cache_response(line: &str, request_id: u64) -> Option<Option<f64>> {
let value: serde_json::Value = serde_json::from_str(line.trim()).ok()?;
let id = value.get("request_id")?.as_u64()?;
if id != request_id {
return None;
}
if value.get("error").and_then(|e| e.as_str()) != Some("success") {
return Some(None);
}
Some(value.get("data").and_then(|d| d.as_f64()))
}
/// Drive one mpv viewer's playback rate over its JSON IPC socket.
///
/// Runs until mpv exits (the socket dies), so it is spawned detached alongside
/// the player and needs no shutdown signal. Every failure path just ends the
/// task: catch-up is an optimization, and a viewer that never gets it still
/// plays, exactly as before this existed.
#[cfg(unix)]
pub async fn drive(socket: PathBuf) {
use tokio::io::{AsyncBufReadExt, AsyncWriteExt, BufReader};
use tokio::net::UnixStream;
// mpv creates the socket a moment after exec, so the first connects race it.
let mut stream = None;
for _ in 0..40 {
match UnixStream::connect(&socket).await {
Ok(s) => {
stream = Some(s);
break;
}
Err(_) => tokio::time::sleep(Duration::from_millis(250)).await,
}
}
let Some(stream) = stream else {
crate::log_msg("livesync: mpv IPC socket never appeared; catch-up disabled");
return;
};
let (read_half, mut write_half) = stream.into_split();
let mut lines = BufReader::new(read_half).lines();
let mut request_id: u64 = 0;
let mut speed = NORMAL_SPEED;
loop {
tokio::time::sleep(POLL_INTERVAL).await;
request_id += 1;
let query = format!("{}\n", get_cache_request(request_id));
if write_half.write_all(query.as_bytes()).await.is_err() {
break;
}
// Skip event lines until our reply arrives.
let cache = loop {
match lines.next_line().await {
Ok(Some(line)) => {
if let Some(value) = parse_cache_response(&line, request_id) {
break value;
}
}
// Socket closed or unreadable: mpv is gone.
_ => return,
}
};
let cache = cache.unwrap_or(f64::NAN);
let next = catchup_speed(cache, speed);
// A proportional law would otherwise re-send on every wobble of the
// reading; only a change worth hearing is worth a round trip.
if (next - speed).abs() > SPEED_EPSILON {
speed = next;
request_id += 1;
let set = format!("{}\n", set_speed_request(request_id, speed));
if write_half.write_all(set.as_bytes()).await.is_err() {
break;
}
crate::log_msg(&format!(
"livesync: cache {cache:.2}s -> playback speed {speed}x"
));
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn deep_buffer_speeds_up_and_drained_buffer_returns_to_realtime() {
assert!(catchup_speed(1.5, NORMAL_SPEED) > NORMAL_SPEED);
assert_eq!(catchup_speed(0.1, MAX_CATCHUP_SPEED), NORMAL_SPEED);
}
#[test]
fn drain_rate_scales_with_how_far_behind_we_are() {
// The point of the proportional law: a small excursion gets a gentle
// nudge, a deep backlog gets real recovery.
let small = catchup_speed(1.5, NORMAL_SPEED);
let large = catchup_speed(4.0, NORMAL_SPEED);
assert!(
large > small,
"deeper buffer must drain faster: {small} vs {large}"
);
assert!(
(small - 1.05).abs() < 1e-9,
"1.5s buffer -> 1.05x, got {small}"
);
}
#[test]
fn drain_rate_is_capped_so_it_never_sounds_absurd() {
// The ~6 s standing buffer measured on the netem rig, and far worse.
assert_eq!(catchup_speed(6.0, NORMAL_SPEED), MAX_CATCHUP_SPEED);
assert_eq!(catchup_speed(600.0, NORMAL_SPEED), MAX_CATCHUP_SPEED);
}
#[test]
fn hysteresis_band_holds_the_current_speed() {
// Between the marks nothing changes, whichever side we came from —
// this is what stops the rate (and audio pitch) oscillating.
for cache in [CACHE_LOW_S, 0.7, CACHE_HIGH_S] {
assert_eq!(catchup_speed(cache, NORMAL_SPEED), NORMAL_SPEED);
assert_eq!(catchup_speed(cache, MAX_CATCHUP_SPEED), MAX_CATCHUP_SPEED);
}
}
#[test]
fn unknown_cache_holds_the_current_speed() {
assert_eq!(
catchup_speed(f64::NAN, MAX_CATCHUP_SPEED),
MAX_CATCHUP_SPEED
);
assert_eq!(catchup_speed(f64::INFINITY, NORMAL_SPEED), NORMAL_SPEED);
}
#[test]
fn a_full_handover_cycle_drains_then_settles() {
// Buffer grows through a loss burst, then drains as we play faster.
let mut speed = NORMAL_SPEED;
for cache in [0.2, 0.5, 1.2, 3.4, 1.4, 0.9, 0.6, 0.3, 0.2] {
speed = catchup_speed(cache, speed);
}
assert_eq!(
speed, NORMAL_SPEED,
"should be back at realtime once drained"
);
}
#[test]
fn requests_are_valid_json_with_their_ids() {
let get: serde_json::Value = serde_json::from_str(&get_cache_request(7)).unwrap();
assert_eq!(get["request_id"], 7);
assert_eq!(get["command"][0], "get_property");
assert_eq!(get["command"][1], CACHE_PROPERTY);
let set: serde_json::Value = serde_json::from_str(&set_speed_request(8, 1.05)).unwrap();
assert_eq!(set["request_id"], 8);
assert_eq!(set["command"][0], "set_property");
assert_eq!(set["command"][1], "speed");
assert_eq!(set["command"][2], 1.05);
}
#[test]
fn parses_our_reply_only() {
assert_eq!(
parse_cache_response(r#"{"error":"success","data":1.25,"request_id":3}"#, 3),
Some(Some(1.25))
);
// Another command's reply, and an unsolicited event, are not ours.
assert_eq!(
parse_cache_response(r#"{"error":"success","data":1.25,"request_id":4}"#, 3),
None
);
assert_eq!(
parse_cache_response(r#"{"event":"playback-restart"}"#, 3),
None
);
assert_eq!(parse_cache_response("not json", 3), None);
}
#[test]
fn reply_without_a_usable_number_is_ours_but_empty() {
// mpv before playback starts, and while the demuxer has no cache.
assert_eq!(
parse_cache_response(r#"{"error":"success","data":null,"request_id":1}"#, 1),
Some(None)
);
assert_eq!(
parse_cache_response(r#"{"error":"property unavailable","request_id":1}"#, 1),
Some(None)
);
}
#[test]
fn socket_path_is_scoped_to_its_token() {
let a = socket_path(Path::new("/run/user/1000"), 42);
assert_eq!(a, Path::new("/run/user/1000/peerspeak-mpv-42.sock"));
assert_ne!(a, socket_path(Path::new("/run/user/1000"), 43));
}
}
+463 -39
View File
@@ -21,6 +21,12 @@ use std::time::Duration;
use tokio::io::{AsyncBufReadExt, BufReader};
use tokio::process::{Child, Command};
use crate::audio::ownership;
pub mod livesync;
use crate::config::{ScreenShareSettings, ShareBuffering, SharePlayer, ShareQuality};
/// The binary we shell out to. Looked up on `$PATH` unless a config override
/// points elsewhere.
const PIXELPASS_BIN: &str = "pixelpass";
@@ -43,6 +49,12 @@ const MAX_TICKET_LEN: usize = 512;
/// are short ("Firefox", "mpv"); this only guards against a pathological value.
const MAX_APP_NAME_LEN: usize = 256;
/// Ceiling on the viewer's demuxer byte cache in the Low latency posture. The
/// cache is a *byte* budget, so at a given bitrate it sets the worst-case
/// backlog in seconds; keeping it tight is what stops a lossy link parking the
/// viewer seconds behind before [`livesync`] even gets a chance to drain it.
const LOW_LATENCY_CACHE_CAP_MB: u32 = 1;
/// How long to wait for the host to emit its ticket / the viewer to connect
/// before giving up and killing the child. Startup is normally sub-second; this
/// is only a safety net so a hung pixelpass can't wedge the caller forever.
@@ -139,7 +151,11 @@ fn json_u32(v: &serde_json::Value, key: &str) -> u32 {
/// otherwise rejects hyphen-leading option values). The name is locally chosen
/// (our own enumeration / the user's pick), not peer-supplied, but is still
/// sanitized via [`sanitize_app_name`] before reaching here. Pure: no I/O.
pub fn host_args(audio_app: Option<&str>) -> Vec<String> {
pub fn host_args(
audio_app: Option<&str>,
settings: &ScreenShareSettings,
quality: ShareQuality,
) -> Vec<String> {
let mut args = vec![
"--host".to_string(),
"--output".to_string(),
@@ -149,9 +165,44 @@ pub fn host_args(audio_app: Option<&str>) -> Vec<String> {
args.push(format!("--app={name}"));
args.push("--strict-audio".to_string());
}
if quality != ShareQuality::Auto {
args.push(format!("--quality={}", pixelpass_quality(quality)));
}
if let Some(height) = settings.max_height {
args.push(format!("--max-height={height}"));
}
if let Some(mbps) = settings.bitrate_mbps {
args.push(format!("--bitrate={}", mbps.saturating_mul(1000)));
}
if let Some(fps) = settings.framerate {
args.push(format!("--framerate={fps}"));
}
if settings.force_software_encode {
args.push("--no-hwencode".to_string());
}
if let Some(max) = settings.max_viewers {
args.push(format!("--max-viewers={max}"));
}
args.extend(split_extra_args(&settings.extra_host_args));
args
}
fn pixelpass_quality(quality: ShareQuality) -> &'static str {
match quality {
ShareQuality::Auto => "auto",
ShareQuality::Low => "low",
ShareQuality::Medium => "medium",
ShareQuality::High => "high",
ShareQuality::Source => "source",
}
}
/// Split user-supplied advanced argv text into separate tokens. Peerspeak does
/// not depend on a shell lexer, so quoted values are not interpreted here.
fn split_extra_args(raw: &str) -> impl Iterator<Item = String> + '_ {
raw.split_whitespace().map(str::to_string)
}
/// Validate a locally-chosen audio app name before it becomes a `--app` value:
/// trim, reject empty / overlong, and reject names carrying control characters
/// (newlines etc.) that have no place in a real `application.name`. `None` means
@@ -320,15 +371,26 @@ pub fn is_available(config_override: Option<&str>) -> bool {
/// whole desktop sink, which avoids the call-loopback echo (A23). The child keeps
/// running (streaming to viewers) until killed or dropped; remaining stdout is
/// drained in a background task so a full pipe can't stall the host. We do
/// **not** pass `--max-viewers`: pixelpass bandwidth-measures its own safe cap,
/// protecting the sharer's uplink, and refuses extras with `viewer_refused`.
/// not pass encode/viewer overrides unless the local settings explicitly ask for
/// them, so pixelpass keeps its own defaults in the common case.
pub async fn spawn_host(
bin: &Path,
audio_app: Option<&str>,
settings: &ScreenShareSettings,
quality: ShareQuality,
notices: Option<tokio::sync::mpsc::UnboundedSender<PixelpassEvent>>,
) -> std::io::Result<(Child, String)> {
let args = host_args(audio_app, settings, quality);
// Log the exact argv we hand pixelpass so a field log can confirm which
// encode/quality flags (e.g. --bitrate) actually reached the host — these
// are local flags with no ticket/secret, so logging them verbatim is safe.
crate::log_msg(&format!(
"pixelpass host spawn: {} {}",
bin.display(),
args.join(" ")
));
let mut child = Command::new(bin)
.args(host_args(audio_app))
.args(&args)
.stdin(Stdio::null())
.stdout(Stdio::piped())
// Capture stderr (not null): pixelpass prints its startup precondition
@@ -428,10 +490,14 @@ pub fn pixelpass_failure_detail(stderr: &str) -> String {
}
/// Spawn a pixelpass viewer for `ticket`, wait for it to connect, and open the
/// stream in a local player (mpv, falling back to vlc). Returns the live viewer
/// child so the caller can kill it on room-leave; it also self-exits when the
/// player window closes (its tunnel ends).
pub async fn spawn_viewer(bin: &Path, ticket: &str) -> std::io::Result<Child> {
/// stream in a local player (mpv/VLC in the configured order, then fallback).
/// Returns the live viewer child so the caller can kill it on room-leave; it also
/// self-exits when the player window closes (its tunnel ends).
pub async fn spawn_viewer(
bin: &Path,
ticket: &str,
settings: &ScreenShareSettings,
) -> std::io::Result<Child> {
let mut child = Command::new(bin)
.args(viewer_args(ticket))
.stdin(Stdio::null())
@@ -465,7 +531,7 @@ pub async fn spawn_viewer(bin: &Path, ticket: &str) -> std::io::Result<Child> {
}
};
if let Err(e) = launch_player(&url) {
if let Err(e) = launch_player(&url, settings) {
let _ = child.kill().await;
return Err(e);
}
@@ -547,29 +613,68 @@ fn event_for_log(ev: &PixelpassEvent) -> String {
}
}
/// Open the viewer stream URL in a media player. Mirrors pixelpass's own
/// low-latency mpv invocation; falls back to vlc. The player is reaped in a
/// background task so it doesn't linger as a zombie when its window closes.
fn launch_player(url: &str) -> std::io::Result<()> {
const MPV_ARGS: &[&str] = &[
"--profile=low-latency",
"--untimed",
"--hwdec=auto",
"--audio-buffer=0.2",
"--demuxer-max-bytes=2M",
"--demuxer-readahead-secs=0.5",
];
const VLC_ARGS: &[&str] = &["--network-caching=200", "--live-caching=200"];
/// Open the viewer stream URL in a media player, then fall back to vlc. The
/// player is reaped in a background task so it doesn't linger as a zombie when
/// its window closes.
///
/// The buffering posture chooses the latency/A/V-sync tradeoff. Low latency
/// keeps the viewer at the live edge: mpv gets an IPC socket and [`livesync`]
/// drains a lagging buffer by playing slightly fast (pitch-corrected, so A/V
/// sync is preserved). Smooth leaves a deeper buffer alone, trading live
/// latency for immunity to jitter. Hardware decoding remains opt-in: forcing
/// `--hwdec=auto` froze some viewers on frame 1 while audio kept playing.
fn launch_player(url: &str, settings: &ScreenShareSettings) -> std::io::Result<()> {
// One socket per viewer launch, so overlapping shares can't collide on it.
// Unix only: mpv's IPC is a named pipe on Windows, which `livesync` does not
// speak, and an unusable socket path on the argv would help nobody.
#[cfg(unix)]
let ipc_socket = Some(livesync::socket_path(
&livesync::socket_dir(),
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_nanos() as u64)
.unwrap_or(0),
));
#[cfg(not(unix))]
let ipc_socket: Option<PathBuf> = None;
let child = match spawn_player("mpv", MPV_ARGS, url) {
Ok(c) => c,
Err(_) => spawn_player("vlc", VLC_ARGS, url).map_err(|_| {
std::io::Error::new(
std::io::ErrorKind::NotFound,
"no media player found — install mpv or vlc to watch screen shares",
)
})?,
let mpv_args = mpv_args(settings, ipc_socket.as_deref());
let vlc_args = vlc_args(settings);
let first = match settings.player {
SharePlayer::Mpv => ("mpv", &mpv_args),
SharePlayer::Vlc => ("vlc", &vlc_args),
};
let second = match settings.player {
SharePlayer::Mpv => ("vlc", &vlc_args),
SharePlayer::Vlc => ("mpv", &mpv_args),
};
let (launched, child) = match spawn_player(first.0, first.1, url) {
Ok(c) => (first.0, c),
Err(_) => (
second.0,
spawn_player(second.0, second.1, url).map_err(|_| {
std::io::Error::new(
std::io::ErrorKind::NotFound,
"no media player found — install mpv or vlc to watch screen shares",
)
})?,
),
};
// Only when the socket actually reached the argv: mpv (VLC has no
// equivalent IPC) in the Low latency posture. The driver ends by itself when
// the player exits, so it needs no shutdown path.
#[cfg(unix)]
if launched == "mpv"
&& settings.buffering == ShareBuffering::LowLatency
&& let Some(socket) = ipc_socket
{
tokio::spawn(livesync::drive(socket));
}
#[cfg(not(unix))]
let _ = launched;
tokio::spawn(async move {
let mut child = child;
let _ = child.wait().await;
@@ -577,21 +682,154 @@ fn launch_player(url: &str) -> std::io::Result<()> {
Ok(())
}
fn spawn_player(bin: &str, args: &[&str], url: &str) -> std::io::Result<Child> {
Command::new(bin)
/// Build the argv for an mpv viewer.
///
/// `ipc_socket` is where mpv should expose its JSON IPC socket so [`livesync`]
/// can drain a lagging buffer. It is wired up for Low latency only: Smooth
/// deliberately holds a ~2 s readahead, which the catch-up thresholds would
/// fight on every poll.
pub fn mpv_args(settings: &ScreenShareSettings, ipc_socket: Option<&Path>) -> Vec<String> {
let mut args = Vec::new();
match settings.buffering {
ShareBuffering::LowLatency => {
args.push("--profile=low-latency".to_string());
// Pixelpass carries MPEG-TS through reliable ordered QUIC/TCP, so a
// lossy link turns every retransmission into buffered latency that
// nothing trims back. `--untimed` does NOT fix that (measured
// marginally worse: it only unpaces *presentation*, while audio
// still drains at 1x, so the backlog never shrinks) — the viewer
// instead drains it by playing slightly fast, see `livesync`.
args.push("--audio-buffer=0.2".to_string());
args.push("--demuxer-readahead-secs=0.5".to_string());
}
ShareBuffering::Smooth => {
args.push("--cache=yes".to_string());
args.push("--demuxer-readahead-secs=2".to_string());
}
}
// The byte cap is what bounds how far behind a viewer can silently fall:
// a demuxer allowed 2 MiB will happily sit on ~6 s of a 2.5 Mbps share (as
// measured on the netem rig) and call it a buffer. Low latency therefore
// gets a tighter ceiling than the user's Smooth-oriented setting, so the
// catch-up has less to claw back after a bad patch of link.
let cache_mb = match settings.buffering {
ShareBuffering::LowLatency => settings.cache_mb.min(LOW_LATENCY_CACHE_CAP_MB),
ShareBuffering::Smooth => settings.cache_mb,
};
args.push(format!("--demuxer-max-bytes={cache_mb}M"));
if settings.hardware_decode {
args.push("--hwdec=auto".to_string());
}
if let Some(socket) = ipc_socket
&& settings.buffering == ShareBuffering::LowLatency
{
args.push(format!("--input-ipc-server={}", socket.display()));
}
// Extra args stay last so a user override wins over everything above.
args.extend(split_extra_args(&settings.extra_mpv_args));
args
}
/// Build the argv for a VLC viewer. VLC honors the subset of viewer settings
/// that map cleanly onto its option set: the buffering posture (network/live
/// caching, in ms) and hardware decoding. The rest of the viewer knobs are
/// mpv-specific — `cache_mb` is an mpv demuxer *byte* cache (VLC's caching is
/// time-based, already covered by `buffering`) and `extra_mpv_args` is literally
/// mpv flags — so they are deliberately not mapped here; the Settings UI labels
/// them as mpv-only. Pure: no I/O.
///
/// The hardware-decode mapping is the load-bearing one: VLC hardware-decodes by
/// default, so without an explicit `--avcodec-hw=none` a VLC viewer would ignore
/// the (default-off) hardware-decode toggle and could hit the frame-1 freeze
/// that default exists to avoid — the same A-bug that made us drop mpv's forced
/// `--hwdec=auto`.
fn vlc_args(settings: &ScreenShareSettings) -> Vec<String> {
let caching_ms = match settings.buffering {
ShareBuffering::LowLatency => 200,
ShareBuffering::Smooth => 1500,
};
let hw = if settings.hardware_decode {
"--avcodec-hw=any"
} else {
"--avcodec-hw=none"
};
vec![
format!("--network-caching={caching_ms}"),
format!("--live-caching={caching_ms}"),
hw.to_string(),
]
}
fn spawn_player(bin: &str, args: &[String], url: &str) -> std::io::Result<Child> {
// Log the player + its flags (mpv/vlc, incl. hardware-decode: --hwdec /
// --avcodec-hw) so a field log can confirm the viewer settings reached the
// player. The `url` is omitted deliberately — it is the local stream address
// and is not needed to verify the flags. Logged on each attempt, so a
// fallback from the preferred player to the other one is visible too.
crate::log_msg(&format!("player spawn: {bin} {}", args.join(" ")));
let mut command = Command::new(bin);
command
.args(args)
.arg(url)
.stdin(Stdio::null())
.stdout(Stdio::null())
.stderr(Stdio::null())
.kill_on_drop(false)
.spawn()
.kill_on_drop(false);
// Ownership tag (plan §5.1): this player is playing the *incoming*
// screenshare's audio, so it is exactly what must not be fanned back out
// if this machine also starts sharing. The role is the player binary, so
// a `pw-dump` during a field test names which one produced the node.
ownership::tag_child(command.as_std_mut(), bin);
command.spawn()
}
#[cfg(test)]
mod tests {
use super::*;
/// Phase-1 exit gate, player half (impl plan §3): the mpv peerspeak
/// actually spawns produces a live node carrying **both** ownership
/// carriers, tagged with the player's own name as the role.
///
/// ⚠️ Drives the real [`spawn_player`], for the same reason the notify
/// gate does: the plan requires the tag to be shown "landing on a live
/// mpv node, not just in the env". Plays a silent WAV, so it is quiet.
///
/// Live: needs PipeWire, `mpv` and `pw-dump`.
/// `cargo test --lib -- --ignored spawned_player`
#[tokio::test]
#[ignore = "live: requires a running PipeWire daemon, mpv and pw-dump"]
async fn spawned_player_node_carries_both_ownership_carriers() {
use crate::audio::ownership::live_test;
let dir = std::env::temp_dir().join(format!("peerspeak-playertest-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
let path = dir.join("silence.wav");
std::fs::write(&path, live_test::silent_wav(6)).unwrap();
let mut child = spawn_player(
"mpv",
&["--no-video".to_string(), "--really-quiet".to_string()],
path.to_str().unwrap(),
)
.expect("mpv spawns");
// The role is the player binary, so this also pins that the call site
// passes `bin` and not a fixed literal.
let prefix = live_test::expected_prefix("mpv");
let found = live_test::poll_for_owned_node(&prefix, std::time::Duration::from_secs(5));
let _ = child.kill().await;
std::fs::remove_dir_all(&dir).ok();
let (name, owned) =
found.unwrap_or_else(|| panic!("no live node named {prefix:?} appeared within 5s"));
assert!(
name.starts_with(ownership::OWNED_NODE_NAME_PREFIX),
"{name}"
);
assert_eq!(owned.as_deref(), Some(ownership::OWNED_PROP_VALUE));
}
#[test]
fn viewer_args_guard_neutralizes_flag_like_ticket() {
// A malicious "ticket" that looks like a flag must end up positional,
@@ -621,7 +859,11 @@ mod tests {
fn host_args_without_app_shares_whole_desktop() {
// No app selected → no --app flag → pixelpass keeps its default
// (whole-desktop) audio capture.
assert_eq!(host_args(None), vec!["--host", "--output", "json"]);
let settings = ScreenShareSettings::default();
assert_eq!(
host_args(None, &settings, ShareQuality::Auto),
vec!["--host", "--output", "json"]
);
}
#[test]
@@ -629,8 +871,9 @@ mod tests {
// The chosen app rides in the `--app=<name>` single-token form so a
// name beginning with `-` can never be reparsed as a flag (A23), plus
// `--strict-audio` so pixelpass never falls back to whole-desktop audio.
let settings = ScreenShareSettings::default();
assert_eq!(
host_args(Some("Firefox")),
host_args(Some("Firefox"), &settings, ShareQuality::Auto),
vec![
"--host",
"--output",
@@ -641,7 +884,7 @@ mod tests {
);
// The hyphen-leading name is still bound to --app as a single token;
// --strict-audio is the trailing flag.
let args = host_args(Some("-rm -rf"));
let args = host_args(Some("-rm -rf"), &settings, ShareQuality::Auto);
assert_eq!(args[3], "--app=-rm -rf");
assert_eq!(args[4], "--strict-audio");
}
@@ -650,11 +893,192 @@ mod tests {
fn host_args_blank_or_control_app_is_dropped() {
// An empty / whitespace / control-laden selection is sanitized away,
// falling back to whole-desktop capture rather than a broken flag.
assert_eq!(host_args(Some(" ")), vec!["--host", "--output", "json"]);
let settings = ScreenShareSettings::default();
assert_eq!(
host_args(Some("bad\nname")),
host_args(Some(" "), &settings, ShareQuality::Auto),
vec!["--host", "--output", "json"]
);
assert_eq!(
host_args(Some("bad\nname"), &settings, ShareQuality::Auto),
vec!["--host", "--output", "json"]
);
}
#[test]
fn host_args_apply_screen_share_settings_and_extra_args_last() {
let settings = ScreenShareSettings {
bitrate_mbps: Some(5),
framerate: Some(60),
max_height: Some(1080),
max_viewers: Some(4),
force_software_encode: true,
extra_host_args: "--relay https://relay.example --verbose".to_string(),
..ScreenShareSettings::default()
};
assert_eq!(
host_args(Some("Firefox"), &settings, ShareQuality::High),
vec![
"--host",
"--output",
"json",
"--app=Firefox",
"--strict-audio",
"--quality=high",
"--max-height=1080",
"--bitrate=5000",
"--framerate=60",
"--no-hwencode",
"--max-viewers=4",
"--relay",
"https://relay.example",
"--verbose",
]
);
}
#[test]
fn mpv_args_default_matches_low_latency_software_decode() {
assert_eq!(
mpv_args(&ScreenShareSettings::default(), None),
vec![
"--profile=low-latency",
"--audio-buffer=0.2",
"--demuxer-readahead-secs=0.5",
"--demuxer-max-bytes=1M",
]
);
}
#[test]
fn low_latency_gets_the_ipc_socket_for_live_edge_catch_up() {
let args = mpv_args(
&ScreenShareSettings::default(),
Some(Path::new("/run/user/1000/peerspeak-mpv-1.sock")),
);
assert!(
args.contains(&"--input-ipc-server=/run/user/1000/peerspeak-mpv-1.sock".to_string()),
"low latency drains a lagging buffer over mpv IPC: {args:?}"
);
// The flag that used to hold this posture at the live edge measured no
// better than pacing, and cost A/V sync — it must not come back.
assert!(!args.contains(&"--untimed".to_string()));
}
#[test]
fn smooth_keeps_its_deep_buffer_and_gets_no_ipc_socket() {
let settings = ScreenShareSettings {
buffering: ShareBuffering::Smooth,
..ScreenShareSettings::default()
};
let args = mpv_args(
&settings,
Some(Path::new("/run/user/1000/peerspeak-mpv-1.sock")),
);
assert!(
!args.iter().any(|a| a.starts_with("--input-ipc-server")),
"catch-up would fight Smooth's deliberate ~2s readahead: {args:?}"
);
}
#[test]
fn low_latency_caps_the_byte_cache_but_smooth_keeps_the_user_value() {
// The cache is a byte budget, so at a given bitrate it sets the
// worst-case backlog: 2 MiB held ~6 s of a 2.5 Mbps share on the rig.
let generous = ScreenShareSettings {
cache_mb: 32,
..ScreenShareSettings::default()
};
assert!(
mpv_args(&generous, None)
.contains(&format!("--demuxer-max-bytes={LOW_LATENCY_CACHE_CAP_MB}M")),
"low latency must bound how far behind the viewer can silently fall"
);
let smooth = ScreenShareSettings {
cache_mb: 32,
buffering: ShareBuffering::Smooth,
..ScreenShareSettings::default()
};
assert!(
mpv_args(&smooth, None).contains(&"--demuxer-max-bytes=32M".to_string()),
"smooth is the posture where the user asked for a deep buffer"
);
}
#[test]
fn user_extra_args_still_come_last() {
let settings = ScreenShareSettings {
extra_mpv_args: "--no-osc".to_string(),
..ScreenShareSettings::default()
};
let args = mpv_args(&settings, Some(Path::new("/tmp/s.sock")));
assert_eq!(
args.last().map(String::as_str),
Some("--no-osc"),
"a user override has to win over everything we add: {args:?}"
);
}
#[test]
fn mpv_args_smooth_hwdecode_and_extra_args_last() {
let settings = ScreenShareSettings {
hardware_decode: true,
buffering: ShareBuffering::Smooth,
cache_mb: 16,
extra_mpv_args: "--no-osc --vd-lavc-threads=2".to_string(),
..ScreenShareSettings::default()
};
assert_eq!(
mpv_args(&settings, None),
vec![
"--cache=yes",
"--demuxer-readahead-secs=2",
"--demuxer-max-bytes=16M",
"--hwdec=auto",
"--no-osc",
"--vd-lavc-threads=2",
]
);
}
#[test]
fn vlc_args_default_disables_hardware_decode() {
// The A-bug fix default (hardware_decode = false) must reach VLC too:
// VLC hardware-decodes by default, so without an explicit
// `--avcodec-hw=none` a VLC viewer would ignore the toggle and could hit
// the frame-1 freeze. Low-latency buffering keeps the 200 ms caches.
assert_eq!(
vlc_args(&ScreenShareSettings::default()),
vec![
"--network-caching=200",
"--live-caching=200",
"--avcodec-hw=none",
]
);
}
#[test]
fn vlc_args_smooth_buffering_and_hwdecode() {
// Enabling hardware decode flips VLC to `--avcodec-hw=any`; Smooth
// buffering raises the network/live caches. cache_mb / extra_mpv_args are
// mpv-only and must NOT leak into the VLC argv.
let settings = ScreenShareSettings {
hardware_decode: true,
buffering: ShareBuffering::Smooth,
cache_mb: 16,
extra_mpv_args: "--no-osc".to_string(),
..ScreenShareSettings::default()
};
assert_eq!(
vlc_args(&settings),
vec![
"--network-caching=1500",
"--live-caching=1500",
"--avcodec-hw=any",
]
);
}
#[test]
+64 -6
View File
@@ -70,6 +70,13 @@ pub fn paste(value: &str, start: usize, end: usize, clip: &str) -> Edit {
}
}
/// Strip control characters (e.g. a trailing newline on an X11 PRIMARY
/// selection) from clipboard text before it is pasted. Shared by the
/// right-click menu Paste and the middle-click PRIMARY paste.
pub fn sanitize_clip(raw: &str) -> String {
raw.chars().filter(|c| !c.is_control()).collect()
}
pub fn select_all_range(value: &str) -> (usize, usize) {
let value = text_input::Value::new(value);
@@ -362,6 +369,48 @@ where
return;
}
// Middle-click pastes the X11 PRIMARY selection at the cursor. iced's
// base text_input only wires Ctrl+V to the Standard (CLIPBOARD)
// selection, so without this the common "select text, middle-click to
// paste" workflow does nothing on X11.
let middle_click_on_input = matches!(
event,
Event::Mouse(mouse::Event::ButtonPressed(mouse::Button::Middle))
) && cursor.is_over(layout.bounds());
if middle_click_on_input && !self.locked {
let clip = sanitize_clip(&clipboard.read(clipboard::Kind::Primary).unwrap_or_default());
if !clip.is_empty() {
let value = text_input::Value::new(&self.value);
let input_state = tree.children[0]
.state
.downcast_mut::<text_input::State<Renderer::Paragraph>>();
let (start, end) = match input_state.cursor().state(&value) {
text_input::cursor::State::Index(index) => {
let index = index.min(value.len());
(index, index)
}
text_input::cursor::State::Selection { start, end } => {
normalized_range(&value, start, end)
}
};
let edit = paste(&self.value, start, end, &clip);
input_state.move_cursor_to(edit.cursor);
if let Some(on_paste) = &self.on_paste {
shell.publish(on_paste.as_ref()(edit.value));
} else if let Some(on_input) = &self.on_input {
shell.publish(on_input.as_ref()(edit.value));
}
}
shell.capture_event();
shell.request_redraw();
return;
}
Widget::update(
&mut self.input,
&mut tree.children[0],
@@ -717,12 +766,11 @@ where
}
}
MenuAction::Paste => {
let clip = clipboard
.read(clipboard::Kind::Standard)
.unwrap_or_default()
.chars()
.filter(|c| !c.is_control())
.collect::<String>();
let clip = sanitize_clip(
&clipboard
.read(clipboard::Kind::Standard)
.unwrap_or_default(),
);
let edit = paste(self.value, start, end, &clip);
self.publish_paste(edit, shell);
@@ -842,6 +890,16 @@ mod tests {
assert_eq!(clip, None);
}
#[test]
fn sanitize_clip_strips_control_chars_keeps_text() {
// An X11 PRIMARY selection commonly carries a trailing newline.
assert_eq!(sanitize_clip("pixelpassF1:abc\n"), "pixelpassF1:abc");
assert_eq!(sanitize_clip("a\tb\r\nc"), "abc");
// Non-control unicode is preserved.
assert_eq!(sanitize_clip("héllo🦀"), "héllo🦀");
assert_eq!(sanitize_clip(""), "");
}
#[test]
fn paste_replaces_selection_or_inserts_at_cursor() {
assert_eq!(
+1 -1
View File
@@ -21,7 +21,7 @@ const HIT_SEARCH_STEPS: usize = 24;
// `Hit::CharOffset(cursor.index)`, and cosmic-text's `cursor.index` is a byte
// offset WITHIN its buffer line — it discards the line number. That equals the
// global byte offset only when the text is a single logical line. Chat bodies
// satisfy this because `app::sanitize_chat` turns every control char (incl. `\n`
// satisfy this because `sanitize::sanitize_chat` turns every control char (incl. `\n`
// and `\r`) into a space and collapses whitespace, so a stored message can never
// contain a newline. If that sanitizer ever starts preserving newlines, this
// widget's per-line offsets would stop being global and selection/copy across
+42
View File
@@ -0,0 +1,42 @@
# Screenshare audio exclusion — ownership tagging wire contract.
#
# peerspeak PRODUCES these carriers on every audio node it owns; pixelpass
# CONSUMES them as the primary taint root of the exclusion engine. Neither
# repo depends on the other, so this file is the contract: it is committed
# byte-identical in both, and each repo has a test that asserts its own named
# constants (and, on the producer side, the environment a real child Command
# would carry) match these values exactly.
#
# peerspeak/tests/fixtures/ownership-tag-contract.txt
# pixelpass/tests/fixtures/ownership-tag-contract.txt
#
# Pinned by peerspeak docs/screenshare-audio-exclusion-impl-plan.md §3 and
# docs/screenshare-audio-exclusion-plan.md §5.1 (v3.5). Changing a value here
# is a cross-repo breaking change: both repos must land in the same session,
# and the phase 5 matrix must be re-run.
#
# Two carriers, matched as a UNION — a node is peerspeak-owned if EITHER
# matches. Round 8 added the second because a property is invisible to the
# PipeWire registry `global` event and readable only via a node bind, so the
# primary taint root must not rest on one observation mechanism alone.
# Carrier 1 — a node property, matched EXACTLY: `prop_value` below is the
# ONLY spelling the consumer reads as owned. A producer emitting "true", "yes"
# or "" is NOT owned on this carrier, and only carrier 2 would still catch it.
#
# ⚠️ This wording is load-bearing and it CHANGED in round 10. The consumer
# used to accept any value other than "false"/"0", on the theory that leniency
# over-excludes and is therefore safe. It is not: leniency buys false-positive
# exclusion, and it let any process suppress a rival application's audio from
# the share with a property it did not even have to spell right. Fail-closed
# on this feature is about ANCESTRY — an unresolvable graph is not eligible —
# not about parsing.
prop_key=peerspeak.owned
prop_value=1
# Carrier 2 — a `node.name` prefix, announced by the registry without a bind.
# `node.description` is deliberately NOT touched, so mixers still show "mpv".
# Only the prefix is matched; the rest of the name is for diagnostics.
node_name_prefix=peerspeak_owned_
node_name_format=peerspeak_owned_<role>_<pid>
node_name_example=peerspeak_owned_mpv_31284
+36
View File
@@ -201,3 +201,39 @@ async fn rejoin_after_grace_eviction_dials_cleanly() {
"an initial dial after a grace eviction must not be treated as a reconnect"
);
}
/// Chat-hardening Phase 2: a peer mid-reconnect-grace keeps its roster-bound
/// chat name (its chat stays admitted), but a TERMINAL grace-expiry eviction
/// revokes it — after that, only a fresh authenticated Announce (PeerJoined)
/// restores chat authority.
#[tokio::test]
async fn grace_eviction_revokes_chat_roster_entry() {
use peerspeak::core::chatroster::ChatRoster;
let (ui_tx, mut ui_rx) = mpsc::channel(100);
let roster = ChatRoster::default();
let h = make_handler(ui_tx, make_transport().await, GRACE).with_chat_roster(roster.clone());
let peer = fake_peer();
roster.upsert(peer, "Victim");
// Link up, then drop: DURING the grace window the peer is still a member —
// its chat must keep rendering under its roster name.
h.handle(ConnEvent::Connected(peer)).await;
h.handle(ConnEvent::Connecting(peer)).await;
assert_eq!(
roster.name_of(&peer),
Some("Victim".to_string()),
"reconnect grace must NOT revoke chat authority"
);
// Once the grace expires and the eviction fires, chat authority goes too.
assert!(
evicted_within(&mut ui_rx, peer, GRACE * 4).await,
"the outage should evict once the grace window elapses"
);
assert_eq!(
roster.name_of(&peer),
None,
"terminal eviction must revoke the roster-bound chat name"
);
}
+158
View File
@@ -244,6 +244,73 @@ async fn loopback_sequenced_audio_reaches_peer_and_decodes() {
);
}
/// Connection transparency: over a real loopback link, `connection_stats()`
/// must report the peer's selected path as direct (relay disabled here), with
/// an IP remote address and counters that advance while audio flows — and the
/// `connstats::derive` seam must turn two such snapshots into badge info with
/// live rates.
#[tokio::test]
async fn connection_stats_report_a_direct_path_with_live_counters() {
let a = spawn_node().await;
let b = spawn_node().await;
a.lookup.add_endpoint_info(b.endpoint.addr());
b.lookup.add_endpoint_info(a.endpoint.addr());
let a_id = a.endpoint.id();
let b_id = b.endpoint.id();
a.transport.admit_audio_sender(b_id);
b.transport.admit_audio_sender(a_id);
// Keep B's receive path subscribed like production (drained implicitly).
let _b_rx = b.transport.receive_datagrams().await.expect("subscribe B");
a.transport.connect_peer(b.endpoint.addr()).await;
b.transport.connect_peer(a.endpoint.addr()).await;
tokio::time::sleep(Duration::from_millis(500)).await;
let snap = |stats: Vec<(iroh::EndpointId, peerspeak::network::PathSnapshot)>| {
stats
.into_iter()
.find(|(id, _)| *id == b_id)
.map(|(_, s)| s)
.expect("peer B should appear in A's connection stats")
};
let s1 = snap(a.transport.connection_stats());
assert!(!s1.is_relay, "loopback with relay disabled must be direct");
assert!(
s1.remote_addr.parse::<std::net::SocketAddr>().is_ok(),
"direct path address should be ip:port, got {}",
s1.remote_addr
);
// Stream real audio so the path counters move.
let mut enc = OpusEncoder::new(48000, Channels::Mono, Application::Voip).unwrap();
for seq in 0..25u32 {
a.transport.broadcast(packet(&mut enc, seq));
tokio::time::sleep(Duration::from_millis(5)).await;
}
let s2 = snap(a.transport.connection_stats());
assert!(s2.tx_bytes > s1.tx_bytes, "sent bytes should advance");
assert!(
s2.tx_datagrams > s1.tx_datagrams,
"sent datagrams should advance"
);
// The derivation seam turns the two snapshots into live badge info.
let info = peerspeak::core::connstats::derive(Some(&s1), &s2, Duration::from_millis(200));
assert!(!info.relay);
assert_eq!(info.remote_addr, s2.remote_addr);
assert!(info.rtt_ms < 1000, "localhost RTT should be sane");
assert!(
info.up_kbps
.expect("same path + positive window has a rate")
> 0.0,
"audio was flowing, so the upstream rate must be non-zero"
);
}
/// Read datagrams off a raw connection until `target` arrive or the deadline
/// passes, asserting each carries the 4-byte sequence header.
async fn count_audio(conn: &Connection, target: u32, deadline: tokio::time::Instant) -> u32 {
@@ -503,3 +570,94 @@ async fn dialer_connects_and_reconnects_without_an_address_lookup() {
"audio should resume after reconnecting via the retained address; got {received} frames"
);
}
/// A node that also serves the file plane (`FILES_ALPN`), mirroring how core
/// registers the `FileRouter` for a session.
async fn spawn_file_server() -> Node {
let lookup = MemoryLookup::new();
let endpoint = Endpoint::builder(presets::Minimal)
.secret_key(iroh::SecretKey::generate())
.relay_mode(RelayMode::Disabled)
.address_lookup(lookup.clone())
.bind()
.await
.expect("bind endpoint");
let transport = Arc::new(IrohTransport::new(endpoint.clone()));
let audio_router = AudioRouter::new();
audio_router.bind(&transport);
let file_router = peerspeak::network::iroh_impl::FileRouter::new();
file_router.bind(&transport);
let router = Router::builder(endpoint.clone())
.accept(AUDIO_ALPN, audio_router)
.accept(peerspeak::protocol::FILES_ALPN, file_router)
.spawn();
Node {
endpoint,
transport,
_router: router,
lookup,
}
}
/// Phase 3C: a file fetch must deliver EXACTLY the declared size — short,
/// overlong, and unknown-id transfers are all rejected with local errors, and
/// an exact transfer round-trips byte-identically.
#[tokio::test]
async fn file_plane_requires_exact_declared_size() {
let fetcher = spawn_node().await;
let server = spawn_file_server().await;
fetcher.lookup.add_endpoint_info(server.endpoint.addr());
server.lookup.add_endpoint_info(fetcher.endpoint.addr());
let server_id = server.endpoint.id();
// Member gating: the server only serves current room members.
server.transport.admit_audio_sender(fetcher.endpoint.id());
let blob = vec![42u8; 1000];
let id = [7u8; 32];
server
.transport
.serve_attachment(id, Arc::new(blob.clone()));
// Exact declared size: byte-identical round trip.
let got = fetcher
.transport
.fetch_blob(server_id, id, 1000)
.await
.expect("exact-size fetch succeeds");
assert_eq!(got, blob);
// Declared larger than served (short transfer): rejected, not cached as-is.
let err = fetcher
.transport
.fetch_blob(server_id, id, 2000)
.await
.expect_err("short transfer must fail");
assert!(
err.to_string().contains("incomplete transfer"),
"unexpected error: {err}"
);
// Declared smaller than served (overlong transfer): the bounded read
// rejects the stream rather than truncating it into a "valid" result.
let err = fetcher
.transport
.fetch_blob(server_id, id, 500)
.await
.expect_err("overlong transfer must fail");
assert!(
err.to_string().contains("read failed"),
"unexpected error: {err}"
);
// Unknown id: the empty body reads as the sender no longer having it.
let err = fetcher
.transport
.fetch_blob(server_id, [9u8; 32], 1000)
.await
.expect_err("unknown id must fail");
assert!(
err.to_string().contains("no longer has the file"),
"unexpected error: {err}"
);
}