Compare commits

..
Author SHA1 Message Date
mollusk c2d82acf82 docs: prepare PeerSpeak 0.6.7 release
CI / check (push) Failing after 2m46s
2026-08-22 06:10:38 -04:00
mollusk 0823f6173f ux: make safe desktop audio the default 2026-08-22 05:44:08 -04:00
mollusk ff2533c95d fix: reap viewers and clear resolved share warnings 2026-08-22 02:41:31 -04:00
mollusk 140e4e7f73 docs: record Phase 9 rig qualification 2026-08-22 00:40:23 -04:00
mollusk db8459fd03 docs: checkpoint Phase 9 correlation rig 2026-08-21 22:25:19 -04:00
mollusk 260815154f docs: close Phase 8 integration 2026-08-21 22:13:33 -04:00
mollusk 2c2b861516 feat(screenshare): integrate desktop audio exclusion 2026-08-21 22:12:48 -04:00
mollusk ba96e59db0 docs: record completed Phase 7 2026-08-21 21:56:06 -04:00
mollusk d72271bd2e docs: close Phase 6 signal qualification 2026-08-21 21:00:38 -04:00
mollusk bf6d0e47b5 docs: record completed Phase 6 matrix 2026-08-21 19:56:18 -04:00
mollusk 46dc5902d6 docs: record Phase 6 refusal gates 2026-08-21 19:45:33 -04:00
mollusk a1df62ce21 docs: record Phase 6 failure-gate checkpoint 2026-08-21 17:35:19 -04:00
mollusk 30a460420a docs: record capture sink recovery checkpoint 2026-08-21 17:01:05 -04:00
mollusk 61960eb76f docs: record audio exclusion status checkpoint 2026-08-21 16:31:29 -04:00
mollusk 3d1f114fd8 docs: record owned fanout checkpoint 2026-08-21 16:13:41 -04:00
mollusk 9e52acf9d3 docs: advance audio exclusion plan into phase 6 2026-08-21 15:40:12 -04:00
mollusk 9ba42c4cda docs: record S3b containment milestone 2026-08-15 15:31:07 -04:00
mollusk f46b2cacc7 build(appimage): harden thin bundle workflow
Keep host multimedia libraries out of the AppImage, document the exact Ubuntu build inputs, and resolve Rust 1.97 release-gate warnings.
2026-08-14 18:05:00 -04:00
mollusk d023621eee chore: enable automatic Nix development shell
CI / check (push) Successful in 2m27s
2026-08-11 11:16:24 -04:00
molluskandClaude Opus 5 59da73c013 docs: the phase-5 gate passed, so stop telling phase 6 it is blocked
CI / check (push) Canceled after 0s
The plan's two status lines both predated the gate's second run. The header
still read "approved to start Phase 0a" nine phases in, and the DAG note still
claimed the re-run had not happened and that phase 6 waits on a passing results
file. That file has existed since 2026-07-26 —
screenshare-audio-exclusion-phase5-results.md records GATE PASSED on run 2 with
all 13 §5.1 rows, including rows 4 and 5 at the real tagging sites.

Both lines now state what actually blocks phase 6: the 0c -> 0d -> 6 edge, with
0c step 2 still open (S1/S2 merged, S3a built but unmerged, S3b/S4/S5 not
started) and 0d unbuilt. The superseded 2026-07-25 note is kept for the trail
rather than deleted.

Docs only; no code or gate changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 22:15:03 -04:00
mollusk 0e395e5c0c build(nix): pin the Rust toolchain to 1.97.1 via rust-overlay
CI / check (push) Canceled after 0s
nixpkgs 26.05 ships rustc 1.95.0, but this crate was developed and verified on
1.97.1 (what CachyOS had, installed 2026-07-17). Taking the compiler from
oxalica/rust-overlay decouples "which Rust the project targets" from "which
release the audio stack came from", so a nixpkgs bump can no longer move the
compiler under the lint gate as a side effect.

Chosen over rustup, which would also have worked here (nix-ld is enabled, so
its prebuilt binaries run) and would have let one rust-toolchain.toml cover the
packaging distroboxes too. The deciding factor is purity: rustup records
nothing in flake.lock, so a fresh clone or darp5 would resolve whatever it
fetched that day. rust-overlay gives the same exact-version control with the
choice pinned in the lock.

`.default` is the rustup "default" profile — rustc, cargo, rust-std, rustfmt
and clippy — so those are no longer listed individually. rust-src and the
x86_64-pc-windows-gnu target are added for win-cross-build.sh, which needs
`-Z build-std=std,panic_abort`; that script still expects the peerspeak-win
distrobox for the mingw half.

Verified on 1.97.1: 640 lib tests pass, fmt clean.

NOT fixed here, and pre-existing rather than a migration artifact:
`cargo clippy --all-targets -- -D warnings` fails with 15 warnings — 13
`float_literal_f32_fallback` (bare 0.05/0.01 into `.step()`, wants `0.05_f32`),
one `manual implementation of Option::filter`, one `redundant reference in
format!`, all in src/app/mod.rs. The f32 lint is `future_incompatible` and is
slated to become a hard error, so it needs fixing regardless of platform.
CachyOS was already on 1.97.1 well before the 2026-07-31 S2 merge logged as
"clippy clean", so that claim reflects a plain `cargo clippy` run, which exits
0 on warnings. `cargo clippy --fix` applies all 15 automatically.
2026-08-07 14:03:35 -04:00
mollusk 774922c6a9 build(nix): add a devShell so peerspeak builds on NixOS
CI / check (push) Canceled after 0s
The repo assumed a distro with a system-wide Rust, which NixOS does not
provide. This adds a flake devShell carrying the full dependency surface:

- Build: rustc/cargo/clippy/rustfmt, pkg-config, clang (pipewire-sys and
  libspa-sys need a real libclang for bindgen, via LIBCLANG_PATH), cmake
  (audiopus_sys's vendored-libopus fallback), and git (build.rs stamps
  PEERSPEAK_GIT_SHORT from `git rev-parse`).
- Link: alsa-lib, libopus, pipewire.
- dlopen'd at runtime: vulkan-loader, libxkbcommon, wayland and the X11 libs.
  Nothing links these, so they never land in the binary's rpath and are
  reachable only through LD_LIBRARY_PATH. Omitting them builds fine and then
  fails at window creation, which is a confusing way to find out.
- The supply-chain gates CI runs (cargo-deny, cargo-audit) plus cargo-deb.
  These were `cargo install`ed on the CachyOS side, which does not carry
  over — those binaries link that distro's glibc.

It also carries the screen-share tools (GStreamer + plugin search path,
pactl, mpv). Those look like they belong only to pixelpass, but
tests/screenshare_host_fault.rs starts a REAL pixelpass host, which aborts at
its own preflight without them — so they are a dependency of this test suite.
They are duplicated from pixelpass's flake rather than imported: the two
projects are mutually optional by design, and having one flake consume the
other would reintroduce the build-level dependency that rule prevents.

nixpkgs is pinned to nixos-26.05, the same channel the hosts run, so the
libraries here match the running PipeWire daemon and Vulkan ICD.

Verified: 640 lib tests pass, all 4 screenshare_host_fault live gates pass
(real audio backend, real network bind, real pixelpass child), fmt clean.
Known delta: clippy 1.95.0 (nixpkgs 26.05) flags one collapsible_match in
src/widget/selectable_text.rs that clippy 1.97.1 on CachyOS did not.
2026-08-07 13:46:18 -04:00
molluskandClaude Opus 5 63c246d976 docs: record the jitter buffer's unreachable shrink path as a known bug
CI / check (push) Canceled after 0s
`target_delay` grows +1 per disruption to MAX_DELAY_FRAMES (240 ms) but only
shrinks after 250 consecutive clean frames — 5 s of unbroken audio. Two of the
five `clean_run` resets fire on every natural pause in speech (jitter.rs:201
benign underrun, jitter.rs:181 re-prime), and the sender stops transmitting
outright while the gate is closed (core/mod.rs:2041). The AIMD decrease half is
therefore unreachable under conversational voice: one early jitter burst pins
the extra latency for the rest of the session.

Found by code review; not yet reproduced live. Pairs with field-test debt #5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 22:38:26 -04:00
molluskandClaude Fable 5 52d842160d Merge s2-host-fault: the screenshare host-fault path (S2)
A pixelpass host that dies mid-share is now torn down instead of
staying advertised: stdout EOF is synthesized as a terminal fault,
routed back into the core on a dedicated channel behind the reliable
arm of the biased select, gated by the ActiveShare generation so a
reaped child's late EOF is dropped as stale, and handled by retiring
the share — presence ticket removal and ScreenShareStopped ahead of
the reap wait, the explanatory error after.

Reviewed by Gemini (three rounds: branch review, full-range merge
review, fix verification round). Its P2s — Join's early-exit ordering
hole, the reap-then-presence advertising window, and the missing
presence-side gate — are fixed and mutation-verified. 640 lib tests;
three live gates green on the desktop, including a two-process
observer gate that reads the sharer's presence from a second real
node.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 14:53:02 -04:00
molluskandClaude Fable 5 e043810eb0 test: presence gate — a host fault clears the ticket for peers, promptly
The 🟡 S2 gap: presence-side ticket removal on fault was asserted by no
gate, and the "same code path as StopScreenShare" argument turned out
false — the fault handler duplicates the presence statements, so a
mutant deleting them passed all 644 tests. Nothing on the sharer's own
UiEvent channel can witness presence; it is only observable from
another node.

The gate runs a real second core as an observer in a SEPARATE PROCESS
(this test binary re-invoked as `presence_probe_helper`, its own
XDG_CONFIG_HOME): two in-process cores would load the same identity.key
and collapse into one node id, and swapping the env var between spawns
races other threads' getenv. The observer asserts the sharer's
PeerState.sharing goes Some → None on fault.

The fake host is a wedge — valid-shaped ticket (the OBSERVER's gossip
ingest sanitizes peer tickets; a garbage one is nulled to None and the
gate goes vacuous), ~1 s of life, then closes stdout while trapping
SIGINT — so the reap burns the full stop grace and TIME discriminates
the ordering, like the SIGINT gate: presence-first clears in ~1 s, the
old reap-then-presence ordering in ~3 s, asserted < the 2 s grace.
Mutation-verified both ways: presence removal deleted ⇒ observer times
out; old ordering restored ⇒ 3002 ms measured, assert fires.

Standalone (a plain --ignored sweep) the helper no-ops; the probe is
kill_on_drop so a parent panic can't orphan it (Gemini P3).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 14:52:10 -04:00
molluskandClaude Fable 5 5b80a1a010 core: pull the ticket off presence before the reap wait on host fault
Gemini's merge-round review (P2, CERTAIN): the fault handler ran
`stop_host().await` first, and a host that merely closed stdout but
lives on — trapped SIGINT, wedged — makes that call burn the full 2 s
stop grace before the SIGKILL fallback. For that whole window the dead
share stayed advertised: peers could still click Watch on it, and the
sharer's own UI kept saying "sharing".

The handler now retires the share where the fault is decided, not where
the corpse is confirmed: presence ticket removal and ScreenShareStopped
are emitted before the reap wait, and only the explanatory error (which
carries the unconfirmed-reap caveat) waits for `stop_host`. The
Stopped-before-Error contract is unchanged and still gated.

StopScreenShare's identical reap-then-presence ordering predates S2 and
is deliberately left alone (user-initiated stop, lower stakes); recorded
as a follow-up note instead of churning reviewed main-line code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 14:51:59 -04:00
molluskandClaude Fable 5 b9803f93fb core: retire the share at Join's session teardown, not after ticket parse
Gemini's review of the S2 branch found the one hole the harness had not
covered (P2, verified reachable): Join tears the old session down —
deliberately killing the share host — BEFORE validating the ticket, and
an invalid ticket exits the arm early, skipping the late
`current_sharing = None`. The killed host's stdout EOF then passed the
staleness gate and the user got a spurious "Screen share ended
unexpectedly" on top of "invalid room ticket". Pre-S2 the stale value
was toothless on this path; the fault handler gave it teeth.

The share now dies where the session does: cleared unconditionally right
after the teardown block, ahead of every early exit. The live gate grew
a third half — share, Join with a garbage ticket, then require silence
after the ticket error — and the mutant restoring the old placement is
killed by exactly that assertion (spurious re-emitted ScreenShareStopped).

Also Gemini's P3: the test's temp dir is now dropped by a guard, so an
assertion panic no longer leaks the fake-pixelpass scripts in /tmp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 04:05:17 -04:00
molluskandClaude Fable 5 3df2378831 test: the owed Stop Share SIGINT gate, against the real pixelpass
0c half (ii) had never been field-run: SIGINT sent, child exits within
the bound, no fallback kill on the normal path. Now it's a repeatable
live gate instead of a one-off manual check: a real whole-desktop host
(idle — no viewer, so no capture) is stopped and must reach
ScreenShareStopped inside STOP_GRACE. The SIGKILL fallback is
indistinguishable from success in the event stream, so time is the
discriminator: the fallback first waits out the full 2 s grace, while a
host honouring SIGINT exits in milliseconds.

Also holds the SIGINTed host's late stdout EOF to the same staleness
contract as the fake-host gate. Verified green on this desktop; no
stray pixelpass processes after the run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 01:10:11 -04:00
molluskandClaude Fable 5 81c230a09c test: live S2 exit gate — dead host torn down, clean stop stays clean
Drives the real core loop through CoreController with the pixelpass
override pointed at fake shell scripts (a host that emits its ticket and
dies; one that lives until signalled). The command loop has no unit
seam, so this is the only harness reaching the fault handler.

Half 1 pins the whole death path: ScreenShareStopped arrives BEFORE the
"ended unexpectedly" error. Half 2 stops a share deliberately and then
requires silence while the retired host's late stdout EOF lands as a
stale fault. The `exec sleep` in the living host is load-bearing: it
makes SIGINT close stdout so the stale fault actually arrives, keeping
the staleness assertion non-vacuous.

Both core-side mutants verified killed: swallowing the forwarder's fault
times out half 1; disabling the staleness gate panics half 2 with the
spurious re-emitted ScreenShareStopped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 01:08:13 -04:00
molluskandClaude Fable 5 be3740f5f9 screenshare/core: a host that dies is no longer advertised as sharing (S2)
The defect: pixelpass's stdout EOF was silently discarded, the notice
channel existed only for app-audio shares, and nothing cleared the host
from the teardown slot or the ticket from presence — so a crashed host
stayed advertised in the room and the UI kept saying "sharing".

Every share now gets a notice channel. The drain task synthesizes a
terminal HostNotice::Eof when the stream ends (EOF or read error — a
crash can abort across `extern "C"` before any JSON line is written, so
the stream ending is the only reliable death signal). The core's
forwarder turns that into a ScreenShareHostFault scoped to the spawn's
generation; a stale fault (already stopped, or a newer share running) is
dropped. The handler reaps the child through the existing confirmed-reap
path, pulls the ticket off presence, and emits ScreenShareStopped BEFORE
the error, so the UI never shows "sharing" next to the explanation.

The ticket and its generation live in one ActiveShare value on purpose:
they must appear and vanish together, or the staleness gate drifts.

Both drain gates are mutation-verified: swallowing the Eof fails both
tests; skipping it only on the read-error path fails exactly the
error-path test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 01:02:41 -04:00
molluskandClaude Opus 5 0d836d14c2 docs: round 18 — repair moves to libpulse; two reviews, three reusable lessons
Rounds 17c and 4 of the repair review, recorded together because they resolve to
one decision: `pactl`'s text output cannot carry the guarantees repair claims, so
observation and unloading now go through libpulse introspection over a single
verified-local connection. The dependency was taken with the user's sign-off after
vetting (details beside the dep in pixelpass Cargo.toml).

Three lessons that generalise beyond this phase:

- A *prescription* can fail reachability just as a finding can. "Use `pactl -f json
  list modules`" is sound reasoning against an API that does not exist — those
  records carry no module index, and `unload-module` accepts only an index.
- Auditing my own fixes paid a third time: two of the four fixes applied in round
  17a were themselves defective, including a correlation scheme that is unsound
  whenever module names repeat.
- The live field test caught a bug unit tests structurally cannot reach, and it was
  phase 0b's bug one layer down: fields drop in declaration order, the Pulse
  context's teardown frees IO events owned by the mainloop, and declaring the
  mainloop first turned a fully successful repair into SIGABRT and exit 134.

Also recorded: the newline defect needed no adversary and was confirmed on the live
server, and the remaining namespace hole is left open with its trade stated — an
owner token would close it but would make orphans from older builds uncleanable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 01:24:46 -04:00
molluskandClaude Opus 5 76c1a13e11 docs: round 17 — the repair review, the 0c actor review, and 0c's slicing
Two reviews in one round. The repair planner returned changes-requested (no P1s,
four reachable P2s, all applied in pixelpass `9145b2a`); the 0c actor design
returned four blocking issues, all accepted, plus a concession on epoch.

Recorded because three of them generalise:

- A prescribed fix was not implementable as written. "Use `pactl -f json list
  modules`" is sound reasoning against an API that does not exist: on pactl 17
  those records carry no module index, and `unload-module` takes only an index.
  The reachability rule now applies to prescriptions, not just findings.
- The actor's bounded join would have disarmed `Drop` by moving the thread handle
  into `spawn_blocking` — the same defect shape as round 15's, a defence disarmed
  exactly when needed.
- Epoch was over-specified and my vacuity instinct was right: serial equality is
  the entire identity guarantee, so epoch is diagnostic and explicitly not a gate.

Measured on the live graph rather than argued: recorded module arguments are
byte-exact with `@DEFAULT_SINK@` unresolved (both load-bearing for exact-form
matching), and two sinks may share one `node.name` with capture attaching to the
OLDER one in 3 of 3 trials — so a surviving wedged owner silently steals the next
session's capture instead of merely risking a collision.

0c step 2 is sliced into S2–S5 so each lands reviewed. Nothing reopens D6: the
connection-owned-sink design is unchanged, and a material part of the growth is
pre-existing debt 0c forced into the light.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 21:33:57 -04:00
mollusk aa0515af1c Merge phase 0b + the peerspeak half of 0c: teardown ordering, reaping, graceful stop
Two invariants land here, both of which phase 6 depends on:

1. The echo-cancel module cannot unload while a pixelpass host is alive. The
   ordering-critical fields moved into `ScreenshareTeardown` with `echo_cancel`
   declared LAST, and killing is no longer taken for reaping — `kill_on_drop`
   only signals, so `ReapOnDrop` blocks on a bounded poll until the child is
   actually gone. The defect was real, not theoretical: `echo_cancel` sat ahead
   of `screenshare_host` in declaration order, so any unwind unloaded the AEC
   first, and unwind is reachable (no `panic=abort`, many `unwrap()`s).

2. Stop Share asks before it insists — SIGINT, a bounded grace, then SIGKILL —
   so pixelpass runs its own cleanup instead of leaking a null-sink module every
   time. SIGINT specifically: pixelpass installs only a `ctrl_c()` handler.

Reviewed by Codex across two rounds: changes-requested (two blocking findings,
both real, both the same shape — a defence disarmed exactly when it was needed)
then approve-with-follow-ups (five P3s, all applied). The mutation matrix was
revised from five to four after one pinned mutation was proved unreachable by
construction, and teardown was hoisted to one unconditional post-loop site so
every loop exit is covered structurally.

Owed and recorded: the live Stop Share SIGINT gate has never been field-run, and
the hoisted call site's live proof belongs to the phase-9 lifecycle row.

638 lib tests, clippy clean, fmt clean.
2026-07-26 19:43:17 -04:00
molluskandClaude Opus 5 9f06741b99 core/teardown: an unconfirmed stop is not a clean stop
Codex's re-review of the branch returned "approve with follow-ups" — no
blocking findings, five P3s. All five are applied here rather than carried as
debt, since each is a few lines.

The one with user-visible consequences: `stop_host` returned a bare "was
sharing" bool, so the single case where the availability-first policy gives up
(SIGKILL queued, reap never confirmed) still sent `ScreenShareStopped` with
nothing else. The UI would say sharing had ended while pixelpass might still be
alive and fanning out — a claim the user cannot see through. `shutdown` now
returns `StopOutcome`, `stop_host` returns `Option<StopOutcome>`, and an
unconfirmed *user-initiated* stop raises a UI error naming the stray process.
Session and viewer teardown discard the outcome deliberately: nobody is waiting
on an answer there, and the residual risk is already logged.

Also: the three failure diagnoses in `shutdown` (the signal never left, the
child ignored it, the wait itself broke) were collapsed into one log line and
are now distinct — they mean different things to whoever reads the log.

The second cancellation gate is the one worth keeping. The review pointed out
that all cancellation coverage sat in the *graceful* wait, so a mutant that
disarmed the wrapper between the two waits would survive. It was right, with a
wrinkle: the naive mutant does not compile, because the child is borrowed from
`self` for the whole function — the borrow checker is doing real work here. The
restructured form (`self.child.take()` once cooperation has failed) does
compile, and the pre-existing mid-wait test passes it.
`cancelling_shutdown_after_the_kill_leaves_the_fallback_armed` kills it.

Mutation-verified, both new gates: reporting an unconfirmed stop as `Reaped`
fails exactly `a_failed_wait_is_not_treated_as_a_confirmed_reap`; disarming
between the waits fails the new cancellation test (and the failed-wait test,
which also asserts armedness) while leaving the old mid-wait test green — which
is the proof the new test is not redundant. The logging split is diagnostics
only and has no gate; said plainly rather than dressed up as covered.

Docs: the "four mutations" line is now an explicit table naming each target and
its test, with 0c's pair counted under 0c; and the aggregate teardown latency is
recorded as a deferred item with a trigger (a fourth routine child, or a
measured teardown over 5 s) instead of an unwritten known cost.

638 lib tests, clippy clean, fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:43:04 -04:00
molluskandClaude Opus 5 3aa768af52 docs: the 0b DAG row says four mutations, matching §10 round 14
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:11:19 -04:00
molluskandClaude Opus 5 1cfa932fbe docs: record the 0c mechanism probe, and revise 0b's mutation matrix
The 0c mechanism probe was run on this host before any structural work, because
one unverified assumption could have invalidated the whole approach: whether a
hand-created `adapter` node is visible to pipewire-pulse under the name the
capture path depends on. It is. Five gates green, including the two that
mattered — `<node.name>.monitor` is exposed as a Pulse source, and SIGKILL of
the owning connection removes both Pulse-visible names with zero graph residue.
No null-sink module is involved at any point. Every O1 stop condition for 0c is
retired, and the default sink never moved, so the probe is safe on a live
desktop.

The probe also settled the native-sink scope question: it applies to every mode
that owns a capture sink, not only `DesktopExcluding`. That makes the `--repair`
rework load-bearing rather than defensive — repair derives dead PIDs only from
`module-null-sink` entries, so once the sink is native its loopbacks become
undiscoverable orphans.

§10 gains rounds 14 and 15: the 0b matrix drops to four mutations because the
best-effort wake arm is unreachable by construction (the loop owns a sender, and
the biased select would win anyway), and teardown is hoisted to one unconditional
post-loop site instead of being duplicated across one live arm and one dead one.
Round 15 records the two blocking implementation-review findings and the vacuous
gate of my own that the review's test-double critique exposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:11:06 -04:00
molluskandClaude Opus 5 d8b8fd79cf core/teardown: stay armed across the wait, and never call a failed wait a reap
Codex's review of the two commits below returned "changes requested" with two
blocking findings. Both were real.

1. `ReapOnDrop` disarmed itself across the async wait. `shutdown` moved the
   child out of the wrapper with `take()` before the first `.await`, so if that
   future was cancelled or unwound mid-wait, the raw child dropped with nothing
   but `kill_on_drop` (signals, does not reap) while `Drop` found `None` and did
   nothing — the AEC could then unload over a live child. That is precisely the
   hole the type exists to close, left open for the duration of every wait. The
   child now stays owned by `self` across every await and is released only on a
   *confirmed* reap.

2. A failed wait was silently converted into success, and the hard-kill path was
   unbounded. `wait_reaped` discarded `io::Result`, so a wait error made the
   timeout return `Ok` and shutdown returned as though the reap were confirmed;
   meanwhile a process stuck in uninterruptible sleep after SIGKILL could wedge
   the core command loop forever. The trait now preserves the result, both waits
   are bounded, and the conflict case has an explicit written policy: we choose
   availability, leave the child owned so the bounded Drop retry stays armed,
   and log the residual risk rather than hiding it.

Codex also showed the test double was flattering the implementation in four
ways. All four are closed: the fake can now be cancelled mid-wait, can fail its
wait, and can take several polls to die, and the grace is pinned independently.

That last one caught a flaw in my own gate. The elapsed-time assertion compares
against `STOP_GRACE` itself, so setting the constant to zero leaves it vacuously
true — both sides move together. `the_grace_is_a_real_interval` pins the
constant to a band instead, and now kills that mutation directly.

Mutation-verified again, five mutants, each killed by its own gate: disarming
the wrapper (cancellation test), treating a wait error as success (failed-wait
test, exactly one), a zero grace (the new band test), a single poll instead of
the drop loop (delayed-reap test, exactly one), reversed field order (the two
ordering tests).

Also applies the matrix adjudication, which Codex and I reached independently:
teardown moves out of the reliable close arm to ONE unconditional site after the
loop, so every `break` is covered structurally — including any added later —
instead of duplicating teardown across one live arm and one provably dead one.

637 lib tests, clippy clean, fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:09:58 -04:00
molluskandClaude Opus 5 92a64465a4 core/teardown: ask pixelpass to stop before killing it
The peerspeak half of phase 0c (design v3.4 §7.4 item 1). Stop Share and
session teardown both went straight to `Child::kill()`, i.e. SIGKILL, which
skips pixelpass's own cleanup and leaks one null-sink module every time.

`ReapOnDrop::shutdown` now asks first: SIGINT, a bounded 2 s wait, then SIGKILL
only if the child ignored the request. SIGINT specifically, not SIGTERM —
pixelpass installs only a `ctrl_c()` handler, so SIGTERM would take the default
disposition and be indistinguishable from SIGKILL.

Signalling by pid is safe against pid reuse here: we have not reaped the child,
so it is a zombie whose pid the kernel reserves until we wait it, and the pid
cannot name a stranger. (Same reasoning that dismissed pixelpass bug #6.)

The grace is 2 s because it is awaited inline in the core command loop, so it
is also how long a wedged child can delay other commands. A healthy pixelpass
never spends it.

The drop/unwind path deliberately stays a hard kill: `Drop` cannot await, and
there the ordering invariant (§7.1) outranks tidiness. Once the pixelpass half
of 0c lands, the capture sink is connection-owned and that path stops leaking
by construction.

`libc` becomes a direct unix-only dependency, pinned to 0.2.186 — the version
already in the tree via alsa/cpal/tokio — so Cargo.lock gains one line and no
new code enters the build.

Mutation-verified, five mutations, each killing its own gate: no wait (6 fail),
reversed field order (2, reap test green), no reap loop (2, ordering test
green), no SIGINT (6), no SIGKILL fallback (exactly 1 — the wedged-child test).
633 lib tests, clippy clean, fmt clean.

Not yet field-tested: the live Stop Share gate (SIGINT sent, child exits within
the bound, no fallback kill on the normal path) still owes a real run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:26:05 -04:00
molluskandClaude Opus 5 6ba763774d core/teardown: reap the screen-share children before the AEC unloads
Phase 0b, fixes 1-3 of design v3.4 §7.2 (decision D4). The invariant is that
the echo-cancel module must not unload while a pixelpass host is alive and
fanning out; two paths have to honour it and only one is code we get to run.

The explicit path: `ActiveSession::shutdown` now awaits
`ScreenshareTeardown::shutdown_children`, and the reliable command channel's
close arm tears the session down explicitly instead of letting it drop on the
way out of `run_core_loop`.

The drop/unwind path: the ordering-critical fields move out of `ActiveSession`
into `core::teardown::ScreenshareTeardown`, where `echo_cancel` is the LAST
declared field and therefore the last dropped. Previously it was declared
first (`:682`, ahead of `screenshare_host` at `:685`), so an unwind unloaded
the AEC while the host was still live — and unwind is reachable, the core is
full of `unwrap()` and has no `panic=abort` profile.

Killing is not enough. `kill_on_drop(true)` only signals: it hands the child to
the runtime's orphan queue and returns, which an unwinding runtime may never
drain. `ReapOnDrop` blocks on a bounded 250 ms budget until the child is really
gone, because a bounded stall beats unloading the AEC out from under a live
pixelpass.

Everything is generic over a narrow `ChildProcess` trait and over the guard
type, so ordering is unit-testable without spawning processes or loading
PipeWire modules — the seam idiom already used by `replace_viewer_index`.

Mutation-verified, and the plan's demand that mutations 4 and 5 prove
*different* defenses holds: reversing the field order fails only the
AEC-ordering tests and leaves the reap test green; removing the reap loop fails
only the reap tests and leaves the ordering test green. Removing the explicit
wait fails the explicit-path tests. 631 lib tests, clippy clean, fmt clean.

⚠️ Mutations 1 and 2 of the pinned matrix do not both exist: the best-effort
wake arm is unreachable by construction, twice over. Documented at the site;
adjudication owed in the impl plan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 18:19:07 -04:00
molluskandClaude Opus 5 692ad677d2 docs: record F11-1 closed — boundedness needs a resolved Client
Design v3.7 §6.1.1 gains the round-13 box (the rule, the ordering that is
load-bearing in both directions, and why bridging deliberately still uses the
full union); the phase-5 results file records the close with the measurement
the deferral was waiting for; the impl plan's phase-6 gate note drops F11-1.

pixelpass c78eb2d is the implementation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 04:46:25 -04:00
mollusk bf908adbf0 docs: phase 5 matrix PASSED (13/13) — design round 10, results run 2
The §5.1 dry-run audit gate passes. All 13 rows completed with a non-empty
eligible half in every one, and O5 is re-measured on the fixed graph: worst
recompute 67 µs across every run, 10 µs mean under deliberate churn, busy
fraction 0.0006, readiness 1-2 ms with 18 binds against a 2000 ms budget.

Round 10's finding, and it is the third of exactly the same shape: the
pipewire-pulse PID derivation required a SINGLE repeated pipewire.sec.pid.
WirePlumber repeats one too (two Clients, both sec_pid 1747), so the derivation
returned None permanently on a stock desktop, key 4's suppression never fired,
and every Pulse-emulated node fused into one owner. Row 1's CLEAN control
forwarder and Firefox were both excluded. Fixed in pixelpass 91c4ded by deleting
the heuristic: probe every distinct sec_pid and let /proc/<pid>/comm decide.

Three measured rounds now, all at the observation boundary, none in the
architecture -- and all three were fail-closed and silent, caught only because
§5.1 requires asserting what must remain ELIGIBLE. An exclusion-only checklist
would have passed every one of these builds.

Rows 4-6 are closed through peerspeak's REAL tagging sites (call, mpv, notify,
plus clip) with a hand-launched mpv staying eligible, so the cross-repo contract
is proven end to end on live nodes. Row 10 covers the full sticky lifecycle
including retirement; row 11 is provably non-vacuous (the recycled
node.link-group came back byte-identical and did not inherit taint).

Recorded and NOT claimed as passes: substitutions in rows 8, 9 and 13, and two
reporting-only findings (the audit's sticky flag is uninformative; a bridge's
named key is lost when a leg reappears under a new serial).

Phase 6 remains blocked by F11-1, phases 0b/0c/0d and the Stereo Mix design
call -- this file removes one gate, not all of them.
2026-07-26 02:59:43 -04:00
mollusk b68fca689e Merge phase 1: ownership tagging (SPA-JSON carriers via libspa)
peerspeak-side half of phase 1 of the screenshare audio-exclusion work: every
node peerspeak owns carries two registry-visible ownership carriers, so the
taint engine has a primary root that survives the registry's filtered global
event (design v3.5 section 6.7).

Reviewed by Codex over rounds 10-12; all findings verified and dispositioned.
Round 12's F12-1 (rfind('}') spliced carriers inside a trailing comment, a
fail-open) and F12-2 (depth ceiling taken from the consumer,
pw_properties_update_string, not from the spa-json-dump grammar) are fixed and
live-verified through the real ALSA plugin.

623 lib tests green, fmt clean, clippy clean.
2026-07-26 02:14:25 -04:00
21 changed files with 4856 additions and 537 deletions
+1
View File
@@ -0,0 +1 @@
use flake
+6
View File
@@ -11,3 +11,9 @@
# the .iss script and .ico are the tracked sources.
/packaging/windows/peerspeak.exe
/packaging/windows/output/
# Nix: the symlink `nix build` drops, and direnv's local cache. flake.nix and
# flake.lock ARE tracked — the lock is what pins the toolchain.
/result
/result-*
/.direnv/
Generated
+1
View File
@@ -4883,6 +4883,7 @@ dependencies = [
"image",
"iroh",
"iroh-gossip",
"libc",
"opus",
"pipewire",
"rand 0.10.1",
+7
View File
@@ -109,3 +109,10 @@ windows-sys = { version = "0.61", features = [
"Win32_System_Diagnostics_ToolHelp",
"Win32_System_Threading",
] }
# Unix-only. Used for exactly one thing: sending SIGINT to our own
# pixelpass child so it can run its cleanup before we resort to SIGKILL
# (src/core/teardown.rs). Already in the tree via alsa/cpal/tokio, so
# declaring it directly adds no new code to the build.
[target.'cfg(unix)'.dependencies]
libc = "0.2.186"
+29
View File
@@ -137,6 +137,35 @@ covers internals). When you ship a feature, add it here.
---
## Known bugs
Defects found by code review, not yet fixed.
1. **Adaptive playout delay never shrinks back in real conversation**
(`src/core/jitter.rs`) — `target_delay` grows +1 per disruption up to
`MAX_DELAY_FRAMES` (12 frames = 240 ms) but only shrinks after
`CLEAN_RUN_TO_SHRINK` = 250 consecutive cleanly-played frames, i.e. **5 s of
unbroken audio**. `clean_run` is reset in five places; two of them fire on
every natural pause in speech: the benign-underrun branch (`jitter.rs:201`,
talker went quiet) and the subsequent re-prime (`jitter.rs:181`). Because the
sender skips transmitting entirely while the noise gate is closed
(`src/core/mod.rs:2041`), a pause between sentences *always* underruns the
receiver and zeroes the clean run — twice.
Net effect: the controller is a one-way ratchet. A single burst of jitter
early in a call pins up to 240 ms of extra playout latency for the rest of
the session, because no conversational speaker talks for 5 continuous
seconds without the gate closing. The AIMD "decrease" half is effectively
unreachable under the workload the app is built for.
Likely fix: let `clean_run` survive a benign idle→re-prime transition rather
than resetting it. Silence is not evidence the link is bad, so it should not
count against the clean run. Distinguish "talker stopped" (benign) from
"playout broke" (real) at `jitter.rs:195-202`.
Found 2026-07-31 by code review. Not yet reproduced in a live call — pairs
with field-test debt item 5 below.
## Known field-test debt (the 🧪 rows above, collected)
Re-run on a real desktop ↔ dopedart call before calling these done:
+295
View File
@@ -0,0 +1,295 @@
# PeerSpeak 0.6.7 release prep
Prepared: 2026-08-22
Planned work day: 2026-08-23 (confirm the actual date before updating the changelog)
Release target: GitButter release page with a verified Linux x86_64 AppImage
This is a checklist, not release authorization. Pushing `main`, creating or
pushing a tag, creating the GitButter release, and uploading assets remain
explicit approval gates.
## Starting checkpoint
Recheck every value live tomorrow; these are the known-good handoff values from
2026-08-22.
- PeerSpeak: `main` at `0823f617`, 17 commits ahead of `origin/main`.
- PixelPass: `main` at `ce909afc`, 14 commits ahead of `origin/main`.
- Latest published PeerSpeak tag: `v0.6.6`.
- Current Cargo/Windows version: `0.6.6`.
- Current field-test build:
`packaging/appimage/peerspeak-0.6.6-unofficial-20260822-ps0823f617-ppce909afc-fieldtest-x86_64.AppImage`
- Current field-test SHA-256:
`f094a665dbd719929d6def32c8fd3d4943e4b88b6290e371e57b127c4e91ec45`.
- The same build was staged for Lindsay at
`/home/lindsay/Downloads/peerspeak-0.6.6-unofficial-20260822-ps0823f617-ppce909afc-fieldtest-x86_64.AppImage`.
- Preserve the existing user-owned `docs/FEATURES.md` modification and
untracked `.codex/` directory. Do not include either in release-prep commits
unless the user explicitly puts them in scope.
The current one-row audio-picker build passed local AppImage artifact checks,
but it still needs the final two-machine field pass. Earlier two-machine passes
proved that desktop audio was heard, call voices were not echoed, and no warning
was visible. They also exposed the now-fixed stale viewer/warning behavior. Do
not substitute those earlier passes for testing the current build.
## Version decision
The expected release is **0.6.7**.
`VERSIONING.md` defines pre-1.0 MINOR bumps as breaking wire-protocol changes,
not as a measure of feature size. The diff from `v0.6.6` through `0823f617` is a
large screen-share implementation and local UI/lifecycle change, but it does
not change `src/protocol.rs` or bump an audio, friends, files, or gossip protocol
version. Under the repository's policy, that makes this a PATCH release.
Use `0.7.0` only if tomorrow's final source review finds or adds an actual
incompatible wire change. If that happens, identify and bump the affected
Layer-2 protocol constant/domain in the same change and plan a coordinated
upgrade for every peer.
## Critical build-label rule
The first/home screen visibly identifies the binary as:
```text
PeerSpeak v<package version> (<8-character commit hash>)
```
The same label also appears in Settings. The two pieces come from different
places:
- `Cargo.toml` supplies the version through `CARGO_PKG_VERSION` and must be
bumped from `0.6.6` to `0.6.7`.
- `build.rs` runs `git rev-parse --short=8 HEAD` and embeds the result as
`PEERSPEAK_GIT_SHORT`. **Do not edit or bump the hash by hand.**
Consequences for the release order:
1. Make and commit all source, version, changelog, and field-evidence changes.
2. Build the final AppImage from that clean final commit.
3. On first launch, verify the title screen says exactly
`PeerSpeak v0.6.7 (<final HEAD short hash>)`.
4. Verify the displayed hash equals `git rev-parse --short=8 HEAD`.
5. Tag that exact commit as `v0.6.7`.
If any commit is added after an AppImage is built—even a field-evidence or
release-note commit—the embedded hash is now old. Rebuild and revalidate the
AppImage. Never publish an artifact whose visible hash differs from the tag
target.
## Tomorrow's runbook
### 1. Re-establish live state
From `/home/mollusk/git/butter/peerspeak`:
```sh
git status --short --branch
git log -1 --oneline --decorate
git -C ../pixelpass status --short --branch
git -C ../pixelpass log -1 --oneline --decorate
git fetch --prune --tags origin
git ls-remote --heads --tags origin
df -h / /mnt/superjar
pgrep -af 'peerspeak|pixelpass|mpv|gst-launch' || true
```
Confirm that:
- the checkpoint commits above are still the intended source;
- neither repository has unexpected changes;
- the existing `docs/FEATURES.md` and `.codex/` state is preserved;
- no stale call/share processes are running;
- `v0.6.7` does not already exist locally, remotely, or on GitButter;
- there is enough space for one-job release builds and extracted AppImages.
### 2. Complete the current-build two-machine field test
Use the stamped `0823f617` AppImage on this machine and Lindsay's staged copy.
Start with PeerSpeak and any old screen-share helper/player processes closed.
Run both directions, with Lindsay keeping software encode enabled where needed:
1. Mollusk hosts; Lindsay views.
2. Lindsay hosts; Mollusk views.
3. In each direction select the sole visible **All system audio** row.
4. Play desktop audio after the share begins, including starting a new audio
stream/application during the share.
5. Talk from both machines while the desktop audio plays.
6. Stop sharing while the voice call remains active.
7. Start one more share, then have both peers leave the call.
Record all of these results explicitly:
- the picker has one desktop-audio row, not duplicate legacy/safe choices;
- desktop audio is heard by the viewer;
- neither person's call voice loops back through the shared audio;
- no transient or persistent missing-output-port warning appears;
- stopping a share clears the viewer and warning promptly;
- the viewer PixelPass child is reaped within roughly 500 ms after the remote
stream ends;
- leaving the call clears the share and leaves no PeerSpeak-owned PixelPass,
player, GStreamer, capture, or echo-cancellation residue.
If any row fails, stop release preparation. Save the exact visible text and
relevant logs, fix the defect, run the focused regression gates, commit the
fix, rebuild, and repeat the matrix.
### 3. Prepare the 0.6.7 source
After the current build passes:
- Change `[package].version` in `Cargo.toml` to `0.6.7`.
- Refresh the root PeerSpeak package entry in `Cargo.lock`; inspect the diff and
make sure it changes only as intended.
- Change `MyAppVersion` in `packaging/windows/peerspeak.iss` to `0.6.7`, even
though tomorrow's planned public asset is AppImage-only. This prevents the
next Windows installer from silently retaining `0.6.6`.
- Move the relevant `[Unreleased]` material into a dated `0.6.7` section in
`CHANGELOG.md` and add the `v0.6.7` comparison/release link.
- State that the release is wire-compatible with `0.6.6`; do not claim a
protocol bump.
- Summarize user-visible behavior, especially:
- **All system audio** now shares desktop sound without feeding PeerSpeak's
own call audio back to listeners;
- the safe path is the single normal desktop-audio choice;
- unsafe or incomplete audio-routing states fail closed and surface a useful
warning;
- viewer/share cleanup is prompt when a stream stops or a peer leaves;
- the Linux AppImage bundles the matching PixelPass helper.
- Record the completed two-machine evidence in
`docs/screenshare-audio-exclusion-impl-plan.md`.
Commit the scoped version/changelog/evidence work. Keep unrelated user changes
out of the commit.
### 4. Run source gates on the final commit
Use the pinned Rust toolchain and one build job if disk or linker pressure is
tight:
```sh
cargo fmt --all -- --check
CARGO_BUILD_JOBS=1 cargo clippy --locked --all-targets -- -D warnings
CARGO_BUILD_JOBS=1 cargo test --locked --all-targets
CARGO_BUILD_JOBS=1 cargo test --locked --doc
cargo deny --locked check
cargo audit
git diff --check
git status --short --branch
```
The commands mirror the repository CI, with `--locked` added to compilation
and test gates after the intentional lockfile refresh. A gate that cannot run
must be recorded as unverified; do not silently treat it as passed.
Also verify the bundled PixelPass source is still exactly the intended clean
commit and run its release diagnostics before packaging:
```sh
git -C ../pixelpass status --short --branch
git -C ../pixelpass rev-parse --short=8 HEAD
cargo run --manifest-path ../pixelpass/Cargo.toml --locked -- --doctor
```
### 5. Build the final AppImage
Build from the clean, committed PeerSpeak release source and clean PixelPass
source in the Ubuntu 24.04 distrobox described in `packaging/appimage/README.md`.
Keep build caches on `/mnt/superjar` and use one build job if space remains
tight. A representative invocation is:
```sh
distrobox enter peerspeak-appimage -- env \
PATH="$HOME/.rustup/toolchains/1.97.1-x86_64-unknown-linux-gnu/bin:$PATH" \
SYSTEM_DEPS_LIBSPA_INCLUDE=/run/host/usr/include/spa-0.2 \
SYSTEM_DEPS_LIBPIPEWIRE_INCLUDE=/run/host/usr/include/pipewire-0.3:/run/host/usr/include/spa-0.2 \
CARGO_BUILD_JOBS=1 \
PEERSPEAK_APPIMAGE_CACHE=/mnt/superjar/peerspeak-appimage-cache \
./packaging/appimage/build-appimage.sh
```
The official output should be
`packaging/appimage/peerspeak-0.6.7-x86_64.AppImage`. Do not overwrite the
stamped 0.6.6 field-test artifact until the release is complete and verified.
### 6. Validate the exact final artifact
At minimum:
- record `sha256sum` and create a matching `.sha256` sidecar;
- run `--appimage-extract` into an isolated scratch directory;
- inspect `ldd` for both bundled `usr/bin/peerspeak` and `usr/bin/pixelpass` and
require no `not found` entries;
- inspect RUNPATH/RPATH and confirm the bundle has not captured the host's
graphics, PulseAudio, or GStreamer stack contrary to the thin-AppImage policy;
- run the bundled `pixelpass --capabilities` and confirm desktop-audio exclusion
support is advertised;
- run the bundled `pixelpass --doctor` with a fresh temporary GStreamer registry;
- launch the AppImage with an isolated config/data directory for a GUI smoke;
- verify the first/home screen reads
`PeerSpeak v0.6.7 (<final PeerSpeak HEAD short hash>)`;
- verify the same version/hash appears in Settings;
- confirm the bundled PixelPass binary corresponds to `ce909afc` (or the exact
newer commit deliberately selected tomorrow).
Because the final version/evidence commit changes the embedded hash relative to
the current field-test artifact, do one short two-machine confirmation using
the final AppImage: connect, share **All system audio**, confirm desktop sound
and no voice echo/warning, stop the share, and leave cleanly.
### 7. Publication gate
Before any external write, present the final facts to the user:
- final PeerSpeak full and short commit;
- bundled PixelPass full and short commit;
- `v0.6.7` proposed tag target;
- AppImage filename, size, and SHA-256;
- title-screen version/hash observed;
- source, artifact, and two-machine gate results;
- exact release notes/assets to publish.
Then wait for explicit approval to publish.
After approval only:
1. Push `main` without force.
2. Confirm remote `main` resolves to the tested final commit.
3. Create an annotated `v0.6.7` tag on that exact commit and push it.
4. Create the `v0.6.7` GitButter release page from the finalized changelog text.
5. Upload `peerspeak-0.6.7-x86_64.AppImage` and its `.sha256` sidecar.
6. Do not paste or store a GitButter token in the repository, documentation,
shell history, or release artifacts.
The planned release scope is the Linux x86_64 AppImage only. Do not imply that
a Windows installer, Debian package, or separate PixelPass release exists
unless those artifacts are deliberately added and independently verified.
### 8. Verify the public release
Publication is not complete until it is independently read back:
- confirm GitButter shows the correct release title, tag, notes, and both assets;
- confirm the remote tag and remote `main` point to the expected commit;
- download the public AppImage to a fresh temporary path;
- verify its SHA-256 against the published sidecar and the local final hash;
- extract or run the downloaded copy and repeat the version/hash and bundled
PixelPass diagnostic checks;
- save the final release URL and verification result in the implementation plan
or handoff.
## Failure and recovery rules
- A failed gate means no tag and no release—not a waiver.
- Never force-push or move a published tag as an automatic recovery step.
- If the tag is pushed but release creation/upload fails, stop and report the
exact remote state before changing anything.
- Keep failed or candidate artifacts clearly stamped so they cannot be mistaken
for the final asset.
- Remove only exact, rebuildable scratch/extraction directories when reclaiming
disk space; preserve source, user changes, final artifacts, hashes, and field
evidence.
+693 -29
View File
@@ -1,8 +1,34 @@
# Implementation plan: whole-desktop screen-share audio without self-echo
**Status:** 🟢 **v4 — three review rounds applied. Approved to start Phase 0a.**
**Date:** 2026-07-21
**Design of record:** [`screenshare-audio-exclusion-plan.md`](screenshare-audio-exclusion-plan.md) v3.4 (`8768cd2`), converged round 7.
**Status:** 🟢 **v4 — three review rounds applied.** *Progress as of 2026-08-21:* phases 0a, 0b,
0c step 1, 1, 2, 3, 3r, 4 and 5 are merged, and the **phase-5 major gate PASSED on 2026-07-26**
(§1). The two pre-Phase-6 decisions are now **closed** in design v3.8 §6.9: the conservative
same-device hardware bridge is built and validated, and the 2 s readiness budget passed
baseline, inflated-graph and live-churn calibration. The completed **0c step 2** was
sliced S1–S5: S1 and S2 are
merged; **S3a is merged locally** in pixelpass (`15d1374`); and **S3b is built, validated, and
committed locally** (`5d3da8b`). **S4, S5, 0d, round 11's bridge changes, and Phase 6's first
pure channel-planning prerequisite are built, validated, and committed locally in pixelpass
(`781defc`).** The first bounded Phase 6 mutation slice is now built, validated and committed
locally in PixelPass (`98cde2c`): the hidden production path creates, retains, revokes and
crash-cleans exact-channel non-lingering fan-out links. The four versioned Phase 6 status events
are built, validated and committed locally in PixelPass (`5a65f50`). Capture-sink replacement
and successful exact-channel relinking are built, validated and committed locally in PixelPass
(`d09ee9b`). **The Row 1 mutation-edge identity gate and the Row 8d/8e fail-closed
construction gates are built, validated and committed locally in PixelPass (`6be07ef`).**
**The independent Row 6a/6b/6c refusal gates are built, validated and committed locally in
PixelPass (`956534f`).** **The revised Row 9 positive/negative late-arrival partition is built,
validated and committed locally in PixelPass (`7b11827`).** The deterministic Phase 6
link-manager matrix is complete. **The production-path three-arm AEC leak qualification is
built, validated and committed locally in PixelPass (`e027bc6`), completing Phase 6.**
**Phase 7's public selector, AEC wiring and versioned capability response are built, validated
and committed locally in PixelPass (`792f2bd`).** **Phase 8's PeerSpeak capability negotiation,
picker/argv integration and causal status surface are built, validated and committed locally
(`2c2b861`).** **Phase 9's PN correlation/xrun rig is implemented and live-qualified locally in
PixelPass (`98bb78f`, hardened by `360d711`); the field matrix remains open, and nothing has
been released. Phase 9 is still the current front and ship gate.**
**Date:** 2026-07-21 (v4); status line refreshed 2026-08-22
**Design of record:** [`screenshare-audio-exclusion-plan.md`](screenshare-audio-exclusion-plan.md) v3.8, round 11.
**Scope:** *ordering, gates and acceptance criteria only.*
**Reference convention.** `v3.4 §N` = the design doc. `plan §N` = this document. The two
@@ -76,7 +102,7 @@ if it differs, failing closed.
| # | Phase | Repo | Mutates graph? | Exit gate |
| --- | --- | --- | --- | --- |
| 0a | `object.serial` u64 fix | pixelpass | no | boundary parse tests |
| 0b | Explicit teardown + drop ordering | peerspeak | no | **five independent mutations** (plan §2) |
| 0b | Explicit teardown + drop ordering | peerspeak | no | **four independent mutations** (plan §2; revised from five, §10 r14) ✅ built |
| 0c | Graceful stop + connection-owned capture sink | both | sink ownership | two-host SIGKILL live gate **+ SIGINT-first gate** |
| 0d | **Typed capture plan + internal mode input** | pixelpass | no | mode matrix; neither unsafe source nor unsafe sink input constructible |
| 1 | peerspeak ownership tagging | peerspeak | no | tag on live nodes; literal pinned in plan §3 |
@@ -100,13 +126,21 @@ if it differs, failing closed.
1 (r8 carriers) ──────────────────────────────────────► 5 (re-run)
```
⚠️ **Status 2026-07-25 (evening): 3r is BUILT AND MERGED; the re-run has not happened yet.**
The phase-5 gate failed on its first live run and put 3r into the DAG; 3r's own four-part
gate now passes, including the live prop-recovery row on this host. Phase 5's machinery is
built and correct — it is the audit that found the defect, twice — so "5 (re-run)" is a
*re-run of the matrix*, not a rebuild. **Phase 6 still does not start** until a passing
results file exists. **Phase 1 is a hard prerequisite of the re-run for both carriers**
(plan §3).
✅ **Status 2026-07-26: the phase-5 gate PASSED on run 2 — all 13 §5.1 rows completed.**
Record: [`screenshare-audio-exclusion-phase5-results.md`](screenshare-audio-exclusion-phase5-results.md)
(audit build pixelpass `main` @ `91c4ded`, release profile). Phase 1 was the hard prerequisite
of the re-run for both carriers (plan §3) and was satisfied — rows 4 and 5 passed at the real
tagging sites. **Phase 6 is no longer blocked by this gate.** What still blocks it is the rest
of the DAG: `0b → 6` is satisfied and merged; **0c and 0d are built, validated, and committed
locally with S4/S5 and round 11 in pixelpass `781defc`**. **Round 11 now closes the two former
design §6.8 blockers** with the v3.8 §6.9 hardware bridge and readiness calibration; Phase 6
is unblocked and its first pure planning prerequisite has landed.
⚠️ **Superseded, kept for the trail — status 2026-07-25 (evening): "3r is BUILT AND MERGED; the
re-run has not happened yet."** The phase-5 gate failed on its first live run and put 3r into
the DAG; 3r's own four-part gate then passed, including the live prop-recovery row on this host.
Phase 5's machinery was built and correct throughout — it is the audit that found the defect,
twice — so "5 (re-run)" was a *re-run of the matrix*, not a rebuild. Run 2 is that re-run.
⚠️ **A smoke run of the audit against the fixed observer immediately found a second defect
(design v3.6 §6.8): a fail-closed `unresolved-ancestry` mark was being promoted to permanent
@@ -151,7 +185,19 @@ v3.4 §7.2, decision D4. All **three** of v3.4's fixes:
AEC guard unloads. This is the *only* protection on the panic/unwind path, and unwind is
reachable — the core has numerous `unwrap()` sites and no `panic=abort` profile.
⚠️ **Mutation testing: five mutations, each independently breaking a named test.** v1 demanded
> ✅ **0b IMPLEMENTED 2026-07-26** (peerspeak branch `phase-0b-teardown`). The ordering
> defect was live: `echo_cancel` was declared *ahead* of `screenshare_host`, so any unwind
> unloaded the AEC while the host was still fanning out. Fields moved into
> `src/core/teardown.rs` with `echo_cancel` declared last, and `ReapOnDrop` added because
> `kill_on_drop(true)` only *signals* — it hands the child to the runtime's orphan queue,
> which an unwinding runtime may never drain. Matrix revised to four mutations; see §10
> round 14, and round 15 for the two blocking review findings that followed.
>
> **Owed to phase 9:** an explicit lifecycle row — *drop the controller / close the command
> channel while sharing* — which is the live proof for the hoisted teardown call site.
⚠️ **Mutation testing: five mutations, each independently breaking a named test.**
⚠️ **SUPERSEDED by §10 round 14 — mutation 2 is vacuous and the matrix is now four.** v1 demanded
a mutation that targeted the wrong defense; v2 fixed that but bundled two defenses into one
combined mutant, which proves neither. Final form:
@@ -167,6 +213,45 @@ Both channel-close arms get their own test; a single "closes the command channel
exercise one arm and leave the other unsafe.
### 0c. Graceful stop + connection-owned capture sink — both
> ✅ **MECHANISM PROBE PASSED on this host, 2026-07-26** (PipeWire 1.6.8). Run *before* any
> structural work, on the reviewer's insistence, because a single unverified assumption could
> have invalidated the entire approach: whether a hand-created adapter is visible to
> pipewire-pulse under the name pixelpass's capture path depends on. It is.
>
> ```
> pw-cli> create-node adapter factory.name=support.null-audio-sink \
> node.name=pixelpass_probe_<pid> media.class=Audio/Sink \
> audio.channels=2 audio.position=[FL,FR] node.virtual=true \
> monitor.channel-volumes=true object.linger=false
> ```
>
> Five gates, all green:
> 1. `pactl list short sinks` shows the sink under the **exact** `node.name`.
> 2. `pactl list short sources` shows **`<node.name>.monitor`** — the derived monitor name is
> a pipewire-pulse contract, not a property of Pulse-created sinks. This was the one that
> could have sunk the approach.
> 3. `gst-launch-1.0 pulsesrc device=<node.name>.monitor num-buffers=40 ! fakesink` pulled its
> buffers and exited clean, and a real recording stream attached — so the pixelpass capture
> path works against it unchanged.
> 4. **No null-sink module was loaded** (`pactl list short modules | grep -c null-sink` stayed
> at its baseline of 3). It is genuinely not a Pulse module.
> 5. **SIGKILL of the owning connection removed both Pulse-visible names**, with zero residue
> anywhere in `pw-dump`. That is the entire point of 0c, demonstrated on the real graph.
>
> The default sink never moved, so this is also safe to run on a live desktop.
> **Every O1 stop condition listed for 0c is retired.** `object.linger=false` is load-bearing:
> the bundled pipewire-rs example sets `linger=1` for the opposite behaviour.
>
> ⚠️ **`--repair`'s job does not shrink — it BREAKS.** Discovery derives dead host PIDs
> **only** from `module-null-sink` entries (`pixelpass/src/repair.rs`), and only then matches
> loopbacks against that PID set. The native sink is scoped to **every mode that owns a
> capture sink**, not just `DesktopExcluding`, so legacy Pulse loopbacks will coexist with a
> connection-owned sink; when that host dies the sink vanishes automatically and its loopbacks
> become **undiscoverable orphans**. Candidate PIDs must be derived independently from all
> three module shapes (`null-sink sink_name=`, `loopback sink=`, `loopback source=…monitor`),
> with a liveness recheck immediately before each destructive unload. This makes the repair
> rework **load-bearing, not defensive**.
v3.4 §7.4. The **largest hidden cost in Phase 0**: moving the null sink off `pactl load-module`
(`pixelpass/src/host/audio.rs:69`, cleaned up only in `Routing::cleanup` at `:259-260`, which
SIGKILL skips) onto a connection-owned PipeWire object.
@@ -249,6 +334,18 @@ through four phases of active work around them. Guardrails go up before the scaf
capture sink;
- legacy behaviour byte-identical.
> ✅ **BUILT, VALIDATED, AND COMMITTED locally in pixelpass `781defc`, 2026-08-21.** The real hidden
> `--internal-desktop-excluding` host input resolves to a typed `CapturePlan`; only
> `LegacyDesktop` can construct `DefaultMonitor`, while `DesktopExcluding` owns a bare
> connection-owned sink whose type has no legacy loopback API. The complete 16-row mode matrix
> passes, both conflict inputs reject before graph mutation, and the exact legacy GStreamer
> audio tail is unchanged. The serialized live graph assertion passed with neither an incoming
> PipeWire link nor a Pulse module feeding the new sink; all four prior ownership/cleanup live
> regressions also passed and left no PixelPass audio residue. Broad result: pixelpass **310
> passed, 9 ignored**, fresh `--doctor` all green; peerspeak screen-share units **41 passed, 1
> ignored**, plus the real Stop Share/SIGINT compatibility gate passed. This remains an internal
> mode with no fan-out and no public selector; Phase 6 is still the first fan-out mutation.
A constructible-but-not-yet-public variant is acceptable for the interval between 0d and Phase
6 provided it is unit-tested and reachable by the hidden trigger.
@@ -377,7 +474,8 @@ and 4 are unaffected, and the phase-5 audit machinery is already correct.
> mutation-verified (a `device_props` ambiguity test that checked for one live *Device*
> rather than one live *global*, and `device.api` corroborating by presence). Two findings
> left open as design items, both pre-existing — hardware playback-to-capture paths and the
> readiness-budget calibration, both recorded in design §6.8.
> readiness-budget calibration. **Both are now closed by design v3.8 §6.9 / the pre-Phase-6
> round-11 gate below.**
>
> **Added beyond the spec: a second live gate for the Device-side path.** Row 1's
> `session_device` assertion is satisfied by a union, and WirePlumber 0.5.15 copies
@@ -458,17 +556,34 @@ observable in Phase 5 before they gate anything real.
## 5. Phase 5 — dry-run audit mode 🚦 MAJOR GATE
> **🚦 STATUS 2026-07-25: BUILT, RUN, AND THE GATE FAILED.** Results:
> `docs/screenshare-audio-exclusion-phase5-results.md`. The machinery is correct and needs no
> rework — **it found the defect on its first live run**, which is the phase working exactly as
> designed. What failed is the observer beneath it (v3.5 §6.7). **Phase 6 does not start.** The
> matrix re-runs after phase 3r and phase 1's second carrier land; no row was completable
> under the defect, so none of it carries over. O5's numbers do not carry over either.
> **🚦 STATUS 2026-07-26: GATE PASSED on run 2.** Results:
> `docs/screenshare-audio-exclusion-phase5-results.md`. All 13 rows completed, the eligible
> half of every row is non-empty, and O5 is re-measured on the fixed graph (worst recompute
> 67 µs; readiness 1–2 ms with 18 binds). Three rows carry recorded substitutions (8, 9, 13)
> and three findings are recorded as non-blocking.
>
> Read this before re-running: `PIXELPASS_AUDIO_AUDIT_FILE=… PIXELPASS_AUDIO_AUDIT_AEC=off
> pixelpass --audit-audio`. **Every partition row must run with `AEC=off`** — a
> **Run 2 found and fixed a third defect of the F2 class, F13-1:** pipewire-pulse's PID was
> unresolvable on this host *permanently*, because stage 1 of the derivation required exactly
> one repeated `sec_pid` and **WirePlumber repeats one too** (two Clients, one PID). Key 4's
> suppression therefore never fired and every Pulse-emulated node fused into one owner. Fixed
> in pixelpass `91c4ded`: probe every distinct `sec_pid` and let `/proc/<pid>/comm` decide.
> **The eligible half of row 1 is the only thing that exposed it** — the verdict was
> fail-closed and silent.
>
> ⚠️ **Phase 6 is NOT unblocked by this file alone.** F11-1 was the other gate and is now
> **closed** (2026-07-26, pixelpass `c78eb2d`: key 4 bounds an owner only when the node's
> Client resolves; measured cost on the live graph, zero — see the results file). Phases
> 0b/0c/0d are satisfied and committed locally with S4/S5 in pixelpass `781defc`.
> **Round 11 (2026-08-21) closed the remaining hardware playback-to-capture and readiness
> gates; Phase 6 is now unblocked for development.**
>
> Two things to keep when re-running: **every partition row must run with `AEC=off`** (a
> configured-but-unvalidated AEC shuts the fan-out gate and empties the eligible half of every
> row, which reads as a failure that is really a harness error.
> row, which reads as a failure that is really a harness error), and **start the audit BEFORE
> building the fixture**. Fixture-first makes the whole graph arrive as one enumeration burst,
> so every node is first tainted while `graph_ready` is false; that partial-graph taint enters
> sticky state and the keyless sticky reason then wins over the evidence-derived one, so a row
> cannot assert its own key. Read keys at *derivation* (first non-sticky appearance).
**Adds no capability. Its entire purpose is to be wrong loudly and safely.**
@@ -501,7 +616,7 @@ exclude-everything implementation fails the eligible half of every row.
| 6 | peerspeak **notification** sound | that node, reason = tag | — |
| 7 | a **second** pixelpass host's capture sink, **plus a controlled forwarder reading that sink's monitor** | the forwarder's **named output serial** (cycle prevention, v3.4 §6.2) | — |
| 8 | EasyEffects running | combined output leg | EasyEffects stopped ⇒ ordinary streams |
| 9 | Firefox: music only / mic on untainted source / capturing a tainted monitor | the third only (v3.4 §6.1.1) | the first two |
| 9 | Firefox: music only / mic on a **different Device** / mic on the **same Device receiving tainted playback** / capturing a tainted monitor | the third and fourth (`tainted-owner-bridge`) | the first two |
| 10 | sticky taint: tainted input leg removed, output leg lives | still excluded | after full owner teardown + restart |
| 11 | recycled serial/index/link-group after teardown | — | must **not** inherit taint |
| 12 | AEC loaded, then unloaded | four nodes; then `Revoked` | — |
@@ -531,8 +646,32 @@ assumptions, whereas Sunshine is an uncontrived third-party forwarder nobody des
test. It stays as row 1b, **opportunistic and non-gating**, because it cannot be relied on to
be present.
**Any surprise here goes back to the design doc as round 8. Phase 6 does not start until this
results file exists.**
**Any surprise here goes back to the design doc as a new measured round. Phase 6 did not start
until this results file existed; round 11's targeted addendum now pins the same-device rule.**
### 5.4 Pre-Phase-6 round-11 closure — ✅ PASSED 2026-08-21
Design v3.8 §6.9 and the addendum in
[`screenshare-audio-exclusion-phase5-results.md`](screenshare-audio-exclusion-phase5-results.md)
are the durable evidence. PixelPass retains snapshot-local `device.id` on positively classified
session-device nodes and adds a conservative `Sink → Source` taint edge only within that Device.
No ALSA control-name guess is part of the runtime policy.
Exit gates:
- 83 focused taint tests pass, including same-device exclusion, different-device eligibility,
and the accepted same-device-microphone over-exclusion.
- 65 pure observer tests and all three serialized live PipeWire observer tests pass; a live
passive device retains the `device.id` consumed by the engine.
- Full non-GUI suite: 313 passed, 0 failed, 9 ignored. The three live observer tests were then
run explicitly and passed.
- Targeted live audit exact partition: tagged ALC897 playback excluded the same-ALC897
capture/re-emitter as `tainted-owner-bridge`; the Arctis-source control remained eligible.
- Readiness: baseline 30 starts p95/max 10/11 ms; 48-module graph 30 starts 113/114 ms;
20 starts during 250 create/remove cycles 6/8 ms, zero timeouts. Keep the 2 s budget.
All temporary modules were unloaded by their exact module ids, configured audio defaults were
unchanged, and root filesystem free space was 17 GiB after tests and Clippy.
---
@@ -545,6 +684,141 @@ retained for the life of the share, per-port link sets, "captured" only when **e
link is `ACTIVE`, same-epoch revalidation immediately before each creation, proxy drop on
ancestry becoming unsafe.
> **Status 2026-08-21 — first Phase 6 mutation slice built, validated and committed locally.** The
> hidden `DesktopExcluding` production path now launches a same-observer-callback link manager.
> It revalidates every recyclable id against its `object.serial` immediately before mutation,
> creates exact FL/FR (or MONO fan-out) links through `link-factory`, retains the proxies, marks a
> stream captured only when every required link is `ACTIVE`, and drops owned links before the
> capture sink on normal teardown. The links explicitly set `object.linger=false`.
>
> Twelve deterministic controller tests cover channel planning, recycled ids, idempotence,
> partial activation, late arrivals, newly unsafe ancestry, AEC gating/revocation and mutation
> failure. The full non-GUI suite passes with 326 passed, 0 failed and 10 ignored, and strict
> all-target Clippy is clean. The real hidden host path was also tested with a late stereo
> stream: exactly two native links became `ACTIVE`, no Pulse `module-loopback` fed the capture
> sink, and both links were revoked when the stream ended. A separate subprocess gate sent
> PixelPass `SIGKILL`; the connection-owned capture sink and both link object serials disappeared
> without Rust destructors. The three serialized live observer tests and `pixelpass --doctor`
> pass.
>
> **Status-event slice, committed locally as `5a65f50`.** PixelPass now emits
> exact version-1 `stream_unsupported`, `aec_failed`, `aec_revoked` and `foreign_aec_warning`
> records. The PipeWire callback only enqueues owned records; a Tokio-side forwarder performs
> JSON serialization and stdout I/O. Each cause has a direct emission test plus an exact JSON
> golden. Repeated graph ticks do not repeat a sticky failure, every leg of one foreign AEC
> collapses to one warning, an unindexed native AEC still counts as foreign, and `aec_revoked`
> is queued only after owned links are dropped. Full PixelPass: 330 passed, 0 failed, 10 ignored;
> strict all-target Clippy, both serialized Phase 6 live gates and `pixelpass --doctor` pass.
>
> **Capture-sink replacement slice, committed locally as `d09ee9b`.** The
> connection owner now recreates an unexpectedly removed non-lingering sink under the stable
> Pulse name but with a fresh exact serial. A narrow identity channel hands that serial to the
> fan-out observer, whose own main-loop command immediately reconciles its coherent snapshot:
> stale proxies are dropped before new links are created, and the stream returns to `Captured`
> only after every replacement channel is `ACTIVE`. The observer treats only PipeWire's
> asynchronous `-ENOENT` for a resource lost during ordinary graph churn as recoverable; every
> other Core error remains fatal.
>
> The pure recovery gate proves both stale links are dropped and two replacement links activate.
> A live owner-only gate destroys the sink and observes a fresh serial. The real hidden host gate
> destroys an actively linked sink, observes a new sink serial and two fresh `ACTIVE` links,
> proves both old link serials are absent, and leaves no residue. All three serialized Phase 6
> live gates pass together. Full PixelPass: 332 passed, 0 failed, 12 ignored; strict all-target
> Clippy, formatting, `git diff --check`, and `pixelpass --doctor` pass.
>
> This does **not** complete Phase 6. The remaining failure-matrix rows are still open. The public
> mode selector remains Phase 7 and PeerSpeak parsing/UI integration remains Phase 8. The last
> committed PixelPass checkpoint is `6be07ef`.
>
> **Row 1 plus Row 8d/8e slice — built, validated and committed locally as `6be07ef` on
> 2026-08-21.** The production
> link mutator now applies serial revalidation to retained proxies as well as new requests; a
> node/global-id recycle between planning and mutation yields zero `create_link` calls and drops
> any stale retained intent. `CapturePlan` has one injectable production construction seam, and
> separate optimized-release gates prove both capture-sink construction failure and the
> observer's sticky readiness timeout escape `DesktopExcluding` as errors without ever resolving
> the default monitor or constructing legacy `Routing`.
>
> The full suite passes with 335 passed, 0 failed and 12 ignored. All three new gates pass in an
> optimized release build; strict all-target Clippy, formatting and `git diff --check` pass. All
> three serialized Phase 6 live audio-plan gates still pass, `pixelpass --doctor` passes, and the
> final Pulse/PipeWire/process residue scan is empty. Nix remains unavailable, so validation used
> system Rust 1.97.1. The disposable 1.0 GiB release cache was removed afterward, leaving 12 GiB
> free. Next is the independent Row 6a/6b/6c refusal/reason slice; Phase 6 is not complete.
> **Row 6a/6b/6c slice — built, validated and committed locally as PixelPass `956534f` on
> 2026-08-21.**
> The bound Node observer now subscribes to the configured `SPA_PARAM_Format` only for
> `Stream/Output/Audio` nodes that advertise it as readable, and classifies the native libspa
> format as raw, encoded or IEC958. Missing/unparseable format evidence is a per-stream
> fail-closed `format-unknown`, not permission to link. `node.passthrough` stays a separate
> predicate, so it cannot be masked by the format classifier.
>
> Three independent controller fixtures prove Row 6a `port.exclusive`, Row 6b encoded and Row
> 6c IEC958 each make **zero** link-creation attempts and emit exactly their own
> `stream_unsupported` reason (`port-exclusive`, `encoded`, `iec958-passthrough`) once. A fourth
> gate independently exercises explicit `node.passthrough`; another proves the ordinary
> Node-info-before-Format ordering stays fail-closed without emitting a transient false warning,
> then captures once raw PCM arrives. Row 6a remains the accepted injected-graph gate: v1 still
> does not bind Ports, so a live exclusive port reaches the already-gated clean link-failure path.
>
> The first read-only live audit exposed and prevented two observer defects before completion:
> generic Pod deserialization rejected the real Format object, and enumerating Format on every
> driver Node produced expected ENOENT/EIO core errors. The corrected path uses libspa's native
> format parser and the Node's advertised readable-param list. A second live audit observed
> Strawberry, FFXIV and Chromium settle from `format-unknown` to eligible raw PCM within the
> initial callback burst, with no parse or core errors. Full PixelPass validation is 342 passed,
> 0 failed and 12 ignored; all three Row 6 gates pass in an optimized one-job release build;
> strict all-target Clippy, formatting, `git diff --check` and `pixelpass --doctor` pass. The
> 966 MiB disposable release cache was removed and disk space returned. Next is the revised Row
> 9 positive/negative partition; Phase 6 is not complete.
> **Revised Row 9 slice — built, validated and committed locally as PixelPass `7b11827` on
> 2026-08-21.**
> The former single music-only late-arrival fixture is now four independent controller gates
> carrying Phase 5's exact Firefox partition across the Phase 6 mutation boundary. Music-only
> and a Firefox output whose microphone is on a different Device each create exactly two stereo
> links and reach `Captured`. A Firefox output whose input reads the same Device receiving
> tainted playback, and one reading that sink's tainted monitor, each retain
> `tainted-owner-bridge`, make zero link-creation calls and hold no proxies. Every case is
> re-driven unchanged to prove idempotence.
>
> Full PixelPass validation is 345 passed, 0 failed and 12 ignored; all four focused Row 9 gates,
> strict all-target Clippy, formatting and `git diff --check` pass. The three serialized live
> Phase 6 mutation gates — late eligible fan-out/cleanup, sink replacement/relink and SIGKILL
> cleanup — pass, `pixelpass --doctor` passes, and the final Pulse/PipeWire/process residue scan
> is empty. Nix is unavailable, so this slice used system Rust 1.96.1. The deterministic
> link-manager matrix is now complete. Phase 6 remains open on its production-path three-arm AEC
> leak measurement re-run, whose naive positive control must exist only behind a test seam.
> **Production-path three-arm AEC leak qualification — built, validated and committed locally as
> PixelPass `e027bc6` on 2026-08-21. Phase 6 is complete.** A `#[cfg(test)]` policy can re-admit
> only the exact configured `aec-identity` candidate; the unsafe policy, constructor and field do
> not exist in a production build. A focused controller gate proves safe mode retains only the
> ordinary stereo pair while the deliberately naive mode adds exactly the AEC playback pair.
>
> The ignored serialized live gate drives the hidden `DesktopExcluding` selector through the real
> connection-owned sink, registry observer, taint controller, native link manager and
> `<sink>.monitor` Pulse source. It loads one exact-ID WebRTC echo-cancel module per arm, filters
> `media.class=Stream/Output/Audio` before matching `pulse.module.id`, retains and reports every
> child stderr stream, verifies the exact graph links, and records 48 kHz stereo s16le with
> `parec`. The guarded and naive arms inject a PeerSpeak-owned 1500 Hz stream into the real AEC
> sink; all arms retain an ordinary 440 Hz desktop stream.
>
> Two complete runs passed the thresholds declared in the test. Desktop 440 Hz stayed between
> -32.84 and -34.05 dBFS. The guarded arm's 1500 Hz result (-75.76 to -78.03 dBFS) never rose more
> than 3 dB above its run's control floor, while the naive positive control measured -33.80 to
> -33.92 dBFS, comparable to its desktop tone and at least 18 dB above both control and guarded.
> The supported conclusion is deliberately narrow: **no incremental 1500 Hz energy was detectable
> above the control floor at this analysis resolution**; this is not a claim that remote audio is
> absent. Phase 9 still owns the PN/MLS intelligibility rig and field variance.
>
> Final PixelPass validation: 346 passed, 0 failed and 13 ignored; strict all-target Clippy,
> formatting, `git diff --check`, `pixelpass --doctor`, and all four serialized Phase 6 live audio
> gates pass. The final Pulse/PipeWire/process/temp-file residue scan is empty. Nix is unavailable,
> so validation used system Rust 1.96.1. The next implementation front is Phase 7's public mode
> selector and versioned capability advertisement.
Failure ⇒ report the stream unsupported. **Never** fall back to the default monitor — and after
0d that fallback is unconstructible in this mode, by either path.
@@ -620,8 +894,9 @@ shippable runtime override.
### Phase 7 — public mode selector + capability advertisement (pixelpass ships first)
⚠️ **Round-3 P1: nothing in v3 ever promoted the hidden trigger to a public flag.** 0d added an
internal mode input; Phase 7 advertised capability and naming; Phase 8 added `--aec`, the picker
and status. No phase required the actual **mode selector** to exist publicly or to be passed.
internal mode input; Phase 7 advertised capability and naming; Phase 8 added PeerSpeak's
`--aec` emission, the picker and status. No phase required the actual **mode selector** to exist
publicly or to be passed.
The result would be a capability-gated picker entry that, when chosen, still spawns legacy
whole-desktop capture — the feature appearing to ship while doing nothing. Reachable: peerspeak's
host argv has no mode parameter (`screenshare/mod.rs:152`) and pixelpass's `HostOpts` has no mode
@@ -640,6 +915,39 @@ byte-identical, absent `--aec` still accepted.
v3.4 §11 public naming is a **blocking user input at the start of this phase**. Internal typed
variant names (0d) do not block on it.
> **Phase 7 — built, validated and committed locally as PixelPass `792f2bd` on 2026-08-21.** The
> user selected `--audio-mode=desktop-shared|desktop-excluding`. The shared value resolves to the
> byte-identical legacy plan; the excluding value reaches the existing typed
> `DesktopExcluding` plan and requires an explicit `--aec=off|pulse-module:<idx>`. The exact
> Phase-4 parser is now the public CLI parser, malformed values fail before host startup, and the
> selected AEC config is carried through `CapturePlan` into the production graph controller
> instead of being hardcoded to `Off`. Both protocol flags require `--host`; the hidden 0d
> trigger remains test-only and defaults to `Off` for its existing fixtures.
>
> `pixelpass --capabilities` now emits exactly one versioned JSON line:
> `{"schema_version":1,"capabilities":{"strict_app_audio":true,"desktop_audio_exclusion":true}}`.
> The two booleans are independent by construction. The old-PeerSpeak/new-PixelPass golden feeds
> both existing whole-desktop and strict per-app argv into the new parser unchanged and proves
> absent `--audio-mode`/`--aec` still resolves to legacy behavior. The existing byte-exact legacy
> GStreamer audio-tail gate remains green; `--help` probing is retained only for old integrations.
>
> Phase-7 validation also made the Phase-6 signal gate consume the selector-owned AEC config. Its
> first rerun caught the stale test-local input immediately. Two later reruns exposed a separate
> measurement issue: a coherent projection of sub-LSB stochastic noise can land in an unusually
> deep single-bin null, so comparing only control-bin to guarded-bin overstated the rig's
> resolution. Retained raw spectra showed no coherent guarded 1500 Hz peak and comparable nearby
> noise. The committed gate now defines the control resolution as the larger of the exact control
> bin and the 90th percentile of neighboring ±5–25 Hz projections outside the Hann main lobe;
> the 3 dB guarded tolerance and 18 dB naive positive-control margin remain unchanged. This is a
> resolution estimate for the steady-state gross-leak gate, not the Phase-9 PN/MLS upgrade.
>
> Final PixelPass validation: 354 passed, 0 failed and 13 ignored; strict all-target Clippy,
> formatting, `git diff --check`, the exact capability/help probes, `pixelpass --doctor`, and all
> four serialized live audio gates pass. The final Pulse/PipeWire/process/temp-file residue scan
> is empty. Nix is unavailable, so validation used system Rust 1.96.1. Phase 8 is next: bind the
> capability to the resolved PixelPass path, add the capability-gated picker/argv, and carry
> exclusion status causally to the UI.
### Phase 8 — peerspeak integration
- `EchoCancelGuard::module_index()` accessor (currently private; only `source_name()` /
`sink_name()` exist).
@@ -663,14 +971,73 @@ variant names (0d) do not block on it.
- Capability-gated picker entry; wording per v3.4 §11.
- Regression: existing `--app` / `--strict-audio` argv byte-identical to today.
> **Phase 8 — built, validated and committed locally as PeerSpeak `2c2b861` on 2026-08-21.**
> The picker/core boundary now uses one typed selection for legacy desktop, desktop-excluding,
> and strict per-app audio. `System audio except PeerSpeak` appears only when the exact resolved
> PixelPass advertises `desktop_audio_exclusion`; the existing `All system audio` choice remains
> alongside it with its echo warning. The machine-readable schema is primary, `--help` is only
> the old-PixelPass strict-app fallback, and the two capability bits stay independent.
>
> Capability results carry the resolved executable path. A capability-gated share reuses the
> picker probe only for that same path; a changed override or `$PATH` resolution is re-probed at
> start and fails closed when the selected feature is absent. The new mode emits exactly
> `--audio-mode=desktop-excluding` plus `--aec=off|pulse-module:<idx>`. The session-owned
> `EchoCancelGuard` exposes its verified numeric module index without moving the guard out of the
> load-bearing teardown object. Legacy desktop and per-app argv still delegate to the old builder
> and are byte-identical.
>
> All four version-1 exclusion events parse into one shared status type that crosses the actual
> host-notice channel into `UiEvent`; there is no duplicate matching enum. The UI retains an
> explanatory warning even when a status races the share-start acknowledgement, displays it with
> the live sharing badge, clears it on stop, and ignores late events for another mode. Exact
> parser and causal tests cover `stream_unsupported`, `aec_failed`, `aec_revoked`, and
> `foreign_aec_warning`.
>
> Final serialized validation (used after the parallel all-target link exhausted disk) passes:
> 649 unit tests with 7 live-only ignored, 20 integration tests with 4 live-only screen-share
> gates ignored, strict all-target Clippy, formatting, and `git diff --check`. The host-fault
> integration target compiles with the typed command. PixelPass's completed 9.1 GiB target tree
> was cleaned to recover disk; no source or Git state was removed. Phase 9 remains the explicit
> rig-upgrade and two-machine field-test ship gate.
---
## 7. Phase 9 — rig upgrade and field tests 🚦 SHIP GATE
v3.4 §9.2's rig upgrade is **owed before any exclusion claim is published**: two orthogonal
PN/MLS probes, windowed per-channel normalised cross-correlation reporting max per-window
correlation, plus xrun telemetry. Until it exists the only defensible claim is the gross-leak
distinction, in v3.4 §9.2's exact wording.
correlation, plus xrun telemetry. Until it is live-qualified the only defensible claim is the
gross-leak distinction, in v3.4 §9.2's exact wording.
> **Rig implementation checkpoint — PixelPass `98bb78f`, 2026-08-21.** Two pinned independent
> PN probes now drive the same production-path control/guarded/naive topology used by the Phase 6
> qualification. The rig downsamples each captured stereo channel to chips, reports normalized
> maximum cyclic correlation for every non-overlapping 1,024-chip window, verifies the eligible
> probe in every window/channel, bounds the absent/excluded probe, and proves detectability with
> the naive positive control. It also snapshots `pw-top` ERR counters for the involved nodes,
> checks original playback and exact capture routes after recording, and retains the existing
> teardown/residue and audio-health gates.
>
> Thresholds are fixed in the ignored live test before any arm runs: present correlation ≥ 0.25,
> excluded correlation ≤ 0.20, and zero new xruns.
>
> **Controlled live qualification PASSED on 2026-08-22; harness hardening committed as PixelPass
> `360d711`.** The first two safe attempts exposed assumptions in the new rig before they could
> hide field variance: PipeWire returned five matching IDs for four requested node names, proving
> again that `node.name` is not unique, and cold `parec` startup yielded too little audio for three
> windows. The resolver now requires every requested name while monitoring all matching IDs, with
> a duplicate/missing-name regression test. Phase 9 alone records for six seconds from twelve-second
> probes; the proven Phase 6 recorder remains at three seconds and its behavior is unchanged.
>
> The final serialized run produced seven complete windows per channel in every arm. The absent
> remote-probe maximum was 0.1356 in control and 0.1340 when guarded, both below 0.20. The naive
> positive-control remote minimum was 0.5741, and the eligible desktop minimum across all arms was
> 0.6429, both above 0.25. Every arm added zero xruns. Original playback and exact capture routes
> remained intact, and the final module/sink/link/process/temp-file residue scan was empty. Full
> validation passes 356 non-live tests with 14 live-only ignored, strict all-target Clippy,
> formatting, diff checks, and `pixelpass --doctor`. Nix was unavailable, so the run used system
> Rust 1.96.1 with one build job. **The rig upgrade is qualified; the two-machine/real-path matrix
> below remains open and keeps the Phase 9 ship gate closed.**
**Every row gets a declared pass/fail threshold before the run, not after.** Baseline for all
rows: excluded probe ≤ the declared rig criterion; **eligible control audio present**; original
@@ -743,6 +1110,303 @@ mid-share load stays a **synthetic** test until a second `enable` site or hot re
## 10. Adjudication record
**Round 14 (2026-07-26) — 0b's five-mutation matrix is revised to four, and one pinned
mutation is retired as vacuous.** Reached independently by both reviewers, then agreed.
- **Mutation 2 cannot be killed by any test, because its site cannot execute.** The
best-effort wake arm (`core/mod.rs`, the `besteffort_wake_rx` close arm) is unreachable
**by construction, twice over**: (i) `run_core_loop` owns a clone of `besteffort_wake_tx`
— created at `CoreController::new` and used for the `has_more` re-arm inside the loop —
and a tokio `Receiver::recv()` yields `None` only once *every* sender is dropped; (ii)
even without that clone, both `CoreController` and `CoreCommandSender` hold `reliable_tx`
alongside the wake sender, and the `select!` is `biased` with the reliable arm first, so
the reliable arm always wins the race to exit. Writing teardown there would be code that
provably never runs, dressed as a tested path.
- **Mutation 1's site is reachable but not unit-testable.** It sits inside `run_core_loop`,
which builds a real iroh endpoint and loads identity; no unit test can drive it.
- **Decision: (b) + (c).** Teardown is **hoisted to one unconditional site after the loop**,
so every `break` is covered structurally, including any added later — strictly better than
duplicating teardown across one live arm and one dead one. The **seam-level mutation gates
are the real ordering proof**, and the call site's live proof is owed to **phase 9**, which
gains an explicit row: *drop the controller / close the command channel while sharing*.
"UI crash" is not precise enough to serve as that row.
- **Rejected: an integration test built to preserve the number five.** It would pay for a
full iroh core plus test-only observability and prove only that a method was called — not
the ordering invariant, which is the thing that actually breaks.
- The 0b gate is therefore **four mutations**, enumerated exactly (round 16 P3-5 — the earlier
wording said "three plus the 0c pair", which reads as five and blurred what 0b owns):
| # | mutation | killed by | status |
|---|----------|-----------|--------|
| 1 (old gate 3) | remove the wait after the host kill | `explicit_shutdown_reaps_the_host_before_the_aec_can_unload` | killed now |
| 2 (old gate 4) | reverse `ScreenshareTeardown`'s field declaration order | `the_aec_unloads_after_the_children_on_the_drop_path` | killed now |
| 3 (old gate 5) | remove the reap loop from `ReapOnDrop::drop` | `dropping_a_guard_kills_and_then_reaps_the_child` | killed now |
| 4 (old 1) | remove the teardown at the hoisted post-loop call site | — | **deferred to the phase-9 row** *drop the controller / close the command channel while sharing* |
4-vs-5 separation verified: reversing the field order leaves the reap test green, and
removing the reap loop leaves the ordering test green. **0c's own pair (no SIGINT · no
SIGKILL fallback) is counted under 0c, not here**, along with the round-15/16 additions
(disarm the wrapper at entry · disarm it between the waits · treat a wait error as a reap ·
report an unconfirmed stop as clean · zero the grace).
**Round 15 (2026-07-26) — review of the 0b/0c-peerspeak implementation returned two blocking
findings, both accepted.** Recorded because both are the same shape: a defence that existed
but was disarmed exactly when it was needed.
- **The drop fallback was disarmed across its own wait.** `shutdown` took the child out of
the wrapper before the first `.await`; a cancellation or unwind during the wait left the
raw child to drop with `kill_on_drop` (which signals without reaping) while `Drop` found
`None`. The child now stays owned until the reap is **confirmed**.
- **A failed wait was reported as a reap, and the hard-kill wait was unbounded.** The
`io::Result` was discarded, so a wait error returned "reaped"; and a process in
uninterruptible sleep after SIGKILL could wedge the core loop forever. Both waits are now
bounded and the conflict case has a written policy: availability wins, the child stays
owned so the bounded `Drop` retry stays armed, and the residual risk is logged.
- **A gate of mine was vacuous and the review's fourth test-double point caught it.** The
elapsed-time assertion compared against `STOP_GRACE` itself, so zeroing the constant left
it trivially true. `the_grace_is_a_real_interval` now pins the constant to a band.
**Round 16 (2026-07-26) — the re-review of the 0b/0c-peerspeak fixes returned *approve with
follow-ups*: no blocking findings, five P3s, all five applied before the merge.** The two that
carry design content:
- **An unconfirmed stop was reported to the user as a clean one.** `stop_host` returned a bare
"was sharing" bool, so the one case where availability-first gives up (SIGKILL queued, reap
never confirmed) still emitted `ScreenShareStopped` with no warning — the UI would say
sharing ended while pixelpass might still be fanning out. `ReapOnDrop::shutdown` now returns
`StopOutcome`, `stop_host` returns `Option<StopOutcome>`, and an `Unconfirmed` user-initiated
stop raises a UI error naming the stray process. Session/viewer teardown discards the outcome
on purpose: no user is waiting on an answer there and the risk is already logged.
- **Cancellation coverage only reached the graceful wait.** The mid-wait test could not kill a
mutant that disarmed the wrapper *between* the two waits. Verified: the naive form of that
mutant does not compile (the child is borrowed from `self`), but the restructured form —
`self.child.take()` once cooperation has failed — compiles, and the pre-existing test passes
it. `cancelling_shutdown_after_the_kill_leaves_the_fallback_armed` kills it.
**Deferred item — aggregate teardown latency (round 16 P3-4).** Bounds are per child, not per
teardown. Sequential drain gives `2 × STOP_GRACE` per unconfirmed child inline (≈4 s), plus
`REAP_BUDGET` (250 ms) per child on the `Drop` path: three wedged children ≈6 s of command-loop
stall, ≈12.75 s worst case including drop retries. Accepted as-is for 0b — one host plus one or
two viewers is the real shape, and concurrency here would mean detaching children from the
session that owns the AEC's lifetime. **Trigger to revisit: a fourth tracked child becomes
routine, or a measured teardown exceeds 5 s.** The fix, when triggered, is to drain viewers
concurrently while still owned by `shutdown_children` — not to detach them.
**Round 17 (2026-07-26 night) — two reviews: the repair planner (changes-requested, all applied)
and the 0c actor design (four blocking issues, all accepted).** 0c is now sliced, because the
fault-handling surface — not the design — is what grew.
*The repair planner: no P1s, four reachable P2s and a P3, all applied in `9145b2a`.*
- **Only the canonical forms are ours.** `classify` recognised any loopback with one
pixelpass-looking endpoint, so a third party's `module-loopback source=some_mic
sink=pixelpass_capture_4242` was ours to unload once that pid died; and a `sink=` token nested
inside a quoted `sink_input_properties` value could be read as a top-level argument. The whole
recorded argument string must now equal what pixelpass itself writes.
- **The matcher's templates are generated from the loader's own renderers.** Hard-coding
`latency_msec=20` beside a matcher means a loader change silently blinds repair to every module
the new build loads — the fail-closed-and-silent class this project has now been bitten by
three times (F2, F13-1, the sticky-uncertainty inversion). `host/audio.rs` loads through the
same renderers, so drift is a compile-time question. Blindness is also *reported*:
`unrecognised_pixelpass_modules` names anything matching `pixelpass_capture_*` that no
canonical form recognises, so a newer pixelpass's shapes cannot make an older `--repair`
quietly clean up nothing.
- **Ordering is not a licence either.** Planning loopbacks before the sink is necessary and
insufficient: an unload can fail or be skipped, and a loopback can appear after planning. The
sink unload is now gated on `sink_still_referenced` against the fresh snapshot — any other
module naming that sink blocks it, ours or not, because the question is what would break rather
than who owns it.
- **Undecidable is not dead.** `Path::exists()` maps a permission error, a missing `/proc` and a
foreign pid namespace all to `false`, which read here as "dead, unload it". Liveness is now
`Alive | Dead | Unknown` via `try_exists()` behind a `/proc/self/stat` preflight, `Unknown`
behaves exactly like `Alive`, and it is reported separately so holding back is visible.
- **⚠️ One prescribed fix was not implementable as written, and measuring first is what caught
it.** The reviewer's fix for fingerprint fidelity was "use `pactl -f json list modules` and
deserialize the complete `argument`". **On pactl 17.0 those records carry no module index at
all** (`"index": null`), and `unload-module` accepts only an index — JSON alone cannot drive
repair. Replacement: **two listings, correlated positionally and checked** (ids and names from
the short listing, exact arguments from JSON; equal counts and equal names at every position or
the run refuses, with retries for a concurrent load). Verified on this host: both listings
return the same 17 modules in an identical name sequence, from 41 physical lines. The check
also turns the reviewer's fabricated-row attack from exploitable into harmless — a crafted
short-listing line has no JSON counterpart, so the sequences misalign and repair stops instead
of unloading an index inferred from text. *This is the reachability rule applied to a
prescription rather than a finding: the chain was valid, the API it assumed did not exist.*
- **Measured before relying on it (pactl 17.0, live server):** recorded arguments come back
byte-for-byte as passed, joined with single spaces, in order, and **`@DEFAULT_SINK@` is not
resolved** to the concrete device. Both facts are load-bearing for exact matching — had either
been false, the P2 fix would itself have been a silent blinding — so both carry a test.
- Normalisation deleted (P3): within one invocation every snapshot comes from one server, so
re-rendering does not happen and normalising only made different arguments compare equal. The
residual ABA window (planned module vanishes, a byte-identical one takes its index) cannot be
closed through an index-only unload API, and is now stated as a limitation in `Fingerprint`'s
own doc comment instead of implied away.
- Five vacuity gaps closed: a raw-pactl-text-to-plan test (the whole planner suite survived a
parser that dropped every argument), per-pid liveness counters over two pids, non-canonical and
nested-quote cases, and a reference-gate test. **One gap deliberately left open and declared:**
a comparator using only `id + args` cannot be killed by a non-vacuous test, because the module
*name* determines which argument grammar can match at all — that field is enforced structurally
by `classify`, and a test appearing to cover it would be the self-satisfying kind.
- **Field-verified twice on the live graph:** the A/B orphan test still removes exactly the two
orphans with the module table otherwise byte-identical, and a new fixture — a dead pid's legacy
sink plus a *non-canonical* loopback naming it — unloads nothing, reports the unrecognised
module, and reports the sink as still referenced.
*The 0c actor design: four blocking issues, all accepted; the epoch requirement conceded.*
- **A bounded join must not move the OS handle into `spawn_blocking`.** My ladder would have
taken the thread handle out of the guard to poll it; if the close future is then cancelled or
unwinds, `Drop` finds no handle and can neither poison nor fail-stop, while the blocking task
stays wedged forever and can pin runtime shutdown. This is the **same defect shape as round
15's** — a defence disarmed exactly when needed. The handle stays owned across every await;
`is_finished()` is polled and `join()` called only once it reports finished. Same rule for the
event task's handle (`await` through `&mut JoinHandle`).
- **`Commit::UnloadNow(id)` cannot forget the id.** An immediate unload can time out or be
cancelled, and a ledger that never recorded the module cannot retry or reconcile it. Slots
become a state machine — `Vacant | Loading { token, expected } | Loaded { fp } | Unloading
{ fp }` — with **affine** permits carrying a unique token, so two permitted loads for one slot
cannot both commit.
- **`kill_on_drop` does not roll back a server-side mutation.** A bounded `pactl load-module`
killed after the server created the module but before its id was read leaves a module with no
id anywhere. So an ambiguous load requires **bounded reconciliation by fingerprint** — reusing
repair's classification idea inside the live session, never its dead-pid policy — before any
further capture may start. Related: cancellation must never be `select!`ed against
`Command::output()`, or a completed load's id is dropped on the floor.
- **`_exit` is right, but the pre-exit sequence must not be able to block.** Event emission,
stdio flushing and tracing all take locks a wedged thread may hold, so the watchdog able to
`_exit` past a stalled diagnostic has to be **armed before** the wedge is detected, not created
in response to it. And `_exit` skips `CaptureHandle::Drop`, so `gst-launch-1.0` and any
in-flight `pactl` need parent-death/process-group containment or they outlive the host that
reported its own death — with gst still holding screen-capture resources.
- **Epoch conceded, and my vacuity instinct was right.** `object.serial` is unique and never
reused while global ids are, so "the object at this id still has the serial I recorded" is
complete proof of identity; there is no same-core interleaving that serial equality misses.
Epoch is carried for diagnostics and explicitly **not** a gate. It would only become
load-bearing across a daemon incarnation or an actor reconnect, and the design makes core
failure terminal with no reconnect — if that changes, the right answer is a core-incarnation
nonce, not a "something churned" counter that invalidates observations on unrelated traffic.
- **"Unjoinability, not slowness" is not literally implementable** and the wording is corrected:
no bounded observation distinguishes "returns one millisecond later" from "never returns", so
the death condition is *failure to terminate within the post-cancellation policy deadline*.
Two budgets, not one — a running MainLoop quitting is a different question from an
initialisation call returning after cancellation, and the second is normally longer.
- **`GraphCmd::Route(Vec<u32>)` is deleted rather than fixed.** Matching and routing stay inside
the actor's registry callback, where removals are already ordered against routes in-thread, so
the privacy race is not introduced at all. For phase 6 the rule is structural: the only
addressable type is an `ObservedNode { global_id, serial, epoch }` constructible solely from
the actor's own observation, kept private and non-`Copy`, revalidated on serial immediately
before any mutation. A bare id is not addressable.
- **An unacked `ClearRoutes` is not a wedge** (agreed), with one qualification taken: a stream
setting `node.dont-reconnect`/`node.dont-fallback` may be left silent rather than moved back to
the default, so the outcome is surfaced as `ClearRoutesUnconfirmed` rather than treated as
benign. Separately, blindly clearing `target.object` can erase a target the user set manually —
the prior value must be recorded and restored only while it is still pixelpass-owned.
- **One terminal fault needs a coordinator, not an emitter.** If the actor emits `CoreError`
immediately and the subsequent teardown then fails to join, peerspeak never learns the process
is fail-stopping. Actor faults are internal *candidates*; the tokio-side coordinator emits
exactly one final fault, and `Wedged` overrides any earlier candidate. Because a callback panic
can cross `extern "C"` and abort before any event is produced, **peerspeak must treat
unexpected stdout EOF as a synthetic terminal fault** rather than trusting that a JSON line
arrives.
*Measured for the actor argument (3 of 3 trials, live graph):* pipewire-pulse accepts **two sinks
with an identical `node.name`** — no rename, no suffix, no refusal, both visible as `<name>` and
`<name>.monitor` — and `pulsesrc device=<name>.monitor` attached to the **older** one every time.
So a surviving wedged owner does not merely risk a collision: it **silently steals the next
session's capture** while the loopbacks feed the new sink. That retires "detach and carry on" as
an option, and it is the evidence behind rejecting session-unique sink names (which would trade a
fail-stop ownership fault for silent accumulation, and re-open the discovery grammar 0c step 1
just closed and field-proved).
**0c step 2 is therefore sliced, and the slices land and are reviewed independently.** Nothing
here reopens D6 — the connection-owned-sink design is unchanged; what grew is the process-
lifecycle and fault surface, and a material part of it is pre-existing debt 0c forced into the
light (the `abort()` orphan race, the unbounded join, peerspeak advertising a dead share):
| slice | scope | why it can land alone |
|-------|-------|-----------------------|
| S1 | repair planner (`919d5bd` + `9145b2a`) | done; awaiting re-review, then merge |
| S2 | peerspeak host-fault path: always-on notice channel, EOF synthesis, session-scoped fault, clear `is_sharing` + presence ticket, `ScreenShareStopped` then error | fixes a defect **today** — a dead share stays advertised — and is independent of the actor |
| S3 | pixelpass ledger transactions + ambiguous-load reconciliation + child containment + pre-armed watchdog + poison state machine + supervisor health arm | fixes the `abort()` orphan race **today**; no libpipewire work |
| S4 | **built, validated, and committed locally in pixelpass `781defc` (2026-08-20):** the `AudioGraphOwner` actor itself, the readiness handshake, and both measured budgets | the only slice that needs new PipeWire mechanism |
| S5 | **built, validated, and committed locally in pixelpass `781defc` (2026-08-21):** the two live exit gates: two-host ownership/repair and Stop Share SIGINT | needs S4 on the graph |
**S5 live evidence (2026-08-21).** Pixelpass's ignored
`live_two_host_sigkill_and_repair_preserve_the_survivor` gate starts two independent routing
owners, observes a distinct native sink and ownership-tagged loopback for each, SIGKILLs the
first, and proves only its sink disappears. `--repair` then removes exactly the dead host's
loopback while explicitly leaving the second live host alone. The survivor exits through a real
SIGINT with its active graph teardown measured at 40 ms, inside peerspeak's 2 s grace, and leaves
no sink or module residue. Peerspeak's separate
`stop_share_ends_the_real_host_via_sigint_within_the_grace` end-to-end gate also passed against
the freshly built pixelpass binary, proving Stop Share drives that signal path before fallback.
The pre-gate cold review also found and fixed an S4 unwind regression: the actor's emergency
`Stop` path had quit without restoring still-owned `target.object` values. `Stop` now performs
the same ownership-checked restoration and waits for a PipeWire core round-trip before closing
the connection, so constructor cancellation or unwind cannot knowingly strand an app on the
disappearing sink.
**Round 18 (2026-07-26 night) — two more repair review rounds. `--repair` now reads and unloads
through libpulse, and one of the review's own prescriptions had to be replaced after measuring.**
*Round 17c — the re-review of my round-17a fixes found two more blocking P2s. Two of the four
fixes I had applied were themselves defective; this is the third time the "audit your own fixes"
rule has paid.*
- **My two-listing correlation was unsound.** Pairing short-listing indices with JSON arguments by
position breaks whenever module names repeat: another client loading one module and unloading
another *between the two calls* leaves counts and names aligned while every argument has shifted
by one, so a foreign module inherits a canonical fingerprint. The name check cannot see it and
the retry never fires, because correlation "succeeded".
- **My liveness fix still converted invisible-but-alive into dead.** A `/proc/self` preflight
proves nothing: inside a pid namespace — a container, a distrobox — `self` stays visible while
every process in the parent namespace is invisible, and `hidepid` has the same shape.
*Round 18 (round 4) — the fix for both, and a third defect neither of us had reached.*
- **Record boundaries in `pactl list short modules` are unprovable, and this needs no adversary.**
A genuine module whose argument contains a newline renders a first line that is byte-exactly one
of our canonical forms, with the rest dropped as an unparseable continuation — no forged index,
so no duplicate-index check can see it. **Field-confirmed on the live server** with
`…latency_msec=20\nremix=false`, `remix` being a real loopback option. A tab in the same position
is worse: it hides a sink reference from the gate that protects a still-referenced sink.
- **Locality was a guess.** `PULSE_SERVER` is a fallback *list*, so `unix:/missing tcp:remote:4713`
passes any "starts with unix:" test and then connects to another machine, where local pids mean
nothing and a live remote host's modules look dead.
- **Resolution: `src/repair/introspect.rs`, one verified-local connection.** `pa_module_info`
carries index, name and exact argument in a single record; `pa_context_is_local()` answers
locality about the connection actually established; and unloading goes back through that same
connection, so listing and destruction cannot disagree about which server they mean. Bounded
throughout (3 s connect, 3 s per request, non-blocking iteration plus a 2 ms sleep). The layer
holds no policy but "refuse the wrong server" — every decision stays in the pure planner.
- **The dependency was the user's call, taken with sign-off after vetting.** libpulse-binding
2.30.1: MIT/Apache-2.0, 5.5M downloads, 3 new crates total, a build script that only probes
pkg-config, no network or subprocess use in any source, and all three historical RustSec
advisories (2018-0020/0021, 2019-0038) fixed by 2.6.0. Reasoning recorded beside the dep.
- ⚠️ **REUSABLE — a prescription can fail reachability, not just a finding.** The reviewer's
fidelity fix was "use `pactl -f json list modules`". On pactl 17 those records carry **no module
index at all** (`"index": null`) while `unload-module` accepts only an index, so it can never
stand alone. Measuring first is what caught it.
- ⚠️ **REUSABLE — the field test found a bug no unit test could reach, and it was 0b's bug again.**
The first introspection version did its work correctly and then aborted on the way out:
`Assertion '!e->dead' failed at mainloop.c:207, function mainloop_io_free()` — SIGABRT, core
dumped, **exit 134, so a fully successful repair reported failure to its caller**. Rust drops
fields in declaration order and the context's teardown frees IO events living in the mainloop,
which I had declared first. Fixed, then hardened past the fix: `Drop` explicitly takes and
destroys the context before the mainloop, so the ordering no longer depends on where the fields
are written. **Field-order drop hazards are not a peerspeak-specific lesson; they recur wherever
one object's teardown reaches into another's.**
- **Still open, deliberately, and recorded rather than guessed:** closing the namespace hole needs
modules to carry an **owner token** (machine/boot identity plus pid-namespace identity), with
token-less modules treated as `Unknown`. That changes what pixelpass writes into the graph *and
how far back `--repair` can clean up* — orphans from any older build would become uncleanable,
which is a regression in the tool's entire purpose. `NSpid > 1` remains a sound negative signal;
`NSpid == 1` is explicitly **not** proof, since its leftmost value is relative to the procfs that
was mounted.
- **Deferred, now cheap to reconsider:** `host/audio.rs` still loads modules via `pactl` and parses
the index off stdout, which is part of why S3's ambiguous-load problem exists. With libpulse in
the tree, `pa_context_load_module` returns the index through an observable operation.
**Round 1 — 13 items, 12 accepted.** Phase reorder (AEC machine before dry-run); typed capture
plan (accepted, moved *earlier* than proposed); Phase 3 five-part gate; tag-consumption gating;
Phase 6 matrix mandatory; 0b unwind backstop restored **and my mutation test corrected — it
+499 -252
View File
@@ -1,35 +1,477 @@
# Phase 5 — dry-run audit gate: results
**Status: 🚦 GATE FAILED. Phase 6 does not start.** Two defects found, one of them
fatal to the whole mechanism. Both go to the design doc as **round 8** per impl
plan §5.3.
**Status: 🟢 GATE PASSED (run 2, 2026-07-26). All 13 §5.1 rows completed; the
eligible half of every row is non-empty. O5 re-measured on the fixed graph and
stays closed.** One new defect was found and fixed during the run (F13-1); three
findings are recorded as non-blocking, and three rows carry recorded
substitutions. Phase 6 is unblocked **by this file**, and F11-1 — the other gate —
was closed with this data on 2026-07-26 (see "What still blocks phase 6").
- **Run date:** 2026-07-25
- **Run date:** 2026-07-26 (run 1: 2026-07-25, gate FAILED — see history below)
- **Host:** `cazen` — PipeWire 1.6.8, WirePlumber 0.5.15, CachyOS
- **Audit build:** pixelpass branch `phase5-dry-run-audit`, release profile
- **Ambient load during the runs:** FINAL FANTASY XIV playing audio (`client.id`
88, pid 14651), Arctis 1 Wireless as an active sink
The audit itself worked exactly as designed: it observed the live graph, ran
phases 2–4 on every registry event, created no links, and reported a complete
eligible/excluded partition with stable reason codes. **It found the defects on
the first live run.** That is the phase doing its job — §5's argument was that a
fixture proves the code matches my model of PipeWire while only a live run proves
my model matches PipeWire, and my model was wrong.
- **Audit build:** pixelpass `main` @ `91c4ded`, release profile
- **peerspeak build:** `main` @ `b68fca6` (phase 1 merged)
- **Ambient load:** Firefox playing audio throughout (a live, uncontrived
candidate); Sunshine running (pid 3838); Arctis 1 Wireless as active sink
- **Graph size:** 14 Nodes, 4 Devices, 57 Ports, 4 Links, 24 Clients
---
## F1 🔴 FATAL — the registry `global` event delivers only a filtered subset of node properties
## Addendum — pre-Phase-6 hardware/readiness gate (2026-08-21)
**The phase-3 adapter reads eight node properties that the PipeWire registry
never announces.** They are parsed off `obj.props` in the registry `global`
callback (`pixelpass/src/host/observer/adapter.rs`), where they are silently
absent, so every one of them is permanently `None`/`false`.
This addendum does not rewrite the historical 2026-07-26 matrix. Design v3.8 §6.9 adds one
conservative edge the old engine did not have: tainted playback into a positively classified
hardware sink taints passive capture nodes carrying the same snapshot-local `device.id`. It
also closes the readiness-budget calibration that round 9 left open.
### Measured
### Hardware path — targeted live exact partition
The complete set of keys the registry announces for a `Node` global on this host
(union over every node, via `pw-cli ls Node`):
The audit started first and reached readiness. Controlled modules then created three named
candidates:
| candidate | expected | observed |
| --- | --- | --- |
| tagged playback into ALC897 | excluded root | `peerspeak-owned` |
| capture/re-emitter reading the ALC897 source | excluded through hidden same-device hop | `tainted-owner-bridge` |
| identical capture/re-emitter reading the Arctis source | eligible; different Device | eligible, no reason |
The settled record was `graph_ready=true`, epoch `complete`. This revises row 9 for every
future full matrix: music-only and a microphone on a **different Device** remain eligible;
a microphone on the **same Device receiving tainted playback** and a tainted-monitor capture
are excluded. This is deliberate fail-closed over-exclusion because a private hardware or
firmware loopback is not observable as a PipeWire Link.
The host's ALC897 had no `Stereo Mix` capture-source item: `Input Source` offered Rear Mic,
Front Mic and Line. Its separate `Loopback Mixing` control was disabled. Runtime safety does
not depend on either spelling; USB/vendor loopbacks need the same rule.
### Readiness calibration — retain the 2 s sticky deadline
Each measurement used a fresh observer process and its emitted monotonic `at_ms` readiness
timestamp:
| arm | runs | min | p50 | p95 | max | timed out |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| ambient live graph | 30 | 4 ms | 5 ms | 10 ms | 11 ms | 0 |
| 24 temporary null sinks + 24 loopbacks | 30 | 109 ms | 111 ms | 113 ms | 114 ms | 0 |
| 250 null-sink create/remove cycles concurrent with starts | 20 | 4 ms | 5 ms | 6 ms | 8 ms | 0 |
The inflated maximum is 17.5x below the 2 s deadline. All temporary modules were unloaded by
the exact ids returned from `pactl`; no test modules remained, and the configured default
sink/source were unchanged.
### Regression gates
- focused taint engine: 83 passed;
- observer: 65 pure passed, then all 3 serialized live PipeWire tests passed;
- complete non-GUI PixelPass suite: 313 passed, 0 failed, 9 ignored;
- disk before the first build: 20 GiB free; after tests and Clippy: 17 GiB free.
**Verdict:** both former pre-Phase-6 design blockers are closed. S4/S5/0d, round 11, and the
first pure Phase 6 channel-planning prerequisite are committed locally in PixelPass `781defc`;
the first fan-out mutation remains the current work.
---
## What changed since run 1
Run 1 failed on two defects, both fixed before this run:
- **F1** (fatal): the registry `global` event delivers only a filtered subset of
node properties, so eight properties the engine depends on were permanently
absent. Fixed by design round 8 / **phase 3r** — bind every Node and Device
and read properties from `info`.
- **F2**: a machine-wide over-exclusion cascade downstream of F1.
Both are gone: the baseline run (no fixture at all) reports **1 candidate,
eligible, empty taint set**.
### 🔴 F13-1 — FOUND AND FIXED DURING THIS RUN
**Row 1 failed on its first attempt, and the cause was a third defect of exactly
the F2 class from a new source: pipewire-pulse's PID was unresolvable on this
host, permanently.**
`pulse_pid::candidate` returned the single `pipewire.sec.pid` shared by two or
more Clients, on the stated reasoning that "native PipeWire clients carry their
own distinct PID; only the Pulse shim repeats one value". Measured: **WirePlumber
repeats one too.** It holds two Clients — `WirePlumber` and
`WirePlumber [export]` — both `sec_pid` 1747. Two values repeated (1747 and
pipewire-pulse's 2528), the rule called that ambiguous, and returned `None`.
With the daemon PID unknown, `owner::keys_of`'s documented fail-closed asymmetry
takes over: key 4's suppression never fires, every Pulse-emulated node fuses into
one owner, and the cascade follows. Row 1's observed failure:
```
ELIGIBLE (1): r1_plain_app
EXCLUDED: Firefox tainted-owner-bridge key=application.process.id
r1_c_play tainted-owner-bridge <- the CLEAN control half
TAINT: ... + both sound cards, all three sunshine sinks, sunshine itself
```
The rule was wrong in **both** directions, so the prefilter was removed rather
than patched:
- **False ambiguity** — any second process holding two Clients defeats it.
WirePlumber always does, so this was permanent, not a corner case.
- **False absence** — a session where pipewire-pulse holds exactly one Client
(one Pulse app running) repeats nothing, so the candidate is missed and the
same cascade follows.
`comm` was always the authoritative check; repetition was a heuristic standing in
front of it, and it was a guess about other processes' Client counts. Fixed in
pixelpass `91c4ded`: `candidates()` lists every distinct `sec_pid`, `resolve()`
picks the unique one whose `/proc/<pid>/comm` is exactly `pipewire-pulse`, and
several matches still fail closed (a single `Option<u32>` cannot suppress two
daemons — recorded, not approximated). The adapter probes only PIDs *entering*
the candidate set, and `retain_probed_comms` bounds the map to live PIDs so a PID
that leaves and returns is re-probed instead of answered from a stale `comm`.
**This is the §5.1 exact-partition requirement earning its keep for the second
time.** The verdict was fail-closed and silent; only the asserted *eligible* half
exposed it. An exclusion-only checklist would have passed this build too.
---
## §5.1 — the matrix
Every row ran with `PIXELPASS_AUDIO_AUDIT_AEC=off` except row 12. Every row ran
in its **own** audit process, so nothing carries over (sticky taint is
per-process state).
⚠️ **Methodology change from run 1, and it is load-bearing.** Run 1 built each
fixture *before* starting the audit. On this host the entire graph then arrives
as one enumeration burst (~122 events in 1–2 ms), so every node is first tainted
while `graph_ready` is still false, that partial-graph taint is recorded into
sticky state, and on the single ready record the sticky pass raises
`TaintedOwnerBridge { key: None }` before the evidence pass can name a key —
`raise` will not replace a same-rank reason. Verdicts were still correct but rows
could not assert their key. This run starts the audit first, waits for readiness,
then builds the fixture, so taint is derived from real topology *changes* against
a ready graph — which is also the dynamic path §6.3 cares about. Keys are read at
**derivation** (first non-sticky appearance), not from the final record.
| # | scenario | status |
| --- | --- | --- |
| 1 | null-sink + loopback forwarder, owner bridge | ✅ **pass** (after F13-1 fixed) |
| 1b | Sunshine's topology (opportunistic, non-gating) | 🟡 observed, nothing to exclude — see below |
| 2 | gst split clients, tainted input | ✅ **pass**, key 4 named at derivation |
| 3 | two Pulse modules, one tainted | ✅ **pass** |
| 4 | peerspeak native call playback | ✅ **pass** — real tagging site |
| 5 | peerspeak-spawned mpv | ✅ **pass** — real tagging site, hand-launched mpv eligible |
| 6 | peerspeak notification sound | ✅ **pass** — real tagging site |
| 7 | second host's capture sink + forwarder | ✅ **pass**, eligible half non-empty |
| 8 | EasyEffects | 🟡 **pass with substitution** — echo-cancel stood in |
| 9 | Firefox three cases | ✅ **pass** (cases 2–3 via gst; see substitution) |
| 10 | sticky taint across teardown | ✅ **pass**, all four phases incl. retirement |
| 11 | recycled serial / index / link-group | ✅ **pass**, and provably non-vacuous |
| 12 | AEC loaded → unloaded → Revoked | ✅ **pass** |
| 13 | `Audio/Duplex` device | 🟡 **pass with synthetic node** — over-taint confirmed |
### Row 1 — owner bridge, key named
```
ELIGIBLE (3): Firefox · r1_c_play · r1_plain_app
EXCLUDED (2): peerspeak_owned_call_4242 peerspeak-owned
r1_t_play tainted-owner-bridge key=node.link-group
TAINT (5): the tagged producer, r1_t_src, r1_t_cap, r1_t_play, r1_t_dest
```
The clean half is an **identically shaped** forwarder — same module type, same
monitor-read, same re-emit — differing only in whether anything tainted feeds it.
`r1_c_play` eligible is the assertion an exclude-everything build cannot satisfy.
The key is `node.link-group`, a strong key, not a link walk.
### Row 2 — GStreamer split clients, key 4
Measured props confirm the shape is the real refutation: `r2_gst_tainted_src`
(client 188) and `r2_gst_tainted_sink` (client 191) are **different Clients** of
**one process**, pid 235628, with no `link-group` and no `pulse.module.id`. So
`application.process.id` is the only key that can relate them.
Derivation record (seq 209): `r2_gst_tainted_sink` → `tainted-owner-bridge`,
**`owner_key=application.process.id`**. `r2_gst_clean_sink`, reading an untainted
monitor in a second process, is eligible.
### Rows 4–6 — peerspeak's own paths, through the real call sites
Driven by peerspeak's phase-1 live gate tests (`--ignored`), i.e. the real
tagging sites, not a hand-rolled env: "emission alone proves only that peerspeak
talks, not that pixelpass listens" (impl plan §3).
| node | verdict |
| --- | --- |
| `peerspeak_owned_call_238172` | EXCLUDED `peerspeak-owned` |
| `peerspeak_owned_mpv_238196` | EXCLUDED `peerspeak-owned` |
| `peerspeak_owned_notify_238231` | EXCLUDED `peerspeak-owned` |
| `peerspeak_owned_clip_238249` | EXCLUDED `peerspeak-owned` (bonus — chat clips) |
| `mpv` (launched by hand, untagged) | **ELIGIBLE** |
This is the cross-repo contract closed end to end on live nodes.
### Row 9 — the over-exclusion promise
```
ELIGIBLE: Firefox (music only) · r9_mic_out (captures an untainted real device)
EXCLUDED: r9_mon_out tainted-owner-bridge key=application.process.id
```
`r9_mic_out` is the row that defends §6.1.1: an app that captures a real
`session_device` source and also plays audio stays shareable. The device source
itself never entered the taint set.
### Row 10 — the full sticky lifecycle
| phase | topology | verdict |
| --- | --- | --- |
| A | tainted producer + forwarder | `r10_play_out` EXCLUDED, key `node.link-group` |
| B | **tagged producer killed**, forwarder lives | **still EXCLUDED** (sticky) — current topology alone no longer justifies it |
| C | forwarder owner replaced, tainted sink kept | fresh forwarder EXCLUDED — correct: a sink that received call audio is still a hazard while it lives |
| D | **every** tainted object torn down, then restart | taint set **empty** at 16.3 s; `r10_new_out` **ELIGIBLE** at 20.3 s |
Phase B proves stickiness works; phase D proves it is not permanent. Phase C is
worth keeping in mind when reading any future report: partial teardown legitimately
does *not* retire taint, and that is easy to mistake for over-exclusion.
### Row 11 — recycled identifiers, provably non-vacuous
| generation | `node.link-group` | global id (`r11_src`) | `object.serial` (`r11_play`) | pulse module |
| --- | --- | --- | --- | --- |
| 1 (tainted) | `loopback-2528-14` | 168 | 4702 | 536870919 |
| 2 (after teardown) | **`loopback-2528-14`** | **168** | 4746 | 536870920 |
The `node.link-group` came back **byte-identical** — and it is the very key that
carried the taint in generation 1 — and the global id was reused. Generation 2's
`r11_play` is **ELIGIBLE** with an empty taint set. `object.serial` correctly did
not recycle, which is why the model keys everything by it.
### Row 12 — AEC lifecycle
| stage | `aec_state` | `fan_out_permitted` | candidates |
| --- | --- | --- | --- |
| module live, configured | `validated` | `true` | Firefox + `r12_plain_app` ELIGIBLE; `echo-cancel-playback` EXCLUDED `aec-identity` |
| module unloaded | `revoked` | `false` (`gate_reason=aec-revoked`) | every candidate EXCLUDED `aec-revoked` |
All **four** link-group siblings (`sink`, `source`, `capture`, `playback`) carry
`aec-identity`; only `echo-cancel-playback` is a candidate, so it is the only one
in the excluded partition. Ordinary apps staying eligible *while validated* is
what makes "the gate is open" observable rather than inferred.
### Row 13 — `Audio/Duplex` over-taint (known accepted)
No real duplex device exists on this host, so one was synthesised by overriding
`media.class=Audio/Duplex` on a null sink. Its playback side was tainted and its
capture-side consumer was dragged down with it (`r13_dup_play` EXCLUDED), with
the eligible half intact. **Fixture limit, stated plainly:** on a null sink the
capture side *is* the monitor, so this cannot separate the duplex smear from the
ordinary sink→monitor edge. The accepted over-taint is confirmed as *behaviour*;
a real duplex device is still the only way to isolate the mechanism.
### Row 1b — Sunshine (opportunistic, non-gating)
Sunshine ran throughout. Its three null sinks stayed SUSPENDED and it read the
**hardware** monitor instead, exactly as §5.3 warned. It appears consistently and
correctly as `sunshine` / `tainted-upstream` whenever the monitor it reads is
tainted (rows 8, 12, o5). It has **no re-emitting output leg** — it sends over
the network — so it is never a candidate and there is nothing to exclude. Recorded
as observed; the "if a re-emitting leg exists" clause did not apply. A real
third-party forwarder sample remains owed.
---
## §5.2 — O5 re-measured
The run-1 numbers do not carry over: they were measured on the graph F1 degraded,
and phase 3r adds a bind plus an `info` round-trip **per node**, which is new I/O
that run never exercised.
Per-run, across all 13 rows (`recompute` in µs):
| run | events | ev/s | max | mean | emit max | busy fraction | ready@ms |
| --- | --- | --- | --- | --- | --- | --- | --- |
| baseline | 123 | 21.4 | 20 | 3 | 6 | 0.0001 | 1 |
| o5 (churn) | 407 | 44.0 | 32 | 10 | 9 | 0.0006 | 1 |
| row01 | 219 | 41.7 | 53 | 10 | 9 | 0.0006 | 1 |
| row02 | 241 | 45.9 | **67** | 11 | 10 | 0.0006 | 1 |
| row03 | 206 | 48.5 | 54 | 8 | 8 | 0.0005 | 1 |
| row0456 | 185 | 20.0 | 38 | 7 | 10 | 0.0002 | 2 |
| row07 | 184 | 43.3 | 40 | 7 | 9 | 0.0004 | 1 |
| row08 | 172 | 32.6 | 44 | 6 | 7 | 0.0003 | 2 |
| row09 | 224 | 30.9 | 52 | 9 | 9 | 0.0004 | 1 |
| row10 | 332 | 14.3 | 41 | 11 | 15 | 0.0002 | 1 |
| row11 | 298 | 24.1 | 41 | 9 | 11 | 0.0003 | 1 |
| row12 | 188 | 25.9 | 39 | 7 | 8 | 0.0003 | 1 |
| row13 | 193 | 36.8 | 43 | 8 | 7 | 0.0004 | 1 |
The dedicated churn run (five load/unload cycles of null-sink + loopback, the
same shape as run 1's measurement):
```json
{"kind":"metrics","graph_events":407,"tick_events":37,"emitted_records":407,
"span_us":9249639,"graph_events_per_sec":44.0,
"recompute_max_us":32,"recompute_mean_us":10,
"recompute_p50":"<50us","recompute_p90":"<50us","recompute_p99":"<50us",
"recompute_distribution":[["<50us",444]],
"emit_max_us":9,"emit_mean_us":1,
"busy_us":5240,"busy_fraction":0.0006,
"queued_events":292,"queue_threshold_us":100}
```
**O5 stays closed on the real graph.** Worst recompute across every run is
**67 µs**; every single recompute in the churn run finished under 50 µs, against
a 44 Hz event rate under churn heavier than a desktop produces at rest. The
observer thread spent **0.06 %** of wall time working. Node binding roughly
doubled the per-event cost (run 1: 15 µs max / 4 µs mean; now 32 µs / 10 µs on
the same churn shape) and that is the honest cost of the F1 fix — it buys three
orders of magnitude of remaining headroom, not one.
**Readiness with node binds: 1–2 ms**, with ~122 enumeration events and 18 binds
(14 Nodes + 4 Devices), against the 2000 ms budget. `queued_events` is high
(292) for the same benign reason as run 1: PipeWire delivers enumeration and
teardown in bursts, and a 32 µs recompute drains a burst faster than it forms.
`busy_fraction` is the number to trust.
⚠️ **Historical run-2 finding, closed by the 2026-08-21 addendum above:** 1–2 ms
against 2000 ms was three orders of magnitude of slack on *this* host with 18
binds, but not an argument about enumeration volume or instability. The addendum
adds a 48-module graph and concurrent create/remove churn and retains the 2 s
budget from that evidence.
---
## Findings recorded, not blocking
### R2-1 — the audit's `sticky` flag is nearly always true, so it says little
As emitted, `sticky` means "this node is in the remembered set", which
`seed_sticky` populates for any node whose current reason the sticky pass agrees
with — i.e. essentially every currently-tainted node. It does **not** mean
"excluded *only* because remembered", which is what its doc comment implies and
what a reader diagnosing "why is this still excluded?" wants.
The information exists: round 9 already computes a second, **evidence-only** pass
(that is the whole provenance mechanism). Emitting "excluded by memory alone"
would make row 10 phase B assertable from a single record instead of from a
sequence. Not fixed here — it is a reporting change to a merged phase in the
middle of a gate run. Row 10 was asserted behaviourally instead, which is
stronger anyway.
### R2-2 — a bridge key is lost when a leg reappears under a new serial
Row 2 named `application.process.id` at derivation (seq 209), then gst re-created
that node; the sticky owner re-seeded the new serial through `reason_for`, whose
documented fallback is `TaintedOwnerBridge { key: None }`, and `raise` will not
replace a same-rank reason with a better-informed one. The verdict is unaffected;
only the diagnosis degrades. The fallback is honest when the owner has no live
tainted receiver, and stale when it does — which is the case worth improving.
### R2-3 — `owner_key` had to be added to the record to run row 1 at all
Row 1 asserts "reason = owner bridge, **naming the key**", and the record could
not express it: `Reason::code` collapses `TaintedOwnerBridge { key }` to one
string. `OwnerKey::code` already documented itself as ending up in the phase 5
audit output; it was simply never wired to it. Added in pixelpass `d462754`
(read-only, diagnostic-only, mutation-verified test). Worth noting as a gate-spec
lesson: the row could not have been asserted from any previous build's output.
---
## Substitutions, stated so they are not mistaken for passes
| row | asked for | used instead | why |
| --- | --- | --- | --- |
| 8 | EasyEffects | `module-echo-cancel` with `AEC=off` | EasyEffects makes itself the default sink on start and the user had live audio playing. `module-filter-chain` cannot stand in either — it is a PipeWire module, so `pactl load-module` answers "No such entity" (measured). The stand-in produces the same shape (four nodes, one `node.link-group`) and exercises `foreign-echo-cancel` (decision D3), a reason code no other row reaches. |
| 9 | Firefox's mic + monitor capture | `gst-launch` pipelines | Firefox's mic and monitor-capture paths need interactive GUI permission grants. Firefox is present live as case 1 in every row. Case 2 captures the motherboard's **analog input**, not the headset mic the user is wearing — identical to the engine (both `session_device` sources), and nothing of the user is recorded. |
| 13 | a real `Audio/Duplex` device | synthetic `media.class` override | None on this host. See row 13 above for what the fixture cannot show. |
---
## What still blocks phase 6
This file passing removes **one** of the two gates. F11-1, the other, is now
closed. Still outstanding:
1. **Hardware playback-to-capture paths ("Stereo Mix")** defeat `session_device`
and are a real echo path — needs ALSA control inspection; user design call owed.
2. **Phases 0b / 0c / 0d** are untouched and all precede phase 6.
3. **The readiness budget calibration argument** (above).
4. **Owed samples:** a real third-party forwarder (row 1b), EasyEffects (row 8),
a real `Audio/Duplex` device (row 13).
### ✅ F11-1 — closed 2026-07-26, with this matrix's data
The rule now implemented (pixelpass `c78eb2d`, §6.1.2's round-13 box): **key 4 bounds an
owner only when the node's Client resolves** — an unambiguous Client yielding
`Some(pipewire.sec.pid)`, read *before* pipewire-pulse suppression — so a node can no
longer bound itself, and escape `propagate_unresolved_owner`'s sweep, with an
`application.process.id` it invented. Bridging still uses the full union.
Codex's round-12 sharpening was the decisive part: "resolved" must mean a `sec_pid`, not
"a unique Client object exists", and the **unique-but-pid-less** row is the only one that
tells the two apart. All five Client cases are unit tests (absent · ambiguous ·
unique-but-pid-less · resolved-native · resolved-to-pipewire-pulse), plus the recorded
three-step leak path end to end. Mutation-verified: dropping the provenance test fails
four of the six rows and leaves the two no-over-exclusion rows green.
**The cost question the deferral was waiting on, measured on this host:** the before- and
after-binaries audited the *same* live graph simultaneously (both are read-only observers)
— tagged producer into the default sink, `parec` on its monitor as a live tainted reader
so the sweep was genuinely armed, Firefox + `aplay` + `pacat` as bystanders. **181 records
each, the same 14 distinct decision states, none exclusive to either side, no
`unresolved-owner` on either, eligible half non-empty throughout.** O5 unmoved (identical
p50 15 µs and busy fraction 0.0012). Every real app here is native or Pulse-emulated and
**both resolve**; sweeping all 18 live nodes, the only unresolved-Client ones were
`Dummy-Driver` and `Freewheel-Driver`, which carry no pid key to lose.
---
## Reproducing this run
Scripts live in the session scratchpad (not committed — they hard-code paths):
one per row, plus `lib.sh`, `summarize.py` and `keys.py`. The shape of every row:
```sh
audit_start out.jsonl off # start FIRST, wait for graph_ready
... build fixture ... # taint arrives as topology CHANGES
audit_stop # SIGTERM: flushes the O5 summary
python3 summarize.py out.jsonl # final partition + derivations + metrics
```
```
env PIXELPASS_AUDIO_AUDIT_FILE=/path/out.jsonl PIXELPASS_AUDIO_AUDIT_AEC=off \
./target/release/pixelpass --audit-audio
```
Rig notes that cost time:
- A tagged producer: `env PIPEWIRE_ALSA='{ "peerspeak.owned": "1", "node.name":
"peerspeak_owned_call_4242", "target.object": "<sink>" }' aplay -c 2 -r 48000
-f S16_LE -t raw -d 30 /dev/zero`. Both carriers land, and `target.object`
routes it.
- ⚠️ `pactl load-module module-echo-cancel --help` **loads the module** with
`--help` as its argument instead of printing help. It was loaded accidentally
during this session and unloaded again; check `pactl list short modules` after
any such probe.
- ⚠️ `pkill -f <pattern>` matches the harness's own shell command line and kills
the script. Use `pkill -x` or an exact pid.
- ⚠️ Under `set -e`, `kill` on an already-exited pid aborts the row before its
modules are unloaded; and `timeout` exiting 124 is *success* for the audit.
---
## History — run 1 (2026-07-25): GATE FAILED
Kept because the reasoning is still the record of why the observation boundary
was redesigned.
### F1 🔴 FATAL — the registry `global` event delivers only a filtered subset of node properties
The phase-3 adapter read eight node properties the registry never announces.
Parsed off `obj.props` in the registry `global` callback, they were silently
absent, so every one was permanently `None`/`false`.
The complete set the registry announces for a `Node` on this host:
```
application.name client.api client.id device.id factory.id media.class
@@ -37,246 +479,51 @@ node.description node.name node.nick object.path object.serial
priority.driver priority.session
```
Against what the adapter tries to read:
| property | announced? | what dies without it |
| property | announced? | what died without it |
| --- | --- | --- |
| `object.serial` | ✅ | — |
| `node.name` | ✅ | — |
| `media.class` | ✅ | — |
| `client.id` | ✅ | — |
| `device.id` | ✅ | — |
| **`peerspeak.owned`** | ❌ | **the primary taint root (v3.4 §5.1, all of phase 1)** |
| `object.serial`, `node.name`, `media.class`, `client.id`, `device.id` | ✅ | — |
| **`peerspeak.owned`** | ❌ | **the primary taint root (all of phase 1)** |
| **`pulse.module.id`** | ❌ | **AEC identity exclusion + phase 4 validation** |
| **`node.link-group`** | ❌ | the link-group owner key (echo-cancel, EasyEffects, loopback siblings) |
| **`application.process.id`** | ❌ | the process owner key (GStreamer split clients, §5.1 row 2) |
| **`node.passthrough`** | ❌ | the passthrough local exclusion (a second link corrupts an encoded stream) |
| **`device.api`** | ❌ | `session_device` classification |
| **`factory.name`** | ❌ | `session_device` classification — the discriminator itself |
| **`alsa.driver_name`** | ❌ | `session_device` classification (the `snd_aloop` denylist) |
| **`node.link-group`** | ❌ | the link-group owner key |
| **`application.process.id`** | ❌ | the process owner key |
| **`node.passthrough`** | ❌ | the passthrough local exclusion |
| **`device.api`**, **`factory.name`**, **`alsa.driver_name`** | ❌ | `session_device` classification |
Ports and Links are also affected, one materially:
Ports lost `port.exclusive`; Links and Clients were fine — notably
`pipewire.sec.pid` **is** announced, so pulse-PID derivation was reachable.
| object | announced | missing |
| --- | --- | --- |
| Port | `node.id`, `object.serial`, `port.direction`, `port.monitor`, `port.physical`, `port.terminal`, `port.group`, `port.alias`, `port.name`, `port.id`, `audio.channel`, `format.dsp` | **`port.exclusive`** — the `port-exclusive` local exclusion never fires |
| Link | `object.serial`, `link.output.node`, `link.input.node`, `link.output.port`, `link.input.port`, `client.id`, `factory.id` | nothing the engine needs |
| Client | `object.serial`, **`pipewire.sec.pid`**, `application.name`, `module.id`, `pipewire.access`, `pipewire.protocol`, `pipewire.sec.{uid,gid,socket}` | nothing the engine needs |
Demonstrated end to end: a null sink carrying `peerspeak.owned=true` whose
monitor a `module-loopback` re-emitted was reported **eligible** with an **empty
taint set**. In phase 6 that is an echo.
**Links and Clients are fine.** Notably the pulse-PID derivation (v3.4 §6.1.2)
works: `pipewire.sec.pid` is announced. Also notable: the Link endpoint props are
*always* present, which confirms the phase-3 exit-gate worry that the
bind-`LinkInfoRef` fallback is dead code in practice — it is correctness
insurance, never exercised on this host.
The fix became design round 8 (v3.5 §6.7) and phase 3r: bind each Node and read
props off its `info`, exactly how `pw-dump` obtains them. `factory.id` is not a
shortcut (`factory.id=19` resolves to `factory.name = "adapter"`), and
`device.api` is on the *Device* global.
### Demonstrated end to end
A null sink carrying `peerspeak.owned=true`, its monitor read by a
`module-loopback` whose playback leg is a fan-out candidate — the exact shape the
tag exists to exclude:
```
pactl load-module module-null-sink sink_name=ppgate_src \
sink_properties="peerspeak.owned=true"
pactl load-module module-loopback source=ppgate_src.monitor sink=ppgate_dest \
source_output_properties=node.name=ppgate_cap \
sink_input_properties=node.name=ppgate_play
```
Audit verdict:
```json
{"kind":"audit","graph_ready":true,"epoch":"complete","aec_state":"not-configured",
"fan_out_permitted":true,
"candidates":[{"serial":280,"name":"FINAL FANTASY XIV","eligible":true,"sticky":false},
{"serial":309,"name":"ppgate_play","eligible":true,"sticky":false}],
"eligible_count":2,"excluded_count":0,"taint":[]}
```
`ppgate_play` **eligible**, and the `taint` set **empty** — the tagged sink was
not even recognised as a root. In phase 6 this is an echo: peerspeak's own call
playback carries `peerspeak.owned` and would be fanned straight into the share.
The AEC path fails in the other direction. With
`PIXELPASS_AUDIO_AUDIT_AEC=pulse-module:536870918` (a real live module index):
```
aec_state = failed fan_out_permitted = false gate_reason = aec-failed
```
Correct behaviour given its inputs — `pulse.module.id` never arrives, so the
identity can never be observed and the validator times out fail-closed — but it
means **§5.1 row 12 cannot be run as written**, and that with a real AEC
configured phase 6 would refuse to share any audio at all.
### The fix (for round 8)
The full property set *is* reachable: **bind each Node global and read the props
off its `info` event**, which is exactly how `pw-dump` obtains them. Verified on
the same objects that were missing them from the registry:
```
alsa_output.usb-SteelSeries… factory.name = 'api.alsa.pcm.sink'
device.api = 'alsa'
alsa.driver_name = 'snd_usb_audio'
ppgate_src peerspeak.owned = True
pulse.module.id = 536870917
ppgate_play pulse.module.id = 536870918
node.link-group = 'loopback-2528-13'
FINAL FANTASY XIV application.process.id = 14651
```
Two notes for whoever designs that change:
- **The pattern already exists.** Phase 3 built exactly this for Links (bind →
`LinkInfoRef` → `LinkEndpointsResolved`, "the optimisation is the props, the
bind is the correctness path"). Nodes need the same, but as the *only* path
rather than a fallback, and the readiness epoch must hold an obligation per
unbound node — which the model already supports (`withheld` / `pending_links`).
- **`factory.id` is not a shortcut.** The Factory global for `factory.id=19`
(which every ALSA node claims) resolves to `factory.name = "adapter"`, not
`api.alsa.pcm.sink`. The node's own `factory.name` is a different property and
binding is the only way to it.
Also relevant: **`device.api` is announced on the *Device* global** even though it
is absent from the Node. That is the phase-3 review's owed fix ("read the ALSA
driver from the backing Device global, authoritative") — now not merely better
but load-bearing, though `factory.name` and `alsa.driver_name` are absent from
the Device global too, so node binding is still required.
---
## F2 🟠 Machine-wide over-exclusion cascade, downstream of F1
### F2 🟠 Machine-wide over-exclusion cascade, downstream of F1
With F1 in force, `pixelpass_capture_*` (matched on `node.name`, which *is*
announced) is the only taint root that still fires. Running §5.1 row 7 —
a capture sink plus a controlled forwarder reading its monitor:
announced) was the only surviving taint root. Row 7 then excluded every
`Stream/Output/Audio` on the machine: with no strong owner keys, every tainted
capture stream was an **unbounded tainted reader**, tripping phase 2's
fail-closed backstop, while WirePlumber's shared `client.id = 42` fused the
device layer into one owner.
```
candidates:
FINAL FANTASY XIV | eligible: false | reason: unresolved-owner
ppgate7_play | eligible: false | reason: tainted-owner-bridge
taint:
Midi-Bridge | tainted-owner-bridge
bluez_midi.server | tainted-owner-bridge
alsa_output.pci-0000_03_00.1.hdmi-stereo-… | tainted-owner-bridge
alsa_output.usb-SteelSeries_…-analog-stereo | tainted-upstream
alsa_input.usb-SteelSeries_…-mono-fallback | tainted-owner-bridge
alsa_output.pci-0000_10_00.6.analog-stereo | tainted-owner-bridge
alsa_input.pci-0000_10_00.6.analog-stereo | tainted-owner-bridge
FINAL FANTASY XIV | unresolved-owner
ppgate_dest | tainted-upstream
pixelpass_capture_ppgate7 | pixelpass-owned
ppgate7_play | tainted-owner-bridge
ppgate7_cap | tainted-upstream
```
Net live behaviour: exclude everything, always, as soon as pixelpass's own
capture sink existed. Fail-closed, so silence rather than echo — but entirely
non-functional, and non-functional in a way that would have looked like "working
safely" to any test that asserted only exclusions.
Row 7's own assertion held — `ppgate7_play` is excluded via the owner bridge, so
the cycle-prevention mechanism works. But the row **fails the §5.1 exact-partition
requirement**, because the eligible half is empty: FFXIV should have been
eligible and was not.
### What run 1's machinery got right
The mechanism: with `node.link-group`, `application.process.id` and
`pulse.module.id` all absent, no node has a *strong* owner key — `client.id` is
explicitly not one (v3.4 §6.1.3). So every tainted capture stream is an
**unbounded tainted reader**, which trips phase 2's documented fail-closed
backstop (`taint/mod.rs`, `an_unbounded_tainted_reader_excludes_every_output`)
and excludes every `Stream/Output/Audio` on the machine. Every device node
separately keeps its coarse keys (`session_device` is universally false, also from
F1) and they all share WirePlumber's `client.id = 42`, which fuses them into a
single owner and spreads the taint across the whole device layer.
So the engine's *net* live behaviour today is: exclude everything, always, as soon
as pixelpass's own capture sink exists. Fail-closed, so silence rather than echo —
but the feature is entirely non-functional, and it is non-functional in a way that
would have looked like "working safely" to any test that only asserted exclusions.
**This is the §5.1 argument vindicated in the most direct possible way.** The
current build *is* the degenerate exclude-everything implementation the plan
warned about, and it is the eligible half of the partition — asserted, per §5.1 —
that caught it. An exclusion-only checklist would have passed this build.
---
## §5.2 — O5 measurements
Recorded under deliberate churn: five load/unload cycles of
`module-null-sink` + `module-loopback`, 6.5 s wall.
```json
{"kind":"metrics","graph_events":308,"tick_events":26,"emitted_records":308,
"span_us":6499634,"graph_events_per_sec":47.39,
"recompute_max_us":15,"recompute_mean_us":4,
"recompute_p50":"<50us","recompute_p90":"<50us","recompute_p99":"<50us",
"recompute_distribution":[["<50us",334]],
"emit_max_us":12,"emit_mean_us":2,"emit_distribution":[["<50us",308]],
"busy_us":2331,"busy_fraction":0.0004,
"queued_events":198,"queue_threshold_us":100}
```
**O5 is closed: full recompute per graph event has roughly four orders of
magnitude of headroom.** Every one of 334 recomputes finished in under 50 µs, the
worst at 15 µs, against a 47 Hz event rate under churn far heavier than a desktop
produces at rest. The observer thread spent 0.04 % of wall time working.
`queued_events: 198` looks alarming and is not: PipeWire delivers enumeration and
teardown as back-to-back bursts, so most events do begin within 100 µs of the
previous one completing. With a 15 µs worst-case recompute the backlog drains
faster than it forms. `busy_fraction` is the number to trust here — it needs no
inference, and it is 0.0004.
**Caveat, and it is a real one.** These numbers were measured on the *degraded*
graph F1 produces. The recompute cost is over the same node and link count so the
taint-engine figure is representative, but the F1 fix adds a bind and an `info`
round-trip **per node**, which is new I/O this run did not measure at all. O5
should be re-measured after round 8 rather than inherited from here.
---
## Matrix status (§5.1)
| # | scenario | status |
| --- | --- | --- |
| 1 | null-sink + loopback forwarder, owner bridge | ⛔ blocked by F1 — needs a taint root (`peerspeak.owned`) |
| 1b | Sunshine's topology (opportunistic, non-gating) | not attempted |
| 2 | gst split clients, tainted input | ⛔ blocked by F1 — needs `application.process.id` |
| 3 | two Pulse modules, one tainted | ⛔ blocked by F1 |
| 4–6 | peerspeak playback / mpv / notification | ⛔ blocked by F1 — all three are `peerspeak.owned` tags |
| 7 | second host's capture sink + forwarder | 🟠 mechanism verified, **partition fails** (F2) |
| 8 | EasyEffects | ⛔ blocked by F1 — needs `node.link-group` |
| 9 | Firefox three cases | ⛔ blocked by F2 (everything excluded) |
| 10 | sticky taint across teardown | ⛔ blocked by F1 |
| 11 | recycled serial / index / link-group | ⛔ blocked by F1 |
| 12 | AEC loaded → unloaded → Revoked | ⛔ blocked by F1 — `pulse.module.id` never arrives; validator goes `failed` |
| 13 | `Audio/Duplex` device | not attempted (none present on this host) |
**No row can be completed until F1 is fixed.** The matrix is not re-runnable in a
meaningful sense before then — every row's eligible half is empty for the same
reason.
---
## What the audit machinery got right
Worth recording, because none of it needs revisiting in round 8:
None of this needed revisiting:
- Running the recompute **inline on the observer thread**, once per applied
registry event, upholds phase 4's no-coalescing contract and put the cost
exactly where O5 could measure it.
- The **complete-partition record** is what caught F2. A record of only the
interesting nodes would have shown row 7 passing.
- **Reason codes survived the trip** and were immediately diagnostic:
`unresolved-owner` on FFXIV named the backstop, not a symptom, and pointed
straight at the missing strong keys.
- The **`peerspeak.owned` / `pulse.module.id` fixtures were right** — phase 2's
engine does the correct thing when handed correct properties. The defect is
entirely at the observation boundary, which is where phase 5 was designed to
look.
## Next
1. **Design round 8** on F1: node binding in the observer, readiness obligations
per unbound node, and where `session_device` reads its inputs from.
2. Re-run this matrix in full afterwards. Rows 4–6 additionally need peerspeak
running; rows 8, 9 and 1b need EasyEffects, Firefox and Sunshine respectively.
3. Re-measure O5 with node binding in place.
registry event, upheld phase 4's no-coalescing contract and put the cost where
O5 could measure it.
- The **complete-partition record** is what caught F2 — and, in run 2, F13-1.
- **Reason codes survived the trip** and were immediately diagnostic.
- The **`peerspeak.owned` / `pulse.module.id` fixtures were right**: the engine
does the correct thing when handed correct properties. Both failures were at
the observation boundary, which is where phase 5 was designed to look.
+275 -37
View File
@@ -1,11 +1,24 @@
# Design v3: whole-desktop screen-share audio without self-echo
# Design v3.8: whole-desktop screen-share audio without self-echo
**Status:** 🟠 **v3.6 — round 9, opened by a second MEASURED finding, this time from a live
audit run of the *fixed* observer.** v3.4's architecture is still unchanged and converged.
Round 8 revised the **observation boundary** (§6.7); round 9 revises what stickiness is
allowed to remember (new §6.8). Both were found by running code, not by reading it.
**Date:** 2026-07-25 (v1: 07-19 · v2: 07-20 · Option C 07-20 · v3.1 r4 · v3.2 r5 · v3.3 r6 ·
v3.4 r7 · v3.5 r8 · v3.6 r9)
**Status:** 🟢 **v3.8 — round 11: both pre-Phase-6 design gates are closed.** The graph now
models an unobservable playback-to-capture route as a conservative same-`device.id` hardware
edge (§6.9), and the 2 s readiness budget is retained after repeated baseline, inflated-graph,
and live-churn calibration. The targeted live dry-run partition passed. The first pure Phase 6
channel-planning prerequisite is committed locally in PixelPass `781defc`; the first bounded
live fan-out mutation slice is built, validated and committed locally in PixelPass `98cde2c`.
The four versioned causal status events are built, validated and committed locally in PixelPass
`5a65f50`. Capture-sink replacement and successful all-channel relinking are built, validated
and committed locally in PixelPass `d09ee9b`. The Row 1 mutation-edge identity gate and Row
8d/8e fail-closed construction gates are built, validated and committed locally in PixelPass
`6be07ef`. The independent Row 6 refusal gates (`956534f`) and revised Row 9 partition
(`7b11827`) complete the deterministic link-manager matrix. The production-path three-arm AEC
leak qualification is built, validated and committed locally in PixelPass `e027bc6`; **Phase 6
is complete.** The public selector, AEC wiring and versioned capability response are built,
validated and committed locally in PixelPass `792f2bd`; **Phase 7 is complete.** PeerSpeak
`2c2b861` completes the capability-bound picker/argv and causal UI status integration;
**Phase 8 is complete and Phase 9 is now the ship gate.**
**Date:** 2026-08-21 (v1: 07-19 · v2: 07-20 · Option C 07-20 · v3.1 r4 · v3.2 r5 · v3.3 r6 ·
v3.4 r7 · v3.5 r8 · v3.6 r9 · v3.7 r10 · v3.8 r11 2026-08-21)
**Origin:** Joe's suggestion — "whitelist all audio except audio coming from peerspeak."
**Scope:** a new capture mode in pixelpass (`src/host/pipeline.rs`, `src/host/audio.rs`),
playback tagging + AEC-identity export + teardown-ordering invariants in peerspeak.
@@ -26,6 +39,7 @@ v1/v2 remain in git history at `88ad5a0` and `10203e1`.
| fan-out spike | `~/Documents/handoff-docs/Claude/peerspeak/fanout-spike-results-2026-07-20.md` + Codex rounds 3/4 | **Option C adopted**, ratified |
| AEC identity gate | `~/Documents/handoff-docs/Claude/peerspeak/aec-playback-leg-identity-2026-07-20.md` | **🟢 gate passed**, both models agree after 2 adversarial rounds |
| **phase 5 dry-run gate (r8)** | `docs/screenshare-audio-exclusion-phase5-results.md` | **🚦 GATE FAILED** — the observation boundary is wrong (§6.7); architecture unaffected |
| **pre-Phase-6 closure (r11)** | `docs/screenshare-audio-exclusion-phase5-results.md` addendum | **PASSED** — same-device bridge, different-device negative control, and readiness calibration (§6.9) |
---
@@ -502,11 +516,47 @@ derive this itself; it cannot assume a value, and peerspeak can supply only a *h
- Read `pipewire.sec.pid` from the **Client** objects of Pulse-emulated streams. Measured:
it is `2541` for Firefox, Steam, KDE Connect, sunshine and libcanberra alike, while each
node's own `application.process.id` differs (Firefox `11114`, sunshine `4119`).
- Require a **single consistent** value across those clients, and validate it by reading
- ~~Require a **single consistent** value across those clients~~ and validate it by reading
`/proc/<pid>/comm` (or cmdline) and confirming it is `pipewire-pulse`.
- "This PID owns implausibly many unrelated streams" is a **diagnostic**, never
correctness logic.
> #### 🔴 Round 10 (MEASURED, phase-5 run 2): "a repeated `sec_pid`" does not identify pulse
>
> The struck rule above was implemented as *the single `sec_pid` shared by two or more
> Clients*, on the reasoning that native clients each carry their own distinct PID so only the
> Pulse shim repeats a value. **Measured on this host: WirePlumber repeats one too** — it holds
> two Clients, `WirePlumber` and `WirePlumber [export]`, both `sec_pid` 1747. Two values
> repeated, "single consistent" was unsatisfiable, and the derivation returned `None`
> **permanently, on a stock desktop**.
>
> The consequence was not a missing optimisation. With the daemon PID unknown the key-4
> exception never fires, every Pulse-emulated node fuses into one owner, and the result is the
> machine-wide over-exclusion cascade of phase 5's F2 — reached again from a new cause, and
> caught again only by the §5.1 requirement to assert the **eligible** half of a row.
>
> The rule failed in both directions, so the repetition test is **deleted** rather than
> tightened:
>
> - **False ambiguity** — any second process holding two Clients defeats it, and WirePlumber
> always does.
> - **False absence** — a session in which pipewire-pulse holds exactly one Client (one Pulse
> app running) repeats nothing at all, so the candidate is never even considered.
>
> `comm` was always the authoritative check; repetition was a heuristic standing in front of it,
> and what it actually encoded was an assumption about *other* processes' Client counts.
> **The rule is now: every distinct `pipewire.sec.pid` is a candidate; the daemon is the unique
> one whose `/proc/<pid>/comm` is exactly `pipewire-pulse`.** Zero matches ⇒ `None` (nothing we
> can prove to suppress). **Several** matches ⇒ also `None`: two live pipewire-pulse daemons (a
> nested or sandboxed session) cannot both be suppressed by a single `Option<u32>`, and failing
> closed there lands on the over-exclusion side, consistent with the failure-mode paragraph
> below. Suppressing a *set* of daemon PIDs is the real answer if a multi-daemon host ever turns
> up; it is out of v1 and recorded rather than silently approximated.
>
> The lesson generalises past this key: **a property of the objects we are trying to identify is
> evidence; a property of everyone else's object count is a guess.** The `/proc` read was already
> there and already authoritative — the heuristic in front of it only added a way to be wrong.
Failure modes: if pixelpass fails to identify the real pipewire-pulse PID, the result is
broad **over-exclusion** (annoying, safe). If it wrongly suppresses a genuine app PID, the
result is over-exclusion **for that app** — safe *only* because unresolved ancestry is
@@ -542,6 +592,30 @@ Where an owner has a tainted input leg and an output leg with **no** resolvable
(§6.1.2), fail closed and exclude the output leg. This only ever engages for owners
actually reading a tainted monitor, so the blast radius is small.
> **⚠️ Round 13 (F11-1) — "bounded" is not "has a key". A self-claimed pid is not
> provenance.** Which legs this backstop sweeps depends on whether the tainted reader and
> the candidate outputs are *bounded* — i.e. whether we could enumerate their sibling legs
> and be right. Key 4 is a union of the node's `application.process.id` (client-controlled,
> optional) and its Client's `pipewire.sec.pid` (protected), so a node could bound itself
> with a value it invented and escape the sweep while its real sibling was unfindable.
>
> **Rule:** a strong key (`node.link-group`, `pulse.module.id`) bounds an owner on its own;
> key 4 bounds an owner **only when the node's Client resolves** — an unambiguous Client
> yielding `Some(pipewire.sec.pid)`, read *before* pipewire-pulse suppression. The
> ordering is load-bearing in both directions: read after suppression and every
> Pulse-emulated app on the box goes unbounded (§6.1.1 catastrophe, new door); accept "a
> unique Client object exists" instead of a `sec_pid` and a pid-less Client leaves the hole
> open.
>
> **Bridging is unchanged** — it still uses the full union, because a self-claimed pid is
> perfectly good *evidence that two legs are related*, which is the taint-increasing
> direction. Only the permission to declare a differently-keyed output "provably someone
> else" now demands a `pipewire.*` answer to "who is this".
>
> Measured cost on this host: **zero** — before/after binaries audited the same live graph
> simultaneously, same 14 decision states, no `unresolved-owner` on either side, eligible
> half non-empty. Implemented in pixelpass `c78eb2d`; five-case Client matrix in the tests.
### 6.1.3 ⚠️ Taint must be STICKY — current topology is not enough
**Conceded to Codex in round 6; my "conditional bridge" was correct about topology and
@@ -844,21 +918,74 @@ through a teardown, and it is unaffected by this change.
Cost: two fixpoints per graph event. Measured 80 µs worst case against a 47 Hz event rate,
so the O5 headroom absorbs it without argument.
⚠️ **Owed, from the round-9 review (Codex, P1 "worth checking"): hardware
playback-to-capture paths.** A card offering "Stereo Mix" / "Digital Loopback" presents an
⚠️ **Round-9 open item — CLOSED in round 11 (§6.9): hardware playback-to-capture
paths.** A card offering "Stereo Mix" / "Digital Loopback" presents an
ordinary driver name (`snd_hda_intel`), so both its sink and its source classify
`session_device` — and audio written to the sink reappears on the source through a hop the
Link graph cannot see. This is the `snd_aloop` hazard (§6.1.1, phase-3 review finding 2) in
a form the driver denylist cannot detect. It is **not new in round 9** and not introduced by
either recent round; distinguishing it needs ALSA control inspection, a new I/O surface and
therefore a design decision. Until then a card with that path enabled can carry the call
from sink to source untainted, and a capture app reading it can re-emit: **echo**.
either recent round. Round 11 closes it without relying on a driver denylist or control name.
⚠️ **Also owed: a calibration argument for the readiness budget.** The observer times out
⚠️ **Round-9 open item — CLOSED in round 11 (§6.9): calibration of the readiness
budget.** The observer times out
after 2 s and `TimedOut` is sticky by design, so a process that never sees one
obligation-free instant during initial enumeration is silent for its lifetime. Measured on
this host: readiness at ~3 ms with 19 binds. The margin is three orders of magnitude, which
is an argument, but it is one measurement on one idle desktop.
obligation-free instant during initial enumeration is silent for its lifetime. The original
~3 ms observation on one idle desktop was not enough; round 11 adds repeated starts, an
inflated graph, and concurrent graph churn.
### 6.9 🟢 Same-device hardware bridge + readiness calibration (round 11, MEASURED)
**The hardware rule.** For every positively classified passive hardware terminal, retain the
Node's snapshot-local `device.id`. The taint walk adds a directed synthetic edge from an
`Audio/Sink` (or output side of `Audio/Duplex`) to every `Audio/Source` (or input side of
`Audio/Duplex`) carrying the **same** `device.id`:
```
tainted stream -> hardware sink ~[private mixer / firmware]~> same-device source -> reader
```
Both predicates are load-bearing. `session_device=true` limits the rule to the observer's
positive passive-hardware allowlist; `device.id` limits it to one physical Device instead of
fusing every card exported by WirePlumber. The id never enters sticky identity and never
survives its snapshot.
**Why this is unconditional rather than an ALSA-control probe.** Measured 2026-08-21: the
ALC897 exposes `Loopback Mixing` (disabled) and two `Input Source` controls containing Rear
Mic, Front Mic and Line, but no `Stereo Mix`. Linux HDA treats analog loopback monitoring and
the optional `Stereo Mix` capture source as distinct mechanisms (kernel
[`hda_generic.c`](https://code.googlesource.com/linux/torvalds/linux/+/master/sound/pci/hda/hda_generic.c)
and [HDA control documentation](https://cdn.kernel.org/doc/html/latest/sound/hd-audio/controls.html)).
More importantly, ALSA/HDA
control spelling cannot prove the absence of USB, vendor-DSP or firmware loopback paths. A
control-name allowlist would therefore be precise on this card and unsound as a portable
absence proof. The graph rule closes every such hidden same-device hop without a new runtime
ALSA dependency.
**Accepted cost.** An app capturing a microphone from the same Device that is receiving
tainted playback is excluded even when that particular microphone path is clean. A source on
a different Device remains eligible. This is deliberate fail-closed over-exclusion, pinned by
pure exact-partition tests and a live dry-run negative control.
**Live partition, 2026-08-21.** Tagged playback was routed to the ALC897 sink. A controlled
reader/re-emitter on the ALC897 source was excluded `tainted-owner-bridge`; the identical
reader/re-emitter on the Arctis source remained eligible. The audit was graph-ready, all
temporary modules were unloaded by exact module id, and the configured default sink/source
were unchanged.
**Readiness budget — keep 2 s for v1.** Fresh observer startup measurements on the same live
desktop, using the observer's own monotonic `at_ms` clock:
| arm | runs | p50 | p95 | max | timeout |
| --- | ---: | ---: | ---: | ---: | ---: |
| ambient graph | 30 | 5 ms | 10 ms | 11 ms | 0 |
| inflated graph: 24 null sinks + 24 loopbacks | 30 | 111 ms | 113 ms | 114 ms | 0 |
| 250 create/remove cycles concurrent with fresh starts | 20 | 5 ms | 6 ms | 8 ms | 0 |
The deliberately inflated maximum leaves 17.5x headroom to the sticky 2 s deadline. This is
not a universal latency promise; it is a calibration argument that exercises enumeration
volume and graph instability, rather than extrapolating from one idle start. Revisit the
budget if a supported target measures startup p95 above 500 ms or produces a real timeout;
do not weaken `TimedOut`'s fail-closed/sticky semantics to hide one.
## 7. Lifecycle and teardown invariants
@@ -1143,9 +1270,34 @@ Both reviewers agree on all seven. Recorded as decided; reopen only with new evi
- **D7 — no materially simpler design exists** that still meets Joe's ask. The available
simplification is to *narrow v1 scope*, not to change architecture. ✅
## 14. Readiness — 🟠 v3.6 (round 9): architecture converged; observation boundary and sticky provenance REVISED.
## 14. Readiness — 🟢 v3.7 (round 10): the §5.1 matrix PASSED; architecture unchanged.
**Round 9 (2026-07-25, same day).** Phase 3r shipped §6.7 and the audit was re-run
**Round 10 (2026-07-26).** The phase-5 matrix ran in full and **passed all 13 rows with a
non-empty eligible half in every one** — results in
`screenshare-audio-exclusion-phase5-results.md`. It also found a third measured defect on
first contact, and once again the failure was fail-closed and *silent*, exposed only by the
requirement to assert what must remain **eligible**: §6.1.2's pulse-PID derivation returned
`None` permanently on this host, so key 4 fused every Pulse-emulated node into one owner.
| | verdict |
| --- | --- |
| Architecture — Option C, taint as a graph property, owner-key union, sticky taint, AEC identity state machine | **unchanged, three times vindicated** |
| §6.1.2 | **revised** — the derivation heuristic is deleted; `comm` alone decides |
| §5.1 matrix | **PASSED** — 13/13, incl. the full sticky lifecycle (row 10) and provable identifier recycling (row 11) |
| O5 | **closed on the real graph** — worst recompute 67 µs, churn mean 10 µs, busy fraction 0.0006 |
| Phase 5 | **machinery unchanged and correct** — three real defects caught on first contact with the live graph, none of them in the engine |
| Phase 6 | **still blocked** — by F11-1, phases 0b/0c/0d, and the "Stereo Mix" design call, *not* by this matrix. (F11-1 was **closed later the same day** with this matrix's data — §6.1.2's round-13 box; the rest stand.) |
Three rows passed with recorded substitutions (8 EasyEffects, 9 Firefox's own mic/monitor
paths, 13 a real `Audio/Duplex` device) and the third-party samples stay owed.
**Round-11 supersession (2026-08-21):** the row above is the round-10 snapshot, not current
status. F11-1 and phases 0b/0c/0d are satisfied, and §6.9 closes the hardware-path decision
plus readiness calibration. S4/S5/0d, round 11, and Phase 6's first pure channel planner are
committed locally in PixelPass `781defc`; the first live fan-out mutation is committed locally
in `98cde2c`.
**Round 9 (2026-07-25).** Phase 3r shipped §6.7 and the audit was re-run
immediately; it found a *second* measured defect within minutes — a permanent sticky taint
on a hardware sink (§6.8). Both rounds share a shape worth naming: **the architecture was
right and the instrumentation was wrong**, and only running the code against a live daemon
@@ -1204,15 +1356,19 @@ How the blockers closed:
| **9** | **a fail-closed unresolved mark became permanent sticky taint (measured, phase 5 re-run)** | **fixed** — §6.8 evidence-only sticky pass |
| **9** | `device_props` tested for one live *Device* rather than one live *global* on the id (Codex, certain) | **fixed** in phase 3r — stale `session_device` on a contested id is an echo path |
| **9** | `device.api` corroborated by presence, so `v4l2` under an ALSA factory passed (Codex) | **fixed** in phase 3r — the API must equal the allowlist's own |
| **9** | hardware playback-to-capture ("Stereo Mix") defeats the `session_device` classifier (Codex, P1 worth checking) | **OPEN — design decision owed**, §6.8; pre-existing, needs ALSA control inspection |
| **9** | the 2 s readiness budget has no calibration argument (Codex) | **OPEN — measurement owed**, §6.8; ~3 ms observed on this host |
| **9 → 11** | hardware playback-to-capture ("Stereo Mix") defeats the `session_device` classifier (Codex, P1 worth checking) | **CLOSED — §6.9.** Conservative `Sink → Source` edge for passive terminals sharing `device.id`; pure and targeted live exact partitions pass. No runtime control-name guess. |
| **9 → 11** | the 2 s readiness budget has no calibration argument (Codex) | **CLOSED — §6.9.** 30 baseline starts, 30 starts with 48 temporary modules, and 20 starts during 250 create/remove cycles; inflated max 114 ms, zero timeouts. Keep 2 s. |
| **10** | **the pulse-PID derivation required a *single* repeated `sec_pid`; WirePlumber repeats one too, so it returned `None` permanently and key 4's suppression never fired (measured, phase-5 run 2)** | **fixed** — §6.1.2 round-10 box: probe every distinct `sec_pid`, let `/proc/<pid>/comm` decide |
| **10** | the audit's `sticky` flag means "is in the remembered set", so it is true for nearly every tainted node and does not answer "excluded only because remembered" | **OPEN — reporting only**; the evidence-only pass §6.8 already computes what is needed |
| **10** | a bridge's named key is lost when a leg reappears under a new serial (sticky `reason_for` falls back to keyless, and `raise` will not replace a same-rank reason) | **OPEN — reporting only**; verdict unaffected |
### v1 scope — agreed
Option C fan-out · explicit `--aec=off|pulse-module:<idx>` · peerspeak playback and child
tagging · exact AEC module validation · graph taint with the owner-key union and the
pipewire-pulse PID exception · sticky taint · readiness epoch · fail-closed unresolved
ancestry · owned non-lingering links · §10 items 1, 4 and 5 landed first.
ancestry · conservative same-device hardware bridge · owned non-lingering links · §10 items
1, 4 and 5 landed first.
### Deliberately OUT of v1
@@ -1226,21 +1382,103 @@ binding, so `port.exclusive` is never observed and the §6.2 row it guards relie
create failing cleanly · **(r8)** no serial-continuity signal for the AEC validator's
no-coalescing contract.
### Next step (round 9)
### Current front after round 11
Phases 0a, 2, 3, 3r, 4 and 5 are built; §6.7 and §6.8 are implemented and merged. What
remains before phase 6 unblocks:
Phases 0a–5 and 3r are built; the full phase-5 matrix passed; 0b is merged; and 0c/0d plus
the S4/S5 ownership work are built, validated, and committed locally. Round 11 closes the last
two pre-Phase-6 design decisions with pure, observer, live PipeWire and targeted dry-run
evidence. Phase 6's channel-aware pure planner is also committed in PixelPass `781defc`. On top
of that checkpoint, local commit `98cde2c` creates and retains serial-guarded
non-lingering links through the hidden production path, requires all planned links to be
`ACTIVE`, revokes them when eligibility changes, and passed both normal-teardown and `SIGKILL`
cleanup gates. Local commit `5a65f50` adds exact version-1 status records for stream link failure,
AEC validation failure, AEC revocation and foreign AEC detection; PipeWire callbacks enqueue them
for Tokio-side JSON emission. Local commit `d09ee9b` recreates an unexpectedly removed
connection-owned sink with a fresh serial, hands that exact identity to the fan-out observer,
drops stale link proxies, and returns every replacement channel to `ACTIVE`. The pure recovery
gate, an owner-only live sink-destruction gate, and the full hidden
host-path sink-destruction/relink gate pass; the full suite reports 332 passed and 12 ignored, all
three serialized Phase 6 live gates pass together, and strict Clippy plus `pixelpass --doctor`
remain clean.
1. ~~**Revise phase 3** to §6.7~~ — **done**, phase 3r merged, four-part gate passed
including the live prop-recovery row and an added live gate for the Device-side path.
2. **Revise phase 1** to emit both carriers (§5.1), literals pinned in plan §3. Unblocked
and next.
3. **Re-run the whole phase-5 §5.1 matrix** — no row was completable under the round-8
defect, so nothing carries over — and **re-measure O5** with bind I/O *and* the round-9
second fixpoint in it. Phase 6 stays blocked until that results file passes.
4. Decide the two items §6.8 leaves open: hardware playback-to-capture paths (a real echo
path, needs a design call) and the readiness-budget calibration.
PixelPass commit `6be07ef` closes Row 1 and Row 8d/8e. Mutation-edge serial
revalidation now covers retained proxies as well as new links, and a recycled identity produces
zero unsafe creates. Two optimized-release fault-injection gates prove capture-sink construction
failure and readiness timeout both fail `DesktopExcluding` closed without resolving the legacy
default monitor or constructing legacy routing. Full validation reports 335 passed and 12
ignored; strict Clippy, all three serialized live Phase 6 gates, diagnostics and residue checks
pass.
Still owed beyond that, unchanged: the §9.2 rig upgrade before any exclusion claim is
published, and **field-test §12** — nothing in this design has been tested over the real
GStreamer/AAC/network path or on two machines.
PixelPass commit `956534f` closes the three independent Row 6a/6b/6c refusal/reason fixtures.
`Stream/Output/Audio` Nodes advertising a readable configured Format
param are subscribed and parsed through libspa into raw, encoded or IEC958; unknown format is
per-stream fail-closed, while the existing `node.passthrough` property remains an independent
predicate. Separate controller gates prove `port.exclusive`, encoded and IEC958 each make zero
link-create calls and emit only `port-exclusive`, `encoded` and `iec958-passthrough`
respectively. An additional gate proves an explicit passthrough property cannot be masked by a
raw Format, and the initial Node-info-before-Format ordering does not emit a false warning.
A read-only live audit initially caught generic Pod parsing failures and harmful Format queries
against Nodes that did not advertise the param. After switching to libspa's native parser and
gating queries on readable Format support, Strawberry, FFXIV and Chromium all settled to raw
eligible streams with no parse/core errors. Full validation is 342 passed and 12 ignored; the
three Row 6 cases also pass optimized release, strict Clippy and `pixelpass --doctor` are clean.
The `port.exclusive` case remains an injected-graph gate under v1's accepted no-Port-binding
limitation; a live exclusive port still degrades through the separately tested link-failure path.
PixelPass commit `7b11827` carries the revised Row 9 partition through the real fan-out
controller. Four independent late-arrival gates prove music-only and a different-device
microphone each create both stereo links and reach `Captured`, while same-device capture and a
tainted-monitor capture retain `tainted-owner-bridge` and make zero link calls. Full validation
is 345 passed and 12 ignored; strict Clippy, all three serialized live Phase 6 mutation gates,
diagnostics and residue checks pass. This completes the deterministic link-manager matrix, not
Phase 6's separate live signal qualification.
PixelPass commit `e027bc6` completes Phase 6's separate live signal qualification. Its
deliberately unsafe predicate is wholly `#[cfg(test)]` and re-admits only the exact configured
`aec-identity` candidate. The live test drives the hidden selector through the real sink,
observer, taint controller, native link manager and Pulse monitor recording path; filters
`media.class` before the exact module ID; preserves child stderr; and verifies the intended
links in PipeWire before recording with `parec`.
Two complete runs kept the guarded 1500 Hz result at or below the run's control floor within the
declared 3 dB tolerance, while the naive arm captured 1500 Hz at desktop level and at least 18 dB
above control and guarded. The supported result is **no incremental 1500 Hz energy detectable
above the control floor at this analysis resolution**, not proof of absence. Full validation is
346 passed and 13 ignored; strict Clippy, all four serialized Phase 6 live audio gates,
`pixelpass --doctor` and residue checks pass.
PixelPass commit `792f2bd` completes Phase 7. The selected public contract is
`--audio-mode=desktop-shared|desktop-excluding`; excluding requires the explicit
`--aec=off|pulse-module:<idx>` state and carries it into the production fan-out controller.
`pixelpass --capabilities` emits schema version 1 with independent `strict_app_audio` and
`desktop_audio_exclusion` booleans. Exact old-PeerSpeak argv remains legacy and byte-compatible.
Phase-7 validation also corrected the Phase-6 estimator without weakening its declared margins.
Several extra runs showed that one sub-LSB coherent control bin can fall into a stochastic null.
The gate now resolves its control floor from that bin plus the 90th percentile of neighboring
frequencies outside the Hann main lobe; the 3 dB guarded tolerance and 18 dB naive margin are
unchanged. Final validation is 354 passed and 13 ignored; strict Clippy, all four serialized
live gates, exact capability/help probes, `pixelpass --doctor` and residue checks pass.
PeerSpeak commit `2c2b861` completes Phase 8. A typed picker/core selection keeps legacy desktop,
desktop-excluding and per-app capture distinct. The versioned capability result is bound to its
resolved PixelPass path and re-probed on path changes before capability-gated argv is built. The
new picker row emits both `--audio-mode=desktop-excluding` and explicit
`--aec=off|pulse-module:<idx>`; old PixelPass keeps the row absent and receives no new flags.
Legacy desktop and per-app argv remain byte-identical.
The four Phase-6 status values now cross the real host-notice channel into the UI as the same
parsed type and drive a persistent, mode-scoped explanation beside the live sharing badge. Final
serialized validation passes 649 unit tests (7 live-only ignored), 20 integration tests (4
live-only screen-share gates ignored), strict all-target Clippy, formatting and diff checks.
**Phase 9 is current.** PixelPass `98bb78f` implements the two-probe, windowed per-channel
normalized-correlation rig with involved-node xrun telemetry, fixed pre-run thresholds, route
integrity checks, and the control/guarded/naive topology; live-discovered harness hardening is
committed as `360d711`. The controlled live run passed on 2026-08-22: excluded-probe maxima were
0.1356 control / 0.1340 guarded against a 0.20 limit; the naive positive-control minimum was
0.5741 and eligible desktop minimum was 0.6429 against a 0.25 floor; every arm added zero xruns.
The full non-live suite, strict Clippy, `pixelpass --doctor`, route checks, and final residue scan
also pass. **Field-test §12 remains owed** — nothing in this design has yet been tested over the
real GStreamer/AAC/network path or on two machines. No exclusion claim may be published yet.
Generated
+48
View File
@@ -0,0 +1,48 @@
{
"nodes": {
"nixpkgs": {
"locked": {
"lastModified": 1785989512,
"narHash": "sha256-HFQhkQcl5D1hUNoen3SGHCSFCt2Bg6uP+HgbrnA3InQ=",
"owner": "nixos",
"repo": "nixpkgs",
"rev": "445d861c6d31b4af0c79d8d4be2331f762a361d7",
"type": "github"
},
"original": {
"owner": "nixos",
"ref": "nixos-26.05",
"repo": "nixpkgs",
"type": "github"
}
},
"root": {
"inputs": {
"nixpkgs": "nixpkgs",
"rust-overlay": "rust-overlay"
}
},
"rust-overlay": {
"inputs": {
"nixpkgs": [
"nixpkgs"
]
},
"locked": {
"lastModified": 1786076960,
"narHash": "sha256-jfR6OhwurCKn1tREyfOcK/Omxf1Q/DzDDFbnEr1mBLs=",
"owner": "oxalica",
"repo": "rust-overlay",
"rev": "57a23bfaf4f7017267294b161175db1e32eb1c85",
"type": "github"
},
"original": {
"owner": "oxalica",
"repo": "rust-overlay",
"type": "github"
}
}
},
"root": "root",
"version": 7
}
+171
View File
@@ -0,0 +1,171 @@
{
description = "PeerSpeak — decentralized P2P voice chat (Rust/iroh/PipeWire/Opus/iced)";
inputs = {
# Pinned to the same channel the hosts run (nixos-config tracks
# nixos-26.05), so the libraries this shell links and dlopens are built
# against the same release as the PipeWire daemon and Vulkan ICD actually
# running on the machine. Floating to unstable here would reintroduce
# precisely the client/server version skew the pin exists to prevent.
nixpkgs.url = "github:nixos/nixpkgs/nixos-26.05";
# The Rust toolchain is pinned SEPARATELY from the system libraries, and
# deliberately so. nixpkgs 26.05 ships rustc 1.95.0, but this crate was
# developed and verified against 1.97.1 — close enough to build and pass
# every test, but not close enough for clippy, which flags a
# `collapsible_match` on 1.95 that 1.97 does not. Taking the compiler from
# here decouples "which Rust the project targets" from "which release the
# audio stack came from", so a nixpkgs bump can never silently move the
# compiler under the lint gate again.
#
# This is the reproducible alternative to rustup: same exact-version
# control, but the choice is recorded in flake.lock, so darp5 or a fresh
# clone resolves the identical toolchain instead of whatever rustup happens
# to fetch that day.
rust-overlay = {
url = "github:oxalica/rust-overlay";
inputs.nixpkgs.follows = "nixpkgs";
};
};
outputs =
{ nixpkgs, rust-overlay, ... }:
let
system = "x86_64-linux";
pkgs = import nixpkgs {
inherit system;
overlays = [ rust-overlay.overlays.default ];
};
# Matches what CachyOS shipped (rust 1:1.97.1-1), which is the toolchain
# every green result in the handoff was produced with.
#
# `default` is the rustup "default" profile — rustc, cargo, rust-std,
# rustfmt and clippy — so those are NOT listed separately below.
#
# rust-src and the windows-gnu target exist for win-cross-build.sh, which
# needs `-Z build-std=std,panic_abort` for the self-contained .exe. That
# script still expects to run in the peerspeak-win distrobox for the
# mingw toolchain; carrying the target here just means the Rust half is
# already in place if it is ever driven from the host.
rustToolchain = pkgs.rust-bin.stable."1.97.1".default.override {
extensions = [ "rust-src" ];
targets = [ "x86_64-pc-windows-gnu" ];
};
# Libraries that iced/winit/wgpu open with dlopen at RUNTIME rather than
# linking at build time. Because nothing links them, they never land in
# the binary's rpath — under `cargo run` the loader finds them only
# through LD_LIBRARY_PATH. Leaving them out builds fine and then panics
# at window creation, which is a genuinely confusing failure, so they are
# listed explicitly instead of discovered the hard way.
runtimeLibs = with pkgs; [
vulkan-loader # wgpu's Vulkan backend (iced's renderer)
libxkbcommon # winit keyboard handling
wayland # wayland-sys, dlopen'd on a Wayland session
libx11 # x11-dl, dlopen'd on the X11 fallback path
libxcursor
libxrandr
libxi
];
# Screen sharing spawns pixelpass as a CHILD PROCESS, and pixelpass in
# turn drives GStreamer as a subprocess. That makes these tools a
# dependency of peerspeak's own test suite, not just of pixelpass:
# `tests/screenshare_host_fault.rs` starts a real pixelpass host, which
# aborts at its preflight if gst-launch-1.0 is missing.
#
# Deliberately duplicated from pixelpass's flake rather than importing it
# as an input. The two projects are mutually optional by design — neither
# is a dependency of the other, and the coupling is a runtime subprocess
# contract. Making one flake consume the other would quietly reintroduce
# exactly the build-level dependency that rule exists to prevent.
screenshareTools = with pkgs; [
gst_all_1.gstreamer
gst_all_1.gst-plugins-base
gst_all_1.gst-plugins-good
gst_all_1.gst-plugins-bad
gst_all_1.gst-plugins-ugly
gst_all_1.gst-libav
pipewire # pipewiresrc (Wayland capture; ships in this pkg)
];
in
{
devShells.${system}.default = pkgs.mkShell {
nativeBuildInputs = [
rustToolchain
]
++ (with pkgs; [
# The supply-chain gates .gitea/workflows/ci.yml runs, so the same
# checks are reproducible locally before a push. These were `cargo
# install`ed on the CachyOS side, which does not carry over — those
# binaries link that distro's glibc and will not run here.
# cargo-deny reads deny.toml; cargo-audit reads .cargo/audit.toml.
cargo-audit
cargo-deny
# Debian packaging (`cargo deb --no-build`). Note the .deb itself
# should still be built inside a Debian/Ubuntu distrobox so the
# binary links that distro's glibc — see the packaging notes in
# Cargo.toml.
cargo-deb
pkg-config
# pipewire-sys and libspa-sys generate their bindings with bindgen,
# which needs a real libclang present at build time.
clang
# audiopus_sys prefers the system libopus via pkg-config but falls
# back to a vendored CMake build; cmake keeps that fallback working
# rather than failing obscurely inside a build script.
cmake
# build.rs shells out to `git rev-parse --short=8 HEAD` to stamp
# PEERSPEAK_GIT_SHORT into the binary (surfaced in Settings).
git
])
++ screenshareTools
++ [
pkgs.pulseaudio # `pactl`, used by pixelpass's audio routing
pkgs.mpv # the screen-share viewer
];
buildInputs =
with pkgs;
[
alsa-lib # alsa-sys, pulled in by rodio/cpal
libopus # audiopus_sys, linked dynamically
pipewire # pipewire-sys + libspa-sys: the Linux audio backend
]
++ runtimeLibs;
# bindgen finds libclang through this variable specifically — having
# clang on PATH is not sufficient.
LIBCLANG_PATH = "${pkgs.llvmPackages.libclang.lib}/lib";
LD_LIBRARY_PATH = pkgs.lib.makeLibraryPath runtimeLibs;
# NixOS keeps every GStreamer plugin in its own store path, so the
# gst-launch-1.0 that pixelpass spawns discovers them ONLY through this
# search path. Same reasoning as hosts/darp5 and hosts/cazen in
# nixos-config.
GST_PLUGIN_SYSTEM_PATH_1_0 =
pkgs.lib.makeSearchPathOutput "lib" "lib/gstreamer-1.0" screenshareTools;
# Only greet an interactive shell. shellHook also runs under
# `nix develop --command …`, where printing this would interleave the
# banner with the command's own output.
shellHook = ''
if [ -t 1 ]; then
echo "peerspeak — rustc $(rustc --version | cut -d' ' -f2) / cargo $(cargo --version | cut -d' ' -f2)"
echo " cargo build --release build"
echo " cargo test lib tests"
echo " cargo clippy --all-targets -- -D warnings lint"
echo
echo "Screen sharing spawns pixelpass as a child process — it must be"
echo "on PATH. Build it from ../pixelpass and add its target/release."
fi
'';
};
};
}
+18 -5
View File
@@ -24,7 +24,10 @@ The AppImage runs on any reasonably current glibc-based distro that has:
- **A Vulkan-capable GPU + driver** (peerspeak's iced/wgpu renderer). Mesa/RADV
on AMD/Intel or the NVIDIA driver all work.
- **PipeWire** (with the PulseAudio shim, for `pactl`).
- **PipeWire** with the PulseAudio shim, `pactl`, and the host `libpulse.so.0`
client library. The Pulse client stack is deliberately not bundled because
PixelPass launches host GStreamer tools that must keep using the host's
matching multimedia libraries.
- For **screen-share only** — pixelpass shells out to these on the host `PATH`;
it prints the exact package names for your distro if any are missing:
- **GStreamer + plugins** (`gst-launch-1.0`/`gst-inspect-1.0`, base,
@@ -58,13 +61,23 @@ distrobox create --yes --image ubuntu:24.04 --name peerspeak-appimage
distrobox enter peerspeak-appimage -- sudo apt-get update
distrobox enter peerspeak-appimage -- sudo apt-get install -y \
build-essential cmake clang libclang-dev pkg-config \
libpipewire-0.3-dev libspa-0.2-dev libasound2-dev libxcb1-dev \
libpipewire-0.3-dev libspa-0.2-dev libpulse-dev libasound2-dev libxcb1-dev \
curl ca-certificates file patchelf git
rustup toolchain install 1.97.1 --profile default
# Build (the host's ~/.rustup toolchain is glibc-2.17-baseline, so it runs in the
# box; isolated CARGO_TARGET_DIRs keep it off the host target/):
# PeerSpeak's ownership validator needs the modern SPA-JSON parser headers it
# is tested against. Distrobox exposes the host's headers under /run/host;
# pkg-config still links Ubuntu's older ABI-compatible libraries, preserving
# the glibc baseline. This checkout's verified host-header version is 1.6.8:
test -f /run/host/usr/include/spa-0.2/spa/utils/json-core.h
# Build with the repository's pinned Rust version (the host's ~/.rustup
# toolchain is glibc-2.17-baseline, so it runs in the box; isolated
# CARGO_TARGET_DIRs keep it off the host target/):
distrobox enter peerspeak-appimage -- env \
PATH="$HOME/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/bin:$PATH" \
PATH="$HOME/.rustup/toolchains/1.97.1-x86_64-unknown-linux-gnu/bin:$PATH" \
SYSTEM_DEPS_LIBSPA_INCLUDE="/run/host/usr/include/spa-0.2" \
SYSTEM_DEPS_LIBPIPEWIRE_INCLUDE="/run/host/usr/include/pipewire-0.3:/run/host/usr/include/spa-0.2" \
./packaging/appimage/build-appimage.sh
```
+7 -3
View File
@@ -8,9 +8,12 @@
# X11) is dlopen'd at runtime and is on the AppImage excludelist because it must
# match the host driver, and pixelpass's capture/encode tools (gst-launch-1.0,
# pactl, mpv) are expected on the host PATH. So the AppImage carries just the two
# binaries plus their handful of non-excludelisted libs. The custom AppRun
# prepends usr/bin to PATH so peerspeak's own $PATH lookup finds the bundled
# pixelpass, while the host's tools stay reachable.
# binaries plus their handful of non-excludelisted libs. PulseAudio's client
# stack is also excluded: pixelpass shells out to host GStreamer, and letting
# those subprocesses inherit Ubuntu's bundled libsndfile/libmpg123 stack can
# override incompatible host multimedia libraries. The custom AppRun prepends
# usr/bin to PATH so peerspeak's own $PATH lookup finds the bundled pixelpass,
# while the host's tools and matching audio stack stay reachable.
#
# All runtime assets (notification WAVs, avatar presets, window icon, fonts) are
# include_bytes!-embedded in the peerspeak binary, so nothing else is bundled.
@@ -80,6 +83,7 @@ echo ">> running linuxdeploy (bundles libs, builds the AppImage)"
--appdir "$appdir" \
-e "$appdir/usr/bin/peerspeak" \
-e "$appdir/usr/bin/pixelpass" \
--exclude-library 'lib*.so*' \
-d "$repo/packaging/peerspeak.desktop" \
-i "$repo/assets/icons/peerspeak-256.png" \
--icon-filename peerspeak \
+367 -66
View File
@@ -16,6 +16,7 @@ use crate::hotkeys::{HotkeyAction, HotkeyContext, KeyBinding, format_binding};
use crate::network::PeerState;
use crate::notify::{self, Sound};
use crate::presence::PresenceMode;
use crate::screenshare::{AudioExclusionStatus, ShareAudioSelection};
use crate::theme::{AppTheme, Palette};
use crate::widget::context_input::{context_input, locked_value};
use crate::widget::selectable_text::selectable_rich_text;
@@ -901,9 +902,8 @@ pub enum AppMessage {
ToggleScreenShare,
/// Close the screen-share audio picker without sharing.
CloseSharePicker,
/// Select which app's audio to share in the picker: `Some(name)` for one app,
/// `None` for the whole desktop ("All system audio").
SelectShareAudioApp(Option<String>),
/// Select legacy desktop, desktop-excluding, or strict per-app audio.
SelectShareAudio(ShareAudioSelection),
/// Session-only quality preset for the next share start.
SelectShareQualityOverride(ShareQuality),
/// Confirm the picker: start the share with the currently selected audio app.
@@ -1150,9 +1150,10 @@ pub struct AppState {
/// Apps currently producing audio, shown in the share picker. Populated from
/// `UiEvent::AudioAppsListed` after the picker requests an enumeration.
share_audio_apps: Vec<String>,
/// The picker's current selection: `Some(name)` = capture that app's audio,
/// `None` = "All system audio" (whole desktop; may echo the call).
share_audio_selection: Option<String>,
/// The picker's current typed selection. Desktop-excluding is the normal
/// whole-desktop choice; desktop-shared remains an internal compatibility
/// fallback for an older PixelPass that does not advertise exclusion.
share_audio_selection: ShareAudioSelection,
/// Session-only quality override for the next screen-share start.
share_quality_selection: ShareQuality,
/// A share start is in flight: `ConfirmShareScreen` was sent but the core
@@ -1171,11 +1172,20 @@ pub struct AppState {
/// just-killed host can't flip the warning on a new whole-desktop share or
/// after stop (audit P3, unscoped events).
share_audio_app_active: bool,
/// Whether the current share is the desktop-excluding mode. This gates its
/// status events so late notices cannot affect another share mode.
share_desktop_excluding_active: bool,
/// The latest still-active warning from PixelPass while desktop exclusion
/// is active. Keeping the typed status preserves a stream serial so an
/// additive `stream_status_cleared` event can clear only its own warning.
share_audio_exclusion_warning: Option<AudioExclusionStatus>,
/// Whether the resolved pixelpass supports `--strict-audio` (per-app audio).
/// `false` ⇒ the picker offers whole-desktop only, because a per-app share
/// would pass a flag an older pixelpass rejects (audit P2). Optimistic `true`
/// until the core's `AudioAppsListed` reports otherwise.
share_app_audio_supported: bool,
/// Whether the resolved PixelPass advertises desktop audio exclusion.
share_desktop_audio_exclusion_supported: bool,
/// Room-level warning for a validly signed peer whose gossip timestamp falls
/// outside the replay freshness window. The peer is not yet in the roster, so
/// this is not attached to a participant card.
@@ -1272,12 +1282,15 @@ impl AppState {
self.self_sharing = false;
self.share_picker_open = false;
self.share_audio_apps.clear();
self.share_audio_selection = None;
self.share_audio_selection = ShareAudioSelection::DesktopShared;
self.share_quality_selection = self.config.screen_share.quality;
self.share_starting = false;
self.share_audio_dropped = false;
self.share_audio_app_active = false;
self.share_desktop_excluding_active = false;
self.share_audio_exclusion_warning = None;
self.share_app_audio_supported = true;
self.share_desktop_audio_exclusion_supported = false;
self.clock_skew_warning = None;
}
@@ -1580,12 +1593,15 @@ impl Default for AppState {
pixelpass_help_open: false,
share_picker_open: false,
share_audio_apps: Vec::new(),
share_audio_selection: None,
share_audio_selection: ShareAudioSelection::DesktopShared,
share_quality_selection,
share_starting: false,
share_audio_dropped: false,
share_audio_app_active: false,
share_desktop_excluding_active: false,
share_audio_exclusion_warning: None,
share_app_audio_supported: true,
share_desktop_audio_exclusion_supported: false,
clock_skew_warning: None,
drawer_chat_open: false,
playlist_drawer_open: false,
@@ -2172,12 +2188,14 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
// Open the audio picker instead of sharing immediately, so the
// user chooses which app's audio to capture rather than the whole
// desktop (which echoes the call back to viewers, A23). Default
// selection is "All system audio" (None). Kick off a fresh
// to the fail-closed desktop-exclusion contract while the fresh
// capability probe runs; an old PixelPass response replaces it
// with the visibly warned legacy fallback. Kick off a fresh
// enumeration so the list reflects what's playing right now.
// Suppressed while a start is already in flight (`share_starting`)
// so the picker can't be reopened during the startup window.
state.share_picker_open = true;
state.share_audio_selection = None;
state.share_audio_selection = ShareAudioSelection::DesktopExcluding;
// NB: do NOT reset `share_quality_selection` here. It is the
// per-call override set by the inline quality dropdown next to
// the Share button, and the picker has no quality control of its
@@ -2190,8 +2208,8 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
AppMessage::CloseSharePicker => {
state.share_picker_open = false;
}
AppMessage::SelectShareAudioApp(app) => {
state.share_audio_selection = app;
AppMessage::SelectShareAudio(audio) => {
state.share_audio_selection = audio;
}
AppMessage::SelectShareQualityOverride(quality) => {
state.share_quality_selection = quality;
@@ -2203,11 +2221,12 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
if state.share_picker_open && !state.share_starting {
state.share_picker_open = false;
state.share_starting = true;
let audio_app = state.share_audio_selection.clone();
state.share_audio_exclusion_warning = None;
let audio = state.share_audio_selection.clone();
let settings = state.config.screen_share.clone();
let quality = state.share_quality_selection;
let _ = state.controller.send(CoreCommand::StartScreenShare {
audio_app,
audio,
settings,
quality,
});
@@ -2513,25 +2532,52 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
UiEvent::AudioAppsListed {
apps,
app_audio_supported,
desktop_audio_exclusion_supported,
} => {
// Only meaningful while the picker is open; if the user
// already cancelled, drop it.
if state.share_picker_open {
let desktop_fallback = if desktop_audio_exclusion_supported {
ShareAudioSelection::DesktopExcluding
} else {
ShareAudioSelection::DesktopShared
};
state.share_app_audio_supported = app_audio_supported;
state.share_desktop_audio_exclusion_supported =
desktop_audio_exclusion_supported;
if app_audio_supported {
// Keep the current selection if it still exists in the
// refreshed list, else fall back to "All system audio".
if let Some(sel) = &state.share_audio_selection
if let ShareAudioSelection::Application(sel) =
&state.share_audio_selection
&& !apps.iter().any(|a| a == sel)
{
state.share_audio_selection = None;
state.share_audio_selection = desktop_fallback.clone();
}
state.share_audio_apps = apps;
} else {
// Older pixelpass: per-app capture would hard-fail
// (--strict-audio unknown). Force whole-desktop only.
state.share_audio_apps.clear();
state.share_audio_selection = None;
if matches!(
state.share_audio_selection,
ShareAudioSelection::Application(_)
) {
state.share_audio_selection = desktop_fallback.clone();
}
}
match state.share_audio_selection {
ShareAudioSelection::DesktopShared
if desktop_audio_exclusion_supported =>
{
state.share_audio_selection = ShareAudioSelection::DesktopExcluding;
}
ShareAudioSelection::DesktopExcluding
if !desktop_audio_exclusion_supported =>
{
state.share_audio_selection = ShareAudioSelection::DesktopShared;
}
_ => {}
}
}
}
@@ -2541,7 +2587,14 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
state.share_audio_dropped = false;
// Remember whether this share captures a specific app, so we
// only apply `app_audio` warnings to app shares (P3).
state.share_audio_app_active = state.share_audio_selection.is_some();
state.share_audio_app_active = matches!(
state.share_audio_selection,
ShareAudioSelection::Application(_)
);
state.share_desktop_excluding_active = matches!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding
);
// Defensive: ensure no picker lingers across a successful start.
state.share_picker_open = false;
state.status_message = "Sharing your screen".to_string();
@@ -2551,6 +2604,8 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
state.share_starting = false;
state.share_audio_dropped = false;
state.share_audio_app_active = false;
state.share_desktop_excluding_active = false;
state.share_audio_exclusion_warning = None;
state.status_message = "Screen share stopped".to_string();
}
UiEvent::ShareAudioActive(active) => {
@@ -2562,6 +2617,40 @@ fn update(state: &mut AppState, message: AppMessage) -> Task<AppMessage> {
state.share_audio_dropped = !active;
}
}
UiEvent::ShareAudioExclusionStatus(status) => {
// A status can race the start acknowledgement because the
// host drain and core loop use cloned UI senders. Accept it
// during an in-flight excluding start as well as the active
// share, but ignore late notices for other modes.
let excluding_starting = state.share_starting
&& matches!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding
);
if state.share_desktop_excluding_active || excluding_starting {
match status {
AudioExclusionStatus::StreamStatusCleared { stream_serial } => {
let clears_visible = matches!(
state.share_audio_exclusion_warning.as_ref(),
Some(AudioExclusionStatus::StreamUnsupported {
stream_serial: visible_serial,
..
}) if *visible_serial == stream_serial
);
if clears_visible {
state.share_audio_exclusion_warning = None;
state.status_message = "Sharing your screen".to_string();
}
}
warning => {
if let Some(message) = audio_exclusion_status_message(&warning) {
state.status_message = message;
state.share_audio_exclusion_warning = Some(warning);
}
}
}
}
}
UiEvent::ClockSkewWarning {
skew_secs,
peer_ahead,
@@ -5568,7 +5657,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
.size(11)
.color(color_subtext),
slider(0.0..=1.0, state.config.background_dim, AppMessage::SetBackgroundDim)
.step(0.05),
.step(0.05_f32),
text("Set a picture from your computer as the app background. Auto-resized; a dimming overlay keeps text readable. Applies live.")
.size(11)
.color(color_subtext),
@@ -5996,7 +6085,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
).width(iced::Length::Fill),
text(format!("Input Volume (mic): {:.0}%", state.config.input_volume * 100.0)).size(11).color(color_subtext),
slider(0.0..=2.0, state.config.input_volume, AppMessage::InputVolumeChanged)
.step(0.05)
.step(0.05_f32)
.on_release(AppMessage::PersistConfig),
].spacing(8).width(iced::Length::Fill),
column![
@@ -6008,7 +6097,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
).width(iced::Length::Fill),
text(format!("Output Volume: {:.0}%", state.config.output_volume * 100.0)).size(11).color(color_subtext),
slider(0.0..=2.0, state.config.output_volume, AppMessage::OutputVolumeChanged)
.step(0.05)
.step(0.05_f32)
.on_release(AppMessage::PersistConfig),
].spacing(8).width(iced::Length::Fill),
].spacing(20).align_y(iced::alignment::Vertical::Top).width(iced::Length::Fill),
@@ -6618,7 +6707,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
column![
row![
avatar_view(&state.config.avatar, &state.name, &state.self_id, 34.0),
text(format!("{} (You)", &state.name)).size(16).color(color_text),
text(format!("{} (You)", state.name)).size(16).color(color_text),
horizontal_space(),
if state.is_muted {
text("[Muted]").size(14).color(color_red)
@@ -6656,20 +6745,24 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
]
.spacing(6)
.align_y(iced::alignment::Vertical::Center);
let mut details = column![badge].spacing(3);
if state.share_audio_dropped {
column![
badge,
details = details.push(
text(
"⚠ Shared app isn't sending audio — viewers hear silence until it plays"
)
.size(11)
.color(color_yellow),
]
.spacing(3)
.into()
} else {
badge.into()
);
}
if let Some(warning) = state
.share_audio_exclusion_warning
.as_ref()
.and_then(audio_exclusion_status_message)
{
details = details.push(text(warning).size(11).color(color_yellow));
}
details.into()
} else {
iced::widget::Space::new().width(0.0).height(0.0).into()
};
@@ -6933,7 +7026,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
slider(0.0..=2.0, current_vol, move |v| {
AppMessage::PeerVolumeChanged(peer_id_clone, v)
})
.step(0.01)
.step(0.01_f32)
.on_release(AppMessage::PersistConfig)
]
.spacing(8)
@@ -6946,7 +7039,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
slider(-1.0..=1.0, current_pan, move |v| {
AppMessage::PeerPanChanged(peer_id_clone, v)
})
.step(0.05)
.step(0.05_f32)
.on_release(AppMessage::PersistConfig),
]
.spacing(8)
@@ -6961,7 +7054,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
slider(0.0..=METER_MAX, current_gate, move |v| {
AppMessage::PeerGateChanged(peer_id_clone, v)
})
.step(0.001)
.step(0.001_f32)
.on_release(AppMessage::PersistConfig),
]
.spacing(8)
@@ -6980,7 +7073,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
slider(EQ_GAIN_DB_MIN..=EQ_GAIN_DB_MAX, value, move |v| {
AppMessage::PeerEqChanged(peer_id_clone, band, v)
})
.step(0.5)
.step(0.5_f32)
.on_release(AppMessage::PersistConfig),
]
.spacing(8)
@@ -7215,7 +7308,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
clip_progress(music_status.position, music_status.total),
AppMessage::MusicSeek,
)
.step(0.001),
.step(0.001_f32),
text(format!("{elapsed} / {duration}"))
.size(11)
.color(color_subtext),
@@ -7225,7 +7318,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
state.config.music_volume,
AppMessage::MusicSetVolume
)
.step(0.01),
.step(0.01_f32),
button(
text("Browse")
.size(12)
@@ -7319,7 +7412,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
effective_music_volume(state),
AppMessage::MusicSetSourceVolume
)
.step(0.01),
.step(0.01_f32),
]
.spacing(8)
.into()
@@ -7815,7 +7908,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
},
move |fraction| AppMessage::SeekAudio(att.id, fraction),
)
.step(0.001)
.step(0.001_f32)
.width(iced::Length::Fixed(180.0)),
text(format!("{elapsed} / {duration}"))
.size(11)
@@ -7827,7 +7920,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
slider(0.0..=2.0, effective_clip_volume(state, att.id), move |v| {
AppMessage::SetClipVolumeFor(att.id, v)
},)
.step(0.01)
.step(0.01_f32)
.width(iced::Length::Fixed(80.0)),
]
.spacing(8)
@@ -7945,7 +8038,7 @@ fn view(state: &AppState) -> Element<'_, AppMessage> {
state.config.clip_volume,
AppMessage::SetClipVolume
)
.step(0.01)
.step(0.01_f32)
.width(iced::Length::Fixed(110.0)),
]
.spacing(10)
@@ -8757,11 +8850,28 @@ fn with_pixelpass_help<'a>(
.into()
}
fn audio_exclusion_status_message(status: &AudioExclusionStatus) -> Option<String> {
match status {
AudioExclusionStatus::StreamUnsupported { reason, .. } => Some(format!(
"Some system audio could not be shared safely ({reason}); it was left out."
)),
AudioExclusionStatus::StreamStatusCleared { .. } => None,
AudioExclusionStatus::AecFailed { .. } =>
Some("Audio exclusion could not verify PeerSpeak's echo canceller; viewers hear silence."
.to_string()),
AudioExclusionStatus::AecRevoked { .. } =>
Some("Audio exclusion stopped because PeerSpeak's echo canceller disappeared; viewers hear silence."
.to_string()),
AudioExclusionStatus::ForeignAecWarning { .. } =>
Some("Another echo-cancel stream was kept out of the screen share.".to_string()),
}
}
/// Overlay the screen-share **audio picker** when open (A23). Lets the user
/// capture a single app's audio instead of the whole desktop sink — the default
/// whole-desktop capture contains our own call playout, so a viewer would
/// otherwise hear themselves echoed back. "All system audio" keeps the legacy
/// behavior (with a warning); picking an app passes `--app=<name>` to pixelpass.
/// capture all safe system audio or one specific app. With a capable PixelPass,
/// "All system audio" uses desktop exclusion so PeerSpeak playback is never
/// included. An older PixelPass gets the same row backed by legacy capture and
/// an explicit echo warning. Picking an app passes `--app=<name>` to pixelpass.
fn with_share_picker<'a>(
base: Element<'a, AppMessage>,
state: &'a AppState,
@@ -8828,22 +8938,47 @@ fn with_share_picker<'a>(
})
};
// "All system audio" first (the whole-desktop default — carries the echo
// warning), then each currently-playing app.
let mut options = column![opt_row(
state.share_audio_selection.is_none(),
"All system audio".to_string(),
Some("⚠ may echo the call back to viewers"),
AppMessage::SelectShareAudioApp(None),
)]
// Present one whole-desktop row. A current PixelPass backs it with safe
// exclusion; only an older binary sees the legacy implementation/warning.
// Treat the optimistic pre-probe DesktopExcluding selection as the safe row
// too; the core re-probes the exact binary before constructing its argv.
let safe_desktop = state.share_desktop_audio_exclusion_supported
|| matches!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding
);
let mut options = if safe_desktop {
column![opt_row(
matches!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding
),
"All system audio".to_string(),
Some("Excludes call and watched-share playback"),
AppMessage::SelectShareAudio(ShareAudioSelection::DesktopExcluding),
)]
} else {
column![opt_row(
matches!(
state.share_audio_selection,
ShareAudioSelection::DesktopShared
),
"All system audio".to_string(),
Some("⚠ may echo the call back to viewers"),
AppMessage::SelectShareAudio(ShareAudioSelection::DesktopShared),
)]
}
.spacing(4);
for app in &state.share_audio_apps {
let selected = state.share_audio_selection.as_deref() == Some(app.as_str());
let selected = matches!(
&state.share_audio_selection,
ShareAudioSelection::Application(selected) if selected == app
);
options = options.push(opt_row(
selected,
app.clone(),
None,
AppMessage::SelectShareAudioApp(Some(app.clone())),
AppMessage::SelectShareAudio(ShareAudioSelection::Application(app.clone())),
));
}
@@ -9668,13 +9803,14 @@ mod tests {
use super::PendingSend;
use super::sendqueue::{self, LocalSend, SendStatus};
use super::{
AppConfig, AppMessage, AppState, AttachmentCache, AttachmentState,
AppConfig, AppMessage, AppState, AttachmentCache, AttachmentState, AudioExclusionStatus,
CLOCK_SKEW_WARNING_VISIBLE_SECS, ChatEntry, ClockSkewBanner, GateMeter, METER_MAX, Screen,
ScreenBounds, UiEvent, attachment_default_name, clamp_window_position,
clear_expired_clock_skew_warning, format_clock_skew_duration, format_duration,
format_relative_ago, friend_presence_notification, initial_window_position,
now_playing_label, reconnect_attempt_chime, reconnected_chime, selected_wav_path,
set_peer_gate_config, set_peer_volume_config, show_clock_skew_warning, update,
ScreenBounds, ShareAudioSelection, UiEvent, attachment_default_name,
audio_exclusion_status_message, clamp_window_position, clear_expired_clock_skew_warning,
format_clock_skew_duration, format_duration, format_relative_ago,
friend_presence_notification, initial_window_position, now_playing_label,
reconnect_attempt_chime, reconnected_chime, selected_wav_path, set_peer_gate_config,
set_peer_volume_config, show_clock_skew_warning, update,
};
use iroh::SecretKey;
use std::collections::VecDeque;
@@ -9970,11 +10106,15 @@ mod tests {
state.self_sharing = true;
state.share_picker_open = true;
state.share_audio_apps = vec!["Firefox".to_string()];
state.share_audio_selection = Some("Firefox".to_string());
state.share_audio_selection = ShareAudioSelection::Application("Firefox".to_string());
state.share_starting = true;
state.share_audio_dropped = true;
state.share_audio_app_active = true;
state.share_desktop_excluding_active = true;
state.share_audio_exclusion_warning =
Some(AudioExclusionStatus::AecFailed { module_index: 77 });
state.share_app_audio_supported = false;
state.share_desktop_audio_exclusion_supported = true;
state.clock_skew_warning = Some(ClockSkewBanner {
skew_secs: 180,
peer_ahead: true,
@@ -10011,7 +10151,10 @@ mod tests {
assert!(!state.self_sharing);
assert!(!state.share_picker_open);
assert!(state.share_audio_apps.is_empty());
assert!(state.share_audio_selection.is_none());
assert_eq!(
state.share_audio_selection,
ShareAudioSelection::DesktopShared
);
assert_eq!(
state.share_quality_selection,
state.config.screen_share.quality
@@ -10019,10 +10162,13 @@ mod tests {
assert!(!state.share_starting);
assert!(!state.share_audio_dropped);
assert!(!state.share_audio_app_active);
assert!(!state.share_desktop_excluding_active);
assert!(state.share_audio_exclusion_warning.is_none());
assert!(
state.share_app_audio_supported,
"reset is optimistic by default"
);
assert!(!state.share_desktop_audio_exclusion_supported);
assert!(state.clock_skew_warning.is_none());
assert!(state.music_broadcast_id.is_none());
assert!(state.music_broadcast_next.is_none());
@@ -10135,7 +10281,7 @@ mod tests {
// Picker open, user confirms a selection.
let mut state = AppState {
share_picker_open: true,
share_audio_selection: Some("mpv".to_string()),
share_audio_selection: ShareAudioSelection::Application("mpv".to_string()),
..Default::default()
};
let _ = update(&mut state, AppMessage::ConfirmShareScreen);
@@ -10204,6 +10350,20 @@ mod tests {
);
}
#[test]
fn share_picker_defaults_whole_desktop_to_safe_exclusion() {
let mut state = AppState::default();
let _ = update(&mut state, AppMessage::ToggleScreenShare);
assert!(state.share_picker_open);
assert_eq!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding,
"the picker must not expose legacy echoing capture as its normal desktop default"
);
}
#[test]
fn share_start_failure_clears_in_flight_flag() {
// A failed spawn surfaces as UiEvent::Error (not ScreenShareStopped); the
@@ -10233,7 +10393,7 @@ mod tests {
// flag; start and stop both reset it so it can't linger across sessions.
// A specific app was chosen in the picker, so the share is app-specific.
let mut state = AppState {
share_audio_selection: Some("mpv".to_string()),
share_audio_selection: ShareAudioSelection::Application("mpv".to_string()),
..Default::default()
};
@@ -10300,7 +10460,7 @@ mod tests {
// (b) After stop: a straggling event can't resurrect the warning.
let mut state = AppState {
share_audio_selection: Some("mpv".to_string()),
share_audio_selection: ShareAudioSelection::Application("mpv".to_string()),
..Default::default()
};
let _ = update(
@@ -10318,6 +10478,100 @@ mod tests {
assert!(!state.share_audio_dropped, "post-stop event is ignored");
}
#[test]
fn desktop_exclusion_status_reaches_visible_state_and_clears_on_stop() {
let mut state = AppState {
share_starting: true,
share_audio_selection: ShareAudioSelection::DesktopExcluding,
..Default::default()
};
// The status may beat ScreenShareStarted because the host drain uses a
// cloned UI sender. It must still be retained and shown.
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ShareAudioExclusionStatus(
AudioExclusionStatus::AecFailed { module_index: 77 },
)),
);
assert!(
state
.share_audio_exclusion_warning
.as_ref()
.and_then(audio_exclusion_status_message)
.is_some_and(|message| message.contains("viewers hear silence"))
);
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ScreenShareStarted),
);
assert!(state.share_desktop_excluding_active);
assert!(state.share_audio_exclusion_warning.is_some());
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ScreenShareStopped),
);
assert!(!state.share_desktop_excluding_active);
assert!(state.share_audio_exclusion_warning.is_none());
// A stale notice after stop cannot resurrect the warning.
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ShareAudioExclusionStatus(
AudioExclusionStatus::AecRevoked { module_index: 77 },
)),
);
assert!(state.share_audio_exclusion_warning.is_none());
}
#[test]
fn desktop_exclusion_clears_only_the_matching_stream_warning() {
let mut state = AppState {
self_sharing: true,
share_desktop_excluding_active: true,
..Default::default()
};
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ShareAudioExclusionStatus(
AudioExclusionStatus::StreamUnsupported {
stream_serial: 41,
reason: "unidentified-channel".to_string(),
},
)),
);
assert!(matches!(
state.share_audio_exclusion_warning.as_ref(),
Some(AudioExclusionStatus::StreamUnsupported {
stream_serial: 41,
..
})
));
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ShareAudioExclusionStatus(
AudioExclusionStatus::StreamStatusCleared { stream_serial: 99 },
)),
);
assert!(
state.share_audio_exclusion_warning.is_some(),
"another stream's recovery must not clear the visible warning"
);
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::ShareAudioExclusionStatus(
AudioExclusionStatus::StreamStatusCleared { stream_serial: 41 },
)),
);
assert!(state.share_audio_exclusion_warning.is_none());
assert_eq!(state.status_message, "Sharing your screen");
}
#[test]
fn old_pixelpass_picker_offers_whole_desktop_only() {
// P2 (version skew): when the resolved pixelpass lacks --strict-audio, the
@@ -10325,7 +10579,7 @@ mod tests {
// so a per-app share (which would pass the unknown flag) can't be started.
let mut state = AppState {
share_picker_open: true,
share_audio_selection: Some("Firefox".to_string()),
share_audio_selection: ShareAudioSelection::Application("Firefox".to_string()),
share_audio_apps: vec!["Firefox".to_string(), "mpv".to_string()],
..Default::default()
};
@@ -10334,12 +10588,14 @@ mod tests {
AppMessage::UiEventReceived(UiEvent::AudioAppsListed {
apps: vec!["Firefox".to_string(), "mpv".to_string()],
app_audio_supported: false,
desktop_audio_exclusion_supported: false,
}),
);
assert!(!state.share_app_audio_supported);
assert!(!state.share_desktop_audio_exclusion_supported);
assert!(state.share_audio_apps.is_empty(), "no per-app rows offered");
assert!(
state.share_audio_selection.is_none(),
state.share_audio_selection == ShareAudioSelection::DesktopShared,
"forced to whole-desktop"
);
@@ -10349,10 +10605,55 @@ mod tests {
AppMessage::UiEventReceived(UiEvent::AudioAppsListed {
apps: vec!["Firefox".to_string(), "mpv".to_string()],
app_audio_supported: true,
desktop_audio_exclusion_supported: true,
}),
);
assert!(state.share_app_audio_supported);
assert!(state.share_desktop_audio_exclusion_supported);
assert_eq!(state.share_audio_apps.len(), 2);
assert_eq!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding,
"a capable PixelPass must make the single desktop row use exclusion"
);
}
#[test]
fn picker_treats_per_app_and_desktop_exclusion_as_independent_capabilities() {
let mut state = AppState {
share_picker_open: true,
share_audio_selection: ShareAudioSelection::DesktopExcluding,
..Default::default()
};
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::AudioAppsListed {
apps: vec!["Firefox".to_string()],
app_audio_supported: false,
desktop_audio_exclusion_supported: true,
}),
);
assert!(!state.share_app_audio_supported);
assert!(state.share_desktop_audio_exclusion_supported);
assert_eq!(
state.share_audio_selection,
ShareAudioSelection::DesktopExcluding,
"lack of per-app support must not hide the independent exclusion mode"
);
let _ = update(
&mut state,
AppMessage::UiEventReceived(UiEvent::AudioAppsListed {
apps: Vec::new(),
app_audio_supported: false,
desktop_audio_exclusion_supported: false,
}),
);
assert_eq!(
state.share_audio_selection,
ShareAudioSelection::DesktopShared,
"old PixelPass must remove the unavailable exclusion selection"
);
}
#[test]
+11 -6
View File
@@ -33,12 +33,17 @@ const NODE_READY_TIMEOUT: Duration = Duration::from_secs(3);
/// Owns a loaded `module-echo-cancel` instance; unloads it on drop so the virtual
/// nodes never leak past the call that created them.
pub struct EchoCancelGuard {
module_index: String,
module_index: u64,
source_name: String,
sink_name: String,
}
impl EchoCancelGuard {
/// The pactl module identity PixelPass validates in desktop-excluding mode.
pub fn module_index(&self) -> u64 {
self.module_index
}
pub fn source_name(&self) -> &str {
&self.source_name
}
@@ -52,7 +57,7 @@ impl Drop for EchoCancelGuard {
fn drop(&mut self) {
let _ = Command::new("pactl")
.arg("unload-module")
.arg(&self.module_index)
.arg(self.module_index.to_string())
.output();
crate::log_msg(&format!(
"Echo cancel: unloaded module {}",
@@ -103,10 +108,10 @@ pub fn enable(
));
}
let module_index = String::from_utf8_lossy(&out.stdout).trim().to_string();
if module_index.parse::<u64>().is_err() {
return Err(format!("unexpected pactl output: {module_index:?}"));
}
let raw_module_index = String::from_utf8_lossy(&out.stdout).trim().to_string();
let module_index = raw_module_index
.parse::<u64>()
.map_err(|_| format!("unexpected pactl output: {raw_module_index:?}"))?;
let guard = EchoCancelGuard {
module_index,
source_name,
+23 -10
View File
@@ -126,17 +126,26 @@ pub enum CoreCommand {
ListAudioApps,
/// Start sharing our screen: spawn a pixelpass host and announce its ticket
/// on our presence so the room can watch. No-op when not in a call.
/// `audio_app` selects which app's audio to capture: `Some(name)` captures
/// only that app (avoiding the call-loopback echo, A23); `None` shares the
/// whole desktop audio (the legacy behavior).
/// `audio` is typed so legacy whole-desktop, desktop-excluding, and strict
/// per-app capture remain distinct across the UI/core boundary.
StartScreenShare {
audio_app: Option<String>,
audio: crate::screenshare::ShareAudioSelection,
settings: ScreenShareSettings,
quality: ShareQuality,
},
/// Stop sharing our screen: kill the pixelpass host and clear the presence
/// ticket. No-op when not sharing.
StopScreenShare,
/// **Core-internal.** The running pixelpass host's stdout ended — the
/// process died (or its event stream broke), so the share identified by
/// `generation` is over: reap the child, pull the ticket off presence, and
/// tell the user. Synthesized by the core's own notice-forwarder task; the
/// UI never sends it. `generation` scopes the fault to one specific host
/// spawn, so a stale fault (the user already stopped, or started a new
/// share) is ignored rather than tearing down the wrong share.
ScreenShareHostFault {
generation: u64,
},
/// Watch a peer's screen share: spawn a pixelpass viewer for `ticket` and
/// open it in a local player.
ViewShare {
@@ -266,11 +275,12 @@ pub fn delivery_class(cmd: &CoreCommand) -> DeliveryClass {
| CoreCommand::SetPixelpassPath(_)
| CoreCommand::ListAudioApps
| CoreCommand::StartScreenShare {
audio_app: _,
audio: _,
settings: _,
quality: _,
}
| CoreCommand::StopScreenShare
| CoreCommand::ScreenShareHostFault { generation: _ }
| CoreCommand::ViewShare {
ticket: _,
settings: _,
@@ -358,11 +368,12 @@ pub fn coalesce_key(cmd: &CoreCommand) -> Option<CoalesceKey> {
| CoreCommand::SetPixelpassPath(_)
| CoreCommand::ListAudioApps
| CoreCommand::StartScreenShare {
audio_app: _,
audio: _,
settings: _,
quality: _,
}
| CoreCommand::StopScreenShare
| CoreCommand::ScreenShareHostFault { generation: _ }
| CoreCommand::ViewShare {
ticket: _,
settings: _,
@@ -499,13 +510,12 @@ pub enum UiEvent {
},
/// The apps currently producing audio, for the screen-share audio picker
/// (A23). Sorted, deduplicated `application.name`s; empty when nothing is
/// playing or enumeration isn't available. `app_audio_supported` reports
/// whether the resolved pixelpass understands `--strict-audio`: when `false`
/// (an older pixelpass) the picker must offer whole-desktop audio only, since
/// a per-app share would pass a flag that older binary rejects (audit P2).
/// playing or enumeration isn't available. The two support bits are
/// independent and belong to the exact resolved PixelPass binary.
AudioAppsListed {
apps: Vec<String>,
app_audio_supported: bool,
desktop_audio_exclusion_supported: bool,
},
/// Our own screen share started; the UI flips the Share button to "Stop".
ScreenShareStarted,
@@ -516,6 +526,9 @@ pub enum UiEvent {
/// run viewers currently hear silence. The UI shows a transient warning while
/// `false`. Only meaningful while sharing a specific app (not whole-desktop).
ShareAudioActive(bool),
/// One desktop-excluding status parsed from PixelPass and forwarded without
/// translating it into a separately maintained PeerSpeak enum.
ShareAudioExclusionStatus(crate::screenshare::AudioExclusionStatus),
/// A validly signed peer cannot be admitted because its gossip timestamp is
/// outside the replay freshness window. `peer_ahead` describes the peer's
/// sender-stamped timestamp relative to this machine's clock.
+400 -95
View File
@@ -4,6 +4,7 @@ pub mod fetchbudget;
pub mod jitter;
pub mod messages;
mod recovery;
mod teardown;
use crate::audio::eq::{Eq, EqSettings};
use crate::audio::{AudioBackend, PlatformAudioBackend};
@@ -677,31 +678,33 @@ struct ActiveSession {
recovery_terminal_task: tokio::task::JoinHandle<()>,
grace_timers: GraceTimers,
transport: Arc<IrohTransport>,
/// Loaded PipeWire echo-cancel module (if enabled); unloads on drop.
#[cfg(target_os = "linux")]
echo_cancel: Option<crate::audio::echo_cancel::EchoCancelGuard>,
/// Our pixelpass screen-share host child while sharing (`kill_on_drop`, so it
/// also dies if the session is dropped without an explicit stop).
screenshare_host: Option<tokio::process::Child>,
/// pixelpass viewer children we spawned to watch peers' shares, each paired
/// with the share ticket it's viewing so a re-watch of the same share can
/// replace (not stack) its player. Killed on session teardown (each also
/// self-exits when its player window closes).
screenshare_viewers: Vec<(String, tokio::process::Child)>,
/// The screen-share children and the echo-cancel module, held together
/// because their **destruction order** is load-bearing: the AEC module must
/// not unload while a pixelpass host is alive and fanning out (design v3.4
/// §7.1). `teardown` owns that ordering; see `core::teardown`.
teardown: SessionTeardown,
}
/// The session's teardown set, with the echo-cancel guard the platform actually
/// has. On non-Linux there is no AEC module, and `Infallible` makes that
/// structural — the `Option` cannot be `Some`.
#[cfg(target_os = "linux")]
type SessionTeardown = teardown::ScreenshareTeardown<
tokio::process::Child,
crate::audio::echo_cancel::EchoCancelGuard,
>;
#[cfg(not(target_os = "linux"))]
type SessionTeardown =
teardown::ScreenshareTeardown<tokio::process::Child, std::convert::Infallible>;
impl ActiveSession {
async fn shutdown(mut self, audio_backend: Arc<PlatformAudioBackend>) {
crate::log_msg("ActiveSession::shutdown started");
// Tear down any screen-share children first so the host stops streaming
// promptly (kill_on_drop is the backstop, but kill explicitly so viewers
// see the stream end without waiting on drop ordering).
if let Some(mut host) = self.screenshare_host.take() {
let _ = host.kill().await;
}
for (_, mut viewer) in self.screenshare_viewers.drain(..) {
let _ = viewer.kill().await;
}
// promptly, and so they are dead *and reaped* well before the AEC guard
// unloads at the end of this function (design v3.4 §7.1). Drop ordering
// is the backstop for the unwind path; this is the path we control.
self.teardown.shutdown_children().await;
self.datagram_task.abort();
self.mixer_task.abort();
self.event_task.abort();
@@ -726,8 +729,9 @@ impl ActiveSession {
// Unload the echo-cancel module now that the audio streams releasing its
// virtual nodes have stopped. (Dropping the guard runs `pactl unload`.)
#[cfg(target_os = "linux")]
drop(self.echo_cancel);
// The screen-share children were killed *and reaped* at the top of this
// function, so nothing pixelpass-side is alive to see the module vanish.
drop(self.teardown);
crate::log_msg("Leaving room...");
let _ = self.room_state.leave().await;
@@ -1237,6 +1241,11 @@ const PING_INTERVAL: Duration = Duration::from_secs(15);
/// Delay before the FIRST presence pass, so the endpoint's background `online()`
/// has a moment to finish (otherwise the first probes fail and friends flash offline).
const PING_STARTUP_DELAY: Duration = Duration::from_secs(3);
/// How quickly the core polls owned PixelPass viewer children for natural exit.
/// A viewer normally exits when the remote host stops sharing; without this
/// independent tick it remains an unreaped zombie until another Watch click or
/// the entire call ends.
const VIEWER_REAP_INTERVAL: Duration = Duration::from_millis(500);
/// One outbound presence-refresh pass (W7 B2): probe every friend and emit a
/// *definitive* status for each, so the UI self-heals every pass instead of only
@@ -1280,6 +1289,42 @@ async fn probe_friends_once(
}
}
fn ui_event_from_pixelpass_event(event: crate::screenshare::PixelpassEvent) -> Option<UiEvent> {
match event {
crate::screenshare::PixelpassEvent::AppAudioRouted => Some(UiEvent::ShareAudioActive(true)),
crate::screenshare::PixelpassEvent::AppAudioLost => Some(UiEvent::ShareAudioActive(false)),
crate::screenshare::PixelpassEvent::AudioExclusion(status) => {
Some(UiEvent::ShareAudioExclusionStatus(status))
}
_ => None,
}
}
/// Carry the actual parsed PixelPass event value across the notice channel to
/// the UI. `Eof` remains a generation-scoped core fault rather than a UI event.
async fn forward_host_notices(
mut notices: mpsc::UnboundedReceiver<crate::screenshare::HostNotice>,
ui_tx: mpsc::Sender<UiEvent>,
fault_tx: mpsc::UnboundedSender<u64>,
generation: u64,
) {
while let Some(notice) = notices.recv().await {
match notice {
crate::screenshare::HostNotice::Event(event) => {
if let Some(event) = ui_event_from_pixelpass_event(event)
&& ui_tx.send(event).await.is_err()
{
break;
}
}
crate::screenshare::HostNotice::Eof => {
let _ = fault_tx.send(generation);
break;
}
}
}
}
async fn run_core_loop(
mut reliable_rx: mpsc::UnboundedReceiver<CoreCommand>,
coalesce: CoalesceStore,
@@ -1393,10 +1438,29 @@ async fn run_core_loop(
// later opt-in can immediately publish whatever is currently running.
let mut current_game: Option<crate::game::DetectedGame> = None;
let mut network_mode = NetworkMode::default();
// Pixelpass binary override (config), and the ticket of our own active screen
// share (rides our presence so the room — incl. late joiners — can watch).
// Pixelpass binary override (config), and our own active screen share: the
// ticket rides our presence so the room — incl. late joiners — can watch,
// and the generation ties host-fault notices to this specific host spawn
// (see `ScreenShareHostFault`). One variable on purpose: the ticket and the
// generation must appear and vanish together, or a stale fault could tear
// down a share it doesn't belong to.
let mut pixelpass_override: Option<String> = None;
let mut current_sharing: Option<String> = None;
// Capability results are meaningful only for the exact resolved binary
// path that produced them. A changed override/PATH resolution must be
// re-probed before any capability-gated argv is constructed.
let mut pixelpass_capabilities: Option<crate::screenshare::ProbedPixelpassCapabilities> = None;
struct ActiveShare {
generation: u64,
ticket: String,
}
let mut current_sharing: Option<ActiveShare> = None;
// Monotonic per-spawn counter feeding `ActiveShare::generation`.
let mut share_generations: u64 = 0;
// Host faults re-enter the loop here (the notice-forwarder task can't touch
// loop state). The loop keeps `host_fault_tx` to clone into each share's
// forwarder, so this channel never closes — the select arm's `Some` pattern
// is total in practice and a closed-channel branch would be unreachable.
let (host_fault_tx, mut host_fault_rx) = mpsc::unbounded_channel::<u64>();
let mut active_session: Option<ActiveSession> = None;
// Standalone capture-only mic meter, live only when no session exists.
@@ -1506,11 +1570,18 @@ async fn run_core_loop(
PING_INTERVAL,
);
ping_interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
let mut viewer_reap_interval = tokio::time::interval_at(
tokio::time::Instant::now() + VIEWER_REAP_INTERVAL,
VIEWER_REAP_INTERVAL,
);
viewer_reap_interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop {
let cmd = tokio::select! {
biased;
maybe_cmd = reliable_rx.recv() => match maybe_cmd {
Some(cmd) => cmd,
// Every `CoreController`/`CoreCommandSender` is gone — the UI has
// dropped the core. Teardown happens once, after the loop.
None => break,
},
maybe_wake = besteffort_wake_rx.recv() => match maybe_wake {
@@ -1529,8 +1600,33 @@ async fn run_core_loop(
None => continue,
}
}
// ⚠️ UNREACHABLE BY CONSTRUCTION, twice over — do not mistake this
// for a live teardown path (phase 0b finding, 2026-07-26):
// 1. this function owns `besteffort_wake_tx` (cloned at the
// `CoreController::new` spawn site, used just above for the
// `has_more` re-arm), so the channel can never close while
// this loop is running;
// 2. even without that, every holder of a wake sender —
// `CoreController` and `CoreCommandSender` — holds
// `reliable_tx` too, and the `biased` select polls that one
// first, so the reliable arm always wins the race to exit.
// Teardown is hoisted after the loop, so if this arm is ever made
// reachable it is already covered — nothing to add here.
None => break,
},
// A share's notice-forwarder task reported the host's stdout ended.
// The `Some` pattern is total: this loop owns `host_fault_tx` (see
// its declaration), so the channel cannot close — no `None` arm is
// written because one would be unreachable by construction.
Some(generation) = host_fault_rx.recv() => {
CoreCommand::ScreenShareHostFault { generation }
}
_ = viewer_reap_interval.tick() => {
if let Some(session) = &mut active_session {
session.teardown.sweep_exited_viewers();
}
continue;
}
game_change = next_game_change(&mut game_rx) => {
// The detector worker published a new debounced game (or `None`).
let Some(detected) = game_change else {
@@ -1549,7 +1645,7 @@ async fn run_core_loop(
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
net.endpoint.addr(),
current_sharing.clone(),
current_sharing.as_ref().map(|s| s.ticket.clone()),
);
let _ = session.room_state.update_self_state(self_state).await;
}
@@ -1693,6 +1789,13 @@ async fn run_core_loop(
net.file_router.clear();
*current_room.lock().unwrap() = None;
}
// Any advertised share died with that session — deliberately —
// so retire it HERE, before the invalid-ticket early exit below
// can skip it. Left populated, the killed host's stdout EOF
// would pass the ScreenShareHostFault staleness gate and
// surface as a spurious "ended unexpectedly" error on top of
// the ticket error (Gemini review of S2, P2-1).
current_sharing = None;
// If a network-mode / identity change was deferred while a call was
// active, rebuild the persistent stack now — after the old session is
@@ -1788,8 +1891,8 @@ async fn run_core_loop(
secret_key.clone(),
));
// Fresh join starts not sharing; clear any stale share ticket.
current_sharing = None;
// (The share was already retired beside the session teardown
// above; a fresh join starts not sharing.)
let self_state =
presence.to_state(is_muted.load(Ordering::Relaxed), endpoint.addr(), None);
@@ -2726,9 +2829,9 @@ async fn run_core_loop(
grace_timers,
transport: transport.clone(),
#[cfg(target_os = "linux")]
echo_cancel: echo_cancel_guard,
screenshare_host: None,
screenshare_viewers: Vec::<(String, tokio::process::Child)>::new(),
teardown: SessionTeardown::new(echo_cancel_guard),
#[cfg(not(target_os = "linux"))]
teardown: SessionTeardown::new(None),
};
let self_id = endpoint.id().to_string();
@@ -2810,8 +2913,11 @@ async fn run_core_loop(
is_muted.store(new_state, Ordering::Relaxed);
if let Some(session) = &active_session {
let self_state =
presence.to_state(new_state, net.endpoint.addr(), current_sharing.clone());
let self_state = presence.to_state(
new_state,
net.endpoint.addr(),
current_sharing.as_ref().map(|s| s.ticket.clone()),
);
let _ = session.room_state.update_self_state(self_state).await;
}
}
@@ -2824,7 +2930,7 @@ async fn run_core_loop(
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
net.endpoint.addr(),
current_sharing.clone(),
current_sharing.as_ref().map(|s| s.ticket.clone()),
);
let _ = session.room_state.update_self_state(self_state).await;
}
@@ -3137,7 +3243,7 @@ async fn run_core_loop(
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
net.endpoint.addr(),
current_sharing.clone(),
current_sharing.as_ref().map(|s| s.ticket.clone()),
);
let _ = session.room_state.update_self_state(self_state).await;
}
@@ -3343,7 +3449,7 @@ async fn run_core_loop(
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
net.endpoint.addr(),
current_sharing.clone(),
current_sharing.as_ref().map(|s| s.ticket.clone()),
);
let _ = session.room_state.update_self_state(self_state).await;
}
@@ -3363,39 +3469,53 @@ async fn run_core_loop(
CoreCommand::SetPixelpassPath(path) => {
pixelpass_override = path.filter(|p| !p.trim().is_empty());
pixelpass_capabilities = None;
}
CoreCommand::ListAudioApps => {
// Probe whether this pixelpass supports `--strict-audio` before
// offering per-app capture: an older binary would reject the flag
// and hard-fail the share (audit P2). When unsupported (or
// pixelpass is missing), skip enumeration and let the picker show
// whole-desktop audio only — never a best-effort `--app` that
// would reopen the A23 echo.
let app_audio_supported =
// Probe the versioned response from the exact binary selected
// for this picker. The help fallback can recover legacy strict
// per-app support, but never desktop exclusion.
let (capabilities, apps) =
match crate::screenshare::pixelpass_path(pixelpass_override.as_deref()) {
Some(bin) => crate::screenshare::supports_strict_audio(&bin).await,
None => false,
Some(bin) => {
let capabilities =
crate::screenshare::probe_pixelpass_capabilities(&bin).await;
pixelpass_capabilities =
Some(crate::screenshare::ProbedPixelpassCapabilities {
binary: bin,
capabilities,
});
let apps = if capabilities.strict_app_audio {
crate::screenshare::list_audio_apps().await
} else {
Vec::new()
};
(capabilities, apps)
}
None => {
pixelpass_capabilities = None;
(
crate::screenshare::PixelpassCapabilities::default(),
Vec::new(),
)
}
};
let apps = if app_audio_supported {
crate::screenshare::list_audio_apps().await
} else {
Vec::new()
};
let _ = ui_tx
.send(UiEvent::AudioAppsListed {
apps,
app_audio_supported,
app_audio_supported: capabilities.strict_app_audio,
desktop_audio_exclusion_supported: capabilities.desktop_audio_exclusion,
})
.await;
}
CoreCommand::StartScreenShare {
audio_app,
audio,
settings,
quality,
} => {
let Some(session) = &mut active_session else {
let Some(session) = active_session.as_ref() else {
let _ = ui_tx
.send(UiEvent::Error(
"Join a call before sharing your screen".into(),
@@ -3403,7 +3523,7 @@ async fn run_core_loop(
.await;
continue;
};
if session.screenshare_host.is_some() {
if session.teardown.is_sharing() {
continue; // already sharing
}
let bin = match crate::screenshare::pixelpass_path(pixelpass_override.as_deref()) {
@@ -3417,46 +3537,85 @@ async fn run_core_loop(
continue;
}
};
// Forward pixelpass `app_audio` events (only emitted when an app
// is selected) to the UI so it can warn when the chosen app's
// audio drops. The channel closes when the host dies (drain hits
// EOF), ending the forwarder task on its own.
let notices = audio_app.as_deref().map(|_| {
let (tx, mut rx) = tokio::sync::mpsc::unbounded_channel::<
crate::screenshare::PixelpassEvent,
>();
let ui_tx_notices = ui_tx.clone();
tokio::spawn(async move {
while let Some(ev) = rx.recv().await {
let active = match ev {
crate::screenshare::PixelpassEvent::AppAudioRouted => true,
crate::screenshare::PixelpassEvent::AppAudioLost => false,
_ => continue,
};
if ui_tx_notices
.send(UiEvent::ShareAudioActive(active))
.await
.is_err()
{
break;
}
// The picker probe is bound to its resolved binary. If the
// override/PATH now resolves elsewhere, immediately re-probe
// before constructing any capability-gated argv and fail closed
// when the selected feature is absent.
if !matches!(
audio,
crate::screenshare::ShareAudioSelection::DesktopShared
) {
let capabilities = crate::screenshare::capabilities_for_resolved_binary(
&bin,
&mut pixelpass_capabilities,
)
.await;
let unsupported = match &audio {
crate::screenshare::ShareAudioSelection::Application(_)
if !capabilities.strict_app_audio =>
{
Some(
"This PixelPass does not support strict per-app audio. Reopen the picker or update PixelPass.",
)
}
});
tx
crate::screenshare::ShareAudioSelection::DesktopExcluding
if !capabilities.desktop_audio_exclusion =>
{
Some(
"This PixelPass does not support desktop audio exclusion. Update PixelPass or choose another audio source.",
)
}
_ => None,
};
if let Some(message) = unsupported {
let _ = ui_tx.send(UiEvent::Error(message.into())).await;
continue;
}
}
let session = active_session
.as_mut()
.expect("session presence checked before PixelPass probe");
let aec_module_index = matches!(
audio,
crate::screenshare::ShareAudioSelection::DesktopExcluding
)
.then(|| {
session
.teardown
.echo_cancel()
.map(crate::audio::echo_cancel::EchoCancelGuard::module_index)
})
.flatten();
// Every share gets a notice forwarder — not just app-audio ones.
// App-audio and desktop-exclusion events become UI state, and
// the drain's terminal `Eof` becomes a generation-scoped fault.
share_generations += 1;
let generation = share_generations;
let (notices_tx, notices_rx) =
tokio::sync::mpsc::unbounded_channel::<crate::screenshare::HostNotice>();
let ui_tx_notices = ui_tx.clone();
let fault_tx = host_fault_tx.clone();
tokio::spawn(async move {
forward_host_notices(notices_rx, ui_tx_notices, fault_tx, generation).await;
});
match crate::screenshare::spawn_host(
&bin,
audio_app.as_deref(),
&audio,
aec_module_index,
&settings,
quality,
notices,
notices_tx,
)
.await
{
Ok((child, ticket)) => {
crate::log_msg("Screen share host started");
session.screenshare_host = Some(child);
current_sharing = Some(ticket.clone());
session.teardown.set_host(child);
current_sharing = Some(ActiveShare {
generation,
ticket: ticket.clone(),
});
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
net.endpoint.addr(),
@@ -3476,9 +3635,24 @@ async fn run_core_loop(
CoreCommand::StopScreenShare => {
current_sharing = None;
if let Some(session) = &mut active_session {
if let Some(mut child) = session.screenshare_host.take() {
let _ = child.kill().await;
crate::log_msg("Screen share host stopped");
match session.teardown.stop_host().await {
None => {}
Some(teardown::StopOutcome::Reaped) => {
crate::log_msg("Screen share host stopped");
}
// We gave up waiting rather than freeze the client, so
// pixelpass may still be alive and serving. Saying
// "stopped" and nothing else would be a lie the user
// cannot see through (round-16 review, P3-2).
Some(teardown::StopOutcome::Unconfirmed) => {
let _ = ui_tx
.send(UiEvent::Error(
"Couldn't confirm the screen-share process exited — \
it may still be sharing. Check for a stray pixelpass."
.into(),
))
.await;
}
}
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
@@ -3490,6 +3664,66 @@ async fn run_core_loop(
let _ = ui_tx.send(UiEvent::ScreenShareStopped).await;
}
CoreCommand::ScreenShareHostFault { generation } => {
// Stale unless it names the share we are advertising RIGHT NOW.
// Every deliberate end of a share (StopScreenShare, Leave, a
// fresh Join) clears `current_sharing` before or while reaping
// the child, and the reaped child's stdout EOF then arrives
// here late — dropping it is the correct handling, not an edge
// case. A mismatched generation likewise: that fault belongs to
// an older spawn than the share now running.
let stale = current_sharing.as_ref().map(|s| s.generation) != Some(generation);
if stale {
continue;
}
crate::log_msg(
"Screen share host died (stdout EOF with the share still advertised)",
);
current_sharing = None;
// Pull the ticket off presence FIRST, before the reap: if the
// child only closed stdout and lives on, `stop_host` burns the
// full stop grace before the SIGKILL fallback, and for that
// whole window peers would still see (and click Watch on) a
// share whose host is already gone (Gemini S2-merge review,
// P2-1).
if let Some(session) = &mut active_session {
let self_state = presence.to_state(
is_muted.load(Ordering::Relaxed),
net.endpoint.addr(),
None,
);
let _ = session.room_state.update_self_state(self_state).await;
}
// Stopped next — it clears the UI's sharing state — so the
// local UI also stops saying "sharing" before the reap wait,
// and the error explaining why comes only after, so the user
// is never left looking at a "sharing" UI with an error
// beside it.
let _ = ui_tx.send(UiEvent::ScreenShareStopped).await;
let mut unconfirmed = false;
if let Some(session) = &mut active_session {
// The child is usually already dead, so this confirms the
// reap immediately; if it merely closed stdout and lives
// on, this is the SIGINT → grace → SIGKILL path. Either
// way the dead-or-dying child leaves the teardown slot, so
// `is_sharing` stops lying.
unconfirmed = matches!(
session.teardown.stop_host().await,
Some(teardown::StopOutcome::Unconfirmed)
);
}
let detail = if unconfirmed {
" Its process also couldn't be confirmed dead — check for a stray pixelpass."
} else {
""
};
let _ = ui_tx
.send(UiEvent::Error(format!(
"Screen share ended unexpectedly — pixelpass exited.{detail}"
)))
.await;
}
CoreCommand::ViewShare { ticket, settings } => {
let bin = match crate::screenshare::pixelpass_path(pixelpass_override.as_deref()) {
Some(b) => b,
@@ -3505,16 +3739,12 @@ async fn run_core_loop(
if let Some(session) = &mut active_session {
// Drop viewers whose player window has already closed so the
// list only tracks live players.
session
.screenshare_viewers
.retain_mut(|(_, child)| !matches!(child.try_wait(), Ok(Some(_))));
session.teardown.sweep_exited_viewers();
// One player per share: a second Watch click on a share we're
// already viewing is a retry (usually because the first window
// froze), so replace the existing player rather than stacking a
// second mpv — two players would double the shared audio.
if let Some(pos) = replace_viewer_index(&session.screenshare_viewers, &ticket) {
let (_, mut old) = session.screenshare_viewers.remove(pos);
let _ = old.kill().await;
if session.teardown.replace_viewer(&ticket).await {
crate::log_msg("Screen share viewer replaced (re-watch)");
}
}
@@ -3522,7 +3752,7 @@ async fn run_core_loop(
Ok(child) => {
crate::log_msg("Screen share viewer started");
if let Some(session) = &mut active_session {
session.screenshare_viewers.push((ticket, child));
session.teardown.push_viewer(ticket, child);
}
}
Err(e) => {
@@ -3535,6 +3765,21 @@ async fn run_core_loop(
}
}
// The command loop has exited, by any route. Tear the session down
// explicitly rather than letting it drop on the way out of this function:
// an implicit drop unloads the echo-cancel module without first reaping the
// pixelpass host (design v3.4 §7.2, decision D4).
//
// This sits *after* the loop rather than in the close arm on purpose. The
// impl plan pinned one teardown per channel-close arm, but the best-effort
// wake arm is unreachable by construction (see the comment at that arm), so
// that shape would have duplicated teardown to cover one live path and one
// dead one. Here every `break` is covered structurally, including any added
// later. Adjudication: impl plan §10, 2026-07-26.
if let Some(session) = active_session.take() {
session.shutdown(audio_backend.clone()).await;
}
Ok(())
}
@@ -3553,10 +3798,11 @@ mod tests {
KnownPeers, MAX_OPUS_PAYLOAD, MAX_RETAINED_PEERS, MIC_LEVEL_REPORT_SAMPLES, MicLevelMeter,
NetworkMode, PLAYBACK_HANDOFF_QUEUE_FRAMES, PeerSpeakTicket, admit_retained,
apply_peer_volume, apply_volume, audio_datagram_len_ok, coalesce_insert, coalesce_pop,
frame_level, mix_frames, mix_stereo_frames, next_game_change, rebuild_with_fallback,
replace_viewer_index, send_playback_frame, should_auto_fetch, stereo_to_mono,
forward_host_notices, frame_level, mix_frames, mix_stereo_frames, next_game_change,
rebuild_with_fallback, replace_viewer_index, send_playback_frame, should_auto_fetch,
stereo_to_mono,
};
use crate::core::messages::{CoalesceKey, CoreCommand, coalesce_key};
use crate::core::messages::{CoalesceKey, CoreCommand, UiEvent, coalesce_key};
use std::collections::{HashMap, HashSet};
use std::sync::mpsc::sync_channel;
use std::time::Duration;
@@ -3565,6 +3811,65 @@ mod tests {
iroh::SecretKey::generate().public()
}
#[tokio::test]
async fn all_exclusion_events_causally_cross_the_host_notice_channel_to_ui() {
use crate::screenshare::{AudioExclusionStatus, HostNotice, parse_pixelpass_event};
let lines = [
r#"{"event":"stream_unsupported","version":1,"stream_serial":4294967303,"reason":"port-exclusive"}"#,
r#"{"event":"stream_status_cleared","version":1,"stream_serial":4294967303}"#,
r#"{"event":"aec_failed","version":1,"module_index":536870919}"#,
r#"{"event":"aec_revoked","version":1,"module_index":536870919}"#,
r#"{"event":"foreign_aec_warning","version":1,"link_group":"echo-cancel-9999-13"}"#,
];
let (notice_tx, notice_rx) = tokio::sync::mpsc::unbounded_channel();
for line in lines {
notice_tx
.send(HostNotice::Event(
parse_pixelpass_event(line).expect("PixelPass event must parse"),
))
.unwrap();
}
drop(notice_tx);
let (ui_tx, mut ui_rx) = tokio::sync::mpsc::channel(8);
let (fault_tx, mut fault_rx) = tokio::sync::mpsc::unbounded_channel();
forward_host_notices(notice_rx, ui_tx, fault_tx, 17).await;
let mut statuses = Vec::new();
while let Some(event) = ui_rx.recv().await {
match event {
UiEvent::ShareAudioExclusionStatus(status) => statuses.push(status),
other => panic!("unexpected forwarded UI event: {other:?}"),
}
}
assert_eq!(
statuses,
vec![
AudioExclusionStatus::StreamUnsupported {
stream_serial: 4_294_967_303,
reason: "port-exclusive".to_string(),
},
AudioExclusionStatus::StreamStatusCleared {
stream_serial: 4_294_967_303,
},
AudioExclusionStatus::AecFailed {
module_index: 536_870_919,
},
AudioExclusionStatus::AecRevoked {
module_index: 536_870_919,
},
AudioExclusionStatus::ForeignAecWarning {
link_group: "echo-cancel-9999-13".to_string(),
},
]
);
assert!(
fault_rx.try_recv().is_err(),
"ordinary status events must not synthesize a host fault"
);
}
#[test]
fn re_watch_replaces_existing_viewer_for_same_ticket() {
// The value type stands in for a viewer Child; only the ticket matters.
+892
View File
@@ -0,0 +1,892 @@
//! Destruction-order guarantees for the screen-share children and the
//! echo-cancel module (phase 0b of the screenshare audio-exclusion plan;
//! design v3.4 §7.1–§7.2, decision D4).
//!
//! # The invariant
//!
//! > **The echo-cancel module must not unload while a pixelpass host is alive
//! > and fanning out.**
//!
//! If it does, the AEC's virtual nodes vanish from under a live pixelpass that
//! still holds link proxies and a stale module index. Phase 6 makes this sharp
//! — it is the first phase whose objects live only as long as pixelpass does —
//! so the ordering guarantee has to exist *before* it.
//!
//! Two paths have to honour it, and only one of them is code we get to run:
//!
//! 1. **The explicit path** — [`ScreenshareTeardown::shutdown_children`], awaited
//! by `ActiveSession::shutdown` before the guard is dropped.
//! 2. **The drop/unwind path** — nobody calls anything. The core has numerous
//! `unwrap()` sites and no `panic=abort` profile, so unwind is reachable, and
//! on that path the only thing standing between us and a violated invariant
//! is *field declaration order* plus [`ReapOnDrop`].
//!
//! Hence the two structural rules enforced here:
//!
//! - `echo_cancel` is the **last declared field** of [`ScreenshareTeardown`].
//! Rust drops fields in declaration order, so last-declared is last-dropped.
//! This is not a style choice; reversing it reintroduces the bug.
//! - Killing is not enough — a child must be **reaped**. `kill_on_drop(true)`
//! only *signals*; it hands the child to the runtime's orphan queue and
//! returns, which on an unwinding runtime may never be drained. [`ReapOnDrop`]
//! therefore blocks, briefly and boundedly, until the child is actually gone.
//!
//! Everything here is generic over [`ChildProcess`] and over the guard type so
//! the ordering is unit-testable without spawning processes or loading PipeWire
//! modules — the same seam idiom as `replace_viewer_index` and
//! `rebuild_with_fallback` in the parent module.
use std::future::Future;
use std::time::{Duration, Instant};
/// How long [`ReapOnDrop::drop`] will block waiting for a killed child to be
/// reaped before giving up and logging. This runs on the unwind path, so it is
/// a deliberate trade: a bounded stall is preferable to unloading the AEC out
/// from under a live pixelpass, and unbounded blocking in a `Drop` is not.
const REAP_BUDGET: Duration = Duration::from_millis(250);
/// Poll interval while waiting out [`REAP_BUDGET`].
const REAP_POLL: Duration = Duration::from_millis(5);
/// How long a child gets to honour the graceful stop before it is killed.
///
/// A healthy pixelpass exits in well under this, so the normal path never
/// spends it; only a wedged child does. It is awaited inline in the core
/// command loop, so it is also how long a wedged child can delay other
/// commands — hence seconds, not tens of seconds.
const STOP_GRACE: Duration = Duration::from_secs(2);
/// The child-process operations the teardown ordering actually depends on.
///
/// Deliberately narrow, and deliberately not `ExitStatus`-shaped: the ordering
/// rules care only about *whether* a child has been signalled and *whether* it
/// has been reaped, so the test double is a few lines instead of a fabricated
/// exit status.
pub(super) trait ChildProcess {
/// Ask the child to exit **gracefully**, so it can run its own cleanup.
/// Does **not** wait, and is not guaranteed to be honoured.
fn request_stop(&mut self) -> std::io::Result<()>;
/// Signal the child to die. Does **not** wait.
fn start_kill(&mut self) -> std::io::Result<()>;
/// Poll once. `true` once the child has exited **and been reaped**.
fn try_reap(&mut self) -> bool;
/// Wait until the child has exited and been reaped.
///
/// The `io::Result` is load-bearing and must not be discarded by callers:
/// a failed wait is *not* a confirmed reap, and treating it as one is how
/// the AEC ends up unloading over a live child.
fn wait_reaped(&mut self) -> impl Future<Output = std::io::Result<()>> + Send;
}
impl ChildProcess for tokio::process::Child {
/// **SIGINT, not SIGTERM.** pixelpass installs only a `tokio::signal::ctrl_c()`
/// handler (`pixelpass/src/common/signal.rs`), so SIGTERM would be the default
/// disposition — instant death, no cleanup — which is indistinguishable from
/// SIGKILL for our purposes.
///
/// Signalling by pid is safe against pid reuse here because we have not
/// reaped this child: an exited-but-unreaped child is a zombie whose pid the
/// kernel reserves until we `wait` it, so the pid cannot name a stranger.
#[cfg(unix)]
fn request_stop(&mut self) -> std::io::Result<()> {
let Some(pid) = self.id() else {
// Already reaped — nothing to signal.
return Ok(());
};
// SAFETY: `kill` is async-signal-safe and takes no pointers; the pid is
// this process's own unreaped child (see above).
if unsafe { libc::kill(pid as libc::pid_t, libc::SIGINT) } == 0 {
Ok(())
} else {
Err(std::io::Error::last_os_error())
}
}
/// Windows has no SIGINT to send to another process without attaching to its
/// console, so the graceful request degrades to the hard kill and the
/// bounded wait below simply returns early.
#[cfg(not(unix))]
fn request_stop(&mut self) -> std::io::Result<()> {
tokio::process::Child::start_kill(self)
}
fn start_kill(&mut self) -> std::io::Result<()> {
tokio::process::Child::start_kill(self)
}
fn try_reap(&mut self) -> bool {
matches!(self.try_wait(), Ok(Some(_)))
}
async fn wait_reaped(&mut self) -> std::io::Result<()> {
self.wait().await.map(|_| ())
}
}
/// Did the explicit stop path actually confirm the child was reaped?
///
/// The distinction is not cosmetic: on [`Unconfirmed`](Self::Unconfirmed) we
/// deliberately stopped waiting (see [`ReapOnDrop::shutdown`]), so pixelpass may
/// still be alive and fanning out. A user-initiated Stop Share must not report
/// that as a clean stop.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[must_use = "an unconfirmed stop means the child may still be sharing"]
pub(super) enum StopOutcome {
/// The child is gone and has been reaped.
Reaped,
/// We could not confirm the reap within the bound and gave up waiting.
Unconfirmed,
}
/// A child that is killed **and reaped** when it is dropped.
///
/// The explicit path calls [`shutdown`](Self::shutdown), which releases the
/// child only once its reap is *confirmed*, so the `Drop` below is a no-op
/// afterwards but stays armed through every await until then. `Drop` is the
/// last-ditch protection for the panic/unwind/cancellation paths.
pub(super) struct ReapOnDrop<C: ChildProcess> {
/// `None` once the child has been reaped through the explicit path.
child: Option<C>,
/// Names the child in the reap-timeout log line.
label: &'static str,
}
impl<C: ChildProcess> ReapOnDrop<C> {
pub(super) fn new(child: C, label: &'static str) -> Self {
Self {
child: Some(child),
label,
}
}
/// Poll once, without killing. `true` if the child has exited on its own —
/// used to sweep player windows the user has already closed.
pub(super) fn has_exited(&mut self) -> bool {
match &mut self.child {
Some(child) => {
if child.try_reap() {
self.child = None;
true
} else {
false
}
}
// Already reaped through the explicit path.
None => true,
}
}
/// Stop the child gracefully if it will go, and by force if it will not.
/// Waits for it to be reaped either way. Idempotent.
///
/// Ask, then insist (design v3.4 §7.4): a pixelpass host that gets SIGINT
/// unloads its capture sink on the way out, whereas SIGKILL skips that and
/// leaks a null-sink module on every Stop Share.
///
/// The wait is the point: returning after signalling would let the caller
/// proceed to unload the AEC while the child is still running.
///
/// ⚠️ The child stays owned by `self` across every `.await`, and is released
/// **only after a confirmed reap**. Taking it out first would disarm the
/// `Drop` fallback for exactly as long as the wait lasts: cancel or unwind
/// this future at that moment and the raw child would drop with nothing but
/// `kill_on_drop` (which signals without reaping) while `Drop` below found
/// `None` and did nothing — the precise hole this type exists to close.
pub(super) async fn shutdown(&mut self) -> StopOutcome {
let Some(child) = self.child.as_mut() else {
return StopOutcome::Reaped;
};
// Three different things can go wrong here and they want three
// different operator diagnoses: the signal never left (a runtime or
// permission fault), the child ignored it (a wedged pixelpass), or the
// wait itself broke (we no longer know anything about the child).
// Collapsing them into one line was P3-1 of the round-16 review.
if let Err(e) = child.request_stop() {
crate::log_msg(&format!(
"teardown: could not ask {} to stop: {e}",
self.label
));
}
match tokio::time::timeout(STOP_GRACE, child.wait_reaped()).await {
Ok(Ok(())) => {
self.child = None;
return StopOutcome::Reaped;
}
Ok(Err(e)) => crate::log_msg(&format!(
"teardown: waiting for {} failed ({e}); killing it",
self.label
)),
Err(_) => crate::log_msg(&format!(
"teardown: {} ignored the graceful stop within {STOP_GRACE:?}; killing it",
self.label
)),
}
if let Err(e) = child.start_kill() {
crate::log_msg(&format!(
"teardown: {} could not be killed: {e}",
self.label
));
}
// The second wait is bounded too. An unbounded one lets a process stuck
// in uninterruptible sleep wedge the core command loop forever, and a
// permanently frozen app is a worse failure than the risk below.
if let Ok(Ok(())) = tokio::time::timeout(STOP_GRACE, child.wait_reaped()).await {
self.child = None;
return StopOutcome::Reaped;
}
// Explicit policy for the one case where the two guarantees conflict:
// we could not confirm the reap and will NOT block indefinitely, so we
// give up availability-first and leave the child owned — `Drop`'s
// bounded retry stays armed, and the AEC may unload over a child that
// is still somehow alive. That residual risk is logged, not silent —
// and, for a user-initiated stop, reported to the caller rather than
// dressed up as success.
crate::log_msg(&format!(
"teardown: {} could not be confirmed dead; the echo-cancel module \
may unload while it lives",
self.label
));
StopOutcome::Unconfirmed
}
/// Is the `Drop` fallback still armed? Test-only: the arming rule is the
/// whole point of holding the child across the waits.
#[cfg(test)]
fn is_armed(&self) -> bool {
self.child.is_some()
}
}
impl<C: ChildProcess> Drop for ReapOnDrop<C> {
fn drop(&mut self) {
let Some(child) = self.child.as_mut() else {
return;
};
let _ = child.start_kill();
// `Drop` cannot await, so poll on a bounded budget. See `REAP_BUDGET`.
let deadline = Instant::now() + REAP_BUDGET;
loop {
if child.try_reap() {
return;
}
if Instant::now() >= deadline {
crate::log_msg(&format!(
"teardown: {} did not exit within the reap budget; \
continuing (the echo-cancel module may unload while it lives)",
self.label
));
return;
}
std::thread::sleep(REAP_POLL);
}
}
}
/// Everything in an `ActiveSession` whose **destruction order** is load-bearing.
///
/// ⚠️ Field order below **is** the invariant. `echo_cancel` is declared last so
/// it is dropped last, after every screen-share child has been killed and
/// reaped. Do not reorder these fields.
pub(super) struct ScreenshareTeardown<C: ChildProcess, G> {
/// Our pixelpass screen-share host child while sharing.
host: Option<ReapOnDrop<C>>,
/// pixelpass viewer children we spawned to watch peers' shares, each paired
/// with the share ticket it is viewing so a re-watch of the same share can
/// replace (not stack) its player.
viewers: Vec<(String, ReapOnDrop<C>)>,
/// Loaded PipeWire echo-cancel module (if enabled); unloads on drop.
///
/// ⚠️ **LAST FIELD ON PURPOSE** — see the module docs and the struct note.
///
echo_cancel: Option<G>,
}
impl<C: ChildProcess, G> ScreenshareTeardown<C, G> {
pub(super) fn new(echo_cancel: Option<G>) -> Self {
Self {
host: None,
viewers: Vec::new(),
echo_cancel,
}
}
pub(super) fn is_sharing(&self) -> bool {
self.host.is_some()
}
/// Borrow the session-owned AEC guard without disturbing its load-bearing
/// last-field drop order. Phase 8 uses this only to pass the module identity
/// to PixelPass while the guard remains owned here.
pub(super) fn echo_cancel(&self) -> Option<&G> {
self.echo_cancel.as_ref()
}
pub(super) fn set_host(&mut self, child: C) {
self.host = Some(ReapOnDrop::new(child, "screen-share host"));
}
/// Stop sharing: kill the host and wait for it to be reaped. `None` if we
/// were not sharing; otherwise whether the reap was actually confirmed —
/// the caller owns telling the user, since an unconfirmed stop may leave
/// pixelpass fanning out after the UI says sharing ended.
pub(super) async fn stop_host(&mut self) -> Option<StopOutcome> {
let mut host = self.host.take()?;
Some(host.shutdown().await)
}
/// Drop viewers whose player window has already closed, so the list only
/// tracks live players.
pub(super) fn sweep_exited_viewers(&mut self) {
self.viewers.retain_mut(|(_, child)| !child.has_exited());
}
/// Kill and reap the viewer already showing `ticket`, if any, so a re-watch
/// replaces its player instead of stacking a second one.
pub(super) async fn replace_viewer(&mut self, ticket: &str) -> bool {
let Some(pos) = super::replace_viewer_index(&self.viewers, ticket) else {
return false;
};
let (_, mut old) = self.viewers.remove(pos);
// A viewer is our own player window, not the thing peers are watching:
// an unconfirmed reap is already logged, and there is no user decision
// riding on it the way there is for Stop Share.
let _ = old.shutdown().await;
true
}
pub(super) fn push_viewer(&mut self, ticket: String, child: C) {
self.viewers
.push((ticket, ReapOnDrop::new(child, "screen-share viewer")));
}
/// Kill and reap **every** screen-share child, host first so viewers see the
/// stream end promptly.
///
/// The caller must await this before the echo-cancel guard is dropped. On
/// the drop/unwind path nothing calls it and field order carries the
/// invariant instead.
pub(super) async fn shutdown_children(&mut self) {
// Outcomes are discarded on purpose: this runs on the session/teardown
// path, where the policy is already availability-first and the residual
// risk is logged by `shutdown` itself. There is no user still waiting
// on an answer here, unlike `stop_host`.
if let Some(host) = &mut self.host {
let _ = host.shutdown().await;
}
self.host = None;
for (_, viewer) in self.viewers.iter_mut() {
let _ = viewer.shutdown().await;
}
self.viewers.clear();
}
}
#[cfg(test)]
mod tests {
use super::{ChildProcess, ReapOnDrop, STOP_GRACE, ScreenshareTeardown, StopOutcome};
use std::future::Future;
use std::sync::{Arc, Mutex};
use std::time::Duration;
type Log = Arc<Mutex<Vec<String>>>;
fn log() -> Log {
Arc::new(Mutex::new(Vec::new()))
}
fn entries(log: &Log) -> Vec<String> {
log.lock().unwrap().clone()
}
fn position(log: &Log, entry: &str) -> Option<usize> {
entries(log).iter().position(|e| e == entry)
}
/// Records the events the ordering rules turn on. Death is gated on an
/// actual signal, so the double cannot report a reap that nothing caused.
struct FakeChild {
log: Log,
label: &'static str,
interrupted: bool,
killed: bool,
reaped: bool,
/// A well-behaved child exits on SIGINT. A wedged one ignores it and
/// dies only to SIGKILL.
honours_interrupt: bool,
/// When true the child is already dead before anyone signals it — the
/// closed-player-window case that `sweep_exited_viewers` looks for.
exited_on_its_own: bool,
/// Death is not instantaneous: `try_reap` reports the child alive this
/// many more times before it goes.
polls_before_death: u32,
/// `wait` reports an error instead of a reap.
wait_fails: bool,
}
impl FakeChild {
/// A well-behaved child: exits when asked.
fn new(log: &Log, label: &'static str) -> Self {
Self {
log: log.clone(),
label,
interrupted: false,
killed: false,
reaped: false,
honours_interrupt: true,
exited_on_its_own: false,
polls_before_death: 0,
wait_fails: false,
}
}
/// A child that ignores the graceful stop entirely.
fn wedged(log: &Log, label: &'static str) -> Self {
Self {
honours_interrupt: false,
..Self::new(log, label)
}
}
/// A child that does not die the instant it is signalled: `try_reap`
/// reports it alive for `polls` calls first. Without this the `Drop`
/// polling loop could be replaced by a single `try_reap` and no test
/// would notice.
fn reaps_after_polls(log: &Log, label: &'static str, polls: u32) -> Self {
Self {
polls_before_death: polls,
..Self::new(log, label)
}
}
/// A child that ignores SIGINT *and* does not die the instant it is
/// killed — the only shape that lets a test reach the post-SIGKILL
/// wait and still be reaped by the `Drop` poll loop afterwards.
fn wedged_then_dies_after_polls(log: &Log, label: &'static str, polls: u32) -> Self {
Self {
honours_interrupt: false,
polls_before_death: polls,
..Self::new(log, label)
}
}
/// A child whose `wait` fails. A failed wait is not a confirmed reap,
/// so it must not be reported as one.
fn wait_fails(log: &Log, label: &'static str) -> Self {
Self {
wait_fails: true,
..Self::new(log, label)
}
}
fn already_exited(log: &Log, label: &'static str) -> Self {
Self {
exited_on_its_own: true,
..Self::new(log, label)
}
}
/// Has anything actually made this child exit yet? A signalled child
/// still has to burn through `polls_before_death` first.
fn is_dead(&self) -> bool {
let signalled = self.killed
|| self.exited_on_its_own
|| (self.interrupted && self.honours_interrupt);
signalled && self.polls_before_death == 0
}
/// One observation of a dying-but-not-yet-dead child.
fn tick(&mut self) {
self.polls_before_death = self.polls_before_death.saturating_sub(1);
}
fn record(&self, event: &str) {
self.log
.lock()
.unwrap()
.push(format!("{}:{event}", self.label));
}
fn mark_reaped(&mut self) {
if !self.reaped {
self.reaped = true;
self.record("reap");
}
}
}
impl ChildProcess for FakeChild {
fn request_stop(&mut self) -> std::io::Result<()> {
if !self.interrupted {
self.interrupted = true;
self.record("sigint");
}
Ok(())
}
fn start_kill(&mut self) -> std::io::Result<()> {
if !self.killed {
self.killed = true;
self.record("kill");
}
Ok(())
}
fn try_reap(&mut self) -> bool {
if self.is_dead() {
self.mark_reaped();
return true;
}
self.tick();
false
}
/// Pending until something actually kills the child, so a wedged child
/// really does make the caller wait out `STOP_GRACE`. No waker is
/// registered: under `start_paused` the runtime auto-advances its clock
/// when every task is idle, which is exactly what fires the timeout.
fn wait_reaped(&mut self) -> impl Future<Output = std::io::Result<()>> + Send {
std::future::poll_fn(move |_cx| {
if self.wait_fails {
return std::task::Poll::Ready(Err(std::io::Error::other("wait failed")));
}
if self.is_dead() {
self.mark_reaped();
std::task::Poll::Ready(Ok(()))
} else {
std::task::Poll::Pending
}
})
}
}
/// Stands in for `EchoCancelGuard`, whose real `Drop` runs `pactl unload`.
struct FakeAec(Log);
impl Drop for FakeAec {
fn drop(&mut self) {
self.0.lock().unwrap().push("aec:unload".to_string());
}
}
fn teardown(log: &Log) -> ScreenshareTeardown<FakeChild, FakeAec> {
ScreenshareTeardown::new(Some(FakeAec(log.clone())))
}
// --- The drop/unwind path: field order + ReapOnDrop carry the invariant ---
/// Mutation gate #5 (remove the reap loop from `ReapOnDrop::drop`).
///
/// Asserts only that dropping a guard reaps, and reaps *after* killing —
/// deliberately says nothing about the AEC, so reversing the struct's field
/// order leaves this test green and only the ordering test below fails.
#[test]
fn dropping_a_guard_kills_and_then_reaps_the_child() {
let log = log();
drop(ReapOnDrop::new(FakeChild::new(&log, "host"), "host"));
assert_eq!(entries(&log), vec!["host:kill", "host:reap"]);
}
/// Mutation gate #4 (reverse the field order of `ScreenshareTeardown`).
///
/// Asserts only kill-before-unload, so removing the reap loop leaves this
/// test green and only the reap test above fails.
#[test]
fn the_aec_unloads_after_the_children_on_the_drop_path() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
t.push_viewer("ticket-A".to_string(), FakeChild::new(&log, "viewer"));
drop(t);
let unload = position(&log, "aec:unload").expect("the AEC guard must be dropped");
let host_kill = position(&log, "host:kill").expect("the host must be killed");
let viewer_kill = position(&log, "viewer:kill").expect("the viewer must be killed");
assert!(
host_kill < unload,
"the AEC unloaded while the host was alive: {:?}",
entries(&log)
);
assert!(
viewer_kill < unload,
"the AEC unloaded while a viewer was alive: {:?}",
entries(&log)
);
}
/// The whole invariant in one sequence, as documentation.
#[test]
fn the_drop_path_reaps_every_child_before_unloading_the_aec() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
drop(t);
assert_eq!(entries(&log), vec!["host:kill", "host:reap", "aec:unload"]);
}
// --- The explicit path: ask, then insist ---
/// A healthy child must be *asked*, never killed. If Stop Share went
/// straight to SIGKILL, pixelpass would skip its own cleanup and leak a
/// null-sink module every time (design v3.4 §7.4).
#[tokio::test]
async fn a_healthy_child_is_asked_to_stop_and_never_killed() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
assert_eq!(t.stop_host().await, Some(StopOutcome::Reaped));
assert_eq!(entries(&log), vec!["host:sigint", "host:reap"]);
assert!(
!entries(&log).contains(&"host:kill".to_string()),
"a child that honoured the graceful stop must not be killed: {:?}",
entries(&log)
);
}
/// ...but a child that ignores the request must not be able to hold the
/// session open forever: the grace is bounded and SIGKILL follows.
#[tokio::test(start_paused = true)]
async fn a_wedged_child_is_killed_once_the_grace_expires() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::wedged(&log, "host"));
// The outer bound turns "the fallback was removed" into a failure
// rather than a hung test. Under `start_paused` no real time passes.
let start = tokio::time::Instant::now();
tokio::time::timeout(Duration::from_secs(60), t.stop_host())
.await
.expect("a wedged child must not block teardown indefinitely");
assert_eq!(entries(&log), vec!["host:sigint", "host:kill", "host:reap"]);
assert!(
start.elapsed() >= STOP_GRACE,
"the child must actually be given the grace period, waited {:?}",
start.elapsed()
);
}
/// The assertion above compares elapsed time against `STOP_GRACE` itself,
/// so it stays vacuously true if the constant is set to zero — both sides
/// move together. Pin the constant independently: the whole point of the
/// graceful stop is that pixelpass gets a real interval in which to unload
/// its capture sink, and zero is not one.
#[test]
fn the_grace_is_a_real_interval() {
assert!(
STOP_GRACE >= Duration::from_millis(500),
"too short to let pixelpass tear its pipeline down: {STOP_GRACE:?}"
);
// ...and short enough that a wedged child cannot visibly stall the core
// command loop, which awaits this inline.
assert!(
STOP_GRACE <= Duration::from_secs(5),
"long enough to freeze the UI's command handling: {STOP_GRACE:?}"
);
}
/// The hole the whole type exists to close, and the one place the old
/// implementation left open: if `shutdown` is cancelled while waiting, the
/// child must still be owned, so dropping the guard still kills and reaps.
#[tokio::test(start_paused = true)]
async fn cancelling_shutdown_mid_wait_leaves_the_fallback_armed() {
let log = log();
let mut guard = ReapOnDrop::new(FakeChild::wedged(&log, "host"), "host");
// Cancel well inside the grace, while it is still waiting.
assert!(
tokio::time::timeout(STOP_GRACE / 4, guard.shutdown())
.await
.is_err(),
"the wedged child should still have been waiting when we cancelled"
);
assert!(
guard.is_armed(),
"a cancelled shutdown must not disarm the drop fallback"
);
drop(guard);
assert_eq!(entries(&log), vec!["host:sigint", "host:kill", "host:reap"]);
}
/// The test above only ever cancels during the *graceful* wait, so a
/// mutation that disarmed the wrapper between the two waits would survive
/// it (round-16 review, P3-3). This one cancels during the post-SIGKILL
/// wait — the window where we have already given up on cooperation and the
/// `Drop` fallback is the only thing left.
#[tokio::test(start_paused = true)]
async fn cancelling_shutdown_after_the_kill_leaves_the_fallback_armed() {
let log = log();
// Ignores SIGINT, so the grace expires and we reach the kill; then
// survives three polls, so the second wait is still pending when we
// cancel, and the drop loop still gets to reap it.
let mut guard = ReapOnDrop::new(
FakeChild::wedged_then_dies_after_polls(&log, "host", 3),
"host",
);
assert!(
tokio::time::timeout(STOP_GRACE + STOP_GRACE / 4, guard.shutdown())
.await
.is_err(),
"we should have been cancelled inside the post-kill wait"
);
assert_eq!(
entries(&log),
vec!["host:sigint", "host:kill"],
"the graceful stop must have expired and escalated before we cancelled"
);
assert!(
guard.is_armed(),
"cancelling after the kill must not disarm the drop fallback either"
);
drop(guard);
// The fake's `start_kill` is idempotent, so `Drop` re-signalling an
// already-killed child adds no entry; the *reap* is what proves the
// fallback ran to completion after we abandoned the wait.
assert_eq!(
entries(&log),
vec!["host:sigint", "host:kill", "host:reap"],
"Drop must poll until the child is actually gone"
);
}
/// A failed wait is not a reap. Reporting it as one is how the AEC ends up
/// unloading over a child that is still alive.
#[tokio::test(start_paused = true)]
async fn a_failed_wait_is_not_treated_as_a_confirmed_reap() {
let log = log();
let mut guard = ReapOnDrop::new(FakeChild::wait_fails(&log, "host"), "host");
assert_eq!(
guard.shutdown().await,
StopOutcome::Unconfirmed,
"a stop we could not confirm must not be reported as a clean one"
);
assert!(
!entries(&log).contains(&"host:reap".to_string()),
"nothing confirmed the reap: {:?}",
entries(&log)
);
assert!(
entries(&log).contains(&"host:kill".to_string()),
"a child that would not stop must still be escalated: {:?}",
entries(&log)
);
assert!(
guard.is_armed(),
"an unconfirmed reap must leave the drop fallback armed"
);
}
/// Death is not instantaneous, so the drop path has to keep polling. A
/// single `try_reap` in place of the loop must not pass.
#[test]
fn the_drop_path_polls_until_the_child_is_actually_gone() {
let log = log();
drop(ReapOnDrop::new(
FakeChild::reaps_after_polls(&log, "host", 3),
"host",
));
assert_eq!(entries(&log), vec!["host:kill", "host:reap"]);
}
/// Mutation gate #3 (remove the wait after the host kill).
#[tokio::test]
async fn explicit_shutdown_reaps_the_host_before_the_aec_can_unload() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
t.push_viewer("ticket-A".to_string(), FakeChild::new(&log, "viewer"));
t.shutdown_children().await;
// Reaped by the explicit path — before the guard is anywhere near dropped.
assert_eq!(
entries(&log),
vec!["host:sigint", "host:reap", "viewer:sigint", "viewer:reap"],
"children must be stopped and reaped by the explicit path"
);
drop(t);
let unload = position(&log, "aec:unload").expect("the AEC guard must be dropped");
let host_reap = position(&log, "host:reap").expect("the host must be reaped");
assert!(host_reap < unload);
}
#[tokio::test]
async fn explicit_shutdown_is_idempotent_with_the_drop_path() {
let log = log();
let mut t = teardown(&log);
t.set_host(FakeChild::new(&log, "host"));
t.shutdown_children().await;
drop(t);
// Exactly one stop and one reap: the drop path must not re-signal a
// child the explicit path already took.
assert_eq!(
entries(&log),
vec!["host:sigint", "host:reap", "aec:unload"]
);
}
// --- Host/viewer bookkeeping ---
#[tokio::test]
async fn stop_host_reports_whether_it_was_sharing() {
let log = log();
let mut t = teardown(&log);
assert!(!t.is_sharing());
assert_eq!(t.stop_host().await, None, "not sharing: nothing to stop");
t.set_host(FakeChild::new(&log, "host"));
assert!(t.is_sharing());
assert_eq!(t.stop_host().await, Some(StopOutcome::Reaped));
assert!(!t.is_sharing());
assert_eq!(entries(&log), vec!["host:sigint", "host:reap"]);
}
#[test]
fn sweeping_drops_only_the_players_that_already_closed() {
let log = log();
let mut t = teardown(&log);
t.push_viewer(
"closed".to_string(),
FakeChild::already_exited(&log, "closed"),
);
t.push_viewer("live".to_string(), FakeChild::new(&log, "live"));
t.sweep_exited_viewers();
// The live player survives the sweep; only the closed one is dropped,
// and dropping it must not kill anything (it was already gone).
assert_eq!(t.viewers.len(), 1);
assert_eq!(t.viewers[0].0, "live");
assert_eq!(entries(&log), vec!["closed:reap"]);
}
#[tokio::test]
async fn re_watching_a_share_replaces_that_player_only() {
let log = log();
let mut t = teardown(&log);
t.push_viewer("ticket-A".to_string(), FakeChild::new(&log, "a"));
t.push_viewer("ticket-B".to_string(), FakeChild::new(&log, "b"));
assert!(t.replace_viewer("ticket-A").await);
assert_eq!(entries(&log), vec!["a:sigint", "a:reap"]);
assert_eq!(t.viewers.len(), 1);
assert_eq!(t.viewers[0].0, "ticket-B");
// A share we are not watching has nothing to replace.
assert!(!t.replace_viewer("ticket-C").await);
}
}
+519 -27
View File
@@ -60,6 +60,48 @@ const LOW_LATENCY_CACHE_CAP_MB: u32 = 1;
/// is only a safety net so a hung pixelpass can't wedge the caller forever.
const STARTUP_TIMEOUT: Duration = Duration::from_secs(20);
/// The audio source selected for one hosted screen share.
///
/// This is shared by the picker and the core so the new desktop-excluding
/// choice cannot collapse back into the legacy `Option<String>` representation
/// (where `None` could only mean whole-desktop audio).
#[derive(Debug, Clone, PartialEq, Eq, Default)]
pub enum ShareAudioSelection {
/// PixelPass's legacy whole-desktop monitor capture.
#[default]
DesktopShared,
/// Whole-desktop audio with PeerSpeak-owned playback excluded.
DesktopExcluding,
/// Strict capture of one locally selected application.
Application(String),
}
/// One version-1 desktop-audio-exclusion status from PixelPass.
///
/// The parsed value itself crosses the core/UI boundary; PeerSpeak does not
/// define a second matching enum that could silently drift from the wire.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum AudioExclusionStatus {
StreamUnsupported {
stream_serial: u64,
reason: String,
},
/// PixelPass has cleared the previously reported status for this exact
/// stream because it became capturable or left the graph.
StreamStatusCleared {
stream_serial: u64,
},
AecFailed {
module_index: u64,
},
AecRevoked {
module_index: u64,
},
ForeignAecWarning {
link_group: String,
},
}
/// One parsed line from pixelpass's `--output json` stdout stream. Mirrors the
/// `event` tags in pixelpass's `src/common/output.rs`. Recognized-but-unused
/// events collapse to [`PixelpassEvent::Other`]; blank or non-JSON lines parse
@@ -86,10 +128,28 @@ pub enum PixelpassEvent {
/// our `--strict-audio` run this means viewers now hear silence (not the call
/// echo) until the app produces audio again — we surface it as a warning.
AppAudioLost,
/// Host (desktop-excluding audio): a versioned status from the fail-closed
/// fan-out controller.
AudioExclusion(AudioExclusionStatus),
/// A recognized event we don't act on (e.g. `host_info`).
Other,
}
/// What the host's stdout drain forwards to the core over the notice channel.
///
/// `Eof` is **synthesized here**, not parsed: pixelpass has no "I died" event,
/// and a crash can abort across `extern "C"` before any JSON line is written,
/// so the stream ending is the only reliable death signal. A read *error*
/// counts too — either way the event stream is gone and the host must be
/// treated as over.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum HostNotice {
/// A parsed pixelpass event line.
Event(PixelpassEvent),
/// The host's stdout ended (EOF or read error). Terminal: nothing follows.
Eof,
}
/// Parse a single stdout line from pixelpass `--output json`. Pure: no I/O.
pub fn parse_pixelpass_event(line: &str) -> Option<PixelpassEvent> {
let line = line.trim();
@@ -125,6 +185,32 @@ pub fn parse_pixelpass_event(line: &str) -> Option<PixelpassEvent> {
Some("lost") => PixelpassEvent::AppAudioLost,
_ => PixelpassEvent::Other,
},
"stream_unsupported" if json_u64(&v, "version") == Some(1) => {
PixelpassEvent::AudioExclusion(AudioExclusionStatus::StreamUnsupported {
stream_serial: json_u64(&v, "stream_serial")?,
reason: v.get("reason")?.as_str()?.to_string(),
})
}
"stream_status_cleared" if json_u64(&v, "version") == Some(1) => {
PixelpassEvent::AudioExclusion(AudioExclusionStatus::StreamStatusCleared {
stream_serial: json_u64(&v, "stream_serial")?,
})
}
"aec_failed" if json_u64(&v, "version") == Some(1) => {
PixelpassEvent::AudioExclusion(AudioExclusionStatus::AecFailed {
module_index: json_u64(&v, "module_index")?,
})
}
"aec_revoked" if json_u64(&v, "version") == Some(1) => {
PixelpassEvent::AudioExclusion(AudioExclusionStatus::AecRevoked {
module_index: json_u64(&v, "module_index")?,
})
}
"foreign_aec_warning" if json_u64(&v, "version") == Some(1) => {
PixelpassEvent::AudioExclusion(AudioExclusionStatus::ForeignAecWarning {
link_group: v.get("link_group")?.as_str()?.to_string(),
})
}
_ => PixelpassEvent::Other,
};
Some(ev)
@@ -134,6 +220,10 @@ fn json_u32(v: &serde_json::Value, key: &str) -> u32 {
v.get(key).and_then(|x| x.as_u64()).unwrap_or(0) as u32
}
fn json_u64(v: &serde_json::Value, key: &str) -> Option<u64> {
v.get(key).and_then(|x| x.as_u64())
}
/// Build the argv for a pixelpass *host*. Always `--host --output json`; when
/// `audio_app` is `Some`, append `--app=<name> --strict-audio` so pixelpass
/// captures only that app's audio instead of the whole desktop sink monitor
@@ -165,6 +255,48 @@ pub fn host_args(
args.push(format!("--app={name}"));
args.push("--strict-audio".to_string());
}
append_host_settings(&mut args, settings, quality);
args
}
/// Build host argv from the picker's typed audio selection.
///
/// The existing desktop-shared and application arms deliberately delegate to
/// [`host_args`] so their argv stays byte-for-byte compatible. Only the new
/// desktop-excluding arm emits the public PixelPass protocol pair, and it
/// always includes an explicit AEC state: `off` when this PeerSpeak session did
/// not load an echo-cancel module, otherwise the exact pactl module index.
pub fn host_args_for_selection(
audio: &ShareAudioSelection,
aec_module_index: Option<u64>,
settings: &ScreenShareSettings,
quality: ShareQuality,
) -> Vec<String> {
match audio {
ShareAudioSelection::DesktopShared => host_args(None, settings, quality),
ShareAudioSelection::Application(name) => host_args(Some(name), settings, quality),
ShareAudioSelection::DesktopExcluding => {
let mut args = vec![
"--host".to_string(),
"--output".to_string(),
"json".to_string(),
"--audio-mode=desktop-excluding".to_string(),
match aec_module_index {
Some(index) => format!("--aec=pulse-module:{index}"),
None => "--aec=off".to_string(),
},
];
append_host_settings(&mut args, settings, quality);
args
}
}
}
fn append_host_settings(
args: &mut Vec<String>,
settings: &ScreenShareSettings,
quality: ShareQuality,
) {
if quality != ShareQuality::Auto {
args.push(format!("--quality={}", pixelpass_quality(quality)));
}
@@ -184,7 +316,6 @@ pub fn host_args(
args.push(format!("--max-viewers={max}"));
}
args.extend(split_extra_args(&settings.extra_host_args));
args
}
fn pixelpass_quality(quality: ShareQuality) -> &'static str {
@@ -244,22 +375,67 @@ pub async fn list_audio_apps() -> Vec<String> {
}
}
/// Hard cap on the capability probe (`pixelpass --help`). Conservative: a slow or
/// hung pixelpass degrades to "strict audio unsupported" → whole-desktop-only
/// picker (safe), never a stalled core loop.
/// Hard cap on each PixelPass capability probe. Conservative: a slow or hung
/// binary degrades to the legacy capability set, never a stalled core loop.
const HELP_PROBE_TIMEOUT: Duration = Duration::from_secs(2);
/// Whether the resolved pixelpass understands `--strict-audio` (added in pixelpass
/// `85fdebe`). peerspeak only offers per-app audio capture when it does: a per-app
/// share always appends `--strict-audio`, and an **older** pixelpass would have
/// clap reject the unknown flag → the host spawn hard-fails and the share is
/// broken (audit P2, version skew). When unsupported the picker degrades to
/// whole-desktop audio only — we never silently drop to best-effort `--app`, which
/// would reintroduce the call echo (A23).
///
/// Any probe failure/timeout returns `false` (degrade to the safe path). The
/// `--help` child is `kill_on_drop` so a hung pixelpass can't linger.
pub async fn supports_strict_audio(bin: &Path) -> bool {
/// Capabilities PeerSpeak consumes from PixelPass's versioned response.
/// Strict per-app capture and desktop exclusion are independent by contract.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct PixelpassCapabilities {
pub strict_app_audio: bool,
pub desktop_audio_exclusion: bool,
}
/// A capability result tied to the exact resolved executable that produced it.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct ProbedPixelpassCapabilities {
pub binary: PathBuf,
pub capabilities: PixelpassCapabilities,
}
#[derive(serde::Deserialize)]
struct CapabilityResponse {
schema_version: u64,
capabilities: CapabilityFlags,
}
#[derive(serde::Deserialize)]
struct CapabilityFlags {
strict_app_audio: bool,
desktop_audio_exclusion: bool,
}
/// Parse the schema-1 response from `pixelpass --capabilities`. Unsupported
/// schemas and malformed output return `None`, which selects the legacy help
/// fallback rather than guessing at a new protocol.
pub fn parse_pixelpass_capabilities(stdout: &[u8]) -> Option<PixelpassCapabilities> {
let response: CapabilityResponse = serde_json::from_slice(stdout).ok()?;
(response.schema_version == 1).then_some(PixelpassCapabilities {
strict_app_audio: response.capabilities.strict_app_audio,
desktop_audio_exclusion: response.capabilities.desktop_audio_exclusion,
})
}
/// Probe one resolved PixelPass binary. The versioned machine response is the
/// primary contract. `--help` survives only as a compatibility fallback for an
/// older PixelPass that predates `--capabilities`; it can recover strict per-app
/// support but can never advertise desktop exclusion.
pub async fn probe_pixelpass_capabilities(bin: &Path) -> PixelpassCapabilities {
let primary = Command::new(bin)
.arg("--capabilities")
.stdin(Stdio::null())
.stdout(Stdio::piped())
.stderr(Stdio::null())
.kill_on_drop(true)
.output();
if let Ok(Ok(output)) = tokio::time::timeout(HELP_PROBE_TIMEOUT, primary).await
&& output.status.success()
&& let Some(capabilities) = parse_pixelpass_capabilities(&output.stdout)
{
return capabilities;
}
let run = Command::new(bin)
.arg("--help")
.stdin(Stdio::null())
@@ -267,12 +443,42 @@ pub async fn supports_strict_audio(bin: &Path) -> bool {
.stderr(Stdio::null())
.kill_on_drop(true)
.output();
match tokio::time::timeout(HELP_PROBE_TIMEOUT, run).await {
Ok(Ok(o)) => help_mentions_strict_audio(&o.stdout),
let strict_app_audio = match tokio::time::timeout(HELP_PROBE_TIMEOUT, run).await {
Ok(Ok(output)) => help_mentions_strict_audio(&output.stdout),
_ => false,
};
PixelpassCapabilities {
strict_app_audio,
desktop_audio_exclusion: false,
}
}
/// Return capabilities for `bin`, re-probing and replacing `cached` whenever
/// the resolved executable path differs. This is the start-time skew guard:
/// capability-gated argv must never be built from a probe of another binary.
pub async fn capabilities_for_resolved_binary(
bin: &Path,
cached: &mut Option<ProbedPixelpassCapabilities>,
) -> PixelpassCapabilities {
if let Some(probe) = cached.as_ref()
&& probe.binary == bin
{
return probe.capabilities;
}
let capabilities = probe_pixelpass_capabilities(bin).await;
*cached = Some(ProbedPixelpassCapabilities {
binary: bin.to_path_buf(),
capabilities,
});
capabilities
}
/// Compatibility helper retained for callers that only need the pre-Phase-8
/// per-app bit.
pub async fn supports_strict_audio(bin: &Path) -> bool {
probe_pixelpass_capabilities(bin).await.strict_app_audio
}
/// Pure check: does `pixelpass --help` advertise `--strict-audio`? Matches the
/// flag token rather than a whole line, since clap may wrap/realign help text.
pub fn help_mentions_strict_audio(help_stdout: &[u8]) -> bool {
@@ -370,17 +576,21 @@ pub fn is_available(config_override: Option<&str>) -> bool {
/// `audio_app` is `Some`, pixelpass captures only that app's audio instead of the
/// whole desktop sink, which avoids the call-loopback echo (A23). The child keeps
/// running (streaming to viewers) until killed or dropped; remaining stdout is
/// drained in a background task so a full pipe can't stall the host. We do
/// drained in a background task so a full pipe can't stall the host. The drain
/// forwards every parsed event over `notices` and — the part no share may opt
/// out of — a terminal [`HostNotice::Eof`] when the stream ends, which is the
/// caller's only reliable signal that the host died. We do
/// not pass encode/viewer overrides unless the local settings explicitly ask for
/// them, so pixelpass keeps its own defaults in the common case.
pub async fn spawn_host(
bin: &Path,
audio_app: Option<&str>,
audio: &ShareAudioSelection,
aec_module_index: Option<u64>,
settings: &ScreenShareSettings,
quality: ShareQuality,
notices: Option<tokio::sync::mpsc::UnboundedSender<PixelpassEvent>>,
notices: tokio::sync::mpsc::UnboundedSender<HostNotice>,
) -> std::io::Result<(Child, String)> {
let args = host_args(audio_app, settings, quality);
let args = host_args_for_selection(audio, aec_module_index, settings, quality);
// Log the exact argv we hand pixelpass so a field log can confirm which
// encode/quality flags (e.g. --bitrate) actually reached the host — these
// are local flags with no ticket/secret, so logging them verbatim is safe.
@@ -433,7 +643,7 @@ pub async fn spawn_host(
if let Some(stderr) = stderr {
drain_stderr_in_background(stderr);
}
drain_in_background(lines, "host", notices);
drain_in_background(lines, "host", Some(notices));
Ok((child, ticket))
}
@@ -572,13 +782,15 @@ where
/// Keep reading the child's stdout to EOF in the background so a full pipe can't
/// stall it; log notable events for diagnostics. When `notices` is `Some`, each
/// parsed event is also forwarded to the caller (the core, which translates the
/// `app_audio` ones into a UI warning); a send failure (receiver dropped) just
/// stops forwarding, draining continues. The task ends on EOF (child exited).
/// parsed event is also forwarded to the caller (the core), and when the stream
/// ends — EOF or read error, i.e. the child exited or its event stream broke —
/// a final [`HostNotice::Eof`] is sent so the caller learns the child is gone
/// (a host that dies must not stay advertised as sharing). A send failure
/// (receiver dropped) just stops forwarding, draining continues.
fn drain_in_background<R>(
mut lines: tokio::io::Lines<BufReader<R>>,
role: &'static str,
notices: Option<tokio::sync::mpsc::UnboundedSender<PixelpassEvent>>,
notices: Option<tokio::sync::mpsc::UnboundedSender<HostNotice>>,
) where
R: tokio::io::AsyncRead + Unpin + Send + 'static,
{
@@ -587,10 +799,14 @@ fn drain_in_background<R>(
if let Some(ev) = parse_pixelpass_event(&line) {
crate::log_msg(&format!("pixelpass {role}: {}", event_for_log(&ev)));
if let Some(tx) = &notices {
let _ = tx.send(ev);
let _ = tx.send(HostNotice::Event(ev));
}
}
}
if let Some(tx) = &notices {
crate::log_msg(&format!("pixelpass {role}: stdout ended"));
let _ = tx.send(HostNotice::Eof);
}
});
}
@@ -609,6 +825,7 @@ fn event_for_log(ev: &PixelpassEvent) -> String {
PixelpassEvent::CaptureStopped => "capture_stopped".to_string(),
PixelpassEvent::AppAudioRouted => "app_audio_routed".to_string(),
PixelpassEvent::AppAudioLost => "app_audio_lost".to_string(),
PixelpassEvent::AudioExclusion(status) => format!("audio_exclusion {status:?}"),
PixelpassEvent::Other => "other".to_string(),
}
}
@@ -889,6 +1106,64 @@ mod tests {
assert_eq!(args[4], "--strict-audio");
}
#[test]
fn desktop_excluding_argv_requires_both_public_mode_and_explicit_aec() {
let settings = ScreenShareSettings::default();
assert_eq!(
host_args_for_selection(
&ShareAudioSelection::DesktopExcluding,
Some(536_870_919),
&settings,
ShareQuality::Auto,
),
vec![
"--host",
"--output",
"json",
"--audio-mode=desktop-excluding",
"--aec=pulse-module:536870919",
]
);
assert_eq!(
host_args_for_selection(
&ShareAudioSelection::DesktopExcluding,
None,
&settings,
ShareQuality::Auto,
),
vec![
"--host",
"--output",
"json",
"--audio-mode=desktop-excluding",
"--aec=off",
]
);
}
#[test]
fn typed_legacy_selections_keep_existing_argv_byte_identical() {
let settings = ScreenShareSettings::default();
assert_eq!(
host_args_for_selection(
&ShareAudioSelection::DesktopShared,
Some(42),
&settings,
ShareQuality::Auto,
),
host_args(None, &settings, ShareQuality::Auto),
);
assert_eq!(
host_args_for_selection(
&ShareAudioSelection::Application("Firefox".to_string()),
Some(42),
&settings,
ShareQuality::Auto,
),
host_args(Some("Firefox"), &settings, ShareQuality::Auto),
);
}
#[test]
fn host_args_blank_or_control_app_is_dropped() {
// An empty / whitespace / control-laden selection is sanitized away,
@@ -1193,6 +1468,106 @@ Install hint: sudo apt install gstreamer1.0-plugins-bad
assert!(!help_mentions_strict_audio(&[0xff, 0xfe, 0x00]));
}
#[test]
fn capability_schema_keeps_strict_and_desktop_exclusion_independent() {
assert_eq!(
parse_pixelpass_capabilities(
br#"{"schema_version":1,"capabilities":{"strict_app_audio":true,"desktop_audio_exclusion":false}}"#,
),
Some(PixelpassCapabilities {
strict_app_audio: true,
desktop_audio_exclusion: false,
})
);
assert_eq!(
parse_pixelpass_capabilities(
br#"{"schema_version":1,"capabilities":{"strict_app_audio":false,"desktop_audio_exclusion":true}}"#,
),
Some(PixelpassCapabilities {
strict_app_audio: false,
desktop_audio_exclusion: true,
})
);
assert!(
parse_pixelpass_capabilities(
br#"{"schema_version":2,"capabilities":{"strict_app_audio":true,"desktop_audio_exclusion":true}}"#,
)
.is_none(),
"an unknown schema must not advertise the new mode"
);
}
#[cfg(unix)]
fn write_fake_pixelpass(dir: &Path, name: &str, body: &str) -> PathBuf {
use std::os::unix::fs::PermissionsExt;
let path = dir.join(name);
std::fs::write(&path, body).unwrap();
let mut permissions = std::fs::metadata(&path).unwrap().permissions();
permissions.set_mode(0o755);
std::fs::set_permissions(&path, permissions).unwrap();
path
}
#[cfg(unix)]
#[tokio::test]
async fn old_pixelpass_help_fallback_cannot_advertise_desktop_exclusion() {
let dir = std::env::temp_dir().join(format!(
"peerspeak-phase8-old-pixelpass-{}",
std::process::id()
));
std::fs::create_dir_all(&dir).unwrap();
let bin = write_fake_pixelpass(
&dir,
"pixelpass-old",
"#!/bin/sh\nif [ \"$1\" = \"--capabilities\" ]; then exit 2; fi\nprintf '%s\\n' 'Options: --app <APP> --strict-audio --output <OUTPUT>'\n",
);
let capabilities = probe_pixelpass_capabilities(&bin).await;
assert!(capabilities.strict_app_audio);
assert!(!capabilities.desktop_audio_exclusion);
assert_eq!(
host_args_for_selection(
&ShareAudioSelection::DesktopShared,
None,
&ScreenShareSettings::default(),
ShareQuality::Auto,
),
vec!["--host", "--output", "json"],
"old-PixelPass fallback must emit no new flags"
);
std::fs::remove_dir_all(dir).unwrap();
}
#[cfg(unix)]
#[tokio::test]
async fn capability_cache_reprobes_when_the_resolved_binary_changes() {
let dir =
std::env::temp_dir().join(format!("peerspeak-phase8-rebind-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
let new_bin = write_fake_pixelpass(
&dir,
"pixelpass-new",
"#!/bin/sh\nprintf '%s\\n' '{\"schema_version\":1,\"capabilities\":{\"strict_app_audio\":true,\"desktop_audio_exclusion\":true}}'\n",
);
let old_bin = write_fake_pixelpass(
&dir,
"pixelpass-old",
"#!/bin/sh\nif [ \"$1\" = \"--capabilities\" ]; then exit 2; fi\nprintf '%s\\n' 'Options: --output <OUTPUT>'\n",
);
let mut cached = None;
let first = capabilities_for_resolved_binary(&new_bin, &mut cached).await;
assert!(first.desktop_audio_exclusion);
assert_eq!(cached.as_ref().unwrap().binary, new_bin);
let rebound = capabilities_for_resolved_binary(&old_bin, &mut cached).await;
assert!(!rebound.desktop_audio_exclusion);
assert!(!rebound.strict_app_audio);
assert_eq!(cached.as_ref().unwrap().binary, old_bin);
std::fs::remove_dir_all(dir).unwrap();
}
#[test]
fn sanitize_ticket_accepts_pixelpass_endpoint_ticket_shape() {
let ticket = "endpointaabwxjexzensznfvuudiapn5tyzws3angd2merarm";
@@ -1310,6 +1685,64 @@ Install hint: sudo apt install gstreamer1.0-plugins-bad
);
}
#[test]
fn parses_all_version_one_audio_exclusion_statuses_exactly() {
assert_eq!(
parse_pixelpass_event(
r#"{"event":"stream_unsupported","version":1,"stream_serial":4294967303,"reason":"port-exclusive"}"#,
),
Some(PixelpassEvent::AudioExclusion(
AudioExclusionStatus::StreamUnsupported {
stream_serial: 4_294_967_303,
reason: "port-exclusive".to_string(),
}
))
);
assert_eq!(
parse_pixelpass_event(
r#"{"event":"stream_status_cleared","version":1,"stream_serial":4294967303}"#,
),
Some(PixelpassEvent::AudioExclusion(
AudioExclusionStatus::StreamStatusCleared {
stream_serial: 4_294_967_303,
}
))
);
assert_eq!(
parse_pixelpass_event(r#"{"event":"aec_failed","version":1,"module_index":536870919}"#,),
Some(PixelpassEvent::AudioExclusion(
AudioExclusionStatus::AecFailed {
module_index: 536_870_919,
}
))
);
assert_eq!(
parse_pixelpass_event(
r#"{"event":"aec_revoked","version":1,"module_index":536870919}"#,
),
Some(PixelpassEvent::AudioExclusion(
AudioExclusionStatus::AecRevoked {
module_index: 536_870_919,
}
))
);
assert_eq!(
parse_pixelpass_event(
r#"{"event":"foreign_aec_warning","version":1,"link_group":"echo-cancel-9999-13"}"#,
),
Some(PixelpassEvent::AudioExclusion(
AudioExclusionStatus::ForeignAecWarning {
link_group: "echo-cancel-9999-13".to_string(),
}
))
);
assert_eq!(
parse_pixelpass_event(r#"{"event":"aec_failed","version":2,"module_index":536870919}"#,),
Some(PixelpassEvent::Other),
"an unknown wire version must not be misinterpreted as version 1"
);
}
#[test]
fn recognized_but_unused_event_is_other() {
assert_eq!(
@@ -1376,4 +1809,63 @@ Install hint: sudo apt install gstreamer1.0-plugins-bad
#[cfg(not(windows))]
assert_eq!(candidates, vec![dir.join("pixelpass")]);
}
/// The host-fault contract, clean-exit half: events are forwarded in order
/// and the stream ending yields exactly one terminal [`HostNotice::Eof`],
/// after which the drain task drops its sender (the closed channel is what
/// ends the core's forwarder). A host that dies silently — EOF swallowed —
/// is the S2 defect: the dead share stays advertised in presence.
#[tokio::test]
async fn drain_forwards_events_then_synthesizes_eof_when_stdout_ends() {
let (tx, mut rx) = tokio::sync::mpsc::unbounded_channel();
let (read_half, mut write_half) = tokio::io::duplex(1024);
drain_in_background(BufReader::new(read_half).lines(), "test", Some(tx));
use tokio::io::AsyncWriteExt;
write_half
.write_all(b"{\"event\":\"app_audio\",\"state\":\"routed\"}\nnot json\n")
.await
.unwrap();
drop(write_half); // child exited: stdout EOF
assert_eq!(
rx.recv().await,
Some(HostNotice::Event(PixelpassEvent::AppAudioRouted))
);
// The non-JSON line is dropped, not forwarded.
assert_eq!(rx.recv().await, Some(HostNotice::Eof));
assert_eq!(rx.recv().await, None, "task ended and dropped the sender");
}
/// The host-fault contract, broken-stream half: a read *error* (not a tidy
/// EOF) must synthesize the same terminal `Eof` — the event stream is gone
/// either way, and only the drain task can tell the core so.
#[tokio::test]
async fn drain_synthesizes_eof_on_a_read_error_too() {
struct BrokenPipe;
impl tokio::io::AsyncRead for BrokenPipe {
fn poll_read(
self: std::pin::Pin<&mut Self>,
_cx: &mut std::task::Context<'_>,
_buf: &mut tokio::io::ReadBuf<'_>,
) -> std::task::Poll<std::io::Result<()>> {
std::task::Poll::Ready(Err(std::io::Error::other("stream broke")))
}
}
use tokio::io::AsyncReadExt;
// One good event line, then the stream breaks mid-read.
let reader =
std::io::Cursor::new(b"{\"event\":\"capture\",\"state\":\"started\"}\n".to_vec())
.chain(BrokenPipe);
let (tx, mut rx) = tokio::sync::mpsc::unbounded_channel();
drain_in_background(BufReader::new(reader).lines(), "test", Some(tx));
assert_eq!(
rx.recv().await,
Some(HostNotice::Event(PixelpassEvent::CaptureStarted))
);
assert_eq!(rx.recv().await, Some(HostNotice::Eof));
assert_eq!(rx.recv().await, None, "task ended and dropped the sender");
}
}
+4 -7
View File
@@ -435,13 +435,10 @@ where
let was_hovered = self.hovered_link.is_some();
self.hovered_link = local_position.and_then(|position| {
state.paragraph.hit_span(position).and_then(|span| {
if spans.get(span)?.link.is_some() {
Some(span)
} else {
None
}
})
state
.paragraph
.hit_span(position)
.filter(|&span| spans.get(span).is_some_and(|span| span.link.is_some()))
});
if was_hovered != self.hovered_link.is_some() {
+590
View File
@@ -0,0 +1,590 @@
//! S2 exit gate: a pixelpass host that dies mid-share must be torn down —
//! reaped, pulled off presence, `ScreenShareStopped` emitted **before** the
//! explanatory error — and a host stopped *deliberately* must NOT produce that
//! error when its stdout EOF arrives late (the staleness gate).
//!
//! Drives the real core loop end to end through `CoreController`, with the
//! pixelpass override pointed at fake shell scripts: one that emits a ticket
//! and dies, one that emits a ticket and lives until signalled. This is the
//! only harness that reaches the core's fault handler — the command loop has
//! no unit seam — so these two halves are what kill the "forwarder drops the
//! Eof" and "handler ignores the generation" mutants.
//!
//! Live: joins a real (solo) room, so it needs a working audio backend and
//! network access for the endpoint bind.
//! `cargo test --test screenshare_host_fault -- --ignored`
#![cfg(unix)]
use std::os::unix::fs::PermissionsExt;
use std::path::PathBuf;
use std::time::Duration;
use peerspeak::core::CoreController;
use peerspeak::core::messages::{CoreCommand, UiEvent};
use peerspeak::screenshare::ShareAudioSelection;
const EVENT_TIMEOUT: Duration = Duration::from_secs(20);
/// How long to listen for events that must NOT arrive. Comfortably past the
/// fake host's exit plus the drain/forwarder hop, so a stale fault that WOULD
/// be mishandled has arrived by the end of it.
const QUIET_WINDOW: Duration = Duration::from_secs(3);
/// Removes the fake-pixelpass dir even when an assertion panics mid-test
/// (a plain trailing `remove_dir_all` never runs on an unwind).
struct TempDir(PathBuf);
impl Drop for TempDir {
fn drop(&mut self) {
std::fs::remove_dir_all(&self.0).ok();
}
}
fn write_fake_pixelpass(dir: &std::path::Path, name: &str, body: &str) -> PathBuf {
let path = dir.join(name);
std::fs::write(&path, body).expect("write fake pixelpass");
std::fs::set_permissions(&path, std::fs::Permissions::from_mode(0o755))
.expect("chmod fake pixelpass");
path
}
/// Skip events until `pick` matches, panicking after [`EVENT_TIMEOUT`].
/// Unrelated events (identity, presence, chat plumbing) flow on this channel
/// too, so gates scan rather than assert exact sequences.
async fn wait_for<T>(
rx: &mut tokio::sync::mpsc::Receiver<UiEvent>,
what: &str,
mut pick: impl FnMut(&UiEvent) -> Option<T>,
) -> T {
let deadline = tokio::time::Instant::now() + EVENT_TIMEOUT;
loop {
let ev = tokio::time::timeout_at(deadline, rx.recv())
.await
.unwrap_or_else(|_| panic!("timed out waiting for {what}"))
.unwrap_or_else(|| panic!("ui channel closed waiting for {what}"));
if let Some(v) = pick(&ev) {
return v;
}
}
}
#[tokio::test]
#[ignore = "live: joins a real solo room (audio backend + network bind)"]
async fn a_dead_host_is_torn_down_and_a_clean_stop_stays_clean() {
let dir_guard =
TempDir(std::env::temp_dir().join(format!("peerspeak-hostfault-{}", std::process::id())));
let dir = dir_guard.0.clone();
std::fs::create_dir_all(&dir).unwrap();
// Half 1's host: emits its ticket, then dies on its own — the S2 defect
// scenario. Plain `sleep` (no exec) so the shell itself exits and closes
// stdout with no orphan holding the pipe.
let dying_host = write_fake_pixelpass(
&dir,
"pixelpass-dies",
"#!/bin/sh\necho '{\"event\":\"ticket\",\"value\":\"fake-ticket-dies\"}'\nsleep 1\n",
);
// Half 2's host: lives until signalled. `exec` so the SIGINT from Stop
// Share hits the sleep itself — the process dies AND its stdout closes,
// which is exactly what makes the late Eof arrive and exercise the
// staleness gate rather than vacuously never sending a fault.
let living_host = write_fake_pixelpass(
&dir,
"pixelpass-lives",
"#!/bin/sh\necho '{\"event\":\"ticket\",\"value\":\"fake-ticket-lives\"}'\nexec sleep 600\n",
);
let (ui_tx, mut ui_rx) = tokio::sync::mpsc::channel(256);
let controller = CoreController::new(ui_tx);
assert!(controller.send(CoreCommand::SetPixelpassPath(Some(
dying_host.to_string_lossy().into_owned()
))));
assert!(controller.send(CoreCommand::Join {
name: "host-fault-gate".into(),
ticket: "create".into(),
room_name: "s2".into(),
input_device: None,
output_device: None,
echo_cancellation: false,
avatar: Default::default(),
}));
wait_for(&mut ui_rx, "RoomJoined", |ev| match ev {
UiEvent::RoomJoined { .. } => Some(()),
UiEvent::Error(e) => panic!("join failed: {e}"),
_ => None,
})
.await;
// ── Half 1: the host dies mid-share ─────────────────────────────────────
assert!(controller.send(CoreCommand::StartScreenShare {
audio: ShareAudioSelection::DesktopShared,
settings: Default::default(),
quality: Default::default(),
}));
wait_for(
&mut ui_rx,
"ScreenShareStarted (dying host)",
|ev| match ev {
UiEvent::ScreenShareStarted => Some(()),
UiEvent::Error(e) => panic!("share start failed: {e}"),
_ => None,
},
)
.await;
// The fake host exits ~1s in. The contract: ScreenShareStopped FIRST (it
// clears the UI's sharing state), the explanatory error only after.
wait_for(&mut ui_rx, "ScreenShareStopped after host death", |ev| {
match ev {
UiEvent::ScreenShareStopped => Some(()),
// An error arriving first is the exact ordering defect S2 fixes:
// the UI would show "sharing" next to the explanation.
UiEvent::Error(e) => panic!("error arrived before ScreenShareStopped: {e}"),
_ => None,
}
})
.await;
let err = wait_for(&mut ui_rx, "the host-death error", |ev| match ev {
UiEvent::Error(e) => Some(e.clone()),
_ => None,
})
.await;
assert!(
err.contains("unexpectedly"),
"the error should say the share ended unexpectedly, got: {err}"
);
// ── Half 2: a deliberate stop must stay clean ───────────────────────────
assert!(controller.send(CoreCommand::SetPixelpassPath(Some(
living_host.to_string_lossy().into_owned()
))));
assert!(controller.send(CoreCommand::StartScreenShare {
audio: ShareAudioSelection::DesktopShared,
settings: Default::default(),
quality: Default::default(),
}));
wait_for(
&mut ui_rx,
"ScreenShareStarted (living host)",
|ev| match ev {
UiEvent::ScreenShareStarted => Some(()),
UiEvent::Error(e) => panic!("second share start failed: {e}"),
_ => None,
},
)
.await;
assert!(controller.send(CoreCommand::StopScreenShare));
wait_for(
&mut ui_rx,
"ScreenShareStopped after Stop Share",
|ev| match ev {
UiEvent::ScreenShareStopped => Some(()),
UiEvent::Error(e) => panic!("clean stop produced an error: {e}"),
_ => None,
},
)
.await;
// The stopped host's stdout EOF is arriving about now as a *stale* fault
// (its generation was retired when Stop Share cleared the share). Without
// the staleness gate the handler would emit a second ScreenShareStopped
// and a spurious "ended unexpectedly" error — listen long enough for that
// mishandling to have shown up, and require silence.
let deadline = tokio::time::Instant::now() + QUIET_WINDOW;
while let Ok(Some(ev)) = tokio::time::timeout_at(deadline, ui_rx.recv()).await {
match ev {
UiEvent::ScreenShareStopped => {
panic!("stale host fault re-emitted ScreenShareStopped after a clean stop")
}
UiEvent::Error(e) if e.contains("unexpectedly") => {
panic!("stale host fault surfaced as an error after a clean stop: {e}")
}
_ => {}
}
}
// ── Half 3: a failed room switch while sharing must not cry "crash" ─────
// Join tears the old session down (killing the host, deliberately) BEFORE
// it validates the ticket, so an invalid ticket exits the Join arm early.
// The share must be retired at the teardown itself — left advertised, the
// killed host's EOF passes the staleness gate and a spurious "ended
// unexpectedly" lands on top of the ticket error (Gemini review, P2-1).
assert!(controller.send(CoreCommand::StartScreenShare {
audio: ShareAudioSelection::DesktopShared,
settings: Default::default(),
quality: Default::default(),
}));
wait_for(
&mut ui_rx,
"ScreenShareStarted (before failed switch)",
|ev| match ev {
UiEvent::ScreenShareStarted => Some(()),
UiEvent::Error(e) => panic!("third share start failed: {e}"),
_ => None,
},
)
.await;
assert!(controller.send(CoreCommand::Join {
name: "host-fault-gate".into(),
ticket: "definitely-not-a-ticket".into(),
room_name: "s2".into(),
input_device: None,
output_device: None,
echo_cancellation: false,
avatar: Default::default(),
}));
wait_for(&mut ui_rx, "the invalid-ticket error", |ev| match ev {
UiEvent::Error(e) if e.contains("invalid room ticket") => Some(()),
UiEvent::Error(e) => panic!("unexpected error before the ticket error: {e}"),
_ => None,
})
.await;
// The deliberately-killed host's EOF is arriving about now; it must be
// dropped as stale, not reported as a crash.
let deadline = tokio::time::Instant::now() + QUIET_WINDOW;
while let Ok(Some(ev)) = tokio::time::timeout_at(deadline, ui_rx.recv()).await {
match ev {
UiEvent::ScreenShareStopped => {
panic!("failed room switch re-emitted ScreenShareStopped for the torn-down share")
}
UiEvent::Error(e) if e.contains("unexpectedly") => {
panic!("deliberate teardown during a failed room switch reported as a crash: {e}")
}
_ => {}
}
}
}
/// S2 presence gate: a host fault must pull the share ticket off PRESENCE —
/// what remote peers actually see — and must do it BEFORE the reap wait, not
/// after. Nothing on the sharer's own `UiEvent` channel can witness either
/// half (presence is only observable from another node), so this test runs a
/// real second core as an OBSERVER and asserts the sharer's `PeerState.sharing`
/// goes `Some` → `None` on fault.
///
/// The observer runs in a SEPARATE PROCESS (`presence_probe_helper`, this same
/// test binary re-invoked): two in-process cores would load the same
/// `identity.key` and collapse into one node id, and swapping `XDG_CONFIG_HOME`
/// between spawns in-process races other threads' getenv.
///
/// The fake host is a WEDGE — it closes stdout (the fault) but ignores SIGINT
/// and lives until the SIGKILL fallback — so `stop_host` burns the full 2 s
/// grace and TIME becomes the discriminator, exactly like the SIGINT gate:
/// with presence-removal-first the observer sees the ticket clear ~1 s after
/// it appeared (the wedge's pre-fault lifetime); with the old
/// reap-then-presence ordering, only after ~3 s. The bound also makes the
/// "presence removal deleted" mutant fail by timeout instead of passing
/// vacuously.
///
/// Live: two real solo-room cores (audio backend + network bind each).
#[tokio::test]
#[ignore = "live: two real cores in one room (audio backend + network bind), observer subprocess"]
async fn a_host_fault_pulls_the_ticket_off_presence_within_the_grace() {
/// Mirrors `core::teardown::STOP_GRACE` (private): the wait the wedge
/// forces before the SIGKILL fallback reaps it.
const STOP_GRACE_MS: u128 = 2000;
let dir_guard = TempDir(
std::env::temp_dir().join(format!("peerspeak-presence-gate-{}", std::process::id())),
);
let dir = dir_guard.0.clone();
std::fs::create_dir_all(&dir).unwrap();
// Emits its ticket, shares for ~1 s, then closes stdout (the fault) while
// staying alive and ignoring SIGINT, so the reap must wait out the grace.
// The trailing sleep is NOT exec'd on purpose: it forks after stdout is
// closed, so it holds no pipe (the vacuous-staleness trap doesn't apply),
// and it merely idles out after the SIGKILL reaps the shell.
//
// The fake ticket must pass `screenshare::sanitize_ticket` (`endpoint` +
// alphanumerics): the OBSERVER's gossip ingest sanitizes peer-advertised
// tickets, and a garbage one is nulled to `sharing: None` there — the
// probe would never see the share appear and the gate would go vacuous.
let wedged_host = write_fake_pixelpass(
&dir,
"pixelpass-wedges",
"#!/bin/sh\ntrap '' INT\n\
echo '{\"event\":\"ticket\",\"value\":\"endpointaabwxjexzensznfvuudiapn5tyzws3angd2merarm\"}'\n\
sleep 1\nexec 1>&-\nsleep 30\n",
);
let (ui_tx, mut ui_rx) = tokio::sync::mpsc::channel(256);
let controller = CoreController::new(ui_tx);
assert!(controller.send(CoreCommand::SetPixelpassPath(Some(
wedged_host.to_string_lossy().into_owned()
))));
assert!(controller.send(CoreCommand::Join {
name: "presence-gate".into(),
ticket: "create".into(),
room_name: "s2-presence".into(),
input_device: None,
output_device: None,
echo_cancellation: false,
avatar: Default::default(),
}));
let room_ticket = wait_for(&mut ui_rx, "RoomJoined", |ev| match ev {
UiEvent::RoomJoined { ticket, .. } => Some(ticket.clone()),
UiEvent::Error(e) => panic!("join failed: {e}"),
_ => None,
})
.await;
// The observer, in its own process with its own config dir (fresh
// identity). It prints `PROBE …` lines this test parses.
let probe_config = dir.join("probe-config");
std::fs::create_dir_all(&probe_config).unwrap();
let probe = tokio::process::Command::new(std::env::current_exe().unwrap())
.kill_on_drop(true)
.args([
"presence_probe_helper",
"--exact",
"--ignored",
"--nocapture",
])
.env("PEERSPEAK_PROBE_TICKET", &room_ticket)
.env("XDG_CONFIG_HOME", &probe_config)
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()
.expect("spawn the presence probe");
// Only share once the probe is in the room, so it witnesses the ticket
// APPEARING before the fault clears it (otherwise `Some` → `None` could
// both predate its join and the gate would go vacuous).
wait_for(&mut ui_rx, "the probe's PeerJoined", |ev| match ev {
UiEvent::PeerJoined { .. } => Some(()),
UiEvent::Error(e) => panic!("waiting for the probe: {e}"),
_ => None,
})
.await;
assert!(controller.send(CoreCommand::StartScreenShare {
audio: ShareAudioSelection::DesktopShared,
settings: Default::default(),
quality: Default::default(),
}));
wait_for(
&mut ui_rx,
"ScreenShareStarted (wedged host)",
|ev| match ev {
UiEvent::ScreenShareStarted => Some(()),
UiEvent::Error(e) => panic!("share start failed: {e}"),
_ => None,
},
)
.await;
// Sharer-side contract, unchanged by the reorder: Stopped first, the
// explanatory error only after.
wait_for(
&mut ui_rx,
"ScreenShareStopped after the wedge faults",
|ev| match ev {
UiEvent::ScreenShareStopped => Some(()),
UiEvent::Error(e) => panic!("error arrived before ScreenShareStopped: {e}"),
_ => None,
},
)
.await;
let err = wait_for(&mut ui_rx, "the host-death error", |ev| match ev {
UiEvent::Error(e) => Some(e.clone()),
_ => None,
})
.await;
assert!(
err.contains("unexpectedly"),
"the error should say the share ended unexpectedly, got: {err}"
);
let out = tokio::time::timeout(Duration::from_secs(60), probe.wait_with_output())
.await
.expect("probe process outlived its budget")
.expect("probe process wait");
let stdout = String::from_utf8_lossy(&out.stdout);
let stderr = String::from_utf8_lossy(&out.stderr);
assert!(
out.status.success(),
"probe failed ({}).\nstdout:\n{stdout}\nstderr:\n{stderr}",
out.status
);
let cleared_ms: u128 = stdout
.lines()
.find_map(|l| l.strip_prefix("PROBE sharing-cleared "))
.unwrap_or_else(|| {
panic!("probe never saw the ticket clear from presence.\nstdout:\n{stdout}")
})
.trim()
.parse()
.expect("probe delta should be integer millis");
// Presence-removal-first: ~1000 ms (the wedge's pre-fault lifetime).
// Reap-then-presence: ~3000 ms (lifetime + the full stop grace). The
// grace itself splits them with ~1 s of jitter headroom on each side.
assert!(
cleared_ms < STOP_GRACE_MS,
"presence kept advertising the dead share for {cleared_ms} ms after it appeared — \
at or past the wedge lifetime + stop grace, i.e. the ticket was only removed \
AFTER the reap wait instead of before it"
);
assert!(controller.send(CoreCommand::Leave));
}
/// Observer half of `a_host_fault_pulls_the_ticket_off_presence_within_the_grace`,
/// run BY that test as a subprocess. Standalone (no `PEERSPEAK_PROBE_TICKET` in
/// the env — e.g. a plain `--ignored` sweep) it is a no-op pass.
#[tokio::test]
#[ignore = "helper: spawned by the presence gate as a subprocess; standalone it no-ops"]
async fn presence_probe_helper() {
let Ok(room_ticket) = std::env::var("PEERSPEAK_PROBE_TICKET") else {
return;
};
let (ui_tx, mut ui_rx) = tokio::sync::mpsc::channel(256);
let controller = CoreController::new(ui_tx);
assert!(controller.send(CoreCommand::Join {
name: "presence-probe".into(),
ticket: room_ticket,
room_name: String::new(),
input_device: None,
output_device: None,
echo_cancellation: false,
avatar: Default::default(),
}));
wait_for(&mut ui_rx, "RoomJoined (probe)", |ev| match ev {
UiEvent::RoomJoined { .. } => Some(()),
UiEvent::Error(e) => panic!("probe join failed: {e}"),
_ => None,
})
.await;
// Watch the sharer's presence: record when its `sharing` ticket appears,
// report the delta when it clears. Timings on both ends are local-loopback
// arrival times, so the parent's bound compares like with like.
let deadline = tokio::time::Instant::now() + Duration::from_secs(30);
let mut seen_at: Option<std::time::Instant> = None;
loop {
let ev = tokio::time::timeout_at(deadline, ui_rx.recv())
.await
.expect("probe timed out watching for the sharing transition")
.expect("probe ui channel closed");
let sharing = match &ev {
UiEvent::PeerJoined { state, .. } | UiEvent::PeerUpdated { state, .. } => {
state.sharing.is_some()
}
_ => continue,
};
match (&seen_at, sharing) {
(None, true) => {
seen_at = Some(std::time::Instant::now());
println!("PROBE sharing-seen");
}
(Some(t0), false) => {
println!("PROBE sharing-cleared {}", t0.elapsed().as_millis());
break;
}
_ => {}
}
}
assert!(controller.send(CoreCommand::Leave));
}
/// The long-owed Stop Share SIGINT gate (0c half (ii)), against the REAL
/// pixelpass binary: a Stop Share must end the host through the graceful
/// SIGINT path — child exits within [`STOP_GRACE`], no SIGKILL fallback, no
/// "couldn't confirm" warning — because SIGKILL would skip pixelpass's own
/// teardown (it unloads its capture sink on the way out in sink-owning modes).
///
/// The fallback is indistinguishable from success in the event stream (both
/// end in a confirmed reap), so the discriminator is TIME: the fallback path
/// first waits out the full 2 s grace, while a host honouring SIGINT exits in
/// milliseconds. The bound asserts the stop completed inside the grace.
///
/// Live: needs `pixelpass` on `$PATH` plus a real solo room (audio + network).
#[tokio::test]
#[ignore = "live: real pixelpass host + a real solo room (audio backend, network bind)"]
async fn stop_share_ends_the_real_host_via_sigint_within_the_grace() {
/// Mirrors `core::teardown::STOP_GRACE` (private): the graceful wait
/// before the SIGKILL fallback.
const STOP_GRACE: Duration = Duration::from_secs(2);
let (ui_tx, mut ui_rx) = tokio::sync::mpsc::channel(256);
let controller = CoreController::new(ui_tx);
// No override: resolve the real binary from $PATH.
assert!(controller.send(CoreCommand::SetPixelpassPath(None)));
assert!(controller.send(CoreCommand::Join {
name: "sigint-gate".into(),
ticket: "create".into(),
room_name: "s2".into(),
input_device: None,
output_device: None,
echo_cancellation: false,
avatar: Default::default(),
}));
wait_for(&mut ui_rx, "RoomJoined", |ev| match ev {
UiEvent::RoomJoined { .. } => Some(()),
UiEvent::Error(e) => panic!("join failed: {e}"),
_ => None,
})
.await;
// Whole-desktop share: no viewers ever connect, so the real host sits idle
// after its ticket (capture starts on first viewer) — exactly the state a
// Stop Share most often hits.
assert!(controller.send(CoreCommand::StartScreenShare {
audio: ShareAudioSelection::DesktopShared,
settings: Default::default(),
quality: Default::default(),
}));
wait_for(
&mut ui_rx,
"ScreenShareStarted (real pixelpass)",
|ev| match ev {
UiEvent::ScreenShareStarted => Some(()),
UiEvent::Error(e) => panic!("real pixelpass host failed to start: {e}"),
_ => None,
},
)
.await;
let stop_started = std::time::Instant::now();
assert!(controller.send(CoreCommand::StopScreenShare));
wait_for(
&mut ui_rx,
"ScreenShareStopped (real pixelpass)",
|ev| match ev {
UiEvent::ScreenShareStopped => Some(()),
// An Unconfirmed reap surfaces exactly this way; it means the
// SIGINT AND the SIGKILL both failed to end the host.
UiEvent::Error(e) => panic!("stop of the real host was not clean: {e}"),
_ => None,
},
)
.await;
let elapsed = stop_started.elapsed();
assert!(
elapsed < STOP_GRACE,
"stop took {elapsed:?} — at or past the {STOP_GRACE:?} grace, i.e. the \
SIGKILL fallback fired instead of pixelpass honouring SIGINT"
);
// And the late stdout EOF from the SIGINTed host must stay silent (same
// staleness contract the fake-host half pins).
let deadline = tokio::time::Instant::now() + QUIET_WINDOW;
while let Ok(Some(ev)) = tokio::time::timeout_at(deadline, ui_rx.recv()).await {
match ev {
UiEvent::ScreenShareStopped => {
panic!("stale fault from the SIGINTed real host re-emitted ScreenShareStopped")
}
UiEvent::Error(e) if e.contains("unexpectedly") => {
panic!("stale fault from the SIGINTed real host surfaced as an error: {e}")
}
_ => {}
}
}
assert!(controller.send(CoreCommand::Leave));
}