ADR: Private realtime voice (Desktop Local)¶
| Field | Value |
|---|---|
| Status | Accepted |
| Date | 2026-07-24 |
| Linear | JOE-1096 epic; continuous VAD JOE-1104 |
| Milestone | Private Realtime Voice |
Context¶
Operators want push-to-talk / conversational voice against local OpenCode sessions without sending microphone audio to cloud STT/TTS vendors by default.
Sibling work:
- Aurum — on-device speech-to-text (
aurum-stt/ whisper.cpp; PCM-first library API). Optional remote paths exist but are never the default. - ZephyrFlow — macOS menu-bar dictation product (Whisper, Local Only). Reference UX for PTT; not a TTS engine and not the Open Cowork voice host.
Open Cowork does not ship placeholder “voice replies” controls. Voice must not appear as a half-wired toggle.
Decision¶
1. Product surface¶
| Rule | Decision |
|---|---|
| Default authority | Desktop Local only |
| Cloud Desktop / Cloud Web | Blocked — voice.* support APIs are not_supported |
| Gateway / paired | Blocked until a future ADR |
| Feature flag | features.voice — secondary, default off (progressive disclosure) |
| Public claims | No “private voice shipping” until V2 PTT UI + local STT path are real |
2. Engine split¶
| Role | Owner | Notes |
|---|---|---|
| STT | Aurum (local_only / on-device) | PCM in → text out; no API key by default |
| TTS | Sibling / separate engine | Not Aurum. MVP = OS system speech (system_os, macOS say+afplay); Piper/neural sidecar deferred |
| Orchestration | Open Cowork voice host (Electron main / native side) | Outside Chromium renderer |
3. Architecture boundary¶
Renderer (UI only)
│ IPC: voice:status | voice:session:* | voice:tts:* | voice:partial | voice:final
▼
Voice host (main / native — never Node in renderer)
│ mic capture (OS APIs)
│ STT via Aurum local_only
│ TTS via sibling engine (OS system speech MVP)
▼
OpenCode session prompt / stream (existing session path)
Rules:
- No raw audio bytes on the renderer IPC path by default. Prefer partial/final text and host-owned playback. Status may include capture frame counts only.
- Chromium
getUserMedia/ Electron sessionmediafor the Studio renderer stays denied unless a future ADR chooses an explicit renderer capture mode. - OS microphone permission is owned by the voice host, not by ad-hoc Settings toggles.
- Cloud Web must never request mic for Open Cowork Studio (support matrix + browser matrix).
- Capture (JOE-1097): host accumulates mono 16 kHz f32 PCM in main (
VoicePcmBuffer). Default backend is ffmpeg when available; tests injectFakeVoiceCapture. PCM is cleared on stop/cancel and is never sent to the renderer. - STT (JOE-1101): on release/stop, host runs Aurum with local provider only (
--provider local, cleanup via rules). Default modeltiny-q5_1. local_only fail-closed unlessOPEN_COWORK_AURUM_ALLOW_DOWNLOAD=1. OpenRouter/cloud ASR is never on the default path. Final text is emitted as avoice:eventfinal payload (text only). - PTT UI (JOE-1105): Chat and Home composers show a mic control when
features.voiceis on and the workspace support matrix allowsvoice.capture+voice.stt(Desktop Local). Click-to-toggle is the shipped interaction (start → Listening → click again → Transcribing → inject text into the composer). Control is hidden on Cloud Web / unsupported authorities. - Partials during PTT (JOE-1102): While listening, the host runs a PartialClock (min ~1s audio, ~15s rolling window, ~1.2s interval, RMS energy gate) and emits
voice:eventpartial payloads (text only). This is not a continuous Whisper stream — the host decides when to call STT. The composer snapshots a baseline at PTT start; partials/finals replace the dictation segment after the baseline (cancel/error restores baseline). - PTT hotkey (JOE-1110, JOE-1188): Default accelerator
CmdOrCtrl+Shift+Space(Edit menu “Toggle Voice Dictation”). Scope is app-focused only — not OS-wide Accessibility paste into other apps. Settings → Privacy validates a configurable Electron accelerator whenfeatures.voiceis on; a valid saved value updates both the focused-window matcher and Edit menu immediately and is restored on restart. Product shortcut conflicts such as the command palette (CmdOrCtrl+Shift+P) are rejected before save. - Local TTS (JOE-1108): Sibling of Aurum STT. Decision: MVP uses OS system voices (macOS
saysynthesize to temp AIFF +afplayplayback). Not Aurum, not cloud TTS, not ChromiumspeechSynthesisin the renderer. Host owns synthesize + playback; IPC carries text only (voice:tts:speak/cancel/voices). Linux/Windows OS backends and Piper/neural packaging are explicit follow-ups (no download on default path). Claim boundary: “local OS speech when available” — not “neural private TTS GA”. - Read-aloud (JOE-1103): Per-message Read aloud on completed assistant bubbles when
features.voice+voice.ttsauthority + host TTS ready. Default off (no auto-read of streaming tokens). Streaming strategy: wait for complete message only (live placeholders have no actions). Stop cancels host playback immediately; optional skip drains the next queued item. Starting PTT callsstopReadAloud(barge-in prep). Markdown is stripped to plain text before speak. - Conversation controller (JOE-1107): Pure state machine
Idle → Listening → FinalizingSTT → Prompting → Streaming → Speaking → Idledrives PTT-gated voice turns when the user enables conversation mode (default off). Release mic → STT final →session.prompt→ wait for generation idle → local TTS of the latest assistant message. Cancel / barge-in stops TTS, cancels listen, and aborts generation. - Continuous VAD + barge-in (JOE-1104): Opt-in energy VAD (RMS gate, not neural) on the voice host when conversation mode and continuous listen are both on (default off — never silent always-on). Host auto-finalizes on end-of-utterance silence or max-listen timeout; UI shows a privacy mic-armed indicator. After
SPEAK_DONEwith continuous on, the machine re-arms listening. During TTS, host monitors mic energy and emitsvadbarge_in(cancels local speak; renderer aborts gen + re-listens). Still text/status/vad IPC only — no raw audio to the renderer. - First-run assets (JOE-1109): STT models are local files under OC
userData/voice/aurum(or system Aurum cache). Default local_only fail-closed when missing — no silent network download. Integrity uses size floor + optional sibling.sha256.OPEN_COWORK_AURUM_ALLOW_DOWNLOAD=1is explicit operator opt-in for a file fetch residual (never audio/transcript upload). Settings → Privacy shows offline-ready status + Ensure local model (copy from system cache when present). TTS readiness remains OS speech probe. - Packaging (JOE-1106): Optional sidecars under packaged
resources/voice/(aurum/ffmpeg); resolution prefers env → packaged path → PATH. CI packages ship the folder + README without pre-bundled model weights or aurum-ffi dylibs (fail closed when missing). macOS supported when tools present; Windows/Linux best-effort with residual TTS backends. Codesign any drop-in binaries with the release pipeline. - Accessibility (JOE-1112): Mic control exposes
aria-pressed/aria-busy; phase chrome uses a single polite live region; disabled reasons render as a status region (permission denied / model missing / unsupported workspace), not a mute one-liner.
4. Workspace support APIs¶
| API | Desktop Local | Cloud / browser / remote |
|---|---|---|
voice.capture | supported (authority) | not_supported |
voice.stt | supported (authority) | not_supported |
voice.tts | supported (authority) | not_supported |
voice.conversation | supported (authority) | not_supported |
“Supported” means this authority may host voice, not that every UI control is complete. Runtime readiness is reported via voice:status (ready / deferred / unavailable). UI stays behind features.voice.
5. Progressive disclosure¶
- Omit or set
features.voice: falsein public default config. - Soft enablement warning via
desktopFeatureEnablementWarnings(local-only, Aurum STT, sibling TTS, host not renderer). - Do not market voice until release checklist evidence is green.
Non-goals (this milestone)¶
- Cloud / multi-tenant voice
- Using Aurum as TTS
- Shipping ZephyrFlow inside Open Cowork
- Renderer-owned continuous listening without PTT policy
- Replacing chat text input as the only interaction mode
Security audit (JOE-1111)¶
Greppable claim: private voice is private-by-construction on the default Desktop Local path.
| Control | Rule | Evidence |
|---|---|---|
| Logs | Lengths + engine metadata only (sttLogMeta / ttsLogMeta) — never transcript text or PCM | apps/desktop/src/main/voice-stt.ts, voice-security.ts, tests/voice-security.test.ts |
| Network STT | Aurum --provider local only; OpenRouter key cleared on spawn; model missing → fail-closed when local_only | voice-stt.ts, security tests |
| Network TTS | OS system speech sibling; no cloud TTS vendor on default path | voice-tts.ts |
| IPC | Status / partial / final text / vad / assets — no raw samples | voice-handlers.ts, preload channel list |
| Renderer mic | getUserMedia / session media denied when captureMode is voice_host | voice-permission-policy.ts |
| Support matrix | Cloud Web / non-local: all voice.* APIs not_supported | browserCloudWorkspaceSupport, workspace support store |
| Model download | Default off; OPEN_COWORK_AURUM_ALLOW_DOWNLOAD=1 = file weights only, never audio upload | voice-assets.ts |
Residual risks (accepted)¶
| ID | Risk | Mitigation |
|---|---|---|
| R-VOICE-01 | Partial/final IPC carries transcript text (product UX) | No audio on IPC; do not ship adoption telemetry of free-text transcripts |
| R-VOICE-02 | Short-lived temp WAV for Aurum CLI | OS temp dir; rmSync in finally after each transcribe |
| R-VOICE-03 | OS TTS may write temp AIFF/WAV for playback | Local host only; cancel best-effort cleanup |
| R-VOICE-04 | Opt-in model download env | Default off; Settings copy; never audio |
| R-VOICE-05 | getLastTranscript() host memory for tests | Not on preload/IPC |
| R-VOICE-06 | Aurum/ffmpeg not pre-bundled in CI packages | Packaging README + fail-closed status; optional drop-in |
| R-VOICE-07 | Windows/Linux OS TTS residual | Best-effort claim; unavailable status when tools missing |
Automated gates: tests/voice-security.test.ts, tests/voice-packaging.test.ts (plus existing STT/TTS/scaffold tests). Dogfood and release-claim checks live in voice-private-dogfood.md.
Consequences¶
- Shared package grows
voiceIPC types andvoice.*workspace support keys. - Electron permission guards stay fail-closed for renderer media; docs state host ownership.
- Residual purity risk for incomplete secondaries remains soft-warn only (same pattern as other Studio flags).
Value review (2026-10)¶
Voice is the largest recent subsystem (~6k LOC plus STT/TTS asset packaging) whose value is asserted from dogfood rather than demonstrated by sustained usage. Review by 2026-10-25 (three months post-close-out): if dogfood / downstream usage has not stuck, the subsystem is the primary deletion candidate under the no-bloat policy; if it has, invest in conversation-mode polish next. Record the outcome here.
Related¶
- Progressive disclosure
- Product contract
- Release checklist
- Aurum: https://github.com/joe-broadhead/aurum
- ZephyrFlow: https://github.com/joe-broadhead/zephyr-flow