TTS listening scorecard (JOE-1735)¶
Not MOS. Use this form for bias-controlled human comparison of Kitten, Kokoro, and future adapters. Record method, listener count, and date.
Session metadata¶
| Field | Value |
|---|---|
| Date | |
| Aurum version / commit | |
| Adapters / models / voices | |
| Listener count | |
| Blinding (yes/no + method) | |
| Hardware |
Items (score 1–5)¶
| Fixture id | Adapter | Intelligibility | Naturalness | Numbers/abbrev | Join artifacts | Notes |
|---|---|---|---|---|---|---|
| tts_short | ||||||
| tts_numbers | ||||||
| tts_long_join |
Disposition¶
- Production-quality candidate? Y/N
- Blocking issues:
- Follow-ups:
Objective PCM checks remain separate (aurum_core::eval duration/PCM guards).
Round 001 (recorded)¶
See repository path evals/reports/listening/listening-round-001.json and the summary in evidence-v004.md. Round 001 is a single-listener, non-blinded pilot — not MOS.