Evaluate the entire voice experience under realistic conditions.
1. Latency breakdown
Measure end-of-speech detection, recognition, model response and first audio separately. Optimizing the fastest stage may have little effect on perceived delay.
2. Accessible alternatives
Provide readable transcripts, touch controls and a text input path. Voice-only interaction excludes users and fails in noisy or private settings.
3. Reliability matrix
Test headset changes, phone interruptions, offline behavior and repeated sessions. Verify capture and playback resources are released after every exit path.
Worked scenario
A fast model still feels slow because turn detection waits too long after the user stops speaking. The measurements reveal the true bottleneck.
Apply it
Publish a latency table and a device test matrix including interruption and text fallback.
Check your understanding
Your voice feature remains usable when speech input or output is unavailable. Explain the decision and show evidence from your implementation or design. If you cannot demonstrate it yet, revisit the relevant section before continuing.