Voice Playback Switching for Noise-Robust Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Background noise and other audio signals can hinder the accuracy of automatic speech recognition (ASR) in voice-controlled devices, making it difficult for them to identify user voice commands effectively.

Innovation Solution

A voice-controlled device alters its audio output by attenuating, pausing, or modifying the audio signal based on predefined phrases spoken by the user, increasing the signal-to-noise ratio and improving speech recognition accuracy by reducing noise interference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio output is maintained at normal levels, then entertainment functionality is preserved, but speech recognition accuracy deteriorates due to background noise interference

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidbackground noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system preemptively reduces or pauses audio output when it detects that speech recognition is about to occur (based on wake words or contextual cues). This preliminary anti-action prevents the audio from interfering with the recognition process before the interference can degrade accuracy, thereby resolving the contradiction between maintaining entertainment functionality and ensuring recognition accuracy.

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The audio output level is made dynamic rather than static. The system continuously adjusts the audio output level based on the operational context - maintaining normal levels during entertainment playback but reducing or pausing levels during speech recognition events. This dynamic adaptation allows the system to satisfy both requirements at different times, resolving the fundamental contradiction.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If audio output is attenuated or paused during speech recognition, then speech recognition accuracy improves, but entertainment continuity deteriorates

Engineering Contradiction:
Improvevoice command detection accuracyVSAvoidaudio playback continuity
Core Design Contradiction:
Measurement precisionVSDuration of action of stationary object

Solution Approach 1:

The system implements periodic monitoring of speech patterns and audio context to determine when to pause or resume entertainment playback. By using periodic detection of wake words, speech patterns, or silence periods, the system can intelligently interrupt and resume audio playback, minimizing disruption to entertainment continuity while ensuring accurate voice command detection during recognition events.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system uses feedback from the speech recognition process to control audio output. When speech is detected or a wake word is recognized, the system provides feedback to pause or attenuate entertainment audio. After recognition completes successfully, the feedback loop triggers resumption of audio playback. This closed-loop control balances recognition accuracy with entertainment continuity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9916830B1Altering audio to improve automatic speech recognition
Publication Date: 2018.03.13 AMAZON TECH INC
  • US9916830B1 patent drawing
  • US9916830B1 patent drawing
  • US9916830B1 patent drawing

AI summary

Techniques for altering audio being output by a voice-controlled device, or another device, to enable more accurate automatic speech recognition (ASR) by the voice-controlled device. For instance, a voice-controlled device may output audio within an environment using a speaker of the device. While outputting the audio, a microphone of the device may capture sound within the environment and may generate an audio signal based on the captured sound. The device may then analyze the audio signal to identify speech of a user within the signal, with the speech indicating that the user is going to provide a subsequent command to the device. Thereafter, the device may alter the output of the audio (e.g., attenuate the audio, pause the audio, switch from stereo to mono, etc.) to facilitate speech recognition of the user's subsequent command.