Audio Controlled Device Non-Speech Command Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition (ASR) in devices is hindered by background noise, making it difficult for devices to accurately interpret voice commands, especially when audio is being played loudly, requiring users to yell to be heard.
Innovation Solution
An audio-controlled device identifies a non-speech command from the user, such as clapping, to alter its audio output, reducing noise and increasing the signal-to-noise ratio, allowing for more accurate speech recognition by attenuating or modifying the audio output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If audio is played loudly in the environment, then the audio output quality is improved, but the speech recognition accuracy deteriorates due to increased background noise
Solution Approach 1:
The system performs preliminary detection of non-speech commands (clapping, snapping, whistling) before processing speech recognition. When such commands are detected, the system proactively adjusts the audio output state (attenuation, pause, or stop) in advance, creating optimal conditions for subsequent speech recognition without requiring users to yell over loud audio
Solution Approach 2:
The audio output system transitions from a static state to a dynamic, adjustable state based on detected user commands. The system can dynamically switch between different output states (normal playback, attenuated volume, paused, or stopped) depending on the detected non-speech command, allowing flexible adaptation to user needs while maintaining speech recognition accuracy
2Ease of operation
If users speak at normal volume, then the ease of operation is improved, but the speech recognition accuracy deteriorates when audio is playing loudly
Solution Approach 1:
The system implements a feedback mechanism where non-speech commands (clapping, snapping, whistling) serve as user feedback signals. When detected, these commands trigger automatic adjustments to the audio output state, creating a feedback loop that allows users to maintain normal speaking volumes while ensuring accurate speech recognition through appropriate audio level management
Data Source
AI summary
Techniques for altering audio being output by an audio-controlled device, or another device, to enable more accurate automatic speech recognition (ASR) by the audio-controlled device. For instance, an audio-controlled device may output audio within an environment using a speaker of the device. While outputting the audio, a microphone of the device may capture sound within the environment and may generate an audio signal based on the captured sound. The device may then analyze the audio signal to identify a predefined non-speech command issued by a user within the environment. In response to identifying the predefined non-speech command, the device may somehow alter the output of the audio for the purpose of reducing the amount of noise within subsequently captured sound.


