Audio Controlled Device Non-Speech Command Noise Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition (ASR) in devices is hindered by background noise, making it difficult for devices to accurately interpret voice commands, especially when audio is being played loudly, requiring users to yell to be heard.

Innovation Solution

An audio-controlled device identifies a non-speech command from the user, such as clapping, to alter its audio output, reducing noise and increasing the signal-to-noise ratio, allowing for more accurate speech recognition by attenuating or modifying the audio output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If audio is played loudly in the environment, then the audio output quality is improved, but the speech recognition accuracy deteriorates due to increased background noise

Engineering Contradiction:
Improveaudio output powerVSAvoidspeech recognition accuracy
Core Design Contradiction:
PowerVSMeasurement precision

Solution Approach 1:

The system performs preliminary detection of non-speech commands (clapping, snapping, whistling) before processing speech recognition. When such commands are detected, the system proactively adjusts the audio output state (attenuation, pause, or stop) in advance, creating optimal conditions for subsequent speech recognition without requiring users to yell over loud audio

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The audio output system transitions from a static state to a dynamic, adjustable state based on detected user commands. The system can dynamically switch between different output states (normal playback, attenuated volume, paused, or stopped) depending on the detected non-speech command, allowing flexible adaptation to user needs while maintaining speech recognition accuracy

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If users speak at normal volume, then the ease of operation is improved, but the speech recognition accuracy deteriorates when audio is playing loudly

Engineering Contradiction:
Improveuser speaking comfortVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system implements a feedback mechanism where non-speech commands (clapping, snapping, whistling) serve as user feedback signals. When detected, these commands trigger automatic adjustments to the audio output state, creating a feedback loop that allows users to maintain normal speaking volumes while ensuring accurate speech recognition through appropriate audio level management

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9087520B1Altering audio based on non-speech commands
Publication Date: 2015.07.21 AMAZON TECH INC
  • US9087520B1 patent drawing
  • US9087520B1 patent drawing
  • US9087520B1 patent drawing

AI summary

Techniques for altering audio being output by an audio-controlled device, or another device, to enable more accurate automatic speech recognition (ASR) by the audio-controlled device. For instance, an audio-controlled device may output audio within an environment using a speaker of the device. While outputting the audio, a microphone of the device may capture sound within the environment and may generate an audio signal based on the captured sound. The device may then analyze the audio signal to identify a predefined non-speech command issued by a user within the environment. In response to identifying the predefined non-speech command, the device may somehow alter the output of the audio for the purpose of reducing the amount of noise within subsequently captured sound.