Voice Audio Output Adjustment for Accurate Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Background noise and other audio signals can hinder the accuracy of automatic speech recognition (ASR) in voice-controlled devices, making it difficult for them to identify voice commands effectively.

Innovation Solution

A voice-controlled device alters its audio output by attenuating, pausing, or modifying the audio signal based on predefined phrases spoken by the user, increasing the signal-to-noise ratio (SNR) to improve speech recognition accuracy. This is achieved by determining user proximity, direction, and preferences, and adjusting the audio output accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio output is maintained at normal levels, then entertainment functionality is preserved, but speech recognition accuracy deteriorates due to background noise interference

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidbackground noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary detection of user presence and intent before the speech recognition process begins. By detecting predefined phrases or user presence in advance, the system proactively adjusts audio output levels to create optimal conditions for accurate speech recognition, preventing noise interference before it degrades recognition accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The audio output system dynamically adjusts its characteristics (volume, pause, modification) based on real-time detection of user presence and speech patterns. This dynamic adaptation allows the system to optimize speech recognition conditions while maintaining entertainment functionality when users are not present, resolving the contradiction between reliable recognition and noise interference

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If audio output is attenuated or paused to improve speech recognition, then speech recognition accuracy improves, but entertainment experience deteriorates

Engineering Contradiction:
Improvevoice command identification accuracyVSAvoidentertainment functionality
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system continuously monitors audio signals for predefined phrases and user presence, using this feedback to intelligently control audio output adjustments. When user speech is detected, the system provides feedback to pause or attenuate entertainment audio, ensuring accurate voice command recognition while automatically resuming entertainment functionality when users are not present

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system changes audio output parameters (volume level, pause state, frequency modification) based on detected user presence and speech characteristics. These parameter adjustments are temporary and context-dependent, allowing the system to improve voice command identification accuracy while minimizing disruption to the overall entertainment experience

Inventive Principle:
Principle #35Parameter changes

3Reliability

If users speak louder to overcome background noise, then speech recognition accuracy improves, but user convenience deteriorates

Engineering Contradiction:
Improvevoice command decoding reliabilityVSAvoiduser interaction comfort
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system applies preliminary anti-action by attenuating or pausing background entertainment audio before the user needs to speak. This preemptive noise reduction eliminates the need for users to increase their voice volume, maintaining both high speech recognition reliability and user interaction comfort without requiring users to yell over background noise

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentUS10354649B2Altering audio to improve automatic speech recognition
Publication Date: 2019.07.16 AMAZON TECH INC
  • US10354649B2 patent drawing
  • US10354649B2 patent drawing
  • US10354649B2 patent drawing

AI summary

Techniques for altering audio being output by a voice-controlled device, or another device, to enable more accurate automatic speech recognition (ASR) by the voice-controlled device. For instance, a voice-controlled device may output audio within an environment using a speaker of the device. While outputting the audio, a microphone of the device may capture sound within the environment and may generate an audio signal based on the captured sound. The device may then analyze the audio signal to identify speech of a user within the signal, with the speech indicating that the user is going to provide a subsequent command to the device. Thereafter, the device may alter the output of the audio (e.g., attenuate the audio, pause the audio, switch from stereo to mono, etc.) to facilitate speech recognition of the user's subsequent command.