Variable-Step Acoustic Echo Cancellation for Playback Wake-Word Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled devices face challenges in distinguishing between user-spoken wake-words and wake-words present in audio playback, leading to false triggers that pause or attenuate audio playback unexpectedly, causing user inconvenience and missed audio content.
Innovation Solution
Utilizing the variable step size (Vss) of an acoustic echo cancellation (AEC) unit to differentiate between user-spoken wake-words and wake-words in downlink audio by monitoring the convergence rate of the AEC's adaptive filter, allowing for efficient suppression or acceptance of detected wake-words without interrupting playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If wake-word detection is implemented during audio playback, then the device can respond to user commands, but false triggers occur when playback contains wake-word-like audio causing unexpected pauses
Solution Approach 1:
The patent introduces an acoustic echo cancellation (AEC) unit as an intermediary component between the audio playback system and wake-word detection system. The AEC unit processes the audio stream and provides cleaned audio signals to the wake-word detector, enabling it to distinguish between wake-words from playback and wake-words from user speech. This intermediary processing layer resolves the conflict by improving detection reliability without disrupting playback continuity.
Solution Approach 2:
The patent utilizes variable step size (Vss) parameter of the AEC unit to dynamically adjust the convergence rate of the adaptive filter. By changing this parameter, the system can optimize the balance between echo cancellation performance and computational efficiency, thereby improving wake-word detection accuracy during playback without causing false triggers that would interrupt audio continuity.
2Reliability
If a secondary wake-word engine is used to detect wake-words in playback audio, then false triggers can be prevented, but computational resources and power consumption increase
Solution Approach 1:
The patent makes the existing AEC unit perform a dual function: its primary function of echo cancellation is enhanced with a secondary function of wake-word detection assistance. By utilizing the Vss parameter and decision flags from the AEC unit, the system achieves false trigger prevention without adding a separate wake-word engine, thereby maintaining energy efficiency while improving reliability.
Solution Approach 2:
The AEC unit generates decision flags that indicate whether detected wake-words are likely from playback or user speech. This self-generated information allows the wake-word detection system to make accurate decisions without requiring additional computational resources for a separate analysis engine, reducing power consumption while maintaining false trigger prevention.
3Extent of automation
If wake-word detection is activated during audio playback, then user commands can be recognized, but audio playback is paused or attenuated causing user inconvenience
Solution Approach 1:
The AEC unit acts as a mediator that provides contextual information about the audio source to the wake-word detection system. By analyzing the audio signal characteristics through the AEC processing, the system can determine whether a detected wake-word originates from playback or user speech, enabling command recognition to proceed without pausing or attenuating the audio playback experience.
Data Source
AI summary
Devices and techniques are generally described for wake word suppression using variable step size of an acoustic echo cancellation (AEC) unit. A reference signal representing an audio stream may be sent to an acoustic echo cancellation (AEC) unit. A microphone may receive an input audio signal and send the input audio signal to the AEC unit. The AEC unit may determine a first set of variable step size (Vss) values over the first time period. Vss values may define a rate at which the AEC unit determines a transfer function between the reference signal and the first input audio signal. A wake-word may be detected during the first time period. A determination may be made that the wake-word is part of the audio output by the loudspeaker based at least in part on the first set of Vss values.


