Wake Suppression for Audio Devices via Voice Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech-enabled devices often falsely trigger when they detect the wake word in audio streams from other devices, leading to unwanted activation and potential security breaches, especially when audio-playing devices play content that includes the wake word.
Innovation Solution
Implementing a system where the listening device uses a voice verification module to differentiate between live and reproduced audio, suppressing wake-up when the wake word is detected in audio streams from audio-playing devices, either through acoustic analysis or network communication, to prevent unnecessary activation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the listening device continuously monitors audio for wake words to enable responsive operation, then the device can quickly respond to user commands, but the device falsely triggers when detecting wake words in audio streams from other devices
Solution Approach 1:
The listening device performs preliminary actions by analyzing audio characteristics before triggering wake-up. The system compares acoustic features such as spectral content, temporal patterns, and signal origin to determine whether the detected wake word comes from a nearby user or from an audio-playing device, thereby preventing false triggers while maintaining responsive operation
Solution Approach 2:
The system introduces an intermediary verification process between audio detection and wake-up triggering. This intermediary layer analyzes the audio signal characteristics and determines the source of the wake word before allowing state transition, effectively filtering out false triggers from audio streams while preserving legitimate user commands
2Reliability
If the device uses acoustic analysis to distinguish live and reproduced audio, then false triggers are reduced, but the device complexity increases
Solution Approach 1:
The audio analysis process is segmented into multiple independent feature extraction modules, each analyzing specific acoustic characteristics such as spectral content, temporal patterns, and signal origin. This segmentation allows the system to handle complex audio differentiation through modular processing, reducing overall system complexity while maintaining high wake suppression accuracy
3Use of energy by stationary object
If the listening device remains in idle state to conserve energy, then power consumption is reduced, but the device cannot respond to wake words in audio streams
Solution Approach 1:
The listening device performs partial audio analysis actions while remaining in idle state. Instead of fully processing all audio streams for wake words, the system performs selective monitoring that identifies audio streams from playing devices and suppresses wake-up triggers only for those streams, thereby maintaining energy efficiency while achieving appropriate adaptability
Data Source
AI summary
A system and method are disclosed for ignoring a wakeword received at a speech-enabled listening device when it is determined the wakeword is reproduced audio from an audio-playing device. Determination can be by detecting audio distortions, by an ignore flag sent locally between an audio-playing device and speech-enabled device, by and ignore flag sent from a server, by comparison of received audio played audio to a wakeword within an audio-playing device or a speech-enabled device, and other means.


