Voice Device Wake Detection Using Directional Audio Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled audio devices struggle to differentiate between user-uttered and device-generated wake expressions, leading to unintended activation due to omnidirectional sound reflections and acoustic complexities in environments.
Innovation Solution
The audio device employs a microphone array with beamforming capabilities to generate directional audio signals, analyzing the number and pattern of these signals to determine if a wake expression is user-generated or device-generated, using machine learning techniques to learn and ignore self-generated expressions, and considering parameters like speaker output, echo characteristics, and loudness to make this distinction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the audio device uses omnidirectional microphone detection to capture wake expressions from all directions, then the device can detect user-uttered wake expressions from any position, but it also detects device-generated wake expressions due to sound reflections and acoustic echoes
Solution Approach 1:
The patent segments the audio signal detection by creating multiple directional audio signals from the omnidirectional microphone input. Each directional signal focuses on a specific spatial sector, allowing the system to distinguish between sounds coming from different directions. This segmentation enables the device to identify whether a wake expression originates from the speaker direction (device-generated) or from other directions (user-uttered).
Solution Approach 2:
The patent introduces directional audio signals as an intermediary layer between the omnidirectional microphone detection and the wake expression recognition. These directional signals act as mediators that process the raw audio input through spatial filtering, providing directional information that helps distinguish between user-uttered and device-generated wake expressions without requiring additional physical microphones.
2Ease of operation
If the device speaks the wake expression aloud to provide feedback, then the user can confirm the device heard them, but the device may mistakenly detect its own spoken wake expression as a new user command
Solution Approach 1:
The patent applies preliminary action by analyzing the directional characteristics of the audio signal before triggering wake expression recognition. The system determines the direction of the detected wake expression and compares it with the speaker's orientation in advance. If the wake expression is detected coming from the speaker direction, the system preemptively prevents false activation by ignoring the detection or marking it as device-generated, thus avoiding the unintended activation problem before it occurs.
3Device complexity
If the device uses simple wake expression detection without directional analysis, then the device complexity remains low, but the device cannot distinguish between user-uttered and device-generated wake expressions
Solution Approach 1:
The patent adds a spatial dimension to the wake expression detection by creating directional audio signals from omnidirectional microphone input. Instead of simply detecting the presence of a wake expression, the system analyzes the directional characteristics of the audio signal, effectively adding spatial information as another dimension to the detection process. This approach enables origin identification without requiring a complex array of multiple microphones, thus balancing measurement precision with device complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Effectively reduces false activations by accurately identifying user-uttered wake expressions, enhancing the reliability and precision of voice interaction systems in various environments.
Implementation Method 1
a microphone array that generates a plurality of directional audio signals
Implementation Method 2
The audio beamformer generates a plurality of directional audio signals based on the input audio
Implementation Method 3
omnidirectional sound reflections and acoustic complexities in environments
Data Source
AI summary
A speech-based audio device may be configured to detect a user-uttered wake expression. For example, the audio device may generate a parameter indicating whether output audio is currently being produced by an audio speaker, whether the output audio contains speech, whether the output audio contains a predefined expression, loudness of the output audio, loudness of input audio, and/or an echo characteristic. Based on the parameter, the audio device may determine whether an occurrence of the predefined expression in the input audio is a result of an utterance of the predefined expression by a user.


