Headphone Imposter Rejection via Accelerometer VAD
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio systems struggle to accurately determine whether a key-phrase is spoken by the wearer of headphones, especially at fast speech rates, leading to false negatives and improper activation of virtual personal assistants.
Innovation Solution
The system uses both a microphone and an accelerometer in the headphones to continuously receive signals, generating a voice activity detection (VAD) signal and score to differentiate between the wearer and imposters, ensuring accurate key-phrase detection regardless of speech rate.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If only microphone signal is used for key-phrase detection, then the system is simple to operate, but the reliability of determining whether the wearer spoke the key-phrase deteriorates
Solution Approach 1:
The patent combines microphone signal processing with accelerometer signal processing to create a more reliable determination system. The microphone detects the key-phrase while the accelerometer detects bone conduction vibrations, and both signals are merged through signal processing to generate a confidence score that reliably determines whether the wearer spoke the key-phrase.
Solution Approach 2:
The patent introduces signal processing algorithms as intermediaries that process both microphone and accelerometer signals. These algorithms include bandpass filtering, magnitude calculation, and confidence score generation, acting as mediators that transform raw sensor data into reliable determination about wearer speech.
2Reliability
If accelerometer signal processing is added to improve reliability, then the reliability of key-phrase detection improves, but the device complexity increases
Solution Approach 1:
The patent segments the signal processing into distinct functional blocks: microphone signal processing, accelerometer signal processing, signal combination, and confidence score generation. Each segment handles a specific aspect of the determination process, making the overall complex system more manageable and implementable through modular processing stages.
3Measurement precision
If continuous signal processing is used to detect fast speech rates, then the measurement precision improves, but the use of energy increases
Solution Approach 1:
The patent implements periodic action through frame-based processing where signals are divided into discrete time frames (e.g., 20ms frames). This periodic processing approach maintains measurement precision for fast speech rates while reducing energy consumption compared to continuous processing, as the system processes signals in periodic intervals rather than continuously.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively rejects imposters and ensures proper activation of virtual personal assistants by determining whether the key-phrase is spoken by the wearer, regardless of speech rate, thereby improving the reliability of voice command recognition.
Implementation Method 1
the accelerometer signal may be different when an imposter speaks versus when the wearer speaks (e.g., the accelerometer signal may have an energy level that is less than an accelerometer energy threshold, and/or the accelerometer signal may not be correlated with a microphone signal produced by the microphone during times at which the imposter is speaking but the wearer is not)
Implementation Method 2
a microphone signal is received from at least one microphone in the headphone
Data Source
AI summary
A signal processing method to determine whether or not a detected key-phrase is spoken by a wearer of a headphone. The method receives an accelerometer signal from an accelerometer in a headphone and receives a microphone signal from at least one microphone in the headphone. The method detects a key-phrase using the microphone signal and generates a voice activity detection (VAD) signal based on the accelerometer signal. The method determines whether the VAD signal indicates that the detected key-phrase is spoken by a wearer of the headphone. Responsive to determining that the VAD signal indicates that the detected key-phrase is spoken by the wearer of the headphone, triggering a virtual personal assistant (VPA).


