Dual Transducer Speech Authentication via Signal Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speaker recognition systems face challenges in reliably authenticating speech signals, particularly in text-independent scenarios where the spoken command does not match the enrolled user's voice pattern, leading to potential unauthorized actions.
Innovation Solution
A method and system utilizing dual transducers, including a microphone and an accelerometer, to perform voice biometrics by correlating speech signals from both transducers, determining if the speech is from an enrolled user and ensuring correlations satisfy a predetermined condition for authentication, thereby distinguishing between the enrolled user and subsequent speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If text-independent speaker recognition is used for commands following the trigger phrase, then the system can handle arbitrary user commands, but the reliability of speaker authentication deteriorates
Solution Approach 1:
The speech signal is divided into multiple segments: the trigger phrase portion and the command portion. Different authentication methods are applied to each segment - text-dependent verification for the trigger phrase and text-independent verification for the command, allowing the system to maintain high reliability while handling arbitrary commands
Solution Approach 2:
The system transitions from single-microphone audio analysis to multi-dimensional analysis by incorporating both audio signals from the microphone and vibration signals from the accelerometer, creating additional verification dimensions that improve authentication reliability for text-independent commands
2Device complexity
If only a single microphone is used for speech recognition, then the device structure is simple, but the ability to distinguish between enrolled user and subsequent speakers deteriorates
Solution Approach 1:
The system merges two different types of transducers - a microphone for capturing acoustic speech signals and an accelerometer for detecting bone-conducted vibrations. By combining these complementary sensing modalities, the system achieves superior speaker identification accuracy while maintaining reasonable device complexity
Solution Approach 2:
The accelerometer acts as an intermediary sensor that captures bone-conducted vibrations which provide additional biometric information about the speaker. This intermediary measurement channel enhances speaker distinction capability without requiring multiple microphones
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances the reliability of speech signal authentication, reducing the risk of unauthorized commands by confirming the speaker's identity through correlations between ambient and bone-conducted sound signals, even after the trigger phrase is spoken.
Implementation Method 1
distinguishing between the enrolled user and subsequent speakers through correlations between ambient and bone-conducted sound signals
Data Source
AI summary
A speech signal is received by a device comprising first and second transducers, and the first transducer comprises a microphone. A method comprises performing a first voice biometric process on speech contained in a first part of a signal received by the microphone, in order to determine whether the speech is the speech of an enrolled user. A first correlation is determined, between said first part of the signal received by the microphone and a corresponding part of the signal received by the second transducer. A second correlation is determined, between said second part of the signal received by the microphone and the corresponding part of the signal received by the second transducer. It is then determined whether the first correlation and the second correlation satisfy a predetermined condition. If it is determined that the speech contained in the first part of the received signal is the speech of an enrolled user and that the first correlation and the second correlation satisfy the predetermined condition, the received speech signal is authenticated.


