Bone-Conduction Biometric Speaker Recognition in Noisy Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition systems face challenges in achieving high reliability and efficiency, particularly in noisy environments, as they rely heavily on computational power-intensive methods like MFCC extraction, which may not be suitable for all scenarios.
Innovation Solution
The method employs a bone-conduction sensor to capture fundamental frequency and harmonic distributions of speech, combining these with acoustic conditions to perform biometric processes like enrolment and authentication, using a cumulative distribution function or probability distribution for enhanced noise resilience and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If MFCC extraction is used for speaker recognition, then measurement precision is improved, but use of energy increases and device complexity increases
Solution Approach 1:
The patent extracts only the essential acoustic features (fundamental frequency F0 and spectral characteristics) needed for speaker recognition, discarding the computationally intensive MFCC extraction process. This selective extraction of critical features maintains recognition accuracy while significantly reducing computational requirements and energy consumption.
Solution Approach 2:
The invention employs simpler, less computationally expensive feature extraction methods that can be quickly computed and discarded, replacing the expensive MFCC process. The system uses lightweight acoustic features that require minimal processing power, enabling efficient speaker recognition in resource-constrained environments.
2Measurement precision
If MFCC extraction is used for speaker recognition, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent removes the complex MFCC extraction pipeline and retains only the essential acoustic feature extraction (fundamental frequency and spectral characteristics). This simplification maintains speaker recognition functionality while dramatically reducing system complexity and making the solution suitable for embedded devices.
Solution Approach 2:
The invention replaces the complex mechanical/computational MFCC extraction system with a simpler acoustic feature extraction approach. By substituting the elaborate signal processing chain with direct fundamental frequency and spectral analysis, the system achieves comparable recognition accuracy with significantly reduced complexity.
3Use of energy by moving object
If alternative speaker recognition methods are used, then use of energy decreases, but measurement precision worsens
Solution Approach 1:
The patent changes the parameter set used for speaker recognition from MFCC coefficients to fundamental frequency (F0) and spectral characteristics. This parameter transformation enables the system to operate with lower computational power while maintaining sufficient accuracy for many speaker recognition applications, particularly in noisy environments where F0 provides robust speaker identification.
4Reliability
If speaker recognition is performed in noisy environments, then reliability worsens, but the patent aims to maintain reliability
Solution Approach 1:
The patent converts the harmful effect of ambient noise into a beneficial feature by using fundamental frequency extraction, which is inherently more robust to noise than MFCC methods. The F0-based approach can actually distinguish speaker characteristics more reliably in noisy conditions, turning the noisy environment from a disadvantage into an area where the alternative method excels.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach provides a more accurate and efficient biometric solution that is resilient to ambient noise, allowing for reliable speaker recognition in various acoustic conditions without the need for extensive computational resources.
Implementation Method 1
receiving, from a bone-conduction sensor, a first audio signal representing bone-conducted speech of the user
Data Source
AI summary
A method of modelling speech of a user of a headset comprising a microphone, the method comprising: receiving a first sample, from a bone-conduction sensor, representing bone-conducted speech of the user; obtaining a measure of fundamental frequency of the bone-conducted speech in each of a plurality of speech frames of the first sample; obtaining a first distribution of the fundamental frequencies of the bone-conducted speech over the plurality of speech frames; receiving, from the microphone, a second sample; determining a first acoustic condition at the headset based on the second signal; performing a biometric process based on the first distribution of fundamental frequencies and the first acoustic condition.


