Bone-Conduction Biometric Speaker Recognition in Noisy Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition systems face challenges in achieving high reliability and efficiency, particularly in noisy environments, as they rely heavily on computational power-intensive methods like MFCC extraction, which may not be suitable for all scenarios.

Innovation Solution

The method employs a bone-conduction sensor to capture fundamental frequency and harmonic distributions of speech, combining these with acoustic conditions to perform biometric processes like enrolment and authentication, using a cumulative distribution function or probability distribution for enhanced noise resilience and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If MFCC extraction is used for speaker recognition, then measurement precision is improved, but use of energy increases and device complexity increases

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidcomputational power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential acoustic features (fundamental frequency F0 and spectral characteristics) needed for speaker recognition, discarding the computationally intensive MFCC extraction process. This selective extraction of critical features maintains recognition accuracy while significantly reducing computational requirements and energy consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention employs simpler, less computationally expensive feature extraction methods that can be quickly computed and discarded, replacing the expensive MFCC process. The system uses lightweight acoustic features that require minimal processing power, enabling efficient speaker recognition in resource-constrained environments.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If MFCC extraction is used for speaker recognition, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent removes the complex MFCC extraction pipeline and retains only the essential acoustic feature extraction (fundamental frequency and spectral characteristics). This simplification maintains speaker recognition functionality while dramatically reducing system complexity and making the solution suitable for embedded devices.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention replaces the complex mechanical/computational MFCC extraction system with a simpler acoustic feature extraction approach. By substituting the elaborate signal processing chain with direct fundamental frequency and spectral analysis, the system achieves comparable recognition accuracy with significantly reduced complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Use of energy by moving object

If alternative speaker recognition methods are used, then use of energy decreases, but measurement precision worsens

Engineering Contradiction:
Improvecomputational power consumptionVSAvoidspeaker recognition accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent changes the parameter set used for speaker recognition from MFCC coefficients to fundamental frequency (F0) and spectral characteristics. This parameter transformation enables the system to operate with lower computational power while maintaining sufficient accuracy for many speaker recognition applications, particularly in noisy environments where F0 provides robust speaker identification.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If speaker recognition is performed in noisy environments, then reliability worsens, but the patent aims to maintain reliability

Engineering Contradiction:
Improvespeaker recognition reliabilityVSAvoidambient noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent converts the harmful effect of ambient noise into a beneficial feature by using fundamental frequency extraction, which is inherently more robust to noise than MFCC methods. The F0-based approach can actually distinguish speaker characteristics more reliably in noisy conditions, turning the noisy environment from a disadvantage into an area where the alternative method excels.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach provides a more accurate and efficient biometric solution that is resilient to ambient noise, allowing for reliable speaker recognition in various acoustic conditions without the need for extensive computational resources.

Implementation Method 1

receiving, from a bone-conduction sensor, a first audio signal representing bone-conducted speech of the user

Methodology Applied
Scientific EffectBone conduction:

Data Source

PatentUS11710475B2Methods and apparatus for obtaining biometric data
Publication Date: 2023.07.25 CIRRUS LOGIC INC
  • US11710475B2 patent drawing
  • US11710475B2 patent drawing
  • US11710475B2 patent drawing

AI summary

A method of modelling speech of a user of a headset comprising a microphone, the method comprising: receiving a first sample, from a bone-conduction sensor, representing bone-conducted speech of the user; obtaining a measure of fundamental frequency of the bone-conducted speech in each of a plurality of speech frames of the first sample; obtaining a first distribution of the fundamental frequencies of the bone-conducted speech over the plurality of speech frames; receiving, from the microphone, a second sample; determining a first acoustic condition at the headset based on the second signal; performing a biometric process based on the first distribution of fundamental frequencies and the first acoustic condition.