Headphone Speaker Recognition Using Inside Microphone

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital signal processing techniques for speaker recognition in headsets and headphones are ineffective in noisy environments and may not accurately differentiate between the registered owner and an imposter, especially when using outside microphones that are prone to ambient noise and lack unique ear canal characteristics.

Innovation Solution

The method utilizes the inside acoustic microphone of a headphone to capture speech signals, which are processed to create a unique model reflecting vocal, nasal tract, and ear canal characteristics, enhancing speaker recognition accuracy by employing a speaker recognition algorithm trained with these signals, and optionally using a ratio of inside to outside microphone signals for feature extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If outside microphones are used for speaker recognition, then speech can be captured, but ambient noise reduces recognition accuracy

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidambient noise
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio signal into two distinct components: speech signal captured by the outside microphone and ambient noise captured by the inside microphone. This segmentation allows the system to process and separate the desired speech from the harmful ambient noise, improving speaker recognition accuracy in noisy environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The inside microphone acts as an intermediary that captures ambient noise which would otherwise contaminate the speech signal from the outside microphone. By using the inside microphone signal as a reference for noise cancellation, the system can subtract the ambient noise component from the outside microphone signal, isolating the speech signal for accurate speaker recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If outside microphones are used for speaker recognition, then speech capture is possible, but unique ear canal characteristics are not obtained

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidear canal characteristics
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent merges the signals from two microphones positioned at different locations: the outside microphone captures speech while the inside microphone captures both speech and ambient noise. By combining these signals through spectral subtraction, the system extracts speech with enhanced characteristics that include both vocal tract information and unique ear canal acoustic properties, improving speaker recognition precision.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If inside microphone is used alone, then ear canal characteristics are captured, but speech signal quality is reduced due to ambient noise

Engineering Contradiction:
Improveear canal characteristic captureVSAvoidambient noise
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The outside microphone signal serves as an intermediary that provides a clean reference of the speech signal without ambient noise contamination. This reference signal is used in spectral subtraction to remove the ambient noise component from the inside microphone signal, preserving the unique ear canal characteristics while eliminating the harmful ambient noise effect.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10896682B1Speaker recognition based on an inside microphone of a headphone
Publication Date: 2021.01.19 APPLE INC
  • US10896682B1 patent drawing
  • US10896682B1 patent drawing
  • US10896682B1 patent drawing

AI summary

A speaker recognition algorithm is trained (one or more of its models are tuned) with samples of a microphone signal produced by an inside microphone of a headphone, while the headphone is worn by a speaker. The trained speaker recognition algorithm then tests other samples of the inside microphone signal and produces multiple speaker identification scores for its given models, or a single speaker verification likelihood score for a single given model. Other embodiments are also described and claimed.