Hierarchical Real-Time Speaker Recognition for Biometric VoIP Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker recognition systems face challenges in accurately identifying speakers in real-time, especially when the number of registered targets is large, and struggle with variations in speech, background noise, and language changes, limiting their scalability and accuracy.

Innovation Solution

A real-time speaker recognition system using a hierarchical architecture that extracts Mel-Frequency Cepstral Coefficients (MFCC) and Gaussian Mixture Model (GMM) components from speech data, allowing for scalable identification of up to millions of users, and operates in both verification and identification modes with low computational complexity, independent of content and language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speaker recognition systems are used to identify speakers in real-time, then the system can provide speaker identification, but the accuracy deteriorates when the number of registered targets is large and under variations in speech, background noise, and language changes

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidrobustness to speech variations, noise, and language changes
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speaker recognition process into multiple hierarchical levels: phoneme-level features are extracted first, then clustered into speaker-specific patterns, and finally combined for identification. This segmentation allows the system to handle large numbers of speakers while maintaining accuracy under various conditions by processing information in manageable stages rather than attempting global matching.

Inventive Principle:
Principle #1Segmentation

2Reliability

If the system monitors all VoIP activities to identify suspects, then the government can detect suspect communications, but privacy violations occur for non-suspects

Engineering Contradiction:
Improveability to detect suspect communicationsVSAvoidprivacy violations for non-suspects
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary biometric feature (voice print) from the communication data for identification purposes, rather than monitoring or storing the entire communication content. This extraction approach enables reliable suspect detection through voice matching while minimizing privacy intrusion by processing only the essential identifying characteristic.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If the government monitors a specific VoIP phone number, then surveillance of the suspect is possible, but the system fails when the suspect uses a new phone number

Engineering Contradiction:
Improvesurveillance effectivenessVSAvoidability to track suspects across different phone numbers
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces voice print biometrics as an intermediary identifier that bridges different communication channels. Instead of tracking specific phone numbers, the system uses the speaker's unique voice characteristics as a mediator to identify the suspect across multiple phone numbers and VoIP services, maintaining surveillance effectiveness regardless of number changes.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If existing speaker recognition algorithms are used, then verification can be performed, but computational complexity increases significantly when scaling to millions of users

Engineering Contradiction:
Improvespeaker verification capabilityVSAvoidcomputational complexity for large-scale deployment
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the large-scale speaker verification problem into hierarchical levels: phoneme feature extraction, speaker-specific clustering, and final identification. This segmentation reduces computational complexity by processing features in stages rather than performing exhaustive comparisons across all registered speakers, enabling scaling to millions of users.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary clustering of phoneme features into speaker-specific patterns during the training phase, creating compact speaker models. This preliminary action reduces the computational burden during real-time verification by replacing complex global matching with efficient comparisons against pre-computed speaker clusters.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8160877B1Hierarchical real-time speaker recognition for biometric VoIP verification and targeting
Publication Date: 2012.04.17 THE BOEING CO
  • US8160877B1 patent drawing
  • US8160877B1 patent drawing
  • US8160877B1 patent drawing

AI summary

A method for real-time speaker recognition including obtaining speech data of a speaker, extracting, using a processor of a computer, a coarse feature of the speaker from the speech data, identifying the speaker as belonging to a pre-determined speaker cluster based on the coarse feature of the speaker, extracting, using the processor of the computer, a plurality of Mel-Frequency Cepstral Coefficients (MFCC) and a plurality of Gaussian Mixture Model (GMM) components from the speech data, determining a biometric signature of the speaker based on the plurality of MFCC and the plurality of GMM components, and determining in real time, using the processor of the computer, an identity of the speaker by comparing the biometric signature of the speaker to one of a plurality of biometric signature libraries associated with the pre-determined speaker cluster.