Speaker Recognition Using Low-Dimensional P-Vector Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice biometric verification systems, particularly those using the i-vector approach, face challenges with high computational requirements and poor performance on short utterances, as they require complex multi-iterative procedures and have low dimensional dimensions, limiting their accuracy and efficiency.

Innovation Solution

The proposed solution combines the strengths of i-vector and GMM-MAP methods by simulating voice using GMM-MAP, transforming the GMM model into a low-dimensional 'p-vector' through dimension reduction, and employing a channel-compensation module based on discriminant analysis, which improves speaker modeling and recognition accuracy, and reduces hardware costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If i-vector approach is used for speaker verification, then computational power requirements increase due to complex multi-iterative factor analysis procedures, but recognition accuracy on short utterances deteriorates due to low dimensional dimensions

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the high-dimensional GMM speaker model into a low-dimensional p-vector by projecting onto a subspace defined by discriminant eigenvectors. This parameter transformation reduces computational complexity while maintaining recognition accuracy by retaining only the most discriminative dimensions for speaker identification.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent extracts the essential speaker information from the high-dimensional GMM model by identifying and retaining only the most discriminative features through discriminant analysis. The p-vector captures the critical speaker characteristics while discarding redundant information, thereby reducing computational requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If i-vector approach is used for speaker verification, then computational power requirements increase, but performance on short utterances deteriorates

Engineering Contradiction:
Improveverification performanceVSAvoidcomputational power consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent changes the dimensional parameter of the speaker model from high-dimensional i-vectors to low-dimensional p-vectors. This parameter change reduces the computational power required for verification while improving reliability on short utterances by focusing on the most discriminative features.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs dimensionality reduction by transforming the speaker model from high-dimensional space to low-dimensional space through projection onto discriminant eigenvectors. This dimensional change maintains verification performance while significantly reducing computational power consumption.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If GMM-MAP approach is used for speaker simulation, then recognition accuracy improves, but hardware costs increase due to complex modeling requirements

Engineering Contradiction:
Improverecognition accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential speaker information from the complex GMM-MAP model into a compact p-vector representation. This extraction maintains recognition accuracy while reducing hardware complexity by storing and processing only the essential low-dimensional features rather than the full complex model.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified copy of the speaker model in the form of a p-vector that captures the essential characteristics without requiring the full GMM-MAP complexity. This copy enables accurate recognition while significantly reducing hardware requirements for storage and processing.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10909991B2System for text-dependent speaker recognition and method thereof
Publication Date: 2021.02.02 ID R&D INC
  • US10909991B2 patent drawing
  • US10909991B2 patent drawing
  • US10909991B2 patent drawing

AI summary

A computer-implemented method for verifying identity of a speaker is proposed. A low dimensional p-vector based on a speech of the speaker is extracted from the generated high dimensional speaker model and is then compared with the stored specific speaker's p-vector obtained previously during the enrollment process. The resulting biometric score is then used to determine whether to verify the speaker, or not.