Voiceprint Recognition Normalization Matrix Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voiceprint recognition systems face accuracy issues due to distortion caused by channel and environment volatility, leading to limited improvement in recognition accuracy when using Linear Discriminant Analysis (LDA)-processed identity vectors.

Innovation Solution

A method for training a voiceprint recognition system involves obtaining a voice training set, determining identity vectors, categorizing them, normalizing using a normalization matrix that maximizes the sum of similarity degrees within categories, and updating the matrix to improve similarity between voice segments of the same user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If Linear Discriminant Analysis (LDA) is used to process identity vectors, then spatial distribution compensation is achieved, but recognition accuracy improvement is limited because actual voice segment distribution does not match multi-dimensional Gaussian distribution assumptions

Engineering Contradiction:
Improvevoiceprint recognition accuracyVSAvoiddistribution assumption validity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the normalization approach from LDA-based (assuming multi-dimensional Gaussian distribution) to a new normalization matrix method that maximizes intra-class similarity without relying on distribution assumptions. This parameter change in the mathematical model resolves the contradiction by adapting to actual voice segment distribution characteristics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent substitutes the LDA mechanical system (which enforces Gaussian distribution assumptions) with a new normalization system based on similarity degree optimization. This replacement eliminates the reliability issue caused by invalid distribution assumptions while maintaining accuracy improvement.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If channel and environment volatility are present, then voice identity vectors become distorted, but traditional normalization methods cannot sufficiently compensate for this distortion

Engineering Contradiction:
Improveidentity vector accuracyVSAvoidvoice distortion
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a feedback mechanism where the normalization matrix is trained using similarity degrees between voice segments. The system continuously optimizes the normalization parameters based on measured similarity, creating a closed-loop compensation system that adapts to channel and environment volatility effects.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary normalization matrix training using a voice training set before actual recognition. This preliminary action pre-compensates for expected distortions from channel and environment volatility, improving identity vector accuracy before recognition occurs.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3477639B1Training a voiceprint recognition system
Publication Date: 2021.06.23 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3477639B1 patent drawingFigure 1
  • EP3477639B1 patent drawingFigure 2A
  • EP3477639B1 patent drawingFigure 2B

AI summary

The present disclosure discloses a method and an apparatus for training a voiceprint recognition system, and belongs to the field of voiceprint recognition technologies. The method includes: obtaining a voice training set; determining identity vectors of all voice segments in the voice training set; recognizing identity vectors of a plurality of voice segments of a same user in the determined identity vectors; putting the recognized identity vectors of the same user in the plurality of users into one of a plurality of user categories; determining any identity vector in the user category as a first identity vector; normalizing the first identity vector by using a normalization matrix; and training the normalization matrix, and outputting a training value of the normalization matrix when the normalization matrix maximizes a sum of first values of all the user categories. The problem of small accuracy improvement in voiceprint recognition using a linear discriminant analysis (LDA)-processed identity vector in the related technology is resolved, and accuracy of voiceprint recognition is improved.