Identity Vector Correction for Short Speech Authentication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Short speech duration and sparse speech data degrade the performance of identity vector generation in speaker authentication systems, leading to reduced authentication accuracy.
Innovation Solution
The method involves calculating posterior probabilities of acoustic features belonging to each Gaussian distribution component in a speaker background model, mapping these statistics to a statistic space, and correcting the statistics using a reference statistic to generate an identity vector, which compensates for the offset caused by short speech duration and sparse data, thereby improving authentication performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional identity vector generation methods are used with short speech duration or sparse speech data, then the processing speed is fast, but the identity authentication performance degrades
Solution Approach 1:
The method performs preliminary actions by pre-calculating and storing the mean vector and covariance matrix from the universal background model before actual identity vector generation. This preprocessing enables the system to handle short and sparse speech data effectively by having reference statistics ready for comparison and correction, thus improving authentication performance without requiring large amounts of input speech data.
Solution Approach 2:
The method changes parameters by introducing correction terms based on the difference between the calculated statistics from short speech data and the reference statistics from the universal background model. This parameter adjustment compensates for the insufficient statistical information in short speech samples, thereby maintaining reliable identity authentication performance even with limited speech data quantity.
2Measurement precision
If traditional identity vector generation methods are used with short speech duration, then the system is simple to operate, but the authentication accuracy reduces
Solution Approach 1:
The method introduces an intermediary mechanism by using the universal background model's statistical parameters (mean vector and covariance matrix) as a reference framework. This intermediary reference enables the system to correct and enhance the statistics derived from short speech durations, thereby achieving high authentication accuracy without requiring long speech samples.
Solution Approach 2:
The method applies parameter changes by adjusting the statistical parameters (mean and covariance) of short speech data through correction terms derived from the universal background model. This parameter transformation allows the system to maintain measurement precision in authentication accuracy even when the speech duration is limited.
3Reliability
If correction based on universal background model is applied, then the identity authentication performance improves, but the computational complexity increases
Solution Approach 1:
The method extracts only the essential statistical parameters (mean vector and covariance matrix) from the universal background model for correction purposes, rather than using the complete model structure. This extraction approach maintains improved authentication performance while reducing computational complexity by focusing only on the necessary correction elements.
Solution Approach 2:
The method applies parameter changes in a computationally efficient manner by only adjusting the mean and covariance parameters using simple correction terms based on the universal background model. This selective parameter modification achieves performance improvement without requiring complex computational operations, thus balancing reliability enhancement with manageable device complexity.
Data Source
Figure 1~2A
Figure 2B
Figure 3
AI summary
An identity vector generation method is disclosed, including: obtaining to-be-processed speech data; extracting corresponding acoustic features from the to-be-processed speech data; calculating a posterior probability that each of the acoustic features belongs to each Gaussian distribution component in a speaker background model to obtain a statistic; mapping the statistic to a statistic space to obtain a reference statistic, the statistic space being built according to a statistic corresponding to a speech sample exceeding preset speech duration; determining a corrected statistic according to the calculated statistic and the reference statistic; and generating an identity vector according to the corrected statistic.