Method for improving voiceprint recognition accuracy of compressed coding voice
The SEDC-UBM general background model, which compensates for speech coding distortion, solves the problem of reduced voiceprint recognition accuracy caused by speech compression coding, and improves the accuracy of voiceprint recognition.
Patent Information
- Application Number
- CN202511271553.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-08
AI Technical Summary
In mobile communication networks, voice compression coding leads to a decrease in the accuracy of voiceprint recognition.
A general background model with speech coding distortion compensation (SEDC-UBM) is adopted. By evaluating the speech coding distortion model, the general background model is compensated to reduce the impact of speech coding on the performance of voiceprint recognition.
It improves the accuracy of voiceprint recognition in compressed-coded speech and reduces the impact of speech coding distortion.
Smart Images

Figure CN121122289A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of voiceprint recognition, and particularly relates to a method for improving the accuracy of compressed and encoded voiceprint recognition. BACKGROUND
[0002] The voiceprint recognition technology has multiple advantages in the identity recognition application in the mobile communication network. First, the voiceprint recognition has the characteristics of convenience, and the user only needs to complete the identity verification through the voice without needing to remember the complex password, so that the identity authentication process is more convenient and fast. Second, the voiceprint recognition has the advantages of uniqueness and security. The voiceprint, as one of the biological characteristics, must be collected by the living body and is difficult to be copied or forged, thereby providing higher level of security. Finally, the voiceprint recognition has the characteristics of non-contact. The voiceprint recognition does not need the contact type device, and only needs to complete the identity authentication through the voice interaction. This non-contact authentication method can realize the remote voiceprint collection and recognition as long as there is the microphone and the network connection, and is suitable for the remote identity authentication in the mobile communication network.
[0003] However, the application of the voiceprint recognition technology in the mobile communication network also has certain limitations. In the mobile communication network, in order to meet the needs of storage and transmission, the voice needs to be processed through the voice compression encoding. The compression and encoding of the voice will cause the distortion of the voice feature and will affect the accuracy of the voiceprint recognition. SUMMARY
[0004] In view of the above analysis, the application aims to disclose a method for improving the accuracy of compressed and encoded voiceprint recognition, so as to improve the problem of low voiceprint recognition accuracy caused by the distortion of the voice encoding compression in the mobile communication.
[0005] The application discloses a method for improving the accuracy of compressed and encoded voiceprint recognition, which comprises the following steps:
[0006] S1, respectively extracting the voice features of the uncompressed voice and the voice after the voice encoding compression;
[0007] S2, respectively utilizing the uncompressed voice features and the compressed voice features for adaptive training based on the general background model, so as to obtain the corresponding uncompressed voice model and the compressed voice model;
[0008] S3, determining the distortion parameter introduced by the encoding process according to the difference between the uncompressed voice model and the compressed voice model;
[0009] S4, compensating the uncompressed voice model based on the distortion parameter, so as to obtain the general background model compensated by the voice encoding distortion;
[0010] S5, correcting the compressed speech using the general background model compensated by the speech coding distortion, and performing voiceprint recognition based on the corrected features to improve the recognition accuracy.
[0011] Further, in S2, the process of obtaining the compressed speech model by adaptive training using the compressed speech features includes:
[0012] 1) calculating a general background model;
[0013] 2) calculating the posterior probability of the compressed speech features under the current model; the current model is initially the general background model;
[0014] 3) updating the prior probability of the sound list according to the calculated posterior probability;
[0015] 4) updating the mean and covariance matrix of the general background model according to the prior probability of the sound list;
[0016] 5) determining whether to stop iteration; if yes, taking the updated general background model as the compressed speech model; if no, returning to step 2) to continue calculating the posterior probability using the updated general background model;
[0017] The condition for stopping iteration is that the log-likelihood probability of the compressed speech features in the updated general background model no longer increases or the number of iterations exceeds a set threshold.
[0018] Further, in step 3), the prior probability of the sound list is updated according to the calculated posterior probability as follows:
[0019]
[0020] where T is the total number of frames of the speech, is the feature vector of the t-th frame after compression by speech coding;
[0021]
[0022] represents the n-th dimension feature of the t-th frame, which satisfies Gaussian distribution, and N is the dimension of the feature vector;
[0023] is the posterior probability of the current model under a certain sound class .
[0024]
[0025] where λ is the parameter of the current model, including M Gaussian components, λ = {w i , μ i,∑ i},i=1,2,...,M,w i denotes the weight of the i-th Gaussian component of the model, μ i denotes the mean of the i-th Gaussian component of the model, Σ i denotes the variance of the i-th Gaussian component of the model. is the probability density function value of the i-th Gaussian component.
[0026] Further, the mean in the updated parameter of the general background model is
[0027]
[0028] where, 1≤i≤M, 1≤n≤N, μ i,n is the mean of the n-th dimension of the i-th Gaussian component of the UBM model before updating, Δμ n The calculation formula of Δμ
[0029]
[0030] Further, the variance in the updated parameter of the general background model is
[0031]
[0032] Further, the feature vector is calculated The log-likelihood probability in the updated model is:
[0033]
[0034] where,
[0035]
[0036] represents the n-th dimension feature of the t-th frame of the unencoded speech.
[0037] Further, in S3, the mean distortion parameter in the distortion parameter introduced by the encoding process is:
[0038]
[0039] where, represents the n-th dimension mean of the i-th Gaussian component of the general background model of the unencoded speech, represents the n-th dimension mean of the mean of the codec distortion model, μ i,n represents the n-th dimension mean of the i-th Gaussian component of the updated UBM;
[0040] The general background model of the unencoded speech is updated by the feature vector of the unencoded speech.
[0041] Further, in S3, among the distortion parameters introduced by the encoding process, the variance distortion parameter is:
[0042]
[0043] wherein, represents the variance of the i-th Gaussian component of the general background model of the unencoded speech in the n-th dimension, represents the variance of the mean of the codec distortion model in the n-th dimension, and i,n represents the variance of the i-th Gaussian component of the updated UBM in the n-th dimension.
[0044] Further, the parameter of the general background model of the speech coding distortion compensation is
[0045]
[0046] Further, in S1, the MFCC feature extraction method is used to extract the speech features of the unencoded speech and the speech compressed by the speech coding, respectively.
[0047] The present application can achieve the following beneficial effects:
[0048] The method for improving the accuracy of compressed and encoded speech voiceprint recognition provided by the present application compensates for the distortion caused by speech compression and coding by using the general background model of speech coding distortion compensation SEDC-UBM. The key step of this new model is to calculate the distortion model of speech coding after evaluating the speech coding and decoding method, and to perform corresponding compensation processing on the general background model based on the speech coding distortion model. This can effectively reduce the influence of speech coding on voiceprint recognition performance, thereby achieving the purpose of improving the accuracy of voiceprint recognition. BRIEF DESCRIPTION OF DRAWINGS
[0049] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:
[0050] Figure 1 The flowchart of the method for improving the accuracy of compressed and encoded speech voiceprint recognition in the embodiments of the present application is shown in FIG. 1.
[0051] Figure 2 The flowchart of the extraction of the MFCC features of the speech in the embodiments of the present application is shown in FIG. 2. DETAILED DESCRIPTION
[0052] Preferred embodiments of the present application are described in detail below with reference to the attached drawing figures, wherein the drawing figures form a part of the specification. The present application is described in connection with these preferred embodiments, but it will be apparent to those skilled in the art that variations to the preferred embodiments can be made in view of the teachings herein without departing from the spirit of the present application.
[0053] One embodiment of the present application discloses a method for improving the accuracy of compressed speech voiceprint recognition, as shown in the accompanying drawings, comprising: Figure 1
[0054] S1, respectively extracting the speech features of uncompressed speech and compressed speech after speech coding;
[0055] S2, based on the general background model, respectively using the uncompressed speech features and the compressed speech features for adaptive training to obtain the corresponding uncompressed speech model and the compressed speech model;
[0056] S3, determining the distortion parameters introduced by the coding process according to the difference between the uncompressed speech model and the compressed speech model;
[0057] S4, compensating the uncompressed speech model based on the distortion parameters to obtain the general background model compensated for speech coding distortion;
[0058] S5, correcting the compressed speech using the general background model compensated for speech coding distortion, and performing voiceprint recognition based on the corrected features to improve the recognition accuracy.
[0059] Specifically, in S1, the MFCC feature extraction method is used to extract the speech features of uncompressed speech and compressed speech after speech coding.
[0060] The extraction process of the MFCC features of the speech is as shown in Figure 2 . The extracted speech features are wherein T represents the total number of frames of the speech, represents the feature vector of 1 frame, represents the feature vector of the tth frame.
[0061] The speech feature vectors of the uncompressed speech and the compressed speech after speech coding are X 0 and X d
[0062] The feature vector of the tth frame of the speech feature vector X 0 of the uncompressed speech:
[0063]
[0064] wherein represents the feature of the nth dimension of the tth frame of the uncompressed speech, which satisfies the Gaussian distribution.
[0065] speech feature vector X of the compressed speech d the feature vector of the t-th frame of the compressed speech:
[0066]
[0067] wherein, represents the n-dimensional feature of the t-th frame of the compressed speech, satisfying Gaussian distribution.
[0068] Specifically, in S2, the process of obtaining the compressed speech model by using the compressed speech feature for adaptive training includes:
[0069] 1) calculating a universal background model UBM;
[0070] 2) calculating the posterior probability of the compressed speech feature under the current model; the current model is initially the universal background model;
[0071] 3) updating the prior probability of the sound list according to the calculated posterior probability;
[0072] 4) updating the mean and covariance matrix of the universal background model according to the prior probability of the sound list;
[0073] 5) judging whether to stop iteration; if yes, taking the updated universal background model as the compressed speech model; if no, returning to step 2) to continue calculating the posterior probability by using the updated universal background model;
[0074] The condition for stopping iteration is that the log-likelihood probability of the compressed speech feature in the updated universal background model no longer increases or the number of iterations exceeds a set threshold.
[0075] In S2, the process of obtaining the uncompressed speech model is the same as the process of obtaining the compressed speech model described above, and the data used is the uncompressed speech feature.
[0076] More specifically, in step 1), the parameters of the UBM model are λ = {w i , μ i , ∑ i}, i = 1, 2,..., M
[0077] The UBM model is composed of M Gaussian components, wherein w i represents the weight of the i-th Gaussian component of the model, μ i represents the mean of the i-th Gaussian component of the model, and Σ i represents the variance of the i-th Gaussian component of the model.
[0078] More specifically, in step 2), the feature vector of the t-th frame of the compressed speech feature under the current model is calculated as follows: For a certain sound class The posterior probability of the i-th Gaussian component is:
[0079]
[0080] is the probability density function value of the i-th Gaussian component,
[0081] More specifically, the prior probability of the sound list in step 3) is updated according to the calculated posterior probability is:
[0082]
[0083] More specifically, the mean in the updated parameter of the universal background model in step 4) is is:
[0084]
[0085] wherein, 1≤i≤M, 1≤n≤N, μ i,n is the mean of the i-th Gaussian component of the UBM model before updating, Δμ n The calculation formula of Δμ
[0086]
[0087] More specifically, the variance in the updated parameter of the universal background model in step 4) is is:
[0088]
[0089] More specifically, in step 5), the stopping iteration condition is
[0090] The feature vector is calculated The log-likelihood probability in the updated model is:
[0091]
[0092] wherein,
[0093]
[0094] represents the feature of the n-th dimension of the t-th frame of the unencoded speech.
[0095] The iteration is stopped when the log-likelihood probability no longer increases, or the iteration number does not exceed 10.
[0096] To solve the influence of speech coding, a model of coding distortion needs to be calculated, and after the model of coding distortion is calculated, the original uncoded speech can be compensated according to the distortion model. This compensation method can match the test speech and the training speech, so as to achieve the purpose of providing accurate voiceprint.
[0097] Specifically, in S3, among the distortion parameters introduced by the encoding process, the mean distortion parameter is:
[0098]
[0099] wherein, represents the mean of the i-th Gaussian component of the n-dimensional general background model of uncoded speech, represents the mean of the n-dimensional mean of the codec distortion model, μ i,n represents the mean of the n-dimensional i-th Gaussian component of the updated UBM;
[0100] The general background model of uncoded speech is obtained by updating the general background model with the feature vector of uncoded speech.
[0101] Specifically, in S3, among the distortion parameters introduced by the encoding process, the variance distortion parameter is:
[0102]
[0103] wherein, represents the variance of the n-dimensional i-th Gaussian component of the general background model of uncoded speech, represents the n-dimensional variance of the mean of the codec distortion model, ∑ i,n represents the n-dimensional variance of the i-th Gaussian component of the updated UBM.
[0104] Specifically, in S4, the general background model SEDC-UBM model for speech coding distortion compensation is calculated as: As can be seen from the formula, the key to obtaining this model is to calculate the parameters and the parameter The calculation formula of the two parameters is as follows:
[0105]
[0106] Specifically, the compressed speech is corrected by using the general background model SEDC-UBM for speech coding distortion compensation, and voiceprint recognition is performed based on the corrected features, so as to improve the recognition accuracy.
[0107] To sum up, the method for improving the accuracy of compressed speech voiceprint recognition in the embodiment of the application compensates the distortion caused by speech compression coding by using a general background model SEDC-UBM. The key step of the new model is to calculate the distortion model of speech coding after evaluating the speech coding and decoding mode, and to compensate the general background model based on the speech coding distortion model. The influence of speech coding on voiceprint recognition performance can be effectively reduced, so as to improve the accuracy of voiceprint recognition.
[0108] The above merely describes the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A method for improving the accuracy of voiceprint recognition in compressed coded speech, characterized in that, include: S1. Extract speech features from uncompressed speech and speech-coded compressed speech respectively; S2. Based on the general background model, adaptive training is performed using uncompressed speech features and compressed speech features respectively to obtain the corresponding uncompressed speech model and compressed speech model. S3. Based on the differences between the uncompressed speech model and the compressed speech model, determine the distortion parameters introduced in the encoding process; S4. Based on the distortion parameters, the uncompressed speech model is compensated to obtain a general background model for speech coding distortion compensation; S5. The compressed speech is corrected using the general background model for speech coding distortion compensation, and voiceprint recognition is performed based on the corrected features to improve recognition accuracy.
2. The method for improving the accuracy of compressed-coded speech voiceprint recognition according to claim 1, characterized in that, In S2, the process of adaptive training using compressed speech features to obtain the corresponding compressed speech model includes: 1) Calculate the general background model; 2) Calculate the posterior probability of compressed speech features under the current model; the current model is initially a general background model; 3) Update the prior probabilities of the sound list based on the calculated posterior probabilities; 4) Update the mean and covariance matrix of the general background model based on the prior probabilities of the sound list; 5) Determine if the iteration has stopped; if yes, use the updated general background model as the compressed speech model; if no, return to step 2) and continue calculating the posterior probability using the updated general background model. The iteration stops when the log-likelihood probability of the compressed speech features no longer increases in updating the general background model, or when the number of iterations exceeds a set threshold.
3. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 2, characterized in that, In step 3), the prior probabilities of the sound list are updated based on the calculated posterior probabilities. for: Where T is the total number of frames in the speech. is the feature vector of the t-th frame after speech coding compression; This represents the nth dimension of the feature vector in the t-th frame, which follows a Gaussian distribution, where N is the dimension of the feature vector. For the current model For a certain sound category The posterior probability; in, λ represents the parameters of the current model, including M Gaussian components, λ = {w i ,μ i ,∑ i }, i = 1, 2, ..., M, w i μ represents the weight of the i-th Gaussian component of the model. i Σ represents the mean of the i-th Gaussian component of the model. i This represents the variance of the i-th Gaussian component of the model; Let be the probability density function value of the i-th Gaussian component.
4. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 3, characterized in that, The mean of the parameters in the updated general background model for: Where 1≤i≤M, 1≤n≤N, μ i,n It is the mean of the i-th Gaussian component in the n-th dimension of the original UBM model, Δμ n The calculation formula is as follows:
5. The method for improving the accuracy of compressed-coded speech voiceprint recognition according to claim 4, characterized in that, The variance in the parameters of the updated general background model is for:
6. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 5, characterized in that, Calculate the eigenvector The log-likelihood probability in the updated model is: in, This represents the nth dimension of the t-th frame of uncoded speech.
7. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 6, characterized in that, In S3, among the distortion parameters introduced during the encoding process, the mean distortion parameter is: in, This represents the mean of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. μ represents the mean of the nth dimension of the encoding / decoding distortion model. i,n This represents the mean of the i-th Gaussian component in the n-th dimension of the updated UBM; The general background model of uncoded speech is obtained by updating the general background model with the feature vector of uncoded speech.
8. The method for improving the accuracy of compressed-coded speech voiceprint recognition according to claim 7, characterized in that, In S3, among the distortion parameters introduced during the encoding process, the variance distortion parameter is: in, This represents the variance of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. ∑ represents the variance of the mean in the nth dimension of the encoding / decoding distortion model. i,n This represents the variance of the i-th Gaussian component in the n-th dimension of the updated UBM.
9. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 8, characterized in that, The parameters of the general background model for speech coding distortion compensation are: i = 1, ..., M; 10. The method for improving the accuracy of compressed coded speech voiceprint recognition according to any one of claims 1-9, characterized in that, In S1, the MFCC feature extraction method is used to extract speech features from uncompressed speech and speech compressed speech after speech coding, respectively.
Citation Information
Patent Citations
Voiceprint identification method based on multi-type combination characteristic parameters
CN104835498A
Voiceprint recognition method based on pitch period mixed characteristic parameters
CN104900235A
Many-to-many voice conversion method based on beta-variational autoencoder (VAE) and identity feature vector (i-vector)
CN110085254A
Language identification and classification method and device based on noise reduction automatic encoder
CN110858477A
Rapid language identification method based on phoneme log likelihood ratio and sparse representation
CN111462729A