Method for improving the accuracy of compressed encoded speech voiceprint recognition
The SEDC-UBM general background model, which compensates for speech coding distortion, solves the problem of reduced voiceprint recognition accuracy caused by speech compression coding in mobile communication networks, and improves the accuracy of voiceprint recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZHONGCHUANG TELECOM TEST CO LTD
- Filing Date
- 2025-09-08
- Publication Date
- 2026-07-21
AI Technical Summary
In mobile communication networks, voice compression coding leads to a decrease in the accuracy of voiceprint recognition.
A general background model with speech coding distortion compensation (SEDC-UBM) is adopted. By evaluating the speech coding distortion model, the general background model is compensated to reduce the impact of speech coding on voiceprint recognition.
It improves the accuracy of compressed-coded voiceprint recognition and reduces the impact of voice coding on voiceprint recognition performance.
Smart Images

Figure CN121122289B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of voiceprint recognition technology, specifically relating to a method for improving the accuracy of voiceprint recognition in compressed coded speech. Background Technology
[0002] Voiceprint recognition technology offers several advantages in identity verification applications within mobile communication networks. Firstly, it is convenient; users don't need to remember complex passwords, and can complete identity verification simply through voice, making the process faster and more efficient. Secondly, voiceprint recognition boasts uniqueness and security. As a biometric feature, voiceprints must be captured by a living person, making them difficult to copy or forge, thus providing a higher level of security. Finally, voiceprint recognition is contactless; it requires no physical devices and can be completed solely through voice interaction. This contactless authentication method, requiring only a microphone and network connection, enables remote voiceprint capture and recognition, making it suitable for remote identity verification in mobile communication networks.
[0003] However, the application of voiceprint recognition technology in mobile communication networks also has certain limitations. In mobile communication networks, in order to meet the needs of storage and transmission, voice needs to undergo voice compression coding. This compression coding of voice can cause distortion of voice features, which will affect the accuracy of voiceprint recognition. Summary of the Invention
[0004] Based on the above analysis, the present invention aims to disclose a method for improving the accuracy of voiceprint recognition in compressed coded speech, thereby addressing the problem of low voiceprint recognition accuracy caused by speech coding compression distortion in mobile communications.
[0005] This invention discloses a method for improving the accuracy of voiceprint recognition in compressed coded speech, comprising:
[0006] S1. Extract speech features from uncompressed speech and speech-coded compressed speech, respectively;
[0007] S2. Based on the general background model, adaptive training is performed using uncompressed speech features and compressed speech features respectively to obtain the corresponding uncompressed speech model and compressed speech model.
[0008] S3. Based on the differences between the uncompressed speech model and the compressed speech model, determine the distortion parameters introduced in the encoding process;
[0009] S4. Based on the distortion parameters, the uncompressed speech model is compensated to obtain a general background model for speech coding distortion compensation;
[0010] S5. The compressed speech is corrected using the general background model for speech coding distortion compensation, and voiceprint recognition is performed based on the corrected features to improve recognition accuracy.
[0011] Furthermore, in S2, the process of adaptively training using compressed speech features to obtain the corresponding compressed speech model includes:
[0012] 1) Calculate the general background model;
[0013] 2) Calculate the posterior probability of compressed speech features under the current model; the current model is initially a general background model;
[0014] 3) Update the prior probabilities of the sound list based on the calculated posterior probabilities;
[0015] 4) Update the mean and covariance matrix of the general background model based on the prior probabilities of the sound list;
[0016] 5) Determine if the iteration has stopped; if yes, use the updated general background model as the compressed speech model; if no, return to step 2) and continue calculating the posterior probability using the updated general background model.
[0017] The iteration stops when the log-likelihood probability of the compressed speech features no longer increases in updating the general background model, or when the number of iterations exceeds a set threshold.
[0018] Furthermore, in step 3), the prior probabilities of the sound list are updated based on the calculated posterior probabilities. for:
[0019]
[0020] Where T is the total number of frames in the speech. is the feature vector of the t-th frame after speech coding compression;
[0021]
[0022] This represents the nth dimension of the feature vector in the t-th frame, which follows a Gaussian distribution, where N is the dimension of the feature vector.
[0023] For the current model For a certain sound category The posterior probability;
[0024]
[0025] in, λ represents the parameters of the current model, including M Gaussian components, λ = {w i ,μ i,∑ i}, i = 1, 2, ..., M, w i μ represents the weight of the i-th Gaussian component of the model. i Σ represents the mean of the i-th Gaussian component of the model. i This represents the variance of the i-th Gaussian component of the model; Let be the probability density function value of the i-th Gaussian component.
[0026] Furthermore, the mean value in the parameters of the updated general background model for:
[0027]
[0028] Where 1≤i≤M, 1≤n≤N, μ i,n It is the mean of the i-th Gaussian component in the n-th dimension of the original UBM model, Δμ n The calculation formula is as follows:
[0029]
[0030] Furthermore, the variance in the parameters of the updated general background model is for:
[0031]
[0032] Further, calculate the eigenvectors. The log-likelihood probability in the updated model is:
[0033]
[0034] in,
[0035]
[0036] This represents the nth dimension of the t-th frame of uncoded speech.
[0037] Furthermore, in S3, among the distortion parameters introduced during the encoding process, the mean distortion parameter is:
[0038]
[0039] in, This represents the mean of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. μ represents the mean of the nth dimension of the encoding / decoding distortion model. i,n This represents the mean of the i-th Gaussian component in the n-th dimension of the updated UBM;
[0040] The general background model of uncoded speech is obtained by updating the general background model with the feature vector of uncoded speech.
[0041] Furthermore, in S3, among the distortion parameters introduced during the encoding process, the variance distortion parameter is:
[0042]
[0043] in, This represents the variance of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. ∑ represents the variance of the mean in the nth dimension of the encoding / decoding distortion model. i,n This represents the variance of the i-th Gaussian component in the n-th dimension of the updated UBM.
[0044] Furthermore, the parameters of the general background model for speech coding distortion compensation are:
[0045]
[0046] Furthermore, in S1, the MFCC feature extraction method is used to extract speech features from both uncompressed speech and speech compressed by speech coding.
[0047] The present invention can achieve the following beneficial effects:
[0048] The method for improving the accuracy of voiceprint recognition in compressed speech proposed in this invention employs the SEDC-UBM (Self-Defining Common Background Model) to compensate for the distortion caused by speech compression coding. The key step in implementing this new model is to evaluate the speech encoding / decoding method and then calculate the speech coding distortion model. Based on this distortion model, the common background model is compensated accordingly. This effectively reduces the impact of speech coding on voiceprint recognition performance, thereby improving the accuracy of voiceprint recognition. Attached Figure Description
[0049] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0050] Figure 1 This is a flowchart of a method for improving the accuracy of compressed-coded speech voiceprint recognition in an embodiment of the present invention;
[0051] Figure 2 This is a flowchart illustrating the extraction of MFCC features from speech in an embodiment of the present invention. Detailed Implementation
[0052] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and, together with the embodiments of the present invention, serve to illustrate the principles of the present invention.
[0053] One embodiment of the present invention discloses a method for improving the accuracy of compressed coded speech voiceprint recognition, such as... Figure 1 As shown, it includes:
[0054] S1. Extract speech features from uncompressed speech and speech-coded compressed speech, respectively;
[0055] S2. Based on the general background model, adaptive training is performed using uncompressed speech features and compressed speech features respectively to obtain the corresponding uncompressed speech model and compressed speech model.
[0056] S3. Based on the differences between the uncompressed speech model and the compressed speech model, determine the distortion parameters introduced in the encoding process;
[0057] S4. Based on the distortion parameters, the uncompressed speech model is compensated to obtain a general background model for speech coding distortion compensation;
[0058] S5. The compressed speech is corrected using the general background model for speech coding distortion compensation, and voiceprint recognition is performed based on the corrected features to improve recognition accuracy.
[0059] Specifically, in S1, the MFCC feature extraction method is used to extract speech features from both uncompressed speech and speech compressed by speech coding.
[0060] The process of extracting MFCC features from speech is as follows: Figure 2 As shown. The extracted speech features are Where T represents the total number of frames in the speech. The feature vector representing one frame, The feature vector representing the t-th frame.
[0061] The extracted speech feature vectors for the uncompressed speech and the speech-coded compressed speech are X, respectively. 0 and X d
[0062] Speech feature vector X of uncompressed speech 0 The feature vector of the t-th frame:
[0063]
[0064] in, This represents the nth dimension of the uncoded speech in the t-th frame, which follows a Gaussian distribution.
[0065] The speech feature vector X of the speech after speech coding and compression d The feature vector of the t-th frame:
[0066]
[0067] in, This represents the nth dimension of the t-th frame of the speech after speech coding and compression, which follows a Gaussian distribution.
[0068] Specifically, in S2, the process of adaptively training using compressed speech features to obtain the corresponding compressed speech model includes:
[0069] 1) Calculate the Universal Background Model (UBM);
[0070] 2) Calculate the posterior probability of compressed speech features under the current model; the current model is initially a general background model;
[0071] 3) Update the prior probabilities of the sound list based on the calculated posterior probabilities;
[0072] 4) Update the mean and covariance matrix of the general background model based on the prior probabilities of the sound list;
[0073] 5) Determine if the iteration has stopped; if yes, use the updated general background model as the compressed speech model; if no, return to step 2) and continue calculating the posterior probability using the updated general background model.
[0074] The iteration stops when the log-likelihood probability of the compressed speech features no longer increases in updating the general background model, or when the number of iterations exceeds a set threshold.
[0075] In S2, the process of obtaining the uncompressed speech model is the same as that of the compressed speech model described above, and the data used is the uncompressed speech features.
[0076] More specifically, in step 1), the parameters of the UBM model are λ = {w} i ,μ i ,∑ i}, i = 1, 2, ..., M
[0077] The UBM model consists of M Gaussian components, where w i μ represents the weight of the i-th Gaussian component of the model. i Σ represents the mean of the i-th Gaussian component of the model. i Let represent the variance of the i-th Gaussian component of the model.
[0078] More specifically, step 2) calculates the feature vector of the compressed speech features in the current model for the t-th frame. For a certain sound category The posterior probability is:
[0079]
[0080] Let be the probability density function value of the i-th Gaussian component.
[0081] More specifically, in step 3), the prior probabilities of the sound list are updated based on the calculated posterior probabilities. for:
[0082]
[0083] More specifically, the mean value in the parameters of the updated general background model in step 4). for:
[0084]
[0085] Where 1≤i≤M, 1≤n≤N, μ i,n It is the mean of the i-th Gaussian component in the n-th dimension of the original UBM model, Δμ n The calculation formula is as follows:
[0086]
[0087] More specifically, the variance in the parameters of the updated general background model in step 4) is for:
[0088]
[0089] More specifically, in step 5), the stopping iteration condition...
[0090] Calculate the eigenvector The log-likelihood probability in the updated model is:
[0091]
[0092] in,
[0093]
[0094] This represents the nth dimension of the t-th frame of uncoded speech.
[0095] Stop iterating when the log-likelihood probability no longer increases, or when the number of iterations does not exceed 10.
[0096] To address the impact of speech coding, a model of coding distortion needs to be calculated. Once this model is calculated, the original uncoded speech can be compensated according to it. This compensated approach allows for matching of test and training speech, thereby improving speaker recognition accuracy.
[0097] Specifically, in S3, among the distortion parameters introduced during the encoding process, the mean distortion parameter is:
[0098]
[0099] in, This represents the mean of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. μ represents the mean of the nth dimension of the encoding / decoding distortion model. i,n This represents the mean of the i-th Gaussian component in the n-th dimension of the updated UBM;
[0100] The general background model of uncoded speech is obtained by updating the general background model with the feature vector of uncoded speech.
[0101] Specifically, in S3, among the distortion parameters introduced during the encoding process, the variance distortion parameter is:
[0102]
[0103] in, This represents the variance of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. ∑ represents the variance of the mean in the nth dimension of the encoding / decoding distortion model. i,n This represents the variance of the i-th Gaussian component in the n-th dimension of the updated UBM.
[0104] Specifically, in S4, the general background model SEDC-UBM model for calculating speech coding distortion compensation is: As can be seen from the modified formula, the key to obtaining this model is calculating the parameters. and parameters The formulas for calculating these two parameters are as follows:
[0105]
[0106] Specifically, the compressed speech is corrected using the general background model SEDC-UBM for speech coding distortion compensation, and voiceprint recognition is performed based on the corrected features to improve recognition accuracy.
[0107] In summary, the method for improving the accuracy of voiceprint recognition in compressed speech according to embodiments of the present invention employs the SEDC-UBM (Self-Defining Background Model for Speech Coding Distortion Compensation) to compensate for the distortion caused by speech compression coding. A key step in implementing this new model is to evaluate the speech encoding / decoding method and then calculate the speech coding distortion model. Based on this distortion model, the general background model is compensated accordingly. This effectively reduces the impact of speech coding on voiceprint recognition performance, thereby improving the accuracy of voiceprint recognition.
[0108] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for improving the accuracy of voiceprint recognition in compressed coded speech, characterized in that, include: S1. Extract speech features from uncompressed speech and speech-coded compressed speech respectively; S2. Based on the general background model, adaptive training is performed using uncompressed speech features and compressed speech features respectively to obtain the corresponding uncompressed speech model and compressed speech model. S3. Based on the differences between the uncompressed speech model and the compressed speech model, determine the distortion parameters introduced in the encoding process; S4. Based on the distortion parameters, the general background model is compensated to obtain the general background model for speech coding distortion compensation; S5. The compressed speech is corrected using the general background model for speech coding distortion compensation, and voiceprint recognition is performed based on the corrected features to improve recognition accuracy. In S2, the process of adaptive training using compressed speech features to obtain the corresponding compressed speech model includes: 1) Calculate the general background model; 2) Calculate the posterior probability of compressed speech features under the current model; the current model is initially a general background model; 3) Update the prior probabilities of the sound list based on the calculated posterior probabilities; 4) Update the mean and covariance matrix of the general background model based on the prior probabilities of the sound list; 5) Determine if the iteration has stopped; if yes, use the updated general background model as the compressed speech model; otherwise, return to step 2) and continue calculating the posterior probability using the updated general background model. The iteration stops when the log-likelihood probability of the compressed speech features no longer increases in updating the general background model, or when the number of iterations exceeds a set threshold.
2. The method for improving the accuracy of compressed-coded speech voiceprint recognition according to claim 1, characterized in that, In step 3), the prior probabilities of the sound list are updated based on the calculated posterior probabilities. for: ; in, The total number of frames in the speech. The first after speech encoding compression The feature vector of a frame; ; It represents the first The first frame The dimensional features satisfy a Gaussian distribution. The dimension of the feature vector; For the current model For a certain sound category The posterior probability; ; in, , The parameters of the current model include Gaussian components , The model represents the first The weights of the Gaussian components, Let represent the mean of the i-th Gaussian component of the model. Let represent the covariance matrix of the i-th Gaussian component of the model; Let be the probability density function value of the i-th Gaussian component.
3. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 2, characterized in that, The mean of the parameters in the updated Universal Background Model (UBM) for: ; in, , It is the mean of the nth dimension of the i-th Gaussian component of the original Universal Background Model (UBM). The calculation formula is as follows: 。 4. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 3, characterized in that, The variance in the parameters of the updated Universal Background Model (UBM) is: for: 。 5. The method for improving the accuracy of compressed-coded speech voiceprint recognition according to claim 4, characterized in that, Calculate the eigenvector The log-likelihood probability in the updated general background UBM model is: ; in, ; This represents the nth dimension of the t-th frame of uncoded speech.
6. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 5, characterized in that, In S3, among the distortion parameters introduced during the encoding process, the mean distortion parameter is: ; in, This represents the mean of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. This represents the mean of the nth dimension of the encoding / decoding distortion model. This represents the mean of the nth dimension of the i-th Gaussian component in the updated Universal Background Model (UBM). The general background model of uncoded speech is obtained by updating the general background model with the feature vector of uncoded speech.
7. The method for improving the accuracy of compressed coded speech voiceprint recognition according to claim 6, characterized in that, In S3, among the distortion parameters introduced during the encoding process, the variance distortion parameter is: ; in, This represents the variance of the nth dimension of the i-th Gaussian component in the general background model of uncoded speech. This represents the variance of the nth dimension of the encoding / decoding distortion model. This represents the variance of the nth dimension of the i-th Gaussian component in the updated Universal Background Model (UBM).
8. The method for improving the accuracy of compressed-coded speech voiceprint recognition according to claim 7, characterized in that, The parameters of the general background model for speech coding distortion compensation are: ; , ; , 。 9. The method for improving the accuracy of compressed coded speech voiceprint recognition according to any one of claims 1-8, characterized in that, In S1, the MFCC feature extraction method is used to extract speech features from uncompressed speech and speech compressed speech after speech coding, respectively.