Voiceprint Extraction Method, Voiceprint Recognition Method and Related Devices, Equipment and Media
By employing Gaussian Mixture Models to analyze and remove channel noise from voice features, the method enhances voiceprint recognition accuracy across various communication channels.
Patent Information
- Application Number
- CN202210683340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-06-15
AI Technical Summary
When existing voiceprint recognition systems face voice data from multiple channel sources, channel noise leads to a degradation of recognition performance, how to weaken channel noise to improve the accuracy of voiceprint recognition.
By obtaining the difference voiceprint characteristics of the target speech and the reference voiceprint characteristics, the Gaussian mixed model is used to match and strip the channel characteristics, and combining feature fusion technology, the voiceprint characteristics are optimized to weaken channel noise.
It effectively weakens the channel noise in the voiceprint characteristics and improves the accuracy of voiceprint recognition.
Smart Images

Figure CN115223571B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech processing, and particularly to a voiceprint extraction method, a voiceprint recognition method, and related devices, equipment, and media. Background Art
[0002] Modern scientific research shows that voiceprints not only have particularity but also have the characteristic of relative stability. After adulthood, a person's voice can remain relatively stable for a long time. Due to the individual differences in the vocal tract, oral cavity, and nasal cavity of each person, the physiological characteristics of the speaker can be extracted from the speech to measure the diversity among speakers. As a biometric technology, automatic speaker recognition has been widely applied to industry fields such as access control systems, e-commerce, and smart products due to its convenience, reliability, low cost, and other characteristics.
[0003] However, existing voiceprint recognition systems usually only process speech data from a single channel source. When facing speech data from multiple channel sources, due to the diversity of transmission media such as landline phones, mobile phones, satellite calls, and instant messaging, the voice of a person will cause channel noise in the voiceprint information when transmitted through different media, and there will be a problem of channel mismatch in voiceprint recognition for different channels, thus seriously affecting the performance of voiceprint recognition. In view of this, how to weaken the channel noise in the voiceprint features as much as possible to improve the accuracy of voiceprint recognition has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by the present application is to provide a voiceprint extraction method, a voiceprint recognition method, and related devices, equipment, and media, which can weaken the channel noise in the voiceprint features as much as possible to improve the accuracy of voiceprint recognition.
[0005] To solve the above technical problem, in the first aspect of the present application, a voiceprint extraction method is provided, including: obtaining the difference voiceprint features between the initial voiceprint features extracted from each target voice of a target object and the reference voiceprint features respectively; determining, from several Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint features as the target mixture model corresponding to the difference voiceprint features; wherein, several Gaussian mixture models are respectively trained based on different sample feature sets, different sample feature sets are obtained by clustering several sample difference voiceprint features, and several sample difference voiceprint features are respectively obtained from the sample voiceprint features and the reference voiceprint features extracted from each sample speech; analyzing to obtain channel features based on the difference voiceprint features and the target mixture model corresponding to the difference voiceprint features, and stripping the channel features from the initial voiceprint features corresponding to the difference voiceprint features to obtain the optimized voiceprint features corresponding to the difference voiceprint features; performing feature fusion based on the optimized voiceprint features respectively corresponding to each difference voiceprint feature to obtain the final voiceprint feature of the target object.
[0006] To solve the above technical problems, a second aspect of the present application provides a voiceprint recognition method, including: obtaining a plurality of registered voiceprint features and the voiceprint features to be recognized of the object to be recognized; wherein, the plurality of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of a plurality of registered objects and the voice to be recognized of the object to be recognized through the voiceprint extraction method in the first aspect above; and obtaining the recognition result of the object to be recognized based on the similarity between the voiceprint features to be recognized and the plurality of registered voiceprint features respectively.
[0007] To solve the above technical problems, a third aspect of the present application provides a voiceprint extraction device, including: a voiceprint difference module, a model matching module, a feature analysis module, an interference stripping module, and a feature fusion module. The voiceprint difference module is used to obtain the difference voiceprint features between the initial voiceprint features extracted from each target voice of the target object and the reference voiceprint feature respectively; the model matching module is used to determine, from a plurality of Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint features as the target mixture model corresponding to the difference voiceprint features; wherein, the plurality of Gaussian mixture models are respectively trained based on different sample feature sets, and the different sample feature sets are obtained by clustering a plurality of sample difference voiceprint features, and the plurality of sample difference voiceprint features are respectively obtained based on the sample voiceprint features extracted from each sample voice and the reference voiceprint feature; the feature analysis module is used to analyze and obtain the channel features based on the difference voiceprint features and the target mixture model corresponding to the difference voiceprint features; the interference stripping module is used to strip the channel features from the initial voiceprint features corresponding to the difference voiceprint features to obtain the optimized voiceprint features corresponding to the difference voiceprint features; the feature fusion module is used to perform feature fusion based on the optimized voiceprint features corresponding to each difference voiceprint feature to obtain the final voiceprint features of the target object.
[0008] To solve the above technical problems, a fourth aspect of the present application provides a voiceprint extraction device, including: a feature acquisition module and a result acquisition module. The feature acquisition module is used to obtain a plurality of registered voiceprint features and the voiceprint features to be recognized of the object to be recognized; wherein, the plurality of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of a plurality of registered objects and the voice to be recognized of the object to be recognized through the voiceprint extraction device in the third aspect; the result acquisition module is used to obtain the recognition result of the object to be recognized based on the similarity between the voiceprint features to be recognized and the plurality of registered voiceprint features respectively.
[0009] To solve the above technical problems, a fifth aspect of the present application provides an electronic device, including a memory and a processor coupled to each other. The memory stores program instructions, and the processor is used to execute the program instructions to implement the voiceprint extraction method in the first aspect above, or implement the voiceprint recognition method in the second aspect above.
[0010] To solve the above technical problems, a sixth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the voiceprint extraction method of the first aspect above, or implement the voiceprint recognition method of the second aspect above.
[0011] In the above solution, the initial voiceprint features extracted from the target voices of the target object are obtained, and the difference voiceprint features between the initial voiceprint features and the reference voiceprint features are respectively determined. Then, from a number of Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint features is determined as the target mixture model corresponding to the difference voiceprint features. The number of Gaussian mixture models are respectively trained based on different sample feature sets, and different sample feature sets are obtained by clustering a number of sample difference voiceprint features. The number of sample difference voiceprint features are respectively obtained from the sample voiceprint features and the reference voiceprint features extracted from each sample voice. On this basis, based on the difference voiceprint features and the target mixture model corresponding to the difference voiceprint features, the channel features are analyzed, and the channel features are stripped from the initial voiceprint features corresponding to the difference voiceprint features to obtain the optimized voiceprint features corresponding to the difference voiceprint features, and the feature fusion is performed based on the optimized voiceprint features respectively corresponding to each difference voiceprint feature to obtain the final voiceprint features of the target object. Since a number of sample difference voiceprint features are obtained from the sample voiceprint features and the reference voiceprint features extracted from each sample voice in advance, so as to highlight the feature information related to the channel in the sample voiceprint features through the sample difference voiceprint features, and different sample feature sets are obtained by clustering based on this, so that the sample difference voiceprint features corresponding to the same or similar channels are gathered in the same set as much as possible. Based on this, a number of Gaussian mixture models are trained through different sample feature sets, so as to reflect the feature distributions of different channels through the trained Gaussian mixture models. Thus, in the process of voiceprint extraction, through the difference voiceprint features and the Gaussian mixture models trained in advance, the Gaussian mixture model that matches the target voice at the channel level can be determined, and then the channel features are analyzed through the matching Gaussian mixture model and the difference voiceprint features, so as to strip the channel features from the initial voiceprint features. Therefore, the channel noise in the voiceprint features can be weakened as much as possible to improve the accuracy of voiceprint recognition. Description of the Drawings
[0012] Figure 1 is a schematic flowchart of an embodiment of the voiceprint extraction method of the present application;
[0013] Figure 2 is a schematic framework diagram of an embodiment of training a voiceprint extraction network;
[0014] Figure 3 is a schematic flowchart of an embodiment of the voiceprint recognition method of the present application;
[0015] Figure 4 is a schematic framework diagram of an embodiment of the voiceprint extraction device of the present application;
[0016] Figure 5 It is a schematic framework diagram of an embodiment of the voiceprint recognition device of the present application;
[0017] Figure 6 It is a schematic framework diagram of an embodiment of the electronic device of the present application;
[0018] Figure 7 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0019] Next, in combination with the accompanying drawings of the specification, the solutions of the embodiments of the present application will be described in detail.
[0020] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0021] The terms "system" and "network" in this article are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0022] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the voiceprint extraction method of the present application.
[0023] Specifically, it may include the following steps:
[0024] Step S11: Obtain the differential voiceprint features between the initial voiceprint features extracted from the target voices of the target object and the reference voiceprint features respectively.
[0025] In an implementation scenario, the target object can be set according to the stage of voiceprint extraction. Specifically, in the voiceprint registration stage, the target object can be the object for which the voiceprint needs to be registered, that is, the voiceprint of these objects needs to be extracted in this stage. For example, in the access control entry scenario, the target object can include, but is not limited to: newly recruited employees, etc. Or, in the e-commerce scenario, the target object can include, but is not limited to: newly registered users, etc. Other scenarios can be inferred by analogy and will not be exemplified one by one here. In the voiceprint recognition stage, the target object can be the object for which the voiceprint needs to be recognized, that is, the voiceprint of these objects needs to be extracted in this stage. For example, in the scenario of passing through the access control, the target object can include, but is not limited to: any person who needs to enter the access control (such as, internal employees, external personnel); Or, in the e-commerce scenario, the target object can include, but is not limited to: any person who needs to perform operations such as payment with the current account (such as, the owner of the current account). Other scenarios can be inferred by analogy and will not be exemplified one by one here.
[0026] It should be noted that in the voiceprint registration stage, at least one target voice of the target object can be collected, and through the steps in any publicly disclosed embodiment of the voiceprint extraction method, the final voiceprint feature of the target object can be extracted based on at least one target voice and used as the registered voiceprint feature. For example, one target voice, two target voices, three or more than three target voices of the target object can be collected, which is not limited here. In addition, in the case of collecting multiple target voices, the channel categories of the multiple target voices can be exactly the same, completely different, or not completely the same (that is, partially the same and partially different). For example, taking the collection of three target voices as an example, two target voices with the channel category of mobile phone can be collected, and one target voice with the channel category of instant messaging can be collected, which is not limited here. In particular, in order to further improve the accuracy of subsequent voiceprint recognition as much as possible, in the voiceprint registration stage, multiple target voices of the target object can be collected, and the channel categories of each target voice are different from each other. Different from the voiceprint registration stage, in the voiceprint recognition stage, only one target voice of the target object can be collected, and through the steps in any publicly disclosed embodiment of the voiceprint extraction method, the final voiceprint feature of the target object can be extracted based on this target voice and used as the voiceprint feature to be recognized.
[0027] In an implementation scenario, the spectrogram of the target voice can be extracted first, and based on the spectrogram, feature information related to the speaker's speaking characteristics, such as different frequency distributions, different average durations of phonemes, fundamental frequencies, etc., can be extracted to obtain the initial voiceprint feature of the target voice. It should be noted that the above "different frequency distributions", "different average durations of phonemes", and "fundamental frequencies" are only artificial features that may be designed in the actual application process, and do not limit other artificial features because of this. That is, in the actual application process, other artificial features can also be designed based on the spectrogram as part of the initial voiceprint feature.
[0028] In another implementation scenario, different from the aforementioned method for extracting initial voiceprint features, in order to improve the extraction efficiency of initial voiceprint features, a voiceprint extraction network can be pre-trained. The voiceprint extraction network can include, but is not limited to, xvector, etc. The network structure of the voiceprint extraction network is not limited herein. In addition, for the specific extraction process of the voiceprint extraction network, technical details such as xvector can be referred to and will not be elaborated herein. The voiceprint extraction network can be obtained through multi-task joint training based on each sample voice. The multi-tasks at least include a speaker prediction task, a channel category prediction task, and a feature distribution constraint task. Each sample voice includes a first voice and a second voice. The first voice is labeled with the sample speaker, and the second voice is labeled with the sample channel category. Moreover, during the training process, the gradient of the prediction loss of the channel category prediction task is reversed. The feature distribution constraint task is used to constrain the second voices labeled with different sample channel categories to tend to the same feature distribution. In the above manner, by jointly training the voiceprint extraction network through multi-tasks, and the multi-tasks at least include a speaker prediction task, a channel category prediction task, and a feature distribution constraint task, the voiceprint extraction network can be constrained by the speaker prediction task to extract as much feature information beneficial to differentiating different speakers (i.e., voiceprint-related feature information) as possible. And the voiceprint extraction network can be constrained by the channel category prediction task to extract as little feature information beneficial to differentiating different channel categories (i.e., channel-related feature information) as possible. And the voiceprint extraction network can be constrained by the feature distribution constraint task to make the features extracted from different channel voices tend to the same feature distribution. Therefore, the influence of channel interference on voiceprint extraction can be further weakened, the network performance of the voiceprint extraction network can be improved, and the accuracy of the initial voiceprint features can be further improved.
[0029] In a specific implementation scenario, to further ensure the effectiveness of training the voiceprint extraction network based on sample voices, each first voice can be no less than 180 seconds, and the total number of different sample speakers can be no less than 2000 people. In addition, the sample voices of the same sample speaker can cover all sample channel categories or only some sample channel categories, which is not limited herein.
[0030] In a specific implementation scenario, similar to the first sample speech, to further ensure the effectiveness of training the voiceprint extraction network based on the sample speech, for each sample channel category, a total of 500 hours of the second speech of this sample channel category can be collected. Each time the voiceprint extraction network is trained, a batch can be taken from the first speech, and each batch involves K1 different sample speaker objects. For each sample speaker object, K2 first speeches are taken, and each batch has a total of K1 * K2 first speeches. For the convenience of description, K1 * K2 can be denoted as N. At the same time, each time the voiceprint extraction network is trained, a batch can also be taken from the second speech, and each batch can also include N second speeches. Since the number of channel types is relatively limited compared to the speaker objects, the N second speeches can cover all different sample channel categories.
[0031] In a specific implementation scenario, the acoustic features of the sample speech (such as filter bank features, mel cepstral coefficient features, etc.) can be pre-extracted, and then the acoustic features of the sample speech are input into the voiceprint extraction network to obtain the sample voiceprint features. For example, the acoustic features of the first speech can be pre-extracted, and then the acoustic features of the first speech are input into the voiceprint extraction network to obtain the first voiceprint feature of the first speech. At the same time, the acoustic features of the second speech can be pre-extracted, and then the acoustic features of the second speech are input into the voiceprint extraction network to obtain the second voiceprint feature of the second speech. On this basis, the aforementioned multi-task joint training is performed based on the first voiceprint feature and the second voiceprint feature.
[0032] In a specific implementation scenario, after the first voiceprint feature of the first speech and the second voiceprint feature of the second speech are respectively extracted based on the voiceprint extraction network, predictions can be made based on the first voiceprint feature and the second voiceprint feature respectively to obtain the predicted speaker object of the first speech and the predicted channel category of the second speech. Please refer to Figure 2 , Figure 2 is a schematic framework diagram of an embodiment for training the voiceprint extraction network. As Figure 2As shown, the first voiceprint feature is processed by a linear fully connected layer, and the predicted probability values of each sample speaker can be obtained. In particular, the sample speaker corresponding to the maximum predicted probability value can be used as the predicted speaker. Similarly, the second voiceprint feature is processed by a channel classification layer, and the predicted probability values of each sample speaker can be obtained. In particular, the sample channel category corresponding to the maximum predicted probability value can be used as the predicted channel category. The channel classification layer may include, but is not limited to, a fully connected layer, a softmax layer, etc., and is not limited herein. At the same time, the first mean feature of the second voiceprint feature extracted from the second voice labeled with the same sample channel category can also be obtained, and the second mean feature of the first mean features corresponding to various sample channel categories can be obtained. That is to say, for each sample channel category, the second voiceprint feature labeled with this sample channel category can be averaged to obtain the first mean feature. That is to say, for each sample channel category, its corresponding first mean feature can be calculated. On this basis, the first mean features corresponding to various sample channel categories can be further averaged to obtain the second mean feature. On this basis, the network loss can be obtained based on the first difference between the sample speaker and the predicted speaker, the second difference between the sample channel category and the predicted channel category, and the third difference between the first mean feature corresponding to each sample channel category and the second mean feature. The first difference and the third difference are positively correlated with the network loss, and the second difference is negatively correlated with the network loss. Then, based on the network loss, the network parameters of the voiceprint extraction network are adjusted. It should be noted that for the measurement methods of the above first difference and second difference, loss functions such as cross entropy can be referred to, and for the measurement method of the above third difference, loss functions such as mean square error can be referred to, which will not be elaborated herein. The adjustment process of the network parameters can refer to optimization methods such as gradient descent, which will not be elaborated herein. In addition, during the adjustment process of the network parameters, the learning rate can be set to 0.1, 0.15, etc., which is not limited herein.In the above method, the first voiceprint feature of the first voice and the second voiceprint feature of the second voice are respectively extracted based on the voiceprint extraction network. Based on this, predictions are respectively made based on the first voiceprint feature and the second voiceprint feature to obtain the predicted speaker of the first voice and the predicted channel category of the second voice. The first mean feature of the second voiceprint feature extracted from the second voice labeled with the same sample channel category is obtained, and the second mean feature of the first mean features corresponding to various sample channel categories is obtained. Thus, based on the first difference between the sample speaker and the predicted speaker, the second difference between the sample channel category and the predicted channel category, and the third difference between the first mean features corresponding to various sample channel categories and the second mean features, the network loss is obtained. The first difference and the third difference are positively correlated with the network loss, and the second difference is negatively correlated with the network loss. Furthermore, based on the network loss, the network parameters of the voiceprint extraction network are adjusted. Therefore, during the training process, by minimizing the network loss, it is possible to constrain the voiceprint extraction network to extract as much feature information beneficial to identifying different speakers (i.e., voiceprint-related feature information) as possible, extract as little feature information beneficial to identifying different channel categories (i.e., channel-related feature information) as possible, and make the features extracted from different-channel voices tend to the same feature distribution as much as possible. Therefore, it is possible to further weaken the influence of channel interference on voiceprint extraction, improve the network performance of the voiceprint extraction network, and further improve the accuracy of the initial voiceprint features.
[0033] In a specific implementation scenario, in order to further make the voiceprint features extracted from different voices of the same speaker as close as possible, before adjusting the network parameters of the voiceprint extraction network based on the network loss, the sample voices with the same sample speaker as the first voice can be used as positive example voices, and the sample voices with different sample speakers from the first voice can be used as negative example voices. On this basis, the mean value of the first voiceprint features extracted from each positive example voice is taken to obtain the positive example voiceprint features, and the mean value of the first voiceprint features extracted from each negative example voice is taken to obtain the negative example voiceprint features. Then, based on the first distance between the first voiceprint feature of the first voice and the positive example voiceprint feature corresponding to the first voice, and the second distance between the first voiceprint feature of the first voice and the negative example voiceprint feature corresponding to the first voice, a feature comparison pair loss is obtained. The first distance is positively correlated with the feature comparison pair loss, and the second distance is negatively correlated with the feature comparison pair loss. Furthermore, based on the feature comparison pair loss, the network loss is updated, and the feature comparison pair loss is positively correlated with the updated network loss. It should be noted that the measurement methods of the above first distance and second distance can refer to measurement methods such as cosine similarity and will not be elaborated here. In addition, the original network loss and the modulation comparison pair loss can be weighted, added, etc. to update the original network loss. The above method further incorporates the feature comparison pair loss into the network loss. The feature comparison pair loss is measured by the first distance between the first voiceprint feature of the first voice and the positive example voiceprint feature corresponding to the first voice, and the second distance between the first voiceprint feature of the first voice and the negative example voiceprint feature corresponding to the first voice. The first distance is positively correlated with the feature comparison pair loss, and the second distance is negatively correlated with the feature comparison pair loss. Therefore, by minimizing the network loss, it is possible to further force the first voiceprint feature of the first voice to be as close as possible to the positive example voiceprint feature corresponding to the first voice, and force the second distance between the first voiceprint feature of the first voice and the negative example voiceprint feature corresponding to the first voice to be as far apart as possible, thereby further improving the network performance of the voiceprint extraction network to further improve the accuracy and robustness of the initial voiceprint features.
[0034] In a specific implementation scenario, for the sake of description, for the nth first voice in a batch of the first voices, if the sample speaker it is labeled with is the ith among all sample speakers, then this first voice can be denoted as The sample speaker it is labeled with can be denoted as Similarly, for the kth second voice in a batch of the second voices, if the sample channel category it is labeled with is the lth among all sample channel categories, then this second voice can be denoted as The sample channel category it is labeled with can be denoted as In addition, the first voice The predicted probability values belonging to various sample speaker objects can be denoted as where f represents the mathematical function of the voiceprint extraction network, i.e., it represents the first voiceprint feature, g1 represents the linear fully connected layer, and σ represents the softmax function; similarly, for the second voice The predicted probability values belonging to various sample channel categories can be denoted as where g2 represents the channel classification layer. Therefore, the network loss Loss can be expressed as:
[0035]
[0036] In the above formula (1), CELoss represents the cross-entropy loss, represents the first loss value calculated from the first difference between the sample speaker object and the predicted speaker object, DomainLoss represents the channel classification loss, represents the second loss value calculated from the second difference between the sample channel category and the predicted channel category, AMMDLoss represents the feature distribution difference, represents the third loss value calculated from the third difference between the first mean feature corresponding to various sample channel categories and the second mean feature, and S represents the total number of all sample channel categories. In addition, TripletLoss represents the triplet loss, represents the first distance between the first voiceprint feature of the first voice and the positive example voiceprint feature corresponding to the first voice, and the feature comparison pair loss value calculated from the second distance between the first voiceprint feature of the first voice and the negative example voiceprint feature corresponding to the first voice. After obtaining the above first loss value, second loss value, third loss value, and fourth loss value, the above loss values can be weighted to obtain the network loss, where α and β represent the weight coefficients. For the above first loss value, it can be further expressed as:
[0037]
[0038] In addition, the above second loss value can be further expressed as:
[0039]
[0040] In addition, the above third loss value can be further expressed as:
[0041]
[0042]
[0043]
[0044] In the above formulas (4) to (6), represents the second voiceprint feature extracted from the l-th second voice among the second voices of the l-th sample channel category, N l represents the total number of second voices labeled as the l-th sample channel category, average_batch_l represents the first mean feature corresponding to the l-th sample channel category, and average_batch_all represents the second mean feature obtained by taking the mean of the first mean features corresponding to various sample channel categories. In addition, mse represents the mean squared error loss function.
[0045] In addition, the above fourth loss value can be further expressed as:
[0046]
[0047]
[0048] In the above formulas (7) and (8), represents the first distance between the first voiceprint feature of the first voice and the positive example voiceprint feature corresponding to the first voice, represents the second distance between the first voiceprint feature of the first voice and the negative example voiceprint feature corresponding to the first voice. In addition, d(f(x n ), f(x k )) characterizes the calculation method of the feature distance. Specifically, the expression on the right side of the minus sign in formula (8) represents the cosine similarity between f(x n ) and f(x k ), and the feature distance is negatively correlated with the cosine similarity.
[0049] In a specific implementation scenario, after calculating the network loss through the above method, the network parameters of the voiceprint extraction network can be adjusted based on this network loss, and thus one training can be completed. By repeating this process, iterative training of the voiceprint extraction network can be achieved. When the preset conditions are met, the iterative training can be terminated. It should be noted that the preset conditions can include but are not limited to: the network loss is lower than the first threshold, the number of iterations is not less than the second threshold, etc., which are not limited here. On this basis, the voiceprint extraction network with training convergence can be used to extract features from each target voice of the target object, and the initial voiceprint features corresponding to each target voice can be obtained.
[0050] In an implementation scenario, the reference voiceprint feature can be obtained based on the sample voiceprint features extracted from each sample voice. Specifically, after the voiceprint extraction network converges during training, the voiceprint extraction network can be used to extract the sample voiceprint features of each sample voice respectively. On this basis, the average value of the sample voiceprint features of each sample voice can be taken as the reference voiceprint feature. For ease of description, the reference voiceprint feature can be denoted as xvector_average, and the reference voiceprint feature xvector_average can be expressed as:
[0051]
[0052] In the above formula (9), N all represents the total number of each sample voice, and xvector(i) represents the sample voiceprint feature extracted by the converged voiceprint extraction network for the i-th sample voice.
[0053] In an implementation scenario, for the target object i, the difference between the initial voiceprint feature extracted from its l-th target voice and the reference voiceprint feature can be calculated to obtain the corresponding difference voiceprint feature, which can be denoted as Δxvector(spk_i_l) for ease of description.
[0054] Step S12: Determine the Gaussian mixture model that matches the difference voiceprint feature from several Gaussian mixture models as the target mixture model corresponding to the difference voiceprint feature.
[0055] In the embodiments of the present disclosure, several Gaussian mixture models (GMMs) are respectively trained based on different sample feature sets. The different sample feature sets are obtained by clustering several sample difference voiceprint features, and the several sample difference voiceprint features are respectively obtained based on the sample voiceprint features and the reference voiceprint features extracted from each sample voice.
[0056] In an implementation scenario, the sample voiceprint feature extracted from the i-th sample voice can be denoted as xvector(i), and the corresponding sample difference voiceprint feature Δxvector(i) can be expressed as:
[0057] Δxvector(i) = xvector(i) - xvector_average......(10)
[0058] On this basis, the sample difference voiceprint features can be clustered to obtain several sample feature sets. Exemplarily, for the above sample difference voiceprint features, clustering algorithms such as AP (Affinity Propagation) clustering can be used for clustering to obtain several sample feature sets. It should be noted that since the sample difference voiceprint features are obtained by subtracting the reference voiceprint features from the sample voiceprint features, the feature information related to the channel remaining in the sample voiceprint features is highlighted by the sample difference voiceprint features. By clustering in this way, the sample difference voiceprint features from the same or similar channel sources can be gathered in the same set as much as possible.
[0059] In an implementation scenario, after obtaining several sample feature sets, Gaussian mixture models can be trained using each sample feature set respectively. Therefore, if P sample feature sets are obtained by clustering, P Gaussian mixture models can be correspondingly trained, representing the feature distributions of P channel sources. Specifically, the MAP (Maximum A Posteriori estimation) algorithm can be used to train the corresponding Gaussian mixture models based on each of the above sample feature sets respectively. For the specific meaning of the Gaussian mixture model, the technical details of the Gaussian mixture model can be referred to and will not be elaborated here.
[0060] In an implementation scenario, after obtaining the difference voiceprint features corresponding to each target voice of the target object, for each difference voiceprint feature, the Gaussian mixture model that matches it can be determined from the several Gaussian mixture models trained above as the target mixture model corresponding to the difference voiceprint feature.
[0061] In an implementation scenario, for each difference voiceprint feature, the difference voiceprint feature can be input into several Gaussian mixture models respectively, so that each Gaussian mixture model can output the matching probability with the difference voiceprint feature. It should be noted that the matching probability represents the matching degree between the difference voiceprint feature and the channel source corresponding to the Gaussian mixture model. The larger the matching probability, the more the difference voiceprint feature matches the channel source corresponding to the Gaussian mixture model. On the contrary, the smaller the matching probability, the less the difference voiceprint feature matches the channel source corresponding to the Gaussian mixture model. On this basis, the Gaussian mixture model corresponding to the maximum matching probability can be selected as the target mixture model corresponding to the difference voiceprint feature. In addition, for the specific calculation process of the matching probability, the relevant description below can be referred to and will not be elaborated here for the time being.
[0062] In another implementation scenario, different from the foregoing method, in order to improve the accuracy of determining the target mixture model, a Universal Background Model (UBM) can be pre-trained based on different sample feature sets. On this basis, for each differential voiceprint feature, the target mixture model can be determined from several Gaussian mixture models based on the matching likelihood ratios of the differential voiceprint feature on each Gaussian mixture model and the Universal Background Model respectively. The above method, based on the matching likelihood ratios of the differential voiceprint vector on each Gaussian mixture model and the Universal Background Model respectively, determines the target mixture model from several Gaussian mixture models, which can further refine the channel differences and interference factors and help improve the accuracy of determining the target mixture model.
[0063] In a specific implementation scenario, it should be noted that the model structure of the Universal Background Model is similar to that of the foregoing Gaussian mixture model, and the main difference lies in the training data. The Universal Background Model can be trained from the union of different sample feature sets. Exemplarily, the Gaussian mixture coefficients of the Universal Background Model can be denoted as M. The initial Universal Background Model can be obtained by using K-Means clustering, and then the union of different sample feature sets can be iteratively trained by using the EM (Expectation-Maximization) algorithm to obtain the final Universal Background Model. It should be noted that the Gaussian mixture coefficient M of the Universal Background Model can be set to an integer value, such as 1024, 512, 256, etc., which is not limited here. Generally speaking, the more labeled sample voices there are, the larger the Gaussian mixture coefficient M can be set.
[0064] In a specific implementation scenario, as described above, for each differential voiceprint feature, the differential voiceprint feature can be input into several Gaussian mixture models respectively, so that each Gaussian mixture model can output the matching probability with the differential voiceprint feature. At the same time, the differential voiceprint feature can be input into the Universal Background Model, so that the Universal Background Model also outputs the matching probability of the differential voiceprint feature. The meaning can refer to the matching probability output by the foregoing Gaussian mixture model and will not be elaborated here. On this basis, the similarity logarithmic likelihood ratio (likelihood Rate, LLR) between the matching probabilities output by each Gaussian mixture model and the matching probability output by the Universal Background Model can be further calculated as the matching likelihood ratio of the differential voiceprint vector on each Gaussian mixture model and the Universal Background Model respectively. Thus, the Gaussian mixture model corresponding to the maximum matching likelihood ratio can be selected as the target mixture model corresponding to the differential voiceprint feature.
[0065] Step S13: Based on the differential voiceprint features and the target mixture model corresponding to the differential voiceprint features, analyze to obtain channel features, and strip the channel features from the initial voiceprint features corresponding to the differential voiceprint features to obtain the optimized voiceprint features corresponding to the differential voiceprint features.
[0066] In an implementation scenario, for each differential voiceprint feature, the occupancy rate of the differential voiceprint feature on each Gaussian component of its corresponding target mixture model can be obtained, and then based on the average feature of each Gaussian component and the occupancy rate on each Gaussian component, the channel features can be obtained. It should be noted that in the case where the Gaussian mixture coefficient of the Gaussian mixture model is M, it means that the Gaussian mixture model has M Gaussian components. Since the Gaussian mixture model represents the feature distribution of a certain channel source and can be decomposed into M data distributions (such as, normal distribution), the average feature of each Gaussian component represents the average feature of each normal feature distribution. Therefore, by measuring the occupancy rate of the differential voiceprint feature on each Gaussian component and combining the average feature of each Gaussian component, the channel features can be obtained, that is, the feature information related to the channel source remaining in the initial voiceprint features corresponding to the differential voiceprint feature. In the above manner, the occupancy rate of the differential voiceprint feature on each Gaussian component of the target mixture model is obtained, and the channel features are obtained based on the average feature of each Gaussian component and the occupancy rate on each Gaussian component. Therefore, the channel features can be measured at the fine-grained level of the Gaussian components, and the channel differences and interference factors can be further refined, which helps to improve the accuracy of determining the target mixture model.
[0067] In a specific implementation scenario, in order to accurately measure the occupancy rate on each Gaussian component, the distribution probability of the differential voiceprint feature on each Gaussian component can be obtained first, and then based on the weight coefficient of each Gaussian component, the distribution probability on each Gaussian component is weighted respectively to obtain the matching probability of the differential voiceprint feature with each Gaussian component. Thus, based on the proportion of the matching probability of each Gaussian component in the sum of the matching probabilities of each Gaussian component, the occupancy rate on each Gaussian component can be obtained. It should be noted that the "matching probability output by the Gaussian mixture model" in the aforementioned step S12 can be obtained by summing the matching probabilities of each Gaussian component. Similarly, the "matching probability output by the universal background model" in the aforementioned step S12 can also be obtained by summing the matching probabilities of each Gaussian component of the universal background model, which will not be elaborated here. As mentioned above, for the sake of convenience of description, for the target object i, the differential voiceprint feature corresponding to its l-th target voice can be denoted as Δxvector(spk_i_l). Then the distribution probability of the differential voiceprint feature Δxvector(spk_i_l) on the j-th Gaussian component of its corresponding target mixture model can be denoted as N j (Δxvector(spk_i_l)), and the weight coefficient of the j-th Gaussian component can be denoted as wj , the proportion p(j|spk_i_l) on the j-th Gaussian component can be expressed as:
[0068]
[0069] In the above formula (11), M represents the Gaussian mixture coefficient of the target mixture model, that is, the total number of Gaussian components in the target mixture model. By the above method, the distribution probability of the differential voiceprint feature on each Gaussian component is obtained, and based on the weight coefficients of each Gaussian component, the distribution probabilities on each Gaussian component are weighted respectively to obtain the matching probabilities of the differential voiceprint feature with each Gaussian component. Then, based on the proportion of the matching probability of each Gaussian component in the sum of the matching probabilities of each Gaussian component, the occupancy rate on each Gaussian component is obtained. Therefore, the occupancy rate of the differential voiceprint feature on each Gaussian component of the target mixture model can be accurately measured by combining the weight coefficients of each Gaussian component and the distribution probabilities on each Gaussian component.
[0070] In a specific implementation scenario, after obtaining the occupancy rates on each Gaussian component, the average features of each Gaussian component can be weighted respectively based on the occupancy rates on each Gaussian component to obtain the weighted features of each Gaussian component, and the channel features can be obtained by fusing (such as adding) the weighted features of each Gaussian component. As mentioned above, the proportion on the j-th Gaussian component can be denoted as p(j|spk_i_l), then the channel feature domain9spk_i_l) can be expressed as:
[0071]
[0072] In the above formula (12), u j represents the average feature of the j-th Gaussian component, and M represents the Gaussian mixture coefficient of the target mixture model, that is, the total number of Gaussian components in the target mixture model. By the above method, based on the occupancy rates on each Gaussian component, the average features of each Gaussian component are weighted respectively to obtain the weighted features of each Gaussian component, and the channel features are obtained by fusing the weighted features of each Gaussian component. Therefore, the channel features can be accurately measured by combining the occupancy rates on each Gaussian component and the average features of each Gaussian component.
[0073] In another implementation scenario, different from the aforementioned calculation method of channel features, when the accuracy requirement for voiceprint features is relatively loose, in order to simplify the calculation process of channel features, the matching probability of the differential voiceprint feature on its corresponding target mixture model can be obtained first. The specific calculation process can refer to the aforementioned relevant description and will not be elaborated here. Meanwhile, the model features of the target mixture model can be obtained. For example, based on the weight coefficients of each Gaussian component of the target mixture model, the average features of each Gaussian component can be weighted and summed to obtain the model features of the target mixture model. On this basis, the channel features can be obtained based on the matching probability of the differential voiceprint feature on its corresponding target mixture model and the model features of the target mixture model. Exemplarily, the channel features can be obtained by directly multiplying the matching probability by the model features.
[0074] It should be noted that for each differential voiceprint feature, after obtaining the channel features, the initial voiceprint feature corresponding to the differential voiceprint feature can be subtracted by the channel feature corresponding to the differential voiceprint feature to strip the channel feature from the initial voiceprint feature, and then the optimized voiceprint feature corresponding to the differential voiceprint feature can be obtained. For the sake of easy description, for the target object i, the optimized voiceprint feature xvector final (spk_i_l) can be expressed as:
[0075]
[0076]
[0077] Step S14: Perform feature fusion based on the optimized voiceprint features respectively corresponding to each differential voiceprint feature to obtain the final voiceprint feature of the target object.
[0078] Specifically, after obtaining the optimized voiceprint features respectively corresponding to each differential voiceprint feature, the means of these optimized voiceprint features can be taken to achieve feature fusion, so as to obtain the final voiceprint feature of the target object. Exemplarily, the final voiceprint feature xvector final (spk_i) can be expressed as:
[0079]
[0080] In the above formula (14), L represents the total number of target voices of the target object. It should be noted that in the registration stage, voiceprint extraction needs to be performed on the registration voice of the registration object. At this time, L is usually greater than 1, and of course, it can also be equal to 1. While in the test stage, voiceprint extraction needs to be performed on the voice to be recognized of the object to be recognized. At this time, L can be equal to 1.
[0081] In the above solution, since a number of sample difference voiceprint features are obtained based on the sample voiceprint features extracted from each sample voice and the reference voiceprint features in advance, so as to highlight the channel-related feature information in the sample voiceprint features through the sample difference voiceprint features, and different sample feature sets are obtained by clustering based on this, so that the sample difference voiceprint features corresponding to the same or similar channels are gathered in the same set as much as possible. Based on this, a number of Gaussian mixture models are trained through different sample feature sets, so as to reflect the feature distributions of different channels through the trained Gaussian mixture models. Thus, in the process of voiceprint extraction, through the difference voiceprint features and the Gaussian mixture models trained in advance, the Gaussian mixture model that matches the target voice at the channel level can be determined, and then the channel features are analyzed through the matching Gaussian mixture model and the difference voiceprint features, so as to strip the channel features from the initial voiceprint features. Therefore, the channel noise in the voiceprint features can be weakened as much as possible to improve the accuracy of voiceprint recognition.
[0082] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of the voiceprint recognition method of this application.
[0083] Specifically, it may include the following steps:
[0084] Step S31: Obtain a number of registered voiceprint features and the voiceprint features to be recognized of the object to be recognized.
[0085] In the embodiments of the present disclosure, a number of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting through the steps in any of the above-mentioned embodiments of the voiceprint extraction method from the registered voices of a number of registered objects and the voiceprint features to be recognized of the object to be recognized.
[0086] In an implementation scenario, when extracting the registered voiceprint features of a registered object, the registered object can be used as the target object, the registered voice of the registered object can be used as the target voice, and the final voiceprint features of the registered object are obtained by extracting through the steps in any of the above-mentioned embodiments of the voiceprint extraction method, as the registered voiceprint features of the registered object; or, when extracting the voiceprint features to be recognized of the object to be recognized, the object to be recognized can be used as the target object, the voiceprint features to be recognized of the object to be recognized can be used as the target voice, and the final voiceprint features of the object to be recognized are obtained by extracting through the steps in any of the above-mentioned embodiments of the voiceprint extraction method, as the voiceprint features to be recognized of the object to be recognized. For the specific extraction process, please refer to the above-mentioned embodiments of the voiceprint extraction method, which will not be elaborated here.
[0087] Step S32: Obtain the recognition result of the object to be recognized based on the similarity between the voiceprint features to be recognized and a number of registered voiceprint features.
[0088] In an implementation scenario, the similarities between the voiceprint features to be recognized and several registered voiceprint features can be measured by metrics such as cosine similarity, which is not limited herein. It should be noted that the greater the similarity between the voiceprint features to be recognized and the registered voiceprint features, the greater the likelihood that the object to be recognized is the registered object to which the registered voiceprint features belong. Conversely, the smaller the similarity between the voiceprint features to be recognized and the registered voiceprint features, the smaller the likelihood that the object to be recognized is the registered object to which the registered voiceprint features belong.
[0089] In an implementation scenario, it can be determined whether the maximum similarity is greater than a preset threshold. If so, it can be determined that the recognition result of the object to be recognized includes: the object to be recognized is the registered object to which the registered voiceprint feature corresponding to the maximum similarity belongs; conversely, if the maximum similarity is not greater than the preset threshold, it can be determined that the recognition result of the object to be recognized includes: the object to be recognized is not a registered object. It should be noted that the preset threshold can be set according to the actual situation. Exemplarily, if a higher requirement for misrecognition in voiceprint recognition is required, that is, the object to be recognized is actually not a registered object, but the probability of being misrecognized as a registered object should be as low as possible, the preset threshold can be set larger. Or, if a higher requirement for missed recognition in voiceprint recognition is required, that is, the object to be recognized is actually a registered object, but the probability of not being recognized as a registered object should be as low as possible, the preset threshold can be set slightly smaller. Other situations can be inferred by analogy and will not be exemplified one by one here.
[0090] In the above solution, since several registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of several registered objects and the voice to be recognized of the object to be recognized through the steps in any of the publicly disclosed embodiments of the voiceprint extraction method, the channel noise in the voiceprint features can be weakened as much as possible. Thus, based on the similarities between the voiceprint features to be recognized and several registered voiceprint features, the recognition result of the object to be recognized can be obtained. Therefore, the channel interference can be reduced as much as possible, and the accuracy of voiceprint recognition can be improved.
[0091] Please refer to Figure 4 , Figure 4It is a schematic framework diagram of an embodiment of the voiceprint extraction device 40 of the present application. The voiceprint extraction device 40 includes: a voiceprint difference module 41, a model matching module 42, a feature analysis module 43, an interference stripping module 44, and a feature fusion module 45. The voiceprint difference module 41 is used to obtain the difference voiceprint features between the initial voiceprint features extracted from the target voices of the target object and the reference voiceprint features respectively. The model matching module 42 is used to determine, from a number of Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint features as the target mixture model corresponding to the difference voiceprint features. Among them, the number of Gaussian mixture models are respectively trained based on different sample feature sets, and the different sample feature sets are obtained by clustering a number of sample difference voiceprint features, and the number of sample difference voiceprint features are respectively obtained based on the sample voiceprint features and the reference voiceprint features extracted from each sample voice. The feature analysis module 43 is used to analyze and obtain the channel features based on the difference voiceprint features and the target mixture model corresponding to the difference voiceprint features. The interference stripping module 44 is used to strip the channel features from the initial voiceprint features corresponding to the difference voiceprint features to obtain the optimized voiceprint features corresponding to the difference voiceprint features. The feature fusion module 45 is used to perform feature fusion based on the optimized voiceprint features respectively corresponding to each difference voiceprint feature to obtain the final voiceprint features of the target object.
[0092] In the above solution, since a number of sample difference voiceprint features are obtained in advance based on the sample voiceprint features and the reference voiceprint features extracted from each sample voice, so as to highlight the feature information related to the channel in the sample voiceprint features through the sample difference voiceprint features, and different sample feature sets are obtained by clustering based on this, so that the sample difference voiceprint features corresponding to the same or similar channels are gathered in the same set as much as possible. Based on this, a number of Gaussian mixture models are trained through different sample feature sets, so as to reflect the feature distribution of different channels through the trained Gaussian mixture models. Thus, in the process of voiceprint extraction, through the difference voiceprint features and the Gaussian mixture models trained in advance, the Gaussian mixture model that matches the target voice at the channel level can be determined, and then the channel features are analyzed and obtained through the matched Gaussian mixture model and the difference voiceprint features, so as to strip the channel features from the initial voiceprint features. Therefore, it is possible to weaken the channel noise in the voiceprint features as much as possible to improve the accuracy of voiceprint recognition.
[0093] In some disclosed embodiments, the feature analysis module 43 includes: an occupancy rate acquisition sub-module, which is used to obtain the occupancy rate of the difference voiceprint features on each Gaussian component of the target mixture model; the feature analysis module 43 includes: a channel feature acquisition sub-module, which is used to obtain the channel features based on the average features of each Gaussian component and the occupancy rate on each Gaussian component.
[0094] In some disclosed embodiments, the occupancy rate acquisition sub-module includes: a distribution probability acquisition unit, configured to acquire the distribution probabilities of the difference voiceprint features on each Gaussian component; the occupancy rate acquisition sub-module includes: a distribution probability weighting unit, configured to respectively weight the distribution probabilities on each Gaussian component based on the weight coefficients of each Gaussian component, to obtain the matching probabilities of the difference voiceprint features with each Gaussian component; the occupancy rate acquisition sub-module includes: an occupancy rate calculation unit, configured to obtain the occupancy rates on each Gaussian component based on the proportion of the matching probabilities of each Gaussian component in the sum of the matching probabilities of each Gaussian component.
[0095] In some disclosed embodiments, the channel feature acquisition sub-module includes: an average feature weighting unit, configured to respectively weight the average features of each Gaussian component based on the occupancy rates on each Gaussian component, to obtain the weighted features of each Gaussian component; the channel feature acquisition sub-module includes: a weighted feature fusion unit, configured to fuse the weighted features of each Gaussian component to obtain the channel features.
[0096] In some disclosed embodiments, the model matching module 42 is specifically configured to determine a target mixture model from several Gaussian mixture models based on the matching likelihood ratios of the difference voiceprint vectors on each Gaussian mixture model and the universal background model; wherein, the universal background model is trained from the union of different sample feature sets.
[0097] In some disclosed embodiments, the reference voiceprint feature is obtained by averaging the sample voiceprint features extracted from each sample voice.
[0098] In some disclosed embodiments, the initial voiceprint feature and the sample voiceprint feature are extracted by a pre-trained voiceprint extraction network, and the voiceprint extraction network is obtained by multi-task joint training based on each sample voice, and the multi-tasks at least include a speaker prediction task, a channel category prediction task, and a feature distribution constraint task. Each sample voice includes a first voice and a second voice. The first voice is labeled with a sample speaker, and the second voice is labeled with a sample channel category. And during the training process, the prediction loss of the channel category prediction task is subjected to gradient reversal, and the feature distribution constraint task is used to constrain the second voices labeled with different sample channel categories to tend to the same feature distribution.
[0099] In some disclosed embodiments, the voiceprint extraction device 40 further includes a sample voiceprint extraction module, configured to extract a first voiceprint feature of the first voice and a second voiceprint feature of the second voice respectively based on a voiceprint extraction network; the voiceprint extraction device 40 further includes a feature prediction module, configured to perform predictions respectively based on the first voiceprint feature and the second voiceprint feature to obtain a predicted speaker of the first voice and a predicted channel category of the second voice; the voiceprint extraction device 40 further includes a feature statistics module, configured to obtain a first mean feature of the second voiceprint features extracted from the second voices labeled with the same sample channel category, and obtain a second mean feature of the first mean features corresponding to various sample channel categories; the voiceprint extraction device 40 further includes a loss metric module, configured to obtain a network loss based on a first difference between a sample speaker and a predicted speaker, a second difference between a sample channel category and a predicted channel category, and a third difference between the first mean features corresponding to various sample channel categories and the second mean features; wherein, the first difference and the third difference are positively correlated with the network loss, and the second difference is negatively correlated with the network loss; the voiceprint extraction device 40 further includes a parameter adjustment module, configured to adjust network parameters of the voiceprint extraction network based on the network loss.
[0100] In some disclosed embodiments, the voiceprint extraction device 40 further includes a reference voice acquisition module, which uses the sample voice labeled with the same sample speaker as the first voice as a positive example voice, and uses the sample voice labeled with a different sample speaker from the first voice as a negative example voice; the voiceprint extraction device 40 further includes a reference feature acquisition module, configured to obtain a positive example voiceprint feature by taking the mean of the first voiceprint features extracted from each positive example voice, and obtain a negative example voiceprint feature by taking the mean of the first voiceprint features extracted from each negative example voice; the voiceprint extraction device 40 further includes a feature comparison module, configured to obtain a feature comparison sub-loss based on a first distance between the first voiceprint feature of the first voice and the positive example voiceprint feature corresponding to the first voice, and a second distance between the first voiceprint feature of the first voice and the negative example voiceprint feature corresponding to the first voice; wherein, the first distance is positively correlated with the feature comparison sub-loss, and the second distance is negatively correlated with the feature comparison sub-loss; the voiceprint extraction device 40 further includes a loss update module, configured to update the network loss based on the feature comparison sub-loss; wherein, the feature comparison sub-loss is positively correlated with the updated network loss.
[0101] Please refer to Figure 5 , Figure 5It is a schematic framework diagram of an embodiment of the voiceprint recognition device 50 of the present application. The voiceprint recognition device 50 includes: a feature acquisition module 51 and a result acquisition module 52. The feature acquisition module 51 is used to acquire a plurality of registered voiceprint features and the voiceprint features to be recognized of the object to be recognized. And the plurality of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of a plurality of registered objects and the voiceprint features to be recognized of the object to be recognized through the voiceprint extraction device in any of the above-mentioned embodiments of the voiceprint extraction device; the result acquisition module 52 is used to obtain the recognition result of the object to be recognized based on the similarity between the voiceprint features to be recognized and the plurality of registered voiceprint features respectively.
[0102] In the above solution, since the plurality of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of a plurality of registered objects and the voiceprint features to be recognized of the object to be recognized through the steps in any of the above-mentioned public embodiments of the voiceprint extraction method, the channel noise in the voiceprint features can be weakened as much as possible. On this basis, the recognition result of the object to be recognized is obtained based on the similarity between the voiceprint features to be recognized and the plurality of registered voiceprint features respectively. Therefore, the channel interference can be reduced as much as possible, and the accuracy of voiceprint recognition can be improved.
[0103] Please refer to Figure 6 , Figure 6 It is a schematic framework diagram of an embodiment of the electronic device 60 of the present application. The electronic device 60 includes a memory 61 and a processor 62 which are mutually coupled. The memory 61 stores program instructions, and the processor 62 is used to execute the program instructions to implement the steps in any of the above-mentioned embodiments of the voiceprint extraction method, or to implement the steps in any of the above-mentioned embodiments of the voiceprint recognition method. Specifically, the electronic device 60 may include, but is not limited to: desktop computers, laptop computers, servers, mobile phones, tablet computers, etc., which are not limited here.
[0104] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps in any of the above embodiments of the voiceprint extraction method, or to implement the steps in any of the above embodiments of the voiceprint recognition method. The processor 62 may also be referred to as a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 62 may be implemented jointly by integrated circuit chips.
[0105] In the above solution, since a number of sample difference voiceprint features are obtained based on the sample voiceprint features and the reference voiceprint features extracted from each sample voice in advance, so as to highlight the channel-related feature information in the sample voiceprint features through the sample difference voiceprint features, and different sample feature sets are obtained by clustering based on this, so that the sample difference voiceprint features corresponding to the same or similar channels are as much as possible clustered in the same set, and then a number of Gaussian mixture models are trained through different sample feature sets to reflect the feature distributions of different channels through the trained Gaussian mixture models. Thus, in the voiceprint extraction process, through the difference voiceprint features and the Gaussian mixture models pre-trained in advance, the Gaussian mixture model that matches the target voice at the channel level can be determined, and then the channel features are analyzed through the matched Gaussian mixture model and the difference voiceprint features to strip the channel features from the initial voiceprint features. Therefore, the channel noise in the voiceprint features can be weakened as much as possible to improve the accuracy of voiceprint recognition.
[0106] Please refer to Figure 7 , Figure 7 is a schematic framework diagram of an embodiment of the computer-readable storage medium 70 of the present application. The computer-readable storage medium 70 stores program instructions 71 that can be run by a processor. The program instructions 71 are used to implement the steps in any of the above embodiments of the voiceprint extraction method, or to implement the steps in any of the above embodiments of the voiceprint recognition method.
[0107] In the above solution, since a number of sample difference voiceprint features are obtained based on the sample voiceprint features extracted from each sample voice and the reference voiceprint features in advance, so as to highlight the channel-related feature information in the sample voiceprint features through the sample difference voiceprint features, and different sample feature sets are obtained by clustering based on this, so that the sample difference voiceprint features corresponding to the same or similar channels are gathered in the same set as much as possible. Based on this, a number of Gaussian mixture models are trained through different sample feature sets, so as to reflect the feature distributions of different channels through the trained Gaussian mixture models. Thus, in the process of voiceprint extraction, through the difference voiceprint features and the Gaussian mixture models trained in advance, the Gaussian mixture model that matches the target voice at the channel level can be determined, and then the channel features are analyzed through the matched Gaussian mixture model and the difference voiceprint features, so as to strip the channel features from the initial voiceprint features. Therefore, the channel noise in the voiceprint features can be weakened as much as possible to improve the accuracy of voiceprint recognition.
[0108] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0109] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated here.
[0110] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0111] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0112] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0114] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. A voiceprint extraction method, characterized in that, Including: Obtaining the difference voiceprint features between the initial voiceprint features extracted from each target voice of the target object and the reference voiceprint feature respectively; Determining, from a number of Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint feature as the target mixture model corresponding to the difference voiceprint feature; wherein, the number of Gaussian mixture models are respectively trained based on different sample feature sets, the different sample feature sets are obtained by clustering a number of sample difference voiceprint features, and the number of sample difference voiceprint features are respectively obtained based on the sample voiceprint features extracted from each sample voice and the reference voiceprint feature, and the reference voiceprint feature is obtained by averaging the sample voiceprint features extracted from each sample voice; Analyzing to obtain a channel feature based on the difference voiceprint feature and the target mixture model corresponding to the difference voiceprint feature, and stripping the channel feature from the initial voiceprint feature corresponding to the difference voiceprint feature to obtain the optimized voiceprint feature corresponding to the difference voiceprint feature; Performing feature fusion based on the optimized voiceprint features respectively corresponding to each difference voiceprint feature to obtain the final voiceprint feature of the target object.
2. The method according to claim 1, wherein The analyzing to obtain a channel feature based on the difference voiceprint feature and the target mixture model corresponding to the difference voiceprint feature includes: Obtaining the occupancy rate of the difference voiceprint feature on each Gaussian component of the target mixture model; Obtaining the channel feature based on the average feature of each Gaussian component and the occupancy rate on each Gaussian component.
3. The method according to claim 2, wherein The obtaining the occupancy rate of the difference voiceprint feature on each Gaussian component of the target mixture model includes: Obtaining the distribution probability of the difference voiceprint feature on each Gaussian component; Based on the weight coefficients of each Gaussian component, respectively weighting the distribution probabilities on each Gaussian component to obtain the matching probabilities of the difference voiceprint feature with each Gaussian component respectively; Based on the proportion of the matching probability of each Gaussian component in the sum of the matching probabilities of each Gaussian component, obtaining the occupancy rate on each Gaussian component.
4. The method according to claim 2, wherein The obtaining the channel feature based on the average feature of each Gaussian component and the occupancy rate on each Gaussian component includes: Based on the occupancy rate on each Gaussian component, respectively weighting the average feature of each Gaussian component to obtain the weighted feature of each Gaussian component; Fusing the weighted features of each Gaussian component to obtain the channel feature.
5. The method according to claim 1, characterized in that, The determining, from a number of Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint feature as the target mixture model corresponding to the difference voiceprint feature includes: Determining the target mixture model from the number of Gaussian mixture models based on the matching likelihood ratio of the difference voiceprint vector on each Gaussian mixture model and the universal background model respectively; Wherein, the universal background model is trained by the union of the different sample feature sets.
6. The method according to claim 1, wherein The initial voiceprint feature and the sample voiceprint feature are extracted by a pre-trained voiceprint extraction network, which is obtained by multi-task joint training based on each sample speech. The multi-task includes at least a speaker prediction task, a channel category prediction task, and a feature distribution constraint task. Each sample speech includes a first speech and a second speech. The first speech is labeled with a sample speaker, and the second speech is labeled with a sample channel category. During the training process, the gradient of the prediction loss of the channel category prediction task is reversed. The feature distribution constraint task is used to constrain the second speeches labeled with different sample channel categories to tend to the same feature distribution.
7. The method according to claim 6, wherein The training steps of the voiceprint extraction network include: Based on the voiceprint extraction network, respectively extract the first voiceprint feature of the first speech and the second voiceprint feature of the second speech; Based on the first voiceprint feature and the second voiceprint feature respectively, make predictions to obtain the predicted speaker of the first speech and the predicted channel category of the second speech, and obtain the first mean feature of the second voiceprint features extracted from the second speeches labeled with the same sample channel category, and obtain the second mean feature of the first mean features corresponding to various sample channel categories; Based on the first difference between the sample speaker and the predicted speaker, the second difference between the sample channel category and the predicted channel category, and the third difference between the first mean features corresponding to various sample channel categories and the second mean feature, obtain the network loss; wherein, the first difference and the third difference are positively correlated with the network loss, and the second difference is negatively correlated with the network loss; Based on the network loss, adjust the network parameters of the voiceprint extraction network.
8. The method according to claim 7, wherein Before adjusting the network parameters of the voiceprint extraction network based on the network loss, the method further includes: Taking the sample speech labeled with the same sample speaker as the first speech as the positive example speech, and taking the sample speech labeled with a different sample speaker from the first speech as the negative example speech; Based on the first voiceprint features extracted from each positive example speech, take the mean to obtain the positive example voiceprint feature, and based on the first voiceprint features extracted from each negative example speech, take the mean to obtain the negative example voiceprint feature; Based on the first distance between the first voiceprint feature of the first speech and the positive example voiceprint feature corresponding to the first speech, and the second distance between the first voiceprint feature of the first speech and the negative example voiceprint feature corresponding to the first speech, obtain the feature comparison pair loss; wherein, the first distance is positively correlated with the feature comparison pair loss, and the second distance is negatively correlated with the feature comparison pair loss; Based on the feature comparison pair loss, update the network loss; wherein, the feature comparison pair loss is positively correlated with the updated network loss.
9. A voiceprint recognition method, characterized in that, Including: Obtain a plurality of registered voiceprint features and the voiceprint features to be recognized of the object to be recognized; wherein, the plurality of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of a plurality of registered objects and the voice to be recognized of the object to be recognized through the voiceprint extraction method according to any one of claims 1 to 8; Based on the similarities between the voiceprint features to be recognized and the plurality of registered voiceprint features respectively, obtain the recognition result of the object to be recognized.
10. A voiceprint extraction device, characterized in that, Comprising: A voiceprint difference module, configured to obtain the difference voiceprint features between the initial voiceprint features extracted from the respective target voices of the target object and the reference voiceprint features; A model matching module, configured to determine, from a plurality of Gaussian mixture models, the Gaussian mixture model that matches the difference voiceprint features as the target mixture model corresponding to the difference voiceprint features; wherein, the plurality of Gaussian mixture models are respectively trained based on different sample feature sets, the different sample feature sets are obtained by clustering a plurality of sample difference voiceprint features, and the plurality of sample difference voiceprint features are respectively obtained based on the sample voiceprint features extracted from each sample voice and the reference voiceprint features, and the reference voiceprint features are obtained by averaging the sample voiceprint features extracted from each sample voice; A feature analysis module, configured to analyze and obtain channel features based on the difference voiceprint features and the target mixture model corresponding to the difference voiceprint features; An interference stripping module, configured to strip the channel features from the initial voiceprint features corresponding to the difference voiceprint features to obtain the optimized voiceprint features corresponding to the difference voiceprint features; A feature fusion module, configured to perform feature fusion based on the optimized voiceprint features respectively corresponding to each difference voiceprint feature to obtain the final voiceprint features of the target object.
11. A voiceprint recognition device, characterized in that, Comprising: A feature acquisition module, configured to obtain a plurality of registered voiceprint features and the voiceprint features to be recognized of the object to be recognized; wherein, the plurality of registered voiceprint features and the voiceprint features to be recognized are respectively obtained by extracting the registered voices of a plurality of registered objects and the voice to be recognized of the object to be recognized through the voiceprint extraction device according to claim 10; A result acquisition module, configured to obtain the recognition result of the object to be recognized based on the similarities between the voiceprint features to be recognized and the plurality of registered voiceprint features respectively.
12. An electronic device, characterized in that, Comprising a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the voiceprint extraction method according to any one of claims 1 to 8, or to implement the voiceprint recognition method according to claim 9.
13. A computer-readable storage medium, characterized in that, Stored with program instructions that can be run by a processor, the program instructions are used to implement the voiceprint extraction method according to any one of claims 1 to 8, or to implement the voiceprint recognition method according to claim 9.
Citation Information
Patent Citations
Method for training voiceprint representation model and related device
CN110491393A
Voice keyword recognition method based on end-to-end, device thereof and equipment
CN111429887A