Voiceprint Recognition Method, Device and Electronic Device
By adopting high-performance and low-performance voiceprint recognition models in the voiceprint recognition system and ensuring the similarity of the extracted voiceprint features through synchronous training, the problem of high resource requirements in the voiceprint verification process is solved, and more efficient resource utilization and recognition performance is achieved.
Patent Information
- Application Number
- CN202210299467.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The prior art frequently extracts the voiceprint characteristics of voiceprints during voiceprint verification, resulting in high resource requirements for electronic devices and affecting the operation efficiency of the equipment.
Two voiceprint recognition models with different performances are adopted: the high-performance second voiceprint recognition model is used for voiceprint registration, and the low-performance first voiceprint recognition model is used for voiceprint verification. Through synchronous training, the extracted voiceprint features of the two are similar.
On the basis of taking into account the accuracy of voiceprint verification, the requirements for electronic device resources are reduced, the deployment flexibility of voiceprint recognition models is improved, and the power consumption and time-consuming of equipment are reduced.
Smart Images

Figure CN114627880B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technologies, and in particular, to a voiceprint recognition method, apparatus, and electronic device. Background Art
[0002] Voiceprint recognition, also known as speaker recognition, can be used for user identity confirmation or speaker identification.
[0003] Voiceprint recognition needs to use a voiceprint recognition model to extract the voiceprint features of a speech signal. Based on this, a voiceprint recognition model needs to be deployed in an electronic device that needs to perform voiceprint recognition. Voiceprint recognition is divided into two stages: voiceprint registration and voiceprint verification. Generally, a user only needs to register voiceprint features once, while voiceprint verification may be performed frequently at different times. Frequent voiceprint verification means that the voiceprint features of different speech signals need to be frequently extracted using the voiceprint recognition model, which will inevitably affect the operation of the electronic device and naturally requires high device resources for the electronic device. Summary of the Invention
[0004] This application provides a voiceprint recognition method, apparatus, and electronic device.
[0005] Among them, a voiceprint recognition method includes:
[0006] Obtain a speech signal to be recognized;
[0007] Based on a first voiceprint recognition model, determine the first voiceprint feature of the speech signal;
[0008] Based on the first voiceprint feature and the second voiceprint feature stored for user information verification, determine a voiceprint recognition result, where the second voiceprint feature is a voiceprint feature extracted from a reference speech signal based on a second voiceprint recognition model;
[0009] Among them, the recognition performance of the second voiceprint recognition model is higher than that of the first voiceprint recognition model.
[0010] In a possible implementation manner, that the recognition performance of the second voiceprint recognition model is higher than that of the first voiceprint recognition model includes:
[0011] The difference degree between the voiceprint features of different users extracted based on the second voiceprint recognition model is greater than the difference degree between the voiceprint features of different users extracted based on the first voiceprint recognition model.
[0012] In another possible implementation manner, the processing duration required for the first voiceprint recognition model to recognize the voiceprint feature of a speech signal is less than the processing duration required for the second voiceprint recognition model to recognize the voiceprint feature of the speech signal.
[0013] In yet another possible implementation, the first voiceprint recognition model and the second voiceprint recognition model are synchronously trained using a plurality of voice samples, and are trained with the similarity of the voiceprint features recognized by the first voiceprint model and the second voiceprint model for the voice samples meeting the requirements as the training objective.
[0014] In yet another possible implementation, the first voiceprint recognition model and the second voiceprint recognition model are trained as follows:
[0015] Obtain a plurality of voice samples;
[0016] Based on the first voiceprint recognition model to be trained, determine the first sample voiceprint feature of the voice sample;
[0017] Based on the second voiceprint recognition model to be trained, determine the second sample voiceprint feature of the voice sample;
[0018] Determine the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample;
[0019] If the similarity between the first sample voiceprint feature and the second sample voiceprint feature meets the requirements, determine that the current training end condition is satisfied to obtain the trained first voiceprint recognition model and the second voiceprint recognition model; otherwise, adjust the internal parameters of the first voiceprint recognition model and the second voiceprint recognition model, and return to perform the operation of determining the first sample voiceprint feature and the second sample voiceprint feature of the voice sample.
[0020] In yet another possible implementation, the voice samples are labeled with user information of the actual attribution;
[0021] After determining the first sample voiceprint feature and the second sample voiceprint feature, it further includes:
[0022] Based on the first classification model and the first sample voiceprint feature of the voice sample, determine the first predicted user information to which the voice sample belongs;
[0023] Based on the second classification model and the second sample voiceprint feature of the voice sample, determine the second predicted user information to which the voice sample belongs;
[0024] Based on the user information of the actual attribution of the voice sample and the first predicted user information, determine the first predicted loss function value corresponding to the first voiceprint recognition model;
[0025] Based on the user information of the actual attribution of the voice sample and the second predicted user information, determine the second predicted loss function value corresponding to the second voiceprint recognition model;
[0026] After determining the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample, the following steps are further included:
[0027] Based on the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample, determine the similarity loss function value of the voiceprint features between the first voiceprint recognition model and the second voiceprint recognition model;
[0028] If the similarity between the first sample voiceprint feature and the second sample voiceprint feature meets the requirements, it is determined that the current training end condition is satisfied, including:
[0029] Based on the first prediction loss function value, the second prediction loss function value, and the similarity loss function value, determine the comprehensive loss function value;
[0030] If the comprehensive loss function value converges, it is determined that the current training end condition is satisfied.
[0031] In another possible implementation, the determining the comprehensive loss function value based on the first prediction loss function value, the second prediction loss function value, and the similarity loss function value includes:
[0032] Based on the first weight corresponding to the first prediction loss function value, the second weight corresponding to the second prediction loss function value, and the third weight corresponding to the similarity loss function value, perform weighted summation on the first loss function value, the second loss function value, and the similarity loss function value to obtain the comprehensive loss function value.
[0033] Wherein, a voiceprint recognition device includes:
[0034] A signal acquisition unit, configured to acquire a voice signal to be recognized;
[0035] A voiceprint extraction unit, configured to determine the first voiceprint feature of the voice signal based on the first voiceprint recognition model;
[0036] A voiceprint recognition unit, configured to determine a voiceprint recognition result based on the first voiceprint feature and the second voiceprint feature stored for user information verification, where the second voiceprint feature is a voiceprint feature extracted from a reference voice signal based on the second voiceprint recognition model; wherein, the recognition performance of the second voiceprint recognition model is higher than that of the first voiceprint recognition model.
[0037] In another possible implementation, the difference degree between the voiceprint features of different users extracted based on the second voiceprint recognition model is greater than the difference degree between the voiceprint features of different users extracted based on the first voiceprint recognition model.
[0038] Wherein, an electronic device includes at least a memory and a processor;
[0039] Wherein, the processor is configured to execute the voiceprint recognition method described in any one of the above;
[0040] The memory is configured to store programs required for the processor to perform operations.
[0041] As can be seen from the above, in the embodiments of the present application, considering that in order to ensure the accuracy of voiceprint verification, it is necessary to accurately extract the registered second voiceprint features. Therefore, there are relatively high requirements for the voiceprint recognition performance of the second recognition model used to extract the second voiceprint features. Relatively speaking, the voiceprint recognition performance of the first voiceprint recognition model used in the verification stage will be relatively low. Therefore, the present application selects the first voiceprint recognition model with a lower recognition performance than the second voiceprint recognition model as the voiceprint recognition model in the voiceprint verification stage, which can reduce the requirements for the device resources of the electronic device while taking into account the voiceprint verification requirements.
[0042] Moreover, the present application allows the first voiceprint recognition model deployed in the electronic device to be different from the second voiceprint recognition model used to extract the registered voiceprint features, so that the device resources of the electronic device can be considered, and the voiceprint recognition models used in the two stages of voiceprint registration and voiceprint verification can be more reasonably deployed, improving the flexibility of the deployment of the voiceprint recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0044] Figure 1 Shows a schematic flowchart of a voiceprint recognition method provided by an embodiment of the present application;
[0045] Figure 2 Shows a comparative schematic diagram of the voiceprint feature distributions of different users recognized by a low-performance voiceprint recognition model;
[0046] Figure 3 Shows a comparative schematic diagram of the voiceprint feature distributions of different users recognized by a high-performance voiceprint recognition model and a low-performance voiceprint recognition model;
[0047] Figure 4 Shows a schematic flowchart of a voiceprint recognition model training method provided by an embodiment of the present application;
[0048] Figure 5 Shows a schematic diagram of the training principle for training the first voiceprint recognition model and the second voiceprint recognition model provided by an embodiment of the present application;
[0049] Figure 6 Another schematic flowchart of the voiceprint recognition model training method provided by the embodiment of the present application is shown;
[0050] Figure 7 A schematic structural diagram of the composition of the voiceprint recognition device provided by the embodiment of the present application is shown;
[0051] Figure 8 A schematic structural diagram of the composition of the electronic device provided by the embodiment of the present application is shown. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0053] As Figure 1 , a schematic flowchart of the voiceprint recognition method provided by the embodiment of the present application is shown. The method of this embodiment can be applied to an electronic device. The method of this embodiment may include:
[0054] S101, obtain a voice signal to be recognized.
[0055] Wherein, the voice signal is an audio signal collected by the electronic device or transmitted to the electronic device by other electronic devices and requiring voiceprint recognition.
[0056] S102, based on the first voiceprint recognition model, determine the first voiceprint feature of the voice signal.
[0057] Wherein, the first voiceprint recognition model is a voiceprint recognition model configured in the electronic device for extracting voiceprint features from the voice signal to be recognized. The first voiceprint recognition model is also the so-called voiceprint recognition model for voiceprint verification.
[0058] It can be understood that, for the sake of convenience of distinction, the voiceprint feature extracted from the voice signal to be recognized is called the first voiceprint feature.
[0059] S103, based on the first voiceprint feature and the second voiceprint feature stored for user information verification, determine the voiceprint recognition result.
[0060] Wherein, the second voiceprint feature is a voiceprint feature extracted from the reference voice signal based on the second voiceprint recognition model.
[0061] Among them, the reference voice signal is a voice signal used to represent user information obtained in advance before voiceprint verification, that is, the so-called registered user voice signal. Correspondingly, the second voiceprint feature extracted from the reference voice signal is the voiceprint feature extracted in advance for identity verification, that is, the so-called registered voiceprint feature.
[0062] It can be understood that voiceprint recognition involves two stages: voiceprint registration and voiceprint verification. In the voiceprint registration stage, it is necessary to use the voiceprint recognition model to extract the voiceprint feature of the voice signal input by the target user (also called the registered user), and store the extracted voiceprint feature as the registered voiceprint feature. In the voiceprint verification stage, it is necessary to use the voiceprint recognition model to extract the voiceprint feature of the voice signal input by the speaker, and compare the extracted voiceprint feature with the registered voiceprint feature to determine whether the speaker belongs to the target user.
[0063] For example, assume that in the voiceprint registration stage, the voiceprint feature A of the voice signal of user A is extracted by the second voiceprint recognition model. Then, the voiceprint feature A will be stored as the registered voiceprint feature in the electronic device. On this basis, if the electronic device obtains the voice signal B and needs to verify whether the voice signal B belongs to the voice signal of user A, the voiceprint feature of the voice signal B can be extracted by using the first voiceprint recognition model. If the voiceprint feature of the voice signal B matches the registered voiceprint feature A, it can be confirmed that the voice signal B belongs to the voice signal of user A.
[0064] In this application, the voiceprint recognition result can represent whether the first voiceprint feature is consistent with the second voiceprint feature. Whether the first voiceprint feature is consistent with the second voiceprint recognition feature can represent whether the voice signal corresponding to the first voiceprint feature belongs to the voice signal of the user corresponding to the second voiceprint feature. Of course, the voiceprint recognition result can also directly represent whether the first voiceprint feature belongs to the voice signal of the user corresponding to the second voiceprint feature.
[0065] It can be understood that in actual applications, multiple second voiceprint features corresponding to different user information can also be stored in the electronic device. However, for the case where multiple second voiceprint features are stored in the electronic device, the first voiceprint feature can be compared with each second voiceprint feature in turn.
[0066] On this basis, the voiceprint recognition result can represent whether the voice signal corresponding to the first voiceprint feature belongs to the user corresponding to a certain second voiceprint feature; or, when the first voiceprint feature is consistent with a certain second voiceprint feature, the voiceprint recognition result represents which user information corresponding to the second voiceprint feature the voice signal corresponding to the first voiceprint feature specifically belongs to.
[0067] In this application, the first voiceprint recognition model is different from the second voiceprint recognition model. Specifically, the recognition performance of the second voiceprint recognition model is higher than that of the first voiceprint recognition model.
[0068] Through research, it is found that: the second voiceprint recognition model is a voiceprint recognition model for registration. Through the second voiceprint recognition model, the second voiceprint feature for registration needs to be extracted, and the second voiceprint feature for registration needs to be able to accurately reflect the voiceprint feature of the voice signal sent by the registered user. Otherwise, it may affect the accuracy of voiceprint verification. Based on this, there are relatively high requirements for the recognition performance of the second voiceprint recognition model for extracting the second voiceprint feature for registration.
[0069] Moreover, generally, a registered user only needs to register the voiceprint feature once, without the need to register the voiceprint feature frequently. Therefore, the number of times of extracting the second voiceprint feature for registration based on the second voiceprint recognition model is relatively small. Even if a second voiceprint recognition model with relatively high performance is deployed, it will not have too much impact on the operation of the electronic device, and there is no need for relatively high requirements for the device resources such as the hardware of the electronic device.
[0070] In the voiceprint verification stage, compared with extracting the second voiceprint feature for registration by the second voiceprint recognition model, the accuracy requirement for extracting the voiceprint feature by the first voiceprint recognition model is relatively low. Therefore, even if the recognition performance of the first voiceprint recognition model set is slightly lower than that of the second voiceprint recognition model, it will not have a great impact on voiceprint verification.
[0071] Moreover, compared with voiceprint registration, the number of times of using the first voiceprint recognition model to extract the voiceprint feature to be verified is relatively high, and the impact on the operation of the electronic device is relatively high. Therefore, there are also relatively high requirements for the device resources such as the hardware of the electronic device.
[0072] Through research on the existing voiceprint registration and voiceprint verification processes, it is found that: the voiceprint recognition models adopted in the current voiceprint registration stage and voiceprint verification stage belong to a symmetric structure. That is, the voiceprint verification model adopted in the voiceprint registration stage and the voiceprint recognition model adopted in the voiceprint verification stage are the same one, or the same voiceprint recognition model. In order to ensure relatively high voiceprint recognition accuracy in the voiceprint registration stage, it is necessary that the voiceprint recognition models adopted in both the voiceprint registration stage and the voiceprint verification stage have relatively high recognition performance. Frequent use of a voiceprint recognition model with relatively high recognition performance to extract the voiceprint of the voice signal to be verified will inevitably have a relatively high impact on the operation of the electronic device, and thus put forward relatively high requirements for the device resources such as the hardware of the electronic device.
[0073] Based on the above research findings, in this application, the first voiceprint recognition model used in the verification stage is set to be different from the second voiceprint recognition model used in the registration stage. By appropriately reducing the recognition performance of the first voiceprint recognition model, it is possible to reduce the hardware resource requirements of voiceprint verification for the electronic device while taking into account the reliability of voiceprint verification.
[0074] It can be understood that generally, the power consumption required by a voiceprint recognition model with higher recognition performance is also generally higher, although there are some special cases. Based on this, in order to more reliably reduce the power consumption of the electronic device, in this application, the power required for the first voiceprint recognition model to recognize a voiceprint is lower than that required for the second voiceprint recognition model to recognize a voiceprint.
[0075] It can be understood that in actual applications, the voiceprint registration stage and the voiceprint verification stage can be executed on the same electronic device or on different electronic devices. Correspondingly, the electronic device can only deploy the first voiceprint recognition model and deploy the second voiceprint recognition model on the server side; or, the first voiceprint recognition model and the second voiceprint recognition model can also be deployed on the electronic device at the same time.
[0076] For example, for an electronic device such as a smart device in the Internet of Things or some intelligent voice terminals, the second voiceprint recognition model in the registration stage can be deployed on the server side. After the electronic device obtains the reference voice signal for user registration, it can send the reference voice signal to the server. The server side extracts the second voiceprint feature based on the second voiceprint recognition model and sends it to the electronic device for storage. The first voiceprint recognition model can be deployed on the electronic device side. In this way, after the electronic device obtains the voice signal to be recognized, it can compare the voiceprint feature extracted from the voice signal by the first voiceprint recognition model with the locally stored voiceprint feature.
[0077] Another example is that the first voiceprint recognition model and the second voiceprint recognition model can also be deployed on a smart voice terminal or other types of electronic devices at the same time. On this basis, in the voiceprint registration stage, the electronic device obtains the reference voice signal input by the user for registration, and the electronic device extracts the second voiceprint feature of the reference voice signal based on the second voiceprint recognition model and stores it. For the voiceprint verification stage, the electronic device can extract the first voiceprint feature of the voice signal to be recognized based on the first voiceprint recognition model.
[0078] As can be seen from the above, in the embodiments of the present application, considering that in order to ensure the accuracy of voiceprint verification, it is necessary to accurately extract the registered second voiceprint feature. Therefore, there are relatively high requirements for the voiceprint recognition performance of the second recognition model used to extract the second voiceprint feature. Relatively speaking, the voiceprint recognition performance of the first voiceprint recognition model used in the verification stage will be relatively low. Therefore, the present application selects the first voiceprint recognition model with a recognition performance lower than that of the second voiceprint recognition model as the voiceprint recognition model in the voiceprint verification stage, which can reduce the requirements for the device resources of the electronic device while taking into account the voiceprint verification requirements.
[0079] Moreover, the present application allows the first voiceprint recognition model deployed in the electronic device to be different from the second voiceprint recognition model used to extract the registered voiceprint feature, so that the device resources of the electronic device can be considered, and the voiceprint recognition models used in the two stages of voiceprint registration and voiceprint verification can be more reasonably deployed, improving the flexibility of the deployment of the voiceprint recognition model.
[0080] It can be understood that the recognition performance of the first voiceprint recognition model being lower than that of the second voiceprint recognition model in the present application can be reflected from multiple dimensions.
[0081] In a possible situation, the recognition performance can be reflected in the accuracy of recognizing voiceprint features. In this possible situation, the difference degree between the voiceprint features of different users extracted based on the second voiceprint recognition model is greater than the difference degree between the voiceprint features of different users extracted based on the first voiceprint recognition model.
[0082] It can be understood that the greater the difference degree between the voiceprint features of different users extracted by the voiceprint recognition model, the greater the distinguishability of the voiceprint features of different users extracted based on this voiceprint recognition model, and the easier it is to distinguish the voices of different users. Based on this, it can be known that if the difference degree between the voiceprint features of different users extracted by the second voiceprint recognition model is greater than the difference degree between the voiceprint features of different users extracted by the first voiceprint recognition model, it can be characterized that the accuracy of the second voiceprint recognition model in recognizing voiceprints is higher than that of the first voiceprint recognition model in recognizing voiceprints.
[0083] Among them, in order to make the difference degree of the voiceprint features of different users extracted by the second voiceprint recognition model relatively large, the number of layers in each structural block of the second voiceprint recognition model will be relatively large, and the model structure will be relatively more complex. Correspondingly, the number of layers in each structural block of the second voiceprint recognition model will be relatively small, and the model structure will be simpler.
[0084] For the sake of easy understanding, the first voiceprint model and the second voiceprint recognition model are configured as voiceprint recognition models with different recognition performances and can ensure the accuracy of voiceprint verification. The following is combined with the attached Figure 2 and the attached Figure 3A comparative description will be given.
[0085] For the sake of convenience of description, the voiceprint recognition model with higher recognition performance is called a high-performance voiceprint recognition model, while the voiceprint recognition model with relatively lower recognition performance is called a low-performance voiceprint recognition model.
[0086] In Figure 2 a comparative schematic diagram of the voiceprint feature distributions of different users recognized by using a low-performance voiceprint recognition model is shown.
[0087] In Figure 2 and Figure 3 the voiceprint feature distribution 201 is the distribution of the voiceprint features of user A extracted by using a low-performance voiceprint recognition model. The voiceprint feature distribution 202 is the distribution of the voiceprint features of user B extracted by using the same low-performance voiceprint recognition model.
[0088] In Figure 2 the minimum inter-class angle (also called the minimum inter-class distance) between the voiceprint feature distribution 201 and the voiceprint distribution feature 202 is a1.
[0089] This minimum inter-class angle represents the distinguishability of the voices of different users (speakers). The larger this minimum inter-class angle is, the better the distinguishability of the voices between different users is.
[0090] As Figure 2 shown in
[0091] In Figure 2 the maximum intra-class angle (also called the maximum intra-class distance) corresponding to user B is a2.
[0092] The maximum intra-class angle is the similarity of the voices of the same user (speaker). The smaller the maximum intra-class angle is, the higher the similarity of the voices between the same users is. Therefore, the smaller the maximum intra-class angle is, the easier it is to distinguish the voices of the same user.
[0093] In Figure 3 the voiceprint feature distribution 301 is the distribution of the voiceprint features of user B extracted by using a high-performance voiceprint recognition model.
[0094] Among them, in Figure 3 the minimum inter-class angle between the voiceprint distribution feature 201 and the voiceprint distribution feature 301 is a3, and the maximum intra-class angle corresponding to user B is a4.
[0095] Comparing Figure 2 and Figure 3 it can be seen that Figure 3 the minimum inter-class angle a3 in Figure 2The minimum inter-class angle a1. That is, after extracting the voiceprint features of user B using a high-performance voiceprint recognition model, the minimum inter-class angle becomes larger, and the distinguishability between the voiceprint features of different users becomes higher. Therefore, different users' voiceprint features can be better distinguished by using a high-performance voiceprint recognition model.
[0096] At the same time, by comparing Figure 2 and Figure 3 it can be seen that Figure 3 the maximum intra-class angle a4 in Figure 2 is smaller than the maximum intra-class angle a2 in
[0097] That is, after extracting the voiceprint features of user B using a high-performance voiceprint recognition model, the maximum intra-class angle becomes smaller, and the distribution of the voiceprint features of the same user is more compact.
[0097] Based on the above analysis, it can be known that if a high-performance voiceprint recognition model is used in the voiceprint registration stage, then the registered voiceprint features extracted by the high-performance recognition model are more compact and can more prominently and accurately represent the features of the voice signals of the registered users. At the same time, because the distinguishability between the voiceprint features of different users extracted by the high-performance recognition model is relatively high.
[0098] Based on this, even if a low-performance voiceprint recognition model is used in the voiceprint verification stage, if the voice signal to be verified does not belong to the voice signal of the registered user, then the voiceprint features extracted from the voice signal to be verified using the low-performance voiceprint features will also have a large difference from the registered voiceprint features, which can meet the requirement of distinguishing the voiceprint features of different users.
[0099] It can be understood that in most cases, the time required for a voiceprint recognition model with relatively low recognition performance to recognize a voiceprint is also relatively low.
[0100] Based on this, when selecting the first voiceprint recognition model in this application, a voiceprint recognition model with a relatively small processing duration required to recognize voiceprint features can also be selected as the first voiceprint recognition model. That is, the processing duration required for the first voiceprint recognition model to recognize the voiceprint features of a voice signal is less than the processing duration required for the second voiceprint recognition model to recognize the voiceprint features of the voice signal.
[0101] It can be understood that in the voiceprint registration stage, the user may only need to perform voiceprint registration once. Therefore, even if the time taken for the second voiceprint recognition model used for voiceprint registration to recognize voiceprint features is relatively long, it will not have a great impact on the electronic device and will not affect the user experience.
[0102] However, an electronic device may perform voiceprint verification multiple times at different times. Therefore, selecting a first voiceprint recognition model with a relatively short processing duration required for identifying a voiceprint as the voiceprint recognition model in the voiceprint verification stage can effectively reduce the time consumed by the electronic device for performing voiceprint recognition on a voice signal to be recognized, improve the voiceprint recognition efficiency, enhance the user experience, and also reduce the impact on the electronic device.
[0103] In summary, the present application can set the voiceprint recognition models adopted in the voiceprint registration stage and the voiceprint verification stage to different voiceprint recognition models. Thus, considering the different requirements of voiceprint registration and voiceprint verification, the first voiceprint recognition model in the voiceprint verification stage can be set to a voiceprint recognition model with relatively weak computing power, relatively low recognition accuracy, but relatively short voiceprint processing duration, so as to reduce the resource consumption and performance requirements of the first voiceprint recognition model for the electronic device.
[0104] Considering that the number of times of voiceprint recognition in the voiceprint registration stage is relatively small and the voiceprint recognition in the voiceprint registration stage can be performed on the server side, therefore, different from the first voiceprint recognition model in the voiceprint verification stage, the second voiceprint recognition model in the voiceprint registration stage is set to a voiceprint recognition model with relatively strong computing power, relatively high recognition accuracy, but relatively long voiceprint processing duration.
[0105] It can be understood that since the first voiceprint recognition model and the second voiceprint recognition model are not the same voiceprint recognition model, in order to make the voiceprint features extracted based on the first voiceprint recognition model and the second voiceprint recognition model comparable, in the model training stage of the present application, the first voiceprint recognition model and the second voiceprint recognition model can be synchronously trained.
[0106] In a possible implementation manner, the first voiceprint recognition model and the second voiceprint recognition model are synchronously trained using multiple voice samples, and are trained with the similarity of the voiceprint features recognized by the first voiceprint model and the second voiceprint model for the voice samples meeting the requirements as the training objective.
[0107] Among them, the so-called synchronous training means that the first voiceprint recognition model and the second voiceprint recognition model are simultaneously trained in the same model training process until the first voiceprint recognition model and the second voiceprint recognition model simultaneously meet the model training requirements.
[0108] It can be understood that during the synchronous training of the first voiceprint recognition model and the second voiceprint recognition model, the present application also needs to ensure that the similarity between the voiceprint features recognized by the first voiceprint recognition model and the second voiceprint recognition model for the same voice sample meets the requirements, so that the voiceprint features extracted by the trained first voiceprint recognition model and the second voiceprint recognition model from the voice signals of the same user are similar. Based on this, if the voice signal to be recognized belongs to the voice signal of the registered target user, then the first voiceprint feature of this voice signal extracted based on the first voiceprint recognition model will inevitably be similar to the second voiceprint feature of this target user stored, so as to ensure the accuracy of voiceprint verification on the premise that the first voiceprint recognition model and the second voiceprint recognition model are inconsistent.
[0109] It can be understood that there are various ways to synchronously train the first voiceprint recognition model and the second voiceprint recognition model in the present application. For ease of understanding, the following will be described in conjunction with a training method.
[0110] As Figure 4 shown, it shows a schematic flowchart of a method for training a voiceprint recognition model in the present application. The method of this embodiment may include:
[0111] S401, obtain a plurality of voice samples.
[0112] The voice samples are voice signals used as training samples.
[0113] S402, for each voice sample, based on the first voiceprint recognition model to be trained, determine the first sample voiceprint feature of this voice sample.
[0114] S403, for each voice sample, based on the second voiceprint recognition model to be trained, determine the second sample voiceprint feature of the voice sample.
[0115] Among them, the first voiceprint recognition model and the second voiceprint model can select network models with different internal structures, etc., and there is no limit to this.
[0116] It can be understood that in order to make the recognition performance of the second voiceprint recognition model relatively high, the internal structure of the second voiceprint recognition model is more complex than the internal model structure of the first voiceprint recognition model. For example, the number of hidden layers in the second voiceprint recognition model can be more than the number of hidden layers in the first voiceprint recognition model.
[0117] It can be understood that for the sake of easy distinction, the voiceprint feature extracted by the first voiceprint recognition model from the voice sample is called the first sample voiceprint feature, and the voiceprint feature extracted by the second voiceprint recognition model from the voice sample is called the second sample voiceprint feature.
[0118] It is understandable that the order of steps S402 and S403 is not limited to Figure 2 As shown, in actual applications, the order of these two steps can be interchanged or they can be executed synchronously.
[0119] S404. For each voice sample, determine the similarity between the first sample voiceprint feature and the second sample voiceprint feature of this voice sample.
[0120] Among them, there are various calculation methods for the similarity between the first sample voiceprint feature and the second sample voiceprint feature.
[0121] For example, the cosine similarity cosθ between the first sample voiceprint feature e and the second sample voiceprint feature v can be calculated e,v As the similarity, specifically, it can refer to the following formula (1):
[0122]
[0123] Of course, there may be other possibilities for the method of calculating the similarity, and there is no limitation on this.
[0124] S405. Detect whether the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample meets the requirements. If so, determine that the current training end condition is satisfied to obtain the trained first voiceprint recognition model and the second voiceprint recognition model; if not, adjust the internal parameters of the first voiceprint recognition model and the second voiceprint recognition model, and return to step S402 to re - execute the operation of determining the first sample voiceprint feature and the second sample voiceprint feature of the voice sample.
[0125] Among them, there are various possibilities for the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample to meet the requirements.
[0126] For example, it can be that the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the same voice sample is greater than a set threshold, or it can be that the average value of the similarities between the first sample voiceprint features and the second sample voiceprint features of multiple voice samples is greater than the set threshold, etc.
[0127] Another example is that a similarity loss function related to the relative degree can also be set. When this similarity loss function converges, it is determined that the similarity meets the requirements.
[0128] Of course, there may be other possibilities for the requirements that the similarity between the two sample voiceprint features of the same voice sample needs to meet, and there is no limitation on this.
[0129] It can be understood that when the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample meets the requirements, it indicates that for the same voice signal, the voiceprint features extracted by the first voiceprint recognition model and the second voiceprint recognition model are similar, making the voiceprint features extracted by the first voiceprint recognition model and the second voiceprint recognition model comparable.
[0130] It can be understood that in the actual process of training the first voiceprint recognition model and the second voiceprint recognition model, in addition to ensuring that the voiceprint features extracted by the first voiceprint recognition model and the second voiceprint recognition model for the same voice sample are similar, it is also necessary to ensure that both the first voiceprint recognition and the second voiceprint recognition models can accurately extract voiceprint features. Based on this, in the model training process, it is also necessary to comprehensively consider the accuracy of the first voiceprint recognition model and the second voiceprint recognition model in recognizing voiceprint features.
[0131] As Figure 5 shown, it shows a schematic diagram of the training principle for training the first voiceprint recognition model and the second voiceprint recognition model provided by an embodiment of the present application.
[0132] From Figure 5 it can be seen that for each voice sample, it needs to be input into the first voiceprint recognition model and the second voiceprint recognition model respectively.
[0133] On this basis, not only is it necessary to calculate the similarity loss function values of the voiceprint features extracted by the first voiceprint recognition model and the second voiceprint recognition model respectively based on the similarity loss function, but also it is necessary to calculate the loss function value of the classification result corresponding to the voiceprint features extracted by the first voiceprint recognition model, and the loss function value of the classification result corresponding to the voiceprint features extracted by the second voiceprint recognition model. Correspondingly, it is necessary to comprehensively determine whether the model training reaches the end condition by combining these three loss function values.
[0134] The following combines the flowchart to Figure 5 illustrate the model training principle shown. As Figure 6 shown, it shows another schematic flowchart of the training method of the voiceprint recognition model provided by an embodiment of the present application. The method of this embodiment may include:
[0135] S601, obtain a plurality of voice samples, and the voice samples are labeled with the user information of the actual attribution.
[0136] S602, for each voice sample, based on the first voiceprint recognition model to be trained, determine the first sample voiceprint feature of the voice sample.
[0137] S603, for each voice sample, based on the second voiceprint recognition model to be trained, determine the second sample voiceprint feature of the voice sample.
[0138] S604. For each voice sample, determine the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample.
[0139] S605. Based on the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample, determine the similarity loss function value of the voiceprint features between the first voiceprint recognition model and the second voiceprint recognition model.
[0140] Wherein, the similarity loss function value can be calculated based on a similarity loss function, and the similarity loss function can be set as needed.
[0141] For example, in a possible implementation, the purpose of the similarity loss function can be to make the first sample voiceprint feature and the second sample voiceprint feature of the same voice sample as similar as possible, while making the first sample voiceprint feature and the second sample voiceprint feature from different voice samples as different as possible. For example, the similarity loss function value L AP can be calculated through the similarity loss function shown in Formula 2 below:
[0142]
[0143] Wherein, B is the total number of voice samples. Both w and b are set constants, and the specific values can be set as needed.
[0144] cosθ ei,vi represents the similarity between the first sample voiceprint feature ei and the second sample voiceprint feature vi of voice sample i.
[0145] cosθ ei,vj represents the similarity between the first sample voiceprint feature ei of voice sample i and the second sample voiceprint feature vj of voice sample j. Both i and j are natural numbers from 1 to B.
[0146] S606. For each voice sample, based on the first classification model and the first sample voiceprint feature of the voice sample, determine the first predicted user information to which the voice sample belongs.
[0147] Wherein, the first classification model is a classification model for determining the user information to which the first sample voiceprint feature belongs.
[0148] Wherein, each user information represents a user identity. For example, the user information can be a user name or other user identifiers, etc.
[0149] In a possible implementation, the first sample voiceprint feature of the voice sample can be input into the first classification model to obtain the probabilities that the first sample voiceprint feature predicted by the first classification model belongs to multiple different user information. On this basis, the user information with the highest corresponding probability can be determined as the predicted user information to which the first sample voiceprint feature belongs.
[0150] Among them, for the sake of easy distinction, the user information to which the first sample voiceprint feature of the voice sample predicted based on the first classification model belongs is called the first predicted user information, and the user information to which the second sample voiceprint feature of the voice sample predicted based on the second classification model belongs later is called the second predicted user information.
[0151] S607. Based on the user information to which each voice sample actually belongs and the first predicted user information, determine the first predicted loss function value corresponding to the first voiceprint recognition model.
[0152] The first predicted loss function value can be calculated based on the predicted loss function of the first voiceprint recognition model.
[0153] Among them, the first predicted loss function used to train the first voiceprint recognition model can be set as needed. Generally, the set predicted loss function aims to make the first predicted user information to which the predicted voice sample belongs the same as the actually labeled user information, and there is no restriction on the specific form of the first predicted loss function.
[0154] S608. For each voice sample, based on the second classification model and the second sample voiceprint feature of the voice sample, determine the second predicted user information to which the voice sample belongs.
[0155] Among them, the process of determining the second predicted user information is similar to the process of determining the first predicted user information before, and will not be elaborated here.
[0156] S609. Based on the user information to which each voice sample actually belongs and the second predicted user information, determine the second predicted loss function value corresponding to the second voiceprint recognition model.
[0157] Among them, the second predicted loss function value can be calculated based on the second predicted loss function set for training the second voiceprint recognition model. The second predicted loss function can also be set as needed, and there is no restriction on this.
[0158] S610. Based on the first predicted loss function value, the second predicted loss function value and the similarity loss function value, determine the comprehensive loss function value.
[0159] Among them, there can be multiple possible ways to determine the comprehensive loss function value, and there is no restriction on this.
[0160] For example, in a possible implementation, the first loss function value, the second loss function value, and the similarity loss function value may be weighted and summed based on the first weight corresponding to the first prediction loss function value, the second weight corresponding to the second prediction loss function value, and the third weight corresponding to the similarity loss function value to obtain a comprehensive loss function value.
[0161] Among them, the first weight, the second weight, and the third weight can be set as needed. For example, the first weight and the second weight can be 1, while the third weight can be a value less than 1.
[0162] S611. Detect whether the comprehensive loss function value converges. If so, determine that the current training end condition is satisfied, and obtain the trained first voiceprint recognition model and the second voiceprint recognition model; if not, adjust the internal parameters of the first voiceprint recognition model and the second voiceprint recognition model, and return to step S602 to re - execute the operation of determining the first sample voiceprint feature and the second sample voiceprint feature of the voice sample.
[0163] It can be understood that the loss function value is obtained by combining the first prediction loss function value, the second prediction loss function value, and the similarity loss function value, so that the comprehensive loss function value comprehensively considers the voiceprint extraction accuracy of the first voiceprint recognition model, the voiceprint extraction accuracy of the second voiceprint recognition model, and the similarity of the voiceprint features extracted by the first voiceprint recognition model and the second voiceprint recognition model respectively in these three dimensions. Therefore, judging whether the model training end condition is reached based on the comprehensive loss function value can not only effectively ensure the accuracy of the first voiceprint recognition model and the second voiceprint recognition model in recognizing voiceprints, but also ensure the comparability of the voiceprint features extracted by these two voiceprint recognition models.
[0164] To facilitate a clear and intuitive understanding of the advantages of this application, the following will be described in combination with the evaluation results of the combination of different voiceprint recognition models used in the voiceprint registration stage and the voiceprint verification stage.
[0165] As shown in the following table, it shows the combination of voiceprint recognition models that may be used in the voiceprint registration stage (abbreviated as the registration stage, such as registration in the following table) and the voiceprint verification stage (abbreviated as the verification stage, such as verification in the following table), as well as the equal error rate (Equal Error Rate, EER) and the minimum detection cost (Minimum Detection Cost Function, MinDCF) obtained by testing each combination on three groups of test data sets.
[0166]
[0167] In the above table, Model A and Model B represent two different voiceprint recognition models with different recognition performances. Among them, Model A is a high-performance voiceprint recognition model with high recognition performance, while Model B is a low-performance recognition model with relatively low recognition performance.
[0168] In the above superscript, the serial numbers 1 to 4 represent four model combination situations.
[0169] Specifically:
[0170] In the model combination situation corresponding to serial number 1, the voiceprint recognition models used in the registration stage and the verification stage are both Model A with high recognition performance. Model A can be obtained by training alone through multiple voice samples using the current conventional model training method.
[0171] In the model combination situation corresponding to serial number 2, the voiceprint recognition models used in the registration stage and the verification stage are both Model B with low recognition performance. Model B can be obtained by training alone through multiple voice samples using the current conventional model training method.
[0172] In the model combination situation corresponding to serial number 3, the low-performance Model B is used in both the registration stage and the verification stage. The Model B used in the situation of serial number 3 has exactly the same model structure as the Model B used in serial number 1. However, in the situation of serial number 3, it is necessary to use the training method mentioned in this application to synchronously train the Model B in the registration stage and the verification stage.
[0173] The model combination situation corresponding to serial number 4, which is the method protected by this application, means that different voiceprint recognition models are used in the registration stage and the verification stage. Specifically, the high-performance recognition model (Model A) with high recognition performance is used in the registration stage, while the low-performance recognition model (Model B) with low recognition performance is used in the verification stage. At the same time, Model A and Model B are synchronously trained based on multiple voice samples with the training objective that the similarity of the voiceprint features recognized by Model A and Model B for the voice samples meets the requirements.
[0174] It can be understood that the smaller the EER, the lower the error rate of voiceprint recognition. Similarly, the smaller the MinDCF, the better the performance of voiceprint recognition.
[0175] Combined with the test results of different test data sets in Table 1, it can be clearly seen that in the situation of serial number 1, high-performance voiceprint recognition models are used in both the voiceprint registration stage and the voiceprint verification stage, and the EER and MinDCF are the smallest, indicating the highest voiceprint recognition accuracy and the best recognition performance. However, due to the generally high power consumption and high latency of high-performance voiceprint recognition models, they have high requirements for the performance of electronic devices and are not conducive to voiceprint feature recognition on small electronic devices such as mobile phones.
[0176] If low-performance voiceprint recognition models are used in both the voiceprint registration stage and the voiceprint verification stage and they are trained using the existing model training method (the situation corresponding to Serial No. 2), although the power consumption and recognition time can be reduced, the values of EER and MinDCF are too high.
[0177] Relative to the situation of Serial No. 2, if low-performance voiceprint recognition models are used in both the voiceprint registration stage and the voiceprint verification stage, but the models used in the voiceprint registration stage and the voiceprint verification stage are synchronously trained using the solution of the present application (i.e., the situation of Serial No. 3), although the values of EER and MinDCF can be reduced, the performance improvement is not obvious.
[0178] By using the solution of the present application, that is, in the voiceprint registration stage mentioned in Serial No. 4, a high-performance voiceprint recognition model is used, and in the voiceprint verification stage, a high-performance voiceprint recognition model is used, and these two models are synchronously trained, then the values of EER and MinDCF can be significantly reduced. For example, relative to the situation of Serial No. 2, in the test based on the test data set, EER drops from 3.07% to 2.31%, and the relative performance improvement is 25%.
[0179] Moreover, by using the solution of the present application, the low-performance voiceprint recognition model used in the voiceprint verification stage can also help reduce the power consumption and recognition time required by the electronic device during the voiceprint verification and recognition process.
[0180] Corresponding to the voiceprint recognition method provided by the embodiment of the present application, the present application also provides a voiceprint recognition device.
[0181] As Figure 7 shown, it shows a schematic structural diagram of a composition of a voiceprint recognition device of the present application. The device in this embodiment may include:
[0182] A signal acquisition unit 701, configured to acquire a voice signal to be recognized;
[0183] A voiceprint extraction unit 702, configured to determine a first voiceprint feature of the voice signal based on a first voiceprint recognition model;
[0184] A voiceprint recognition unit 703, configured to determine a voiceprint recognition result based on the first voiceprint feature and a second voiceprint feature stored for user information verification, where the second voiceprint feature is a voiceprint feature extracted from a reference voice signal based on a second voiceprint recognition model; wherein, the recognition performance of the second voiceprint recognition model is higher than the recognition performance of the first voiceprint recognition model.
[0185] In a possible implementation manner, the difference degree between the voiceprint features of different users extracted based on the second voiceprint recognition model is greater than the difference degree between the voiceprint features of different users extracted based on the first voiceprint recognition model.
[0186] In yet another possible implementation, the processing duration required for the first voiceprint recognition model to recognize the voiceprint features of a voice signal is less than the processing duration required for the second voiceprint recognition model to recognize the voiceprint features of the voice signal.
[0187] In yet another possible implementation, the first voiceprint recognition model and the second voiceprint recognition model are synchronously trained using a plurality of voice samples, and are trained with the similarity of the voiceprint features recognized by the first voiceprint model and the second voiceprint model for the voice samples meeting the requirements as the training objective.
[0188] In yet another possible implementation, the apparatus further includes:
[0189] A sample acquisition unit, configured to acquire a plurality of voice samples before the signal acquisition unit acquires a voice signal;
[0190] A first feature determination unit, configured to determine first sample voiceprint features of the voice samples based on the first voiceprint recognition model to be trained;
[0191] A second feature determination unit, configured to determine second sample voiceprint features of the voice samples based on the second voiceprint recognition model to be trained;
[0192] A similarity determination unit, configured to determine the similarity between the first sample voiceprint features and the second sample voiceprint features of the voice samples;
[0193] A training control unit, configured to determine that the current training end condition is met if the similarity between the first sample voiceprint features and the second sample voiceprint features meets the requirements, so as to obtain the trained first voiceprint recognition model and the second voiceprint recognition model; otherwise, adjust the internal parameters of the first voiceprint recognition model and the second voiceprint recognition model, and return to execute the operations of the first feature determination unit and the second feature determination unit.
[0194] In yet another possible implementation, the voice samples acquired by the sample acquisition unit are labeled with user information of the actual attribution;
[0195] The apparatus further includes:
[0196] A first user determination unit, configured to determine first predicted user information to which the voice samples belong based on the first classification model and the first sample voiceprint features of the voice samples after the first feature determination unit determines the first sample voiceprint features;
[0197] A second user determination unit, configured to determine second predicted user information to which the voice sample belongs based on a second classification model and the second sample voiceprint feature of the voice sample after the second feature determination unit determines the second sample voiceprint feature;
[0198] A first loss determination unit, configured to determine a first predicted loss function value corresponding to the first voiceprint recognition model based on the user information to which the voice sample actually belongs and the first predicted user information;
[0199] A second loss determination unit, configured to determine a second predicted loss function value corresponding to the second voiceprint recognition model based on the user information to which the voice sample actually belongs and the second predicted user information;
[0200] A similarity loss determination unit, configured to determine a similarity loss function value of voiceprint features between the first voiceprint recognition model and the second voiceprint recognition model based on the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample after the similarity determination unit determines the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice sample;
[0201] The training control unit includes:
[0202] An integrated determination subunit, configured to determine an integrated loss function value based on the first predicted loss function value, the second predicted loss function value, and the similarity loss function value;
[0203] A training control subunit, configured to determine that the current satisfies the training end condition if the integrated loss function value converges.
[0204] In a possible implementation manner, the integrated determination subunit is specifically configured to perform a weighted sum of the first loss function value, the second loss function value, and the similarity loss function value based on a first weight corresponding to the first predicted loss function value, a second weight corresponding to the second predicted loss function value, and a third weight corresponding to the similarity loss function value to obtain an integrated loss function value.
[0205] In another aspect, the present application further provides an electronic device, as Figure 8 shown, which shows a schematic composition structure of the electronic device. The electronic device can be any type of electronic device, and the electronic device at least includes a processor 801 and a memory 802;
[0206] Wherein, the processor 801 is configured to execute the voiceprint recognition method in any one of the above embodiments.
[0207] The memory 802 is configured to store a program required for the processor to perform operations.
[0208] It can be understood that the electronic device may further include a display unit 803 and an input unit 804.
[0209] Of course, the electronic device may also have Figure 8 more or fewer components, which is not limited herein.
[0210] On the other hand, the present application also provides a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, which is loaded and executed by a processor to implement the voiceprint recognition method described in any one of the above embodiments.
[0211] The present application also proposes a computer program including computer instructions stored in a computer-readable storage medium. When the computer program runs on an electronic device, it is used to execute the voiceprint recognition method in any one of the above embodiments.
[0212] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. At the same time, the features described in the embodiments of this specification can be replaced or combined with each other, enabling those skilled in the art to implement or use the present application. For the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial descriptions of the method embodiments.
[0213] Finally, it should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0214] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0215] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A voiceprint recognition method, comprising: Obtaining a voice signal to be recognized; Determining a first voiceprint feature of the voice signal based on a first voiceprint recognition model; Determining a voiceprint recognition result based on the first voiceprint feature and a stored second voiceprint feature for user information verification, where the second voiceprint feature is a voiceprint feature extracted from a reference voice signal based on a second voiceprint recognition model; Wherein, the difference degree between the voiceprint features of different users extracted based on the second voiceprint recognition model is greater than the difference degree between the voiceprint features of different users extracted based on the first voiceprint recognition model.
2. The method according to claim 1, wherein the processing duration required for the first voiceprint recognition model to recognize the voiceprint feature of a voice signal is less than the processing duration required for the second voiceprint recognition model to recognize the voiceprint feature of the voice signal.
3. The method according to claim 1, wherein the first voiceprint recognition model and the second voiceprint recognition model are synchronously trained using a plurality of voice samples, and are trained with the similarity of the voiceprint features recognized by the first voiceprint model and the second voiceprint model for the voice samples meeting the requirements as the training objective.
4. The method according to claim 3, wherein the first voiceprint recognition model and the second voiceprint recognition model are trained through the following method: Obtaining a plurality of voice samples; Determining a first sample voiceprint feature of the voice samples based on a first voiceprint recognition model to be trained; Determining a second sample voiceprint feature of the voice samples based on a second voiceprint recognition model to be trained; Determining the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice samples; If the similarity between the first sample voiceprint feature and the second sample voiceprint feature meets the requirements, determining that the current training end condition is satisfied to obtain the trained first voiceprint recognition model and second voiceprint recognition model; otherwise, adjusting the internal parameters of the first voiceprint recognition model and the second voiceprint recognition model, and returning to execute the operation of determining the first sample voiceprint feature and the second sample voiceprint feature of the voice samples.
5. The method according to claim 4, wherein the voice samples are labeled with actual user information of attribution; After determining the first sample voiceprint feature and the second sample voiceprint feature, further comprising: Determining a first predicted user information to which the voice samples belong based on a first classification model and the first sample voiceprint feature of the voice samples; Determining a second predicted user information to which the voice samples belong based on a second classification model and the second sample voiceprint feature of the voice samples; Determining a first predicted loss function value corresponding to the first voiceprint recognition model based on the actual user information of attribution of the voice samples and the first predicted user information; Determining a second predicted loss function value corresponding to the second voiceprint recognition model based on the actual user information of attribution of the voice samples and the second predicted user information; After determining the similarity between the first sample voiceprint feature and the second sample voiceprint feature of the voice samples, further comprising: Determine the similarity loss function value of the voiceprint features between the first voiceprint recognition model and the second voiceprint recognition model based on the similarity between the first sample voiceprint features and the second sample voiceprint features of the voice samples; If the similarity between the first sample voiceprint features and the second sample voiceprint features meets the requirements, it is determined that the current training end condition is satisfied, including: Determine the comprehensive loss function value based on the first prediction loss function value, the second prediction loss function value, and the similarity loss function value; If the comprehensive loss function value converges, it is determined that the current training end condition is satisfied.
6. The method according to claim 5, wherein the determining the comprehensive loss function value based on the first prediction loss function value, the second prediction loss function value, and the similarity loss function value includes: Perform weighted summation on the first prediction loss function value, the second prediction loss function value, and the similarity loss function value based on the first weight corresponding to the first prediction loss function value, the second weight corresponding to the second prediction loss function value, and the third weight corresponding to the similarity loss function value to obtain the comprehensive loss function value.
7. A voiceprint recognition device, comprising: A signal acquisition unit, configured to acquire a voice signal to be recognized; A voiceprint extraction unit, configured to determine the first voiceprint feature of the voice signal based on the first voiceprint recognition model; A voiceprint recognition unit, configured to determine a voiceprint recognition result based on the first voiceprint feature and the second voiceprint feature stored for user information verification, where the second voiceprint feature is a voiceprint feature extracted from a reference voice signal based on the second voiceprint recognition model; wherein, the difference degree between the voiceprint features of different users extracted based on the second voiceprint recognition model is greater than the difference degree of the voiceprint features of different users extracted based on the first voiceprint recognition model.
8. An electronic device, at least comprising a memory and a processor; Among them, The processor is configured to execute the voiceprint recognition method according to any one of claims 1 to 6 above; The memory is configured to store a program required for the processor to perform operations.
Citation Information
Patent Citations
Voiceprint verification method, voiceprint recognition model training method, device and equipment
CN112802481A
Device awakening method and electronic equipment thereof
CN113870855A