Voiceprint Recognition Model Training Method, Device, Mobile Terminal and Storage Medium
By using multiple fully connected layers and loss calculation layers in the xvector voiceprint recognition model to train feature vectors, the problem of poor effect of xvector model in text semi-related voiceprint recognition is solved, and the accuracy of voiceprint recognition is improved.
Patent Information
- Application Number
- CN202010469636.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-05-28
AI Technical Summary
The existing xvector model has poor effect on text semi-related voiceprint recognition, resulting in a reduced accuracy of text semi-related voiceprint recognition.
By obtaining training data, inputting the xvector vocalprint recognition model for feature extraction, and inputting the extracted feature vectors into multiple fully connected layers for training. The loss calculation layer is used to calculate the loss probability, and the fully connected layer is trained according to the loss probability until the output converges.
The xvector voiceprint recognition model has improved the recognition effect of text semi-related information, so that the model can more accurately perform text-independent, text-related and text-semi-related voiceprint recognition.
Smart Images

Figure CN111783939B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voiceprint recognition, and particularly relates to a method, a device, a mobile terminal and a storage medium for training a voiceprint recognition model. Background Art
[0002] Everyone's voice contains unique biological characteristics. Voiceprint recognition refers to a technical means of identifying a speaker by using the speaker's voice. Like technologies such as fingerprint recognition, voiceprint recognition has high security and reliability and can be applied to all occasions where identity recognition is required, such as in financial fields such as criminal investigation, banking, securities, and insurance. Compared with traditional identity recognition technologies, the advantages of voiceprint recognition are that the voiceprint extraction process is simple, the cost is low, and it has uniqueness and is not easy to forge or counterfeit.
[0003] In the existing voiceprint recognition process, the xvector model has good results in voiceprint recognition. The application scenarios of voiceprints generally include text-independent, text-dependent (fixed passwords), and text semi-dependent (dynamic numbers). However, in the process of using the existing xvector model, the voiceprint recognition effect for text semi-dependent is poor, which reduces the accuracy of text semi-dependent voiceprint recognition. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method, a device, a mobile terminal and a storage medium for training a voiceprint recognition model, aiming to solve the problems of low audio detection efficiency and poor audio detection accuracy in the existing voiceprint recognition model training method.
[0005] The embodiments of the present invention are implemented as follows. A method for training a voiceprint recognition model includes:
[0006] Obtain training data and input the training data into an xvector voiceprint recognition model; wherein, the training data includes preset data and dynamic number data;
[0007] Extract features from the training data based on the xvector voiceprint recognition model to obtain training feature vectors, and input the training feature vectors into a first fully connected layer;
[0008] Perform type recognition on the training feature vectors through the first fully connected layer to obtain preset feature vectors and dynamic number feature vectors;
[0009] Input the preset feature vectors and the dynamic number feature vectors into a second fully connected layer and a third fully connected layer respectively, and both the second fully connected layer and the third fully connected layer correspond to one output;
[0010] Use a loss calculation layer to calculate the losses of the outputs of the second fully connected layer and the third fully connected layer respectively, obtaining a first loss probability and a second loss probability;
[0011] Train the second fully connected layer according to the first loss probability, and train the third fully connected layer according to the second loss probability until the outputs of the second fully connected layer and the third fully connected layer converge.
[0012] Furthermore, the step of extracting features from the training data based on the xvector speaker recognition model includes:
[0013] Input the training data into the TDNN network in the xvector speaker recognition model, and control the TDNN network to extract features from the training data, obtaining training features;
[0014] The TDNN network controls the TDNN network to perform a non-linear transformation on the training features to obtain the training feature vectors.
[0015] Furthermore, the step of using a loss calculation layer to calculate the losses of the outputs of the second fully connected layer and the third fully connected layer respectively, obtaining a first loss probability and a second loss probability includes:
[0016] Calculate the loss of the output of the second fully connected layer according to a preset loss function and the preset feature vector, obtaining a first loss probability;
[0017] Calculate the loss of the output of the third fully connected layer according to the preset loss function and the dynamic digital feature vector, obtaining a second loss probability.
[0018] Furthermore, the step of training the second fully connected layer according to the first loss probability and training the third fully connected layer according to the second loss probability includes:
[0019] Perform backpropagation in the xvector speaker recognition model according to the first loss probability, and perform backpropagation in the xvector speaker recognition model according to the second loss probability.
[0020] Furthermore, before the step of inputting the training feature vectors into the first fully connected layer, the method further includes:
[0021] Perform pooling processing on the training feature vectors output by each TDNN network, and input the pooled training feature vectors into the first fully connected layer.
[0022] Further, the step of performing pooling processing on the training feature vectors output by each of the TDNN networks includes:
[0023] Accumulate the training feature vectors output by each of the TDNN networks, calculate the mean and standard deviation of all the training feature vectors according to the vector accumulation result, and use the mean and the standard deviation as the output after the pooling processing of the training feature vectors.
[0024] Further, the method further includes:
[0025] Obtain the voiceprint data to be recognized, and input the voiceprint data to be recognized into the xvector voiceprint recognition model;
[0026] Control the xvector voiceprint recognition model to recognize the voiceprint data to be recognized, and use the output result of the first fully connected layer as the output vector of the xvector voiceprint recognition model;
[0027] Calculate the matching value between the output vector and the sample vector pre-stored locally according to the Euclidean distance formula, and obtain the number value of the sample vector corresponding to the maximum value in the matching value;
[0028] When it is determined that the number value is greater than the number threshold, it is determined that the voiceprint recognition of the voiceprint data to be recognized is qualified.
[0029] Another object of the embodiments of the present invention is to provide a voiceprint recognition model training device, and the device includes:
[0030] A training data acquisition module, configured to acquire training data and input the training data into an xvector voiceprint recognition model, where the training data includes preset data and dynamic digital data;
[0031] A feature extraction module, configured to extract features from the training data based on the xvector voiceprint recognition model to obtain training feature vectors, and input the training feature vectors into a first fully connected layer;
[0032] A feature type recognition module, configured to perform type recognition on the training feature vectors through the first fully connected layer to obtain a preset feature vector and a dynamic digital feature vector;
[0033] A feature output module, configured to input the preset feature vector and the dynamic digital feature vector into a second fully connected layer and a third fully connected layer respectively, and both the second fully connected layer and the third fully connected layer correspond to one output;
[0034] A loss calculation module, configured to use a loss calculation layer to separately calculate losses for the outputs of the second fully-connected layer and the third fully-connected layer, obtaining a first loss probability and a second loss probability;
[0035] A model training module, configured to train the second fully-connected layer according to the first loss probability, and train the third fully-connected layer according to the second loss probability until the outputs of the second fully-connected layer and the third fully-connected layer converge.
[0036] Another object of the embodiments of the present invention is to provide a mobile terminal, including a storage device and a processor, where the storage device is used to store a computer program, and the processor runs the computer program to enable the mobile terminal to execute the above-mentioned voiceprint recognition model training method.
[0037] Another object of the embodiments of the present invention is to provide a storage medium, which stores the computer program used in the above-mentioned mobile terminal, and when the computer program is executed by a processor, the steps of the above-mentioned voiceprint recognition model training method are implemented.
[0038] In the embodiments of the present invention, through the design of training the second fully-connected layer according to a preset feature vector and controlling the dynamic digital feature vector to train the third fully-connected layer, the recognition effect of the xvector voiceprint recognition model on text semi-correlation after model training is improved, so that the xvector voiceprint recognition model can improve effective voiceprint recognition for text-independent, text-related, and text semi-related, and improve the accuracy of voiceprint recognition. Description of the Drawings
[0039] Figure 1 is a flowchart of the voiceprint recognition model training method provided by the first embodiment of the present invention;
[0040] Figure 2 is a flowchart of the voiceprint recognition model training method provided by the second embodiment of the present invention;
[0041] Figure 3 is a schematic structural diagram of the voiceprint recognition model training device provided by the third embodiment of the present invention;
[0042] Figure 4 is a schematic structural diagram of the mobile terminal provided by the fourth embodiment of the present invention. Detailed Embodiments
[0043] In the following description, specific details such as specific system architectures and technologies are presented for purposes of illustration and not limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0044] It should be understood that when used in the specification and claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0045] It should also be understood that the term "and / or" used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0046] As used in the specification and claims of the present application, the term "if" may be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0047] In addition, in the description of the specification and claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0048] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0049] Example 1
[0050] Please refer to Figure 1, is the flowchart of the voiceprint recognition model training method provided by the first embodiment of the present invention, including the steps:
[0051] Step S10, obtain training data and input the training data into the xvector voiceprint recognition model;
[0052] Among them, the size and information parameters of the training data can be set according to user needs. The training data includes preset data and dynamic digital data. The preset data can be text data, voice data or digital data. For example, the training data includes 100 fixed texts and 100 dynamic numbers, and each fixed text and dynamic number can be randomly generated;
[0053] Preferably, the xvector voiceprint recognition model includes a TDNN network (Time-Delay Neural Network), a pooling layer and multiple fully connected layers. Specifically, the xvector voiceprint recognition model includes a first fully connected layer, a second fully connected layer and a third fully connected layer. Further, users can set the number of fully connected layers according to their own needs;
[0054] Step S20, extract features from the training data based on the xvector voiceprint recognition model to obtain training feature vectors, and input the training feature vectors into the first fully connected layer;
[0055] Among them, the training feature vector can be an MFCC feature vector. After the xvector voiceprint recognition model performs MFCC feature extraction on the training data, the extracted MFCC features are vector-converted to obtain the MFCC feature vector, and the MFCC feature vector is input into the first fully connected layer for convolution;
[0056] Step S30, perform type recognition on the training feature vectors through the first fully connected layer to obtain preset feature vectors and dynamic digital feature vectors;
[0057] Among them, by controlling the first fully connected layer to recognize the vector identifiers in the training feature vectors, and based on the recognition results, the preset feature vectors corresponding to the preset data and the dynamic digital feature vectors corresponding to the dynamic digital data are obtained respectively;
[0058] Preferably, in this step, the sample data in the training data can be set according to user needs, but at least two different types of sample data are included in the sample data to ensure that after the first fully connected layer performs type recognition on the training feature vectors, at least two different types of feature vectors are obtained;
[0059] Step S40, input the preset feature vectors and the dynamic digital feature vectors into the second fully connected layer and the third fully connected layer respectively;
[0060] Among them, both the second fully connected layer and the third fully connected layer correspond to one output. The second fully connected layer and the third fully connected layer share the TDNN network, pooling layer, and first fully connected layer in the xvector voiceprint recognition model;
[0061] Specifically, in this embodiment, by respectively setting a different output for the second fully connected layer and the third fully connected layer, the second fully connected layer and the third fully connected layer can respectively perform model training on different features, so that the first fully connected layer can identify different types of features subsequently, improving the diversity of voiceprint recognition of the first fully connected layer and the xvector voiceprint recognition model;
[0062] Step S50: Use the loss calculation layer to respectively calculate the losses of the outputs of the second fully connected layer and the third fully connected layer to obtain a first loss probability and a second loss probability;
[0063] Among them, through the design of using the loss calculation layer to respectively calculate the losses of the outputs of the second fully connected layer and the third fully connected layer, the posterior probabilities of the second fully connected layer and the third fully connected layer can be effectively calculated, that is, the probability value of the speaker can be calculated;
[0064] Step S60: Train the second fully connected layer according to the first loss probability and train the third fully connected layer according to the second loss probability until the outputs of the second fully connected layer and the third fully connected layer converge;
[0065] Preferably, when the second fully connected layer and the third fully connected layer reach a preset number of iterations, the model training of the xvector voiceprint recognition model is automatically stopped, so that the trained xvector voiceprint recognition model can effectively recognize voiceprint data that is text-independent, text-dependent, and text-semidependent;
[0066] In this embodiment, through the design of training the second fully connected layer according to a preset feature vector and controlling the dynamic digital feature vector to train the third fully connected layer, the recognition effect of the xvector voiceprint recognition model on text-semidependent after model training is improved, so that the xvector voiceprint recognition model can improve effective voiceprint recognition for text-independent, text-dependent, and text-semidependent, and improve the accuracy of voiceprint recognition.
[0067] Example 2
[0068] Please refer to Figure 2 , which is a flowchart of the voiceprint recognition model training method provided by the second embodiment of the present invention, including the steps:
[0069] Step S11, obtain training data and input the training data into the xvector speaker recognition model;
[0070] Among them, the training data includes preset data and dynamic digital data. Further, the training data includes at least two different sample data, and one of the sample data is dynamic digital data to ensure that the trained xvector speaker recognition model can effectively recognize speaker for text semi - related;
[0071] Step S21, input the front - end features of the training data into the TDNN network in the xvector speaker recognition model, and control the TDNN network to extract features from the front - end features to obtain training features;
[0072] Among them, the front - end features can be MFCC features. The TDNN network is used to express the relationship of speaker features in time. Preferably, two TDNN networks are used in the xvector speaker recognition model;
[0073] Step S31, control the TDNN network to perform non - linear transformation on the training features to obtain training feature vectors;
[0074] Among them, the training feature vectors can be MFCC feature vectors. After the TDNN network extracts MFCC features from the training data, non - linear transformation is performed on the extracted MFCC features to achieve the effect of feature vector conversion and obtain the MFCC feature vectors;
[0075] Specifically, in this step, the MFCC feature vectors can be obtained after pre - emphasis, framing, windowing, fast Fourier transform, band - pass filtering, logarithmic operation and discrete cosine transform processing on the training features;
[0076] Step S41, perform pooling processing on the training feature vectors output by each TDNN network, and input the pooled training feature vectors into the first fully - connected layer;
[0077] Among them, pooling processing (Pooling), also known as downsampling or subsampling, is mainly used for feature dimensionality reduction, compressing the amount of data and parameters to reduce overfitting and improve the fault tolerance of the model at the same time;
[0078] Specifically, in this step, the step of performing pooling processing on the training feature vectors output by each TDNN network includes:
[0079] Accumulate the training feature vectors output by each TDNN network, calculate the mean and standard deviation of all the training feature vectors according to the vector accumulation result, and use the mean and the standard deviation as the output after pooling processing of the training feature vectors.
[0080] Step S51, perform type recognition on the training feature vector through the first fully connected layer to obtain a preset feature vector and a dynamic digital feature vector;
[0081] Among them, by controlling the first fully connected layer to recognize the vector identifier in the training feature vector, and correspondingly obtaining the preset feature vector corresponding to the preset data and the dynamic digital feature vector corresponding to the dynamic digital data based on the recognition result;
[0082] Step S61, input the preset feature vector and the dynamic digital feature vector into the second fully connected layer and the third fully connected layer respectively;
[0083] Among them, both the second fully connected layer and the third fully connected layer correspond to one output, and the second fully connected layer and the third fully connected layer share the TDNN network, pooling layer and the first fully connected layer in the xvector voiceprint recognition model;
[0084] Step S71, calculate the loss of the output of the second fully connected layer according to the preset loss function and the preset feature vector to obtain the first loss probability, and calculate the loss of the output of the third fully connected layer according to the preset loss function and the dynamic digital feature vector to obtain the second loss probability;
[0085] Step S81, perform backpropagation in the xvector voiceprint recognition model according to the first loss probability, and perform backpropagation in the xvector voiceprint recognition model according to the second loss probability until the outputs of the second fully connected layer and the third fully connected layer converge;
[0086] Step S91, obtain the voiceprint data to be recognized and input the voiceprint data to be recognized into the xvector voiceprint recognition model;
[0087] Step S101, control the xvector voiceprint recognition model to recognize the voiceprint data to be recognized, and use the output result of the first fully connected layer as the output vector of the xvector voiceprint recognition model;
[0088] Step S111, calculate the matching value between the output vector and the sample vector pre-stored locally according to the Euclidean distance formula, and obtain the number value of the sample vector corresponding to the maximum value in the matching value;
[0089] Among them, the Euclidean distance formula used between the output vector and the sample vector is:
[0090] ;
[0091] Let \(a\) be the output vector and \(b\) be the sample vector. By using the Euclidean distance formula, the current eigenvalue (output vector) is compared with the existing eigenvalues in the voiceprint library (sample vector) for 1:N retrieval scoring to obtain the matching value;
[0092] Specifically, in this embodiment, a number table is pre-stored, and the corresponding relationship between different matching values and number values is stored in the number table. Therefore, by matching the maximum value in the matching value with the number table, the number value is queried;
[0093] Step S121, when it is determined that the number value is greater than the number threshold, it is determined that the voiceprint recognition of the voiceprint data to be recognized is qualified;
[0094] Among them, the size of the queried number value is compared with the number threshold to determine whether the voiceprint recognition of the voiceprint data to be recognized is qualified. Specifically, the number threshold can be set according to requirements. For example, the number threshold can be 0.8, 0.9, or 0.95, etc. The number threshold is used to determine whether the voiceprint features in the voiceprint data to be recognized are consistent with the sample voiceprint features pre-stored locally;
[0095] Further, in this embodiment, when it is determined that the voiceprint recognition of the voiceprint data to be recognized is qualified, the user identifier corresponding to the number value is obtained and the user identifier is output. Among them, the user identifier can be stored in the form of text, numbers, numbers, images, or biometric features. The user identifier is used to point to the corresponding user. For example, when the user identifier is stored in the form of text, the user identifier can be the user name, such as "Zhang San", "Li Si", etc.; when the user identifier is stored in the form of numbers, the user identifier can be the user work number, and when the user identifier is stored in the form of an image, the user identifier is the user's avatar picture;
[0096] In this embodiment, by training the second fully connected layer according to the preset feature vector and controlling the dynamic digital feature vector to train the third fully connected layer, the recognition effect of the xvector voiceprint recognition model on text semi-correlation after model training is improved, so that the xvector voiceprint recognition model can improve the effective voiceprint recognition for text-independent, text-related, and text semi-related, and improve the accuracy of voiceprint recognition.
[0097] Example 3
[0098] Please refer to Figure 3 , which is a schematic structural diagram of the voiceprint recognition model training device 100 provided in the third embodiment of the present invention, including: a training data acquisition module 10, a feature extraction module 11, a feature type recognition module 12, a feature output module 13, and a model training module 16, where:
[0099] The training data acquisition module 10 is used to acquire training data and input the training data into the xvector voiceprint recognition model. The training data includes preset data and dynamic digital data. Wherein, the size and information parameters of the training data can both be set according to user requirements. The training data includes preset data and dynamic digital data;
[0100] Preferably, the xvector voiceprint recognition model includes a TDNN network (Time-Delay Neural Network), a pooling layer, and multiple fully connected layers. Specifically, the xvector voiceprint recognition model includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. Further, the user can set the number of fully connected layers according to their own needs.
[0101] The feature extraction module 11 is used to extract features from the training data based on the xvector voiceprint recognition model to obtain training feature vectors, and input the training feature vectors into the first fully connected layer. Wherein, the training feature vectors can be MFCC feature vectors. After the xvector voiceprint recognition model performs MFCC feature extraction on the training data, it performs vector conversion on the extracted MFCC features to obtain the MFCC feature vectors, and inputs the MFCC feature vectors into the first fully connected layer for convolution.
[0102] Preferably, the feature extraction module 11 is further used to: input the training data into the TDNN network in the xvector voiceprint recognition model, and control the TDNN network to extract features from the training data to obtain training features;
[0103] Control the TDNN network to perform a non-linear transformation on the training features to obtain the training feature vectors.
[0104] The feature type recognition module 12 is used to perform type recognition on the training feature vectors through the first fully connected layer to obtain preset feature vectors and dynamic digital feature vectors. Wherein, by controlling the first fully connected layer to recognize the vector identifiers in the training feature vectors, and based on the recognition results, the preset feature vectors corresponding to the preset data and the dynamic digital feature vectors corresponding to the dynamic digital data are obtained respectively.
[0105] The feature output module 13 is used to input the preset feature vectors and the dynamic digital feature vectors into the second fully connected layer and the third fully connected layer respectively. Both the second fully connected layer and the third fully connected layer have one output. The second fully connected layer and the third fully connected layer share the TDNN network, the pooling layer, and the first fully connected layer in the xvector voiceprint recognition model.
[0106] A loss calculation module 15 is configured to calculate losses for the outputs of the second fully-connected layer and the third fully-connected layer respectively using a loss calculation layer, to obtain a first loss probability and a second loss probability.
[0107] Wherein, the loss calculation module 15 is further configured to: calculate a loss for the output of the second fully-connected layer according to a preset loss function and the preset feature vector, to obtain a first loss probability;
[0108] calculate a loss for the output of the third fully-connected layer according to the preset loss function and the dynamic digital feature vector, to obtain a second loss probability.
[0109] A model training module 16 is configured to train the second fully-connected layer according to the first loss probability, and train the third fully-connected layer according to the second loss probability until the outputs of the second fully-connected layer and the third fully-connected layer converge. Preferably, when the second fully-connected layer and the third fully-connected layer reach a preset number of iterations, the model training of the xvector speaker recognition model is automatically stopped, so that the trained xvector speaker recognition model can effectively recognize speaker data that is text-independent, text-dependent, and semi-text-dependent.
[0110] Wherein, the model training module 16 is further configured to: perform backpropagation in the xvector speaker recognition model according to the first loss probability, and perform backpropagation in the xvector speaker recognition model according to the second loss probability.
[0111] Specifically, in this embodiment, the speaker recognition model training apparatus 100 further includes:
[0112] A feature pooling module 14 is configured to perform pooling processing on the training feature vectors output by each of the TDNN networks, and input the pooled training feature vectors into the first fully-connected layer.
[0113] Preferably, the feature pooling module 14 is further configured to: accumulate the training feature vectors output by each of the TDNN networks, calculate the mean and standard deviation of all the training feature vectors according to the vector accumulation result, and use the mean and the standard deviation as the output after the pooling processing of the training feature vectors.
[0114] In addition, the speaker recognition model training apparatus 100 further includes:
[0115] A speaker recognition module 17 is configured to obtain the to-be-recognized speaker data, and input the to-be-recognized speaker data into the xvector speaker recognition model;
[0116] Control the x-vector voiceprint recognition model to recognize the voiceprint data to be recognized, and use the output result of the first fully-connected layer as the output vector of the x-vector voiceprint recognition model;
[0117] Calculate the matching value between the output vector and the sample vector pre-stored locally according to the Euclidean distance formula, and obtain the number value of the sample vector corresponding to the maximum value in the matching value;
[0118] When it is determined that the number value is greater than the number threshold, it is determined that the voiceprint recognition of the voiceprint data to be recognized is qualified.
[0119] In this embodiment, through the design of training the second fully-connected layer according to the preset feature vector and controlling the dynamic digital feature vector to train the third fully-connected layer, the recognition effect of the x-vector voiceprint recognition model on text semi-correlation after model training is improved, so that the x-vector voiceprint recognition model can improve effective voiceprint recognition for text-independent, text-related, and text semi-related, and improve the accuracy of voiceprint recognition.
[0120] Example 4
[0121] Please refer to Figure 4 , which is the mobile terminal 101 provided by the fourth embodiment of the present invention, including a storage device and a processor. The storage device is used to store a computer program, and the processor runs the computer program to enable the mobile terminal 101 to execute the above-mentioned voiceprint recognition model training method.
[0122] This embodiment also provides a storage medium, on which the computer program used in the above-mentioned mobile terminal 101 is stored. When the program is executed, it includes the following steps:
[0123] Obtain training data and input the training data into the x-vector voiceprint recognition model; wherein, the training data includes preset data and dynamic digital data;
[0124] Based on the x-vector voiceprint recognition model, extract features from the training data to obtain training feature vectors, and input the training feature vectors into the first fully-connected layer;
[0125] Perform type recognition on the training feature vectors through the first fully-connected layer to obtain a preset feature vector and a dynamic digital feature vector;
[0126] Input the preset feature vector and the dynamic digital feature vector into the second fully-connected layer and the third fully-connected layer respectively. Both the second fully-connected layer and the third fully-connected layer correspond to an output;
[0127] Use a loss calculation layer to calculate the losses of the outputs of the second fully connected layer and the third fully connected layer respectively, obtaining a first loss probability and a second loss probability;
[0128] Train the second fully connected layer according to the first loss probability, and train the third fully connected layer according to the second loss probability until the outputs of the second fully connected layer and the third fully connected layer converge. The storage medium, such as: ROM / RAM, magnetic disk, optical disk, etc.
[0129] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units or modules according to needs, that is, the internal structure of the storage device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application.
[0130] Those skilled in the art can understand that Figure 3 the composition structure shown in does not constitute a limitation on the voiceprint recognition model training device of the present invention, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components, while Figures 1-2 the voiceprint recognition model training method in also adopts Figure 3 more or fewer components shown in, or combine some components, or arrange different components to implement. The units, modules, etc. referred to in the present invention refer to a series of computer programs that can be executed by a processor (not shown in the figure) in the target voiceprint recognition model training device and can complete specific functions, and all of them can be stored in the storage device (not shown in the figure) of the target voiceprint recognition model training device.
[0131] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training a voiceprint recognition model, characterized in that, The method includes: Obtain training data and input the training data into the xvector speaker recognition model; wherein, the training data includes preset data and dynamic digital data, and the preset data includes text data or voice data; Extract features from the training data based on the xvector speaker recognition model, the xvector speaker recognition model includes a TDNN network, a pooling layer, a first fully-connected layer, a second fully-connected layer, and a third fully-connected layer, and the second fully-connected layer and the third fully-connected layer share the TDNN network, the pooling layer, and the first fully-connected layer in the xvector speaker recognition model; Obtain training feature vectors and input the training feature vectors into the first fully-connected layer; Perform type recognition on the training feature vectors through the first fully-connected layer to obtain preset feature vectors and dynamic digital feature vectors; Respectively input the preset feature vectors and the dynamic digital feature vectors into the second fully-connected layer and the third fully-connected layer, and each of the second fully-connected layer and the third fully-connected layer corresponds to an output; Use a loss calculation layer to calculate losses for the outputs of the second fully-connected layer and the third fully-connected layer respectively to obtain a first loss probability and a second loss probability; Train the second fully-connected layer according to the first loss probability and train the third fully-connected layer according to the second loss probability until the outputs of the second fully-connected layer and the third fully-connected layer converge.
2. The method for training a voiceprint recognition model according to claim 1, characterized in that, The step of extracting features from the training data based on the xvector speaker recognition model includes: Input the training data into the TDNN network in the xvector speaker recognition model and control the TDNN network to extract features from the training data to obtain training features; Control the TDNN network to perform a non-linear transformation on the training features to obtain the training feature vectors.
3. The method for training a voiceprint recognition model according to claim 1, characterized in that, The step of using a loss calculation layer to calculate losses for the outputs of the second fully-connected layer and the third fully-connected layer respectively includes: Calculate the loss for the output of the second fully-connected layer according to a preset loss function and the preset feature vectors to obtain a first loss probability; Calculate the loss for the output of the third fully-connected layer according to the preset loss function and the dynamic digital feature vectors to obtain a second loss probability.
4. The method for training a voiceprint recognition model according to claim 1, characterized in that, The step of training the second fully-connected layer according to the first loss probability and training the third fully-connected layer according to the second loss probability includes: Perform backpropagation in the xvector speaker recognition model according to the first loss probability and perform backpropagation in the xvector speaker recognition model according to the second loss probability.
5. The method for training a voiceprint recognition model according to claim 2, characterized in that, Before the step of inputting the training feature vectors into the first fully-connected layer, the method further includes: Perform pooling processing on the training feature vectors output by each TDNN network and input the pooled training feature vectors into the first fully-connected layer.
6. The method for training a voiceprint recognition model according to claim 5, characterized in that, The step of performing pooling processing on the training feature vectors output by each of the TDNN networks includes: Accumulating the training feature vectors output by each of the TDNN networks, calculating the mean and standard deviation of all the training feature vectors according to the vector accumulation result, and using the mean and the standard deviation as the output after the pooling processing of the training feature vectors.
7. The method for training a voiceprint recognition model according to claim 1, characterized in that, The method further includes: Obtaining the voiceprint data to be recognized and inputting the voiceprint data to be recognized into the xvector voiceprint recognition model; Controlling the xvector voiceprint recognition model to recognize the voiceprint data to be recognized, and using the output result of the first fully connected layer as the output vector of the xvector voiceprint recognition model; Calculating the matching value between the output vector and the sample vector pre-stored locally according to the Euclidean distance formula, and obtaining the number value of the sample vector corresponding to the maximum value in the matching values; When it is determined that the number value is greater than the number threshold, determining that the voiceprint recognition of the voiceprint data to be recognized is qualified.
8. A device for training a voiceprint recognition model, characterized in that, The device includes: A training data acquisition module, configured to acquire training data and input the training data into an xvector voiceprint recognition model, where the training data includes preset data and dynamic digital data, and the preset data includes text data or voice data; A feature extraction module, configured to extract features from the training data based on the xvector voiceprint recognition model, where the xvector voiceprint recognition model includes a TDNN network, a pooling layer, a first fully connected layer, a second fully connected layer, and a third fully connected layer, and the second fully connected layer and the third fully connected layer share the TDNN network, the pooling layer, and the first fully connected layer in the xvector voiceprint recognition model; Obtaining training feature vectors and inputting the training feature vectors into the first fully connected layer; A feature type recognition module, configured to perform type recognition on the training feature vectors through the first fully connected layer to obtain preset feature vectors and dynamic digital feature vectors; A feature output module, configured to respectively input the preset feature vectors and the dynamic digital feature vectors into the second fully connected layer and the third fully connected layer, and each of the second fully connected layer and the third fully connected layer corresponds to an output; A loss calculation module, configured to respectively perform loss calculation on the outputs of the second fully connected layer and the third fully connected layer using a loss calculation layer to obtain a first loss probability and a second loss probability; A model training module, configured to train the second fully connected layer according to the first loss probability and train the third fully connected layer according to the second loss probability until the outputs of the second fully connected layer and the third fully connected layer converge.
9. A mobile terminal, characterized in that, It includes a storage device and a processor, where the storage device is used to store a computer program, and the processor runs the computer program to enable the mobile terminal to execute the voiceprint recognition model training method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, It stores a computer program used in the mobile terminal described in claim 9, and when the computer program is executed by a processor, it implements the steps of the voiceprint recognition model training method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice print detection method, device and equipment and storage medium based on short texts
CN110010133A
Voiceprint recognition method, system, mobile terminal and storage medium
CN111145758A