Voiceprint recognition model training method, voiceprint recognition method and related equipment

By combining labeled and unlabeled data, the voiceprint recognition model is trained using a semi-supervised learning mode, which solves the problem of difficulty in training voiceprint recognition models in the existing technology, and achieves more efficient and accurate model training.

CN115346534BActive Publication Date: 2025-06-06MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110527175.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-14
Publication Date
2025-06-06
Estimated Expiration
2041-05-14

AI Technical Summary

Technical Problem

The training of existing voiceprint recognition models is difficult, mainly due to the high requirements for the quantity and quality of sample data.

Method used

By combining the labeled first sample data and the unlabeled second sample data, the voiceprint recognition model is trained using a semi-supervised learning mode. The specific steps include: iteratively training the encoded network using labeled data, and then transmit the unlabeled data to the decoding network and the feedforward network for further training through the trained encoded network.

Benefits of technology

The requirements for the quantity and quality of sample data are reduced, the training process of the vocalprint recognition model is simplified, and the efficiency and accuracy of model training are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346534B_ABST
    Figure CN115346534B_ABST
Patent Text Reader

Abstract

The present application provides a voiceprint recognition model training method, a voiceprint recognition method and related equipment, the method comprising: inputting the first sample data with annotations into the encoding network included in the model to be trained, and performing the Nth iteration training; inputting the second sample data without annotations into the decoding network through the encoding network after the Nth iteration training, and performing the N+1th iteration training; inputting the second sample data into the feedforward network, and performing the N+1th iteration training; obtaining the voiceprint recognition model when the mean square error between the first vector and the second vector is less than the first threshold; the first vector is the output of the decoding network after the N+1th iteration training, and the second vector is the output of the feedforward network after the N+1th iteration training, and the voiceprint recognition model includes the encoding network after the Nth iteration training, the decoding network after the N+1th iteration training, and the feedforward network after the N+1th iteration training. In this way, the difficulty of model training can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of voiceprint recognition, and in particular to a voiceprint recognition model training method, a voiceprint recognition method and related equipment. Background Art

[0002] As a reliable voiceprint feature authentication technology, voiceprint recognition has broad application prospects in many fields and scenarios such as identity authentication and security verification. However, voice is easily affected by various external factors such as noise environment, emotions, physical condition and internal factors. Therefore, improving the accuracy of voiceprint recognition is of great practical significance. The currently trained voiceprint recognition model has high requirements for the quantity and quality of sample data, which makes the training of voiceprint recognition model more difficult. Summary of the invention

[0003] The embodiments of the present application provide a voiceprint recognition model training method, a voiceprint recognition method and related equipment to solve the problem that the training of the voiceprint recognition model is difficult.

[0004] In a first aspect, an embodiment of the present application provides a voiceprint recognition model training method, comprising:

[0005] Inputting the labeled first sample data into the encoding network included in the model to be trained, and performing the Nth iteration training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network;

[0006] Inputting the unlabeled second sample data into the decoding network through the encoding network after the Nth iteration training to perform the N+1th iteration training; and inputting the second sample data into the feedforward network to perform the N+1th iteration training;

[0007] When the mean square error between the first vector and the second vector is less than a first threshold, a voiceprint recognition model is obtained; wherein the first vector is output by the decoding network after the N+1th iterative training, the second vector is output by the feedforward network after the N+1th iterative training, and the voiceprint recognition model includes the encoding network after the Nth iterative training, the decoding network after the N+1th iterative training, and the feedforward network after the N+1th iterative training.

[0008] It can be seen that in the embodiment of the present application, the voiceprint recognition model can be trained using the labeled first sample data and the unlabeled second sample data at the same time, which reduces the requirements on the quantity and quality of the sample data, thereby reducing the difficulty of training the voiceprint recognition model; in addition, in the training process of the voiceprint recognition model, the labeled first sample data is first used to perform the Nth iteration training on the encoding network, and then the second sample data is transmitted to the decoding network through the encoding network after the Nth iteration training. At the same time, the second sample data is input into the feedforward network, so that when the second sample data is used to train the decoding network and the feedforward network, the encoding network trained by the first sample data can supervise and guide the training process of the decoding network and the feedforward network, so that the second sample data has a very obvious learning direction, thereby further reducing the difficulty of training the voiceprint recognition model.

[0009] In a second aspect, the embodiment of the present application further provides a voiceprint recognition method, including:

[0010] Obtaining first voiceprint data of the user to be identified;

[0011] Inputting the first voiceprint data into an encoding network included in a voiceprint recognition model, and outputting a first feature vector corresponding to the first voiceprint data;

[0012] The first feature vector and the pre-stored second feature vector are input into a target classifier, and a likelihood distribution value is output; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is output by inputting the second voiceprint data of the target user into the encoding network included in the voiceprint recognition model;

[0013] When the likelihood distribution value is greater than a second threshold, it is determined that the user to be identified and the target user are the same user.

[0014] It can be seen that in the embodiment of the present application, the voiceprint recognition model and the target classifier connected to the voiceprint recognition model are used to determine whether the user to be identified and the target user are the same user, thereby improving the accuracy of the recognition result of the voiceprint data of the user to be identified and reducing the loss caused by the inability to accurately identify the voiceprint data of the user to be identified.

[0015] In a third aspect, an embodiment of the present application provides a voiceprint recognition model training device, including:

[0016] A first input module, used for inputting the labeled first sample data into the encoding network included in the model to be trained, and performing N-th iterative training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network;

[0017] A second input module is used to input the unlabeled second sample data into the decoding network through the encoding network after the Nth iteration training, and perform the N+1th iteration training; and input the second sample data into the feedforward network, and perform the N+1th iteration training;

[0018] The first obtaining module is used to obtain a voiceprint recognition model when the mean square error between the first vector and the second vector is less than a first threshold; wherein the first vector is output by a decoding network after N+1th iterative training, and the second vector is output by a feedforward network after N+1th iterative training, and the voiceprint recognition model includes an encoding network after Nth iterative training, a decoding network after N+1th iterative training, and a feedforward network after N+1th iterative training.

[0019] In a fourth aspect, the present application also provides a voiceprint recognition device, including:

[0020] A first acquisition module, used to acquire first voiceprint data of a user to be identified;

[0021] A third input module, used to input the first voiceprint data into an encoding network included in a voiceprint recognition model, and output a first feature vector corresponding to the first voiceprint data;

[0022] a fourth input module, configured to input the first feature vector and a pre-stored second feature vector into a target classifier, and output a likelihood distribution value; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is output by inputting the second voiceprint data of the target user into the encoding network included in the voiceprint recognition model;

[0023] The first determination module is used to determine that the user to be identified and the target user are the same user when the likelihood distribution value is greater than a second threshold.

[0024] In a fifth aspect, an embodiment of the present application further provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps in the above-mentioned voiceprint recognition model training method when executing the computer program, or implements the steps in the above-mentioned voiceprint recognition method when executing the computer program.

[0025] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned voiceprint recognition model training method are implemented, or when the computer program is executed by a processor, the steps in the above-mentioned voiceprint recognition method are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0027] Figure 1 It is a flow chart of a voiceprint recognition model training method provided in an embodiment of the present application;

[0028] Figure 2 is a structural diagram of a model to be trained provided in an embodiment of the present application;

[0029] Figure 3 is a flow chart of a voiceprint recognition method provided by an embodiment of the present application;

[0030] Figure 4 It is a flow chart of a voiceprint recognition model training method and a voiceprint recognition method provided in an embodiment of the present application;

[0031] Figure 5 It is a structural schematic diagram of a voiceprint recognition model training device provided in an embodiment of the present application;

[0032] Figure 6 It is a structural schematic diagram of a voiceprint recognition device provided in an embodiment of the present application;

[0033] Figure 7 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0035] See also Figure 1 , Figure 1 is a flow chart of a voiceprint recognition model training method provided in an embodiment of the present application. Figure 1 As shown, the following steps are included:

[0036] Step 101: input the labeled first sample data into the encoding network included in the model to be trained, and perform the Nth iterative training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network.

[0037] The model to be trained includes an encoding network, a decoding network and a feedforward network, and the encoding network and the decoding network can also be combined to form an autoencoding network, that is, the autoencoding network includes two parts: the encoding network and the decoding network. Figure 2 , Figure 2 It can be used to represent the structural diagram of the model to be trained. Figure 2 It includes an encoding network 20, a decoding network 21 and a feedforward network 23.

[0038] It should be noted that the above encoding network may also be referred to as an encoder, and the above decoding network may also be referred to as a decoder.

[0039] Among them, both the first sample data and the second sample data can be audio data, so as to facilitate the training of the voiceprint recognition model. In addition, the first sample data can be data in a public data set, and the public data set can be downloaded from a server to obtain the first sample data. The annotations of the data in the public data set are all authenticated with relatively high accuracy. It should be noted that the annotations in the first sample data can also be referred to as annotation information.

[0040] As an optional implementation, the method further includes:

[0041] Acquire first initial data and second initial data from a sample pool, wherein the first initial data is labeled, and the second initial data is unlabeled;

[0042] The first initial data is subjected to data amplification processing to obtain the first sample data; and the second initial data is subjected to data amplification processing to obtain the second sample data.

[0043] The specific method of obtaining the first sample data and the second sample data through data augmentation processing is not limited here. For example, the first initial data after data augmentation processing can be determined as the first sample data, and the second initial data after data augmentation processing can be determined as the second sample data.

[0044] In addition, for example, the first initial data and the first initial data after data augmentation processing are determined as the first sample data, and the second initial data and the second initial data after data augmentation processing are determined as the second sample data.

[0045] In this way, since the amount of the first initial data and the second initial data is small, the first initial data and the second initial data can be amplified to obtain the first sample data and the second sample data. The amount of sample data is increased through the above data amplification processing, thereby facilitating the subsequent training of the voiceprint recognition model.

[0046] It should be noted that there may be a difference between the first initial data and the first initial data after data amplification processing. For example, the first initial data is the voice data of user A, and the first initial data does not contain noise, while the first initial data after data amplification processing is the voice data of user A, but the first initial data after data amplification processing also includes noise. Similarly, there may be a difference between the second initial data and the second initial data after data amplification processing. In this way, the diversity of sample data is increased.

[0047] In addition, the specific method of data augmentation processing is not limited here. For example, as an optional implementation, the data augmentation processing includes at least one of the following processing methods: increasing noise, increasing speech speed, and increasing data disturbance. In this way, the data augmentation processing method can be made more diversified and flexible. Of course, data augmentation processing can also include: lowering speech speed and reducing data disturbance.

[0048] In addition, as an optional implementation, performing data amplification processing on the first initial data to obtain the first sample data; and performing data amplification processing on the second initial data to obtain the second sample data, comprises:

[0049] Performing feature extraction on the first initial data and the first initial data after data amplification processing to obtain a third feature vector; and performing feature extraction on the second initial data and the second initial data after data amplification processing to obtain a fourth feature vector;

[0050] Performing spectrum enhancement on the third eigenvector, and determining the spectrum enhanced third eigenvector as the first sample data; and determining the fourth eigenvector as the second sample data.

[0051] In this way, the third eigenvector can be spectrally enhanced, and the third eigenvector after spectral enhancement can be determined as the first sample data, thereby enhancing the generalization performance of the first sample data and reducing the occurrence of overfitting and other phenomena in the first sample data during the training process; in addition, since the second sample data is unlabeled sample data, if spectral enhancement is used, it will increase the uncertainty in the training process and cause the model to be trained to fail to converge. Therefore, the fourth eigenvector can be determined as the second sample data, thereby improving the training rate of the model to be trained.

[0052] It should be noted that spectral enhancement is an algorithm in the field of voiceprint recognition, which randomly sets the features (i.e., sample data) to 0 in a certain proportion to reduce overfitting of sample data during training.

[0053] Wherein, as an optional implementation, the third eigenvector and the fourth eigenvector are both 80-dimensional filter bank features. The 80-dimensional filter bank feature can also be referred to as an 80Fbank feature or an 80-dimensional Fbank feature. In this way, the 80Fbank feature can effectively map the time domain and frequency domain information of the speech, thereby making the coverage of the speech (which can also be understood as voiceprint data or sample data) higher. Of course, the dimension of the filter bank feature can also be reduced, which is not specifically limited here.

[0054] In addition, the Fbank feature can also be called the log mel spectrum feature, which is a feature widely used in speech emotion recognition, speech recognition, voiceprint recognition, and speech synthesis.

[0055] It should be noted that the Fbank feature obtained through sample data can be described as follows: the audio signal (that is, the sample data can be an audio signal) is first pre-emphasized, framed and windowed, and then each frame of the signal is short-time Fourier transform STFT to obtain a short-time amplitude spectrum, and finally the short-time amplitude spectrum is passed through a Mel filter group to obtain the Fbank feature.

[0056] Step 102: input the unlabeled second sample data into the decoding network through the encoding network after the Nth iterative training, and perform the N+1th iterative training; and input the second sample data into the feedforward network, and perform the N+1th iterative training.

[0057] It should be noted that the second sample data can be input into the decoding network and the feedforward network at the same time. Of course, the second sample data can also be input into the decoding network and the feedforward network successively, that is to say, the time when the second sample data is input into the decoding network and the feedforward network is not the same time, but the two times are relatively close. The specific method is not limited here.

[0058] Among them, the unlabeled second sample data can also be audio data. It should be noted that the amount of second sample data can be much larger than the first sample data. In this way, the demand for the first sample data is lower, thereby further reducing the requirements for the quantity and quality of sample data when training the voiceprint recognition model, thereby reducing the difficulty of training the voiceprint recognition model.

[0059] In addition, since the labeled first sample data and the unlabeled second sample data are used simultaneously in the process of training the voiceprint recognition model in the embodiment of the present application, such a training and learning mode can be called a semi-supervised learning (SSL) mode. When the semi-supervised learning mode is adopted, the number of personnel to do the work will be reduced, and at the same time, the output result of the trained voiceprint recognition model can be made more accurate.

[0060] Among them, the second sample data is input into the decoding network and the feedforward network respectively, so that the decoding network and the feedforward network can be trained; in addition, the encoding network is firstly trained for the Nth iteration using the labeled first sample data, and then the second sample data is transmitted to the decoding network through the encoding network after the Nth iteration training, and at the same time, the second sample data is input into the feedforward network, so that when the decoding network and the feedforward network are trained using the second sample data, the encoding network trained by the first sample data can supervise and guide the training process of the decoding network and the feedforward network, so that the second sample data has a very obvious learning direction, thereby improving the training rate and training accuracy of the decoding network and the feedforward network.

[0061] Step 103, when the mean square error between the first vector and the second vector is less than a first threshold, a voiceprint recognition model is obtained; wherein the first vector is output by the decoding network after the N+1th iterative training, the second vector is output by the feedforward network after the N+1th iterative training, and the voiceprint recognition model includes the encoding network after the Nth iterative training, the decoding network after the N+1th iterative training, and the feedforward network after the N+1th iterative training.

[0062] The above can also be understood as: the model formed by the combination of the encoding network after the Nth iterative training, the decoding network after the N+1th iterative training, and the feedforward network after the N+1th iterative training can be called a voiceprint recognition model.

[0063] It should be noted that the first sample data and the second sample data can both be data in the training set, and there can also be a test set in the sample pool. The test set can be used to test the voiceprint recognition model. The test set data is input into the voiceprint recognition model. When the difference between the output result and the actual result is less than the preset difference, the voiceprint recognition model is determined to be an available model.

[0064] In addition, the data obtained from the sample pool can be divided into data in the training set and data in the test set according to a certain ratio. The specific value of the above ratio is not limited here. For example, the above ratio can be 98:2. Of course, there is no repeated overlap between the users in the training set and the test set. For example, the voice data of user A can only exist in the training set or the test set, but not in both sets. In this way, when the test set is used to test the voiceprint recognition model, the test results can be more accurate.

[0065] It should be noted that the specific structures of the encoding network, decoding network and feedforward network included in the model to be trained in the embodiment of the present application are not limited. For example, the encoding network, decoding network and feedforward network can respectively include multiple convolutional layers, and the convolutional layers of the encoding network, decoding network and feedforward network can correspond one to one.

[0066] For example, the encoding network has the same structure as the feedforward network, which is more conducive to the loss convergence of the network trained with unlabeled data.

[0067] For another example: as an optional implementation, the decoding network includes M first convolutional layers, the feedforward network includes M second convolutional layers, the M first convolutional layers and the M second convolutional layers are connected in a one-to-one correspondence, and M is a positive integer;

[0068] The M first convolutional layers output M first vectors, and the M second convolutional layers output M second vectors; the mean square error between the first vectors and the second vectors is less than a first threshold, including: the sum of the M mean square errors is less than the first threshold, and the M mean square errors are obtained by calculating the mean square errors of the M first vectors and the M second vectors.

[0069] In this way, the voiceprint recognition model is determined only when the sum of the M mean square errors is less than a preset threshold, thereby making the output result of the voiceprint recognition model more accurate.

[0070] Among them, each first convolutional layer can correspond to a second convolutional layer. For example, the decoding network includes the first convolutional layer A, the first convolutional layer B and the first convolutional layer C, and the feedforward network can include the second convolutional layer A, the second convolutional layer B and the third convolutional layer C. The first convolutional layer A and the second convolutional layer A can be connected to each other, the first convolutional layer B and the second convolutional layer B can be connected to each other, and the first convolutional layer C and the second convolutional layer C can be connected to each other. Therefore, the first convolutional layer A corresponds to the second convolutional layer A, the first convolutional layer B corresponds to the second convolutional layer B, and the first convolutional layer C corresponds to the second convolutional layer C.

[0071] In this way, each first convolution layer and the corresponding second convolution layer can output a mean square error, thereby obtaining multiple mean square errors. The voiceprint recognition model is determined only when the sum of the multiple mean square errors is less than a preset threshold, thereby making the output result of the voiceprint recognition model more accurate.

[0072] Among them, the first convolution layer can be called deconvolution, the second convolution layer can be called convolution, and correspondingly, the encoding network can also have a third convolution layer corresponding to the first convolution layer one by one. It should be noted that the first convolution layer, the second convolution layer and the third convolution layer can all adopt a 3x3 structure.

[0073] For example: See Figure 2 , Figure 2 The method includes an encoding network 20, a decoding network 21, a feedforward network 23, and a classifier 24 connected to the encoding network 20, wherein the encoding network 20 may include multiple layers of third convolutional layers 203, the decoding network 21 may include multiple layers of first convolutional layers 201, and the feedforward network 23 may include multiple layers of second convolutional layers 202. Figure 2 The arrows in can be used to indicate the direction of data transmission between the encoding network 20, the decoding network 21 and the feedforward network 23.

[0074] Also, see Figure 2 , Figure 2 A represents the input direction of the first sample data (i.e., labeled data) and the second sample data (i.e., unlabeled data), C represents the input direction of the second sample data (i.e., unlabeled data), and B and D represent the output directions of the output results of the encoding network 20 and the feedforward network 23, respectively.

[0075] In addition, when multiple first convolutional layers and second convolutional layers are included, the calculation formula of the sum of the mean square errors can be referred to as follows:

[0076]

[0077] Here, MSE stands for mean square error, and x and x i One of them can represent the first vector, and the other can represent the second vector, i can be used to represent the number of the first convolutional layer or the second convolutional layer, and n represents the total number of the first convolutional layer or the second convolutional layer.

[0078] In the embodiment of the present application, through steps 101 to 103, the voiceprint recognition model can be trained using the labeled first sample data and the unlabeled second sample data at the same time, which reduces the requirements on the quantity and quality of the sample data, thereby reducing the difficulty of training the voiceprint recognition model; at the same time, in the training process of the voiceprint recognition model, the first sample data provides a very obvious learning direction for the second sample data, effectively utilizes the second sample data, and further reduces the difficulty of training the voiceprint recognition model.

[0079] The present application also provides a voiceprint recognition method, which can be applied to the voiceprint recognition model trained in the above embodiment. Figure 3 , including the following steps:

[0080] Step 301: Obtain the first voiceprint data of the user to be identified.

[0081] Step 302: input the first voiceprint data into the encoding network included in the voiceprint recognition model, and output the first feature vector corresponding to the first voiceprint data.

[0082] Step 303: input the first feature vector and the pre-stored second feature vector into a target classifier, and output a likelihood distribution value; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is the output of the encoding network included in the voiceprint recognition model when the second voiceprint data of the target user is input into the voiceprint recognition model.

[0083] Step 304: When the likelihood distribution value is greater than a second threshold, determine that the user to be identified and the target user are the same user.

[0084] It should be noted that the first eigenvector and the second eigenvector may both be x-vector features, and the x-vector features are neural network features extracted by the deep neural network.

[0085] The second feature vector may be understood as a feature vector corresponding to the second voiceprint data of the target user collected in advance, and may be stored on a server corresponding to the database.

[0086] Among them, the type of target classifier is not limited here. For example, the target classifier can be a plda classifier, that is, the plda algorithm can be run in the plda classifier. In this way, the accuracy of the likelihood distribution value output by the plda classifier is higher, thereby making the accuracy of the judgment result of whether the user to be identified and the target user are the same user higher.

[0087] The likelihood distribution value can also be understood as similarity, that is, the larger the likelihood distribution value is, the higher the possibility that the user to be identified and the target user are the same user.

[0088] In the implementation manner of the present application, a voiceprint recognition model and a target classifier connected to the voiceprint recognition model can be used to determine whether the user to be identified and the target user are the same user, thereby improving the accuracy of the recognition result of the voiceprint data of the user to be identified and reducing the loss caused by the inability to accurately identify the voiceprint data of the user to be identified.

[0089] It should be noted that the target classifier can be pre-trained, and the training process of the classifier can be described as follows:

[0090] As an optional implementation, the training method of the target classifier includes:

[0091] Acquire a target feature vector output by an encoding network included in the voiceprint recognition model, wherein the target feature vector corresponds to the first sample data;

[0092] Input the target feature vector into the classifier to be trained, perform the Nth iteration training, and output the likelihood distribution value after the Nth iteration training; wherein the likelihood distribution value after the Nth iteration training, the target parameter of the Nth iteration training and the target feature vector correspond one to one;

[0093] When the mathematical expectation value of the target parameter of the Nth iteration training converges, the classifier after the Nth iteration training is determined as the target classifier.

[0094] The target feature vector obtained by extracting features from the first sample data may be used to train the classifier to be trained.

[0095] In this way, since the voiceprint recognition model and the target classifier are trained using the same first sample data, the correlation between the target classifier and the voiceprint recognition model can be made stronger, so that the likelihood distribution value output by the target classifier can accurately reflect whether the user to be identified and the target user are the same user, thereby improving the accuracy of the judgment result.

[0096] For example: Take an example to illustrate the training process of the target classifier, from obtaining the user's voiceprint data (it can be the voiceprint data uploaded by the user to the server, or it can be the voiceprint data in the sample pool), and extracting the feature vector of the voiceprint data (such as the x-vector feature vector), the x-vector feature vector is input into the classifier (such as the plda classifier), and the EM algorithm is used to perform full probability posterior estimation. The role of the posterior estimation is to use the EM algorithm to perform probability estimation calculations based on the probability, and obtain the best parameters for each round of classifier training. Through multiple rounds of iterative training, until the classifier finds the optimal feature parameters, the classifier corresponding to the optimal feature parameters can be determined as the final available classifier (i.e., the target classifier).

[0097] The training idea of ​​the EM algorithm is to estimate the value of the model parameters in each round of training through maximum likelihood estimation based on the given observed data; then estimate the value of the missing data based on the parameter value estimated by the previous model, and then re-estimate the parameter value based on the estimated missing data plus the previously observed data, and then iterate repeatedly until the model converges and the iteration ends. In other words: in each iterative training, the input target feature vector, the output likelihood distribution value after the iterative training, and the target parameter after the iterative training are one-to-one corresponding. When the mathematical expectation of the target parameter converges, the classifier after the iterative training can be determined as the target classifier.

[0098] As an optional implementation manner, before inputting the first feature vector and the pre-stored second feature vector into the target classifier and outputting the likelihood distribution value, the method further includes:

[0099] Acquire second voiceprint data of the target user;

[0100] Inputting the second voiceprint data into the encoding network included in the voiceprint recognition model to extract the second feature vector;

[0101] Save the second eigenvector.

[0102] The second voiceprint data of the target user may be voiceprint data collected by a sensor, for example, by a microphone. Of course, the second voiceprint data may also be voiceprint data collected in a multi-person conversation scenario, and all voiceprint data included in the multi-person conversation scenario may be subjected to channel separation to obtain the second voiceprint data.

[0103] It should be noted that in order to ensure the comprehensiveness of the data, the feature vector corresponding to the voiceprint data of each user can be extracted and saved on the server, so that when the voiceprint data of a certain user is received again during use, it can be compared with the feature vector of the pre-stored voiceprint data to quickly determine the identity of the above-mentioned user.

[0104] In the implementation manner of the present application, since the second feature vector of the target user can be saved, when subsequently identifying the voiceprint data of the user to be identified, only the second feature vector needs to be obtained, thereby improving the speed of obtaining the second feature vector of the target user.

[0105] The present application is illustrated below by taking a specific embodiment as an example.

[0106] See also Figure 4 , including the following steps:

[0107] Step 401: 30 hours of accurately labeled voiceprint data of 500 people (i.e., the first sample data with annotations) are amplified by adding noise, speeding up the speech speed, increasing data disturbance, etc. (i.e., data amplification method), and 4000 hours of unlabeled data (i.e., the second sample data without annotations) are amplified in the same way.

[0108] Step 402: extract 80-dimensional Fbank features from each speech file in the training set (ie, the first sample data and the second sample data), use spectrum enhancement, and store them in a feature file.

[0109] The first sample data and the second sample data may both be sample data in a training set. Meanwhile, there may also be test set data for testing the voiceprint recognition model obtained through subsequent training.

[0110] It should be noted that the third eigenvector corresponding to the first initial data and the fourth eigenvector corresponding to the second initial data may be spectrally enhanced respectively, or only the third eigenvector corresponding to the first initial data may be spectrally enhanced. In this embodiment, the third eigenvector and the fourth eigenvector are spectrally enhanced simultaneously.

[0111] Among them, step 401 and step 402 can be referred to as the voiceprint feature extraction stage.

[0112] Step 403: For the labeled data (i.e., the first sample data), read the feature files to be trained in batches to form a feature data combination of data-label (note: 128 files are read each time to form a feature data combination). Only the encoding network of the autoencoder network is sent to perform forward propagation and reverse backpropagation training. (It can be understood that the encoding network after reverse backpropagation training can be used as the initial network for the next training. Training is an iterative process with continuous parameters.)

[0113] Step 404: For the unlabeled data (i.e., the second sample data), the training feature files are directly read in batches, sent to the encoding network in the autoencoder network, and flow through the encoding network and the decoding network (passing through the encoding network and the decoding network of the autoencoder network in sequence); at the same time, the unlabeled data is also sent to the feedforward network;

[0114] The decoding network and the feedforward network are trained using the unlabeled data above. The mean square error is calculated between the output vectors of each hidden layer (i.e., the first convolutional layer and the corresponding second convolutional layer) of the decoding network and the feedforward network, and this mean square error is minimized during the training process (Note: training is completed when the mean square error is minimized).

[0115] Among them, the labeled data in step 403 plays a role of supervised training for the training of the decoding network and the feedforward network with unlabeled data. Under the supervision of the labeled data, the mean square error of the unsupervised data is minimized until the loss converges and the model is saved.

[0116] Among them, step 403 and step 404 can be referred to as the voiceprint confirmation semi-supervised network (ie, voiceprint recognition model) training phase.

[0117] Step 405: Before a user performs voiceprint recognition, the user first records a registration voice to register the user's voiceprint. The encoding network part of the trained and converged autoencoder network extracts the x-vector feature (i.e., the second feature vector corresponding to the second voiceprint data, which is domain-wide knowledge and is a neural network feature extracted by the deep neural network) and stores it in the registration library.

[0118] Among them, step 405 can be called the voiceprint confirmation registration stage.

[0119] Step 406: Obtain the conversation recording between the customer service and the user, separate the audio channels, and separate the customer service and user audio channels.

[0120] Step 407, extract 80-dimensional Fbank features from the recording of the customer's vocal channel (i.e., the first voiceprint data), send the obtained features to the encoding network of the convergent autoencoder network, obtain the feature vector x-vector of the speech to be identified (i.e., the first feature vector), and then input the first feature vector and the pre-stored second feature vector into the plda classifier to obtain the likelihood distribution value, and judge whether the user to be identified corresponding to the first feature vector and the target user corresponding to the second feature vector are the same user according to the likelihood distribution value.

[0121] Among them, step 406 and step 407 can be referred to as a voiceprint confirmation stage or a voiceprint recognition stage.

[0122] In this way, through steps 401 to 407, the training process of the voiceprint recognition model and the recognition process of the voiceprint recognition model on the voiceprint data can be fully reflected.

[0123] See also Figure 5 , Figure 5 is a structural diagram of a voiceprint recognition model training device provided in an embodiment of the present application, such as Figure 5 As shown, the voiceprint recognition model training device 500 includes:

[0124] A first input module 501 is used to input the labeled first sample data into the encoding network included in the model to be trained, and perform the Nth iteration training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network;

[0125] The second input module 502 is used to input the unlabeled second sample data into the decoding network through the encoding network after the Nth iteration training, and perform the N+1th iteration training; and input the second sample data into the feedforward network, and perform the N+1th iteration training;

[0126] The first obtaining module 503 is used to obtain a voiceprint recognition model when the mean square error between the first vector and the second vector is less than a first threshold; wherein the first vector is output by the decoding network after the N+1th iteration training, the second vector is output by the feedforward network after the N+1th iteration training, and the voiceprint recognition model includes an encoding network after the Nth iteration training, a decoding network after the N+1th iteration training, and a feedforward network after the N+1th iteration training. Optionally, the decoding network includes M first convolutional layers, the feedforward network includes M second convolutional layers, the M first convolutional layers and the M second convolutional layers are connected one-to-one, and M is a positive integer;

[0127] The M first convolutional layers output M first vectors, and the M second convolutional layers output M second vectors; the mean square error between the first vectors and the second vectors is less than a first threshold, including: the sum of the M mean square errors is less than the first threshold, and the M mean square errors are obtained by calculating the mean square errors of the M first vectors and the M second vectors.

[0128] Optionally, the voiceprint recognition model training device 500 further includes:

[0129] A second acquisition module is used to acquire first initial data and second initial data from the sample pool, wherein the first initial data is labeled and the second initial data is unlabeled;

[0130] The amplification processing module is used to perform data amplification processing on the first initial data to obtain the first sample data; and to perform data amplification processing on the second initial data to obtain the second sample data.

[0131] Optionally, the amplification processing module includes:

[0132] A feature extraction submodule, configured to perform feature extraction on the first initial data and the first initial data after data amplification processing to obtain a third feature vector; and perform feature extraction on the second initial data and the second initial data after data amplification processing to obtain a fourth feature vector;

[0133] The spectrum enhancement submodule is used to perform spectrum enhancement on the third eigenvector, determine the spectrum enhanced third eigenvector as the first sample data; and determine the fourth eigenvector as the second sample data.

[0134] Optionally, the third eigenvector and the fourth eigenvector are both 80-dimensional filter bank features.

[0135] The voiceprint recognition model training device provided in the embodiment of the present application can achieve Figure 1 In order to avoid repetition, the various processes implemented by the voiceprint recognition model training device in the method embodiment are not described here. In the embodiment of the present application, the voiceprint recognition model can be trained using the labeled first sample data and the unlabeled second sample data at the same time, which reduces the requirements on the quantity and quality of the sample data, thereby reducing the difficulty of voiceprint recognition model training.

[0136] See also Figure 6 , Figure 6 A structural diagram of a voiceprint recognition device provided in an embodiment of the present application, such as Figure 6 As shown, the voiceprint recognition device 600 includes:

[0137] The first acquisition module 601 is used to acquire the first voiceprint data of the user to be identified;

[0138] A third input module 602 is used to input the first voiceprint data into the encoding network included in the voiceprint recognition model, and output a first feature vector corresponding to the first voiceprint data;

[0139] The fourth input module 603 is used to input the first feature vector and the pre-stored second feature vector into a target classifier, and output a likelihood distribution value; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is output by inputting the second voiceprint data of the target user into the encoding network included in the voiceprint recognition model;

[0140] The first determination module 604 is configured to determine that the user to be identified and the target user are the same user when the likelihood distribution value is greater than a second threshold.

[0141] Optionally, the voiceprint recognition device 600 further includes:

[0142] A third acquisition module, used to acquire a target feature vector output by the encoding network included in the voiceprint recognition model, wherein the target feature vector corresponds to the first sample data;

[0143] A fifth input module, used for inputting the target feature vector into the classifier to be trained, performing the Nth iteration training, and outputting the likelihood distribution value after the Nth iteration training; wherein the likelihood distribution value after the Nth iteration training, the target parameter of the Nth iteration training and the target feature vector correspond one to one;

[0144] The second determination module is used to determine the classifier after the Nth iterative training as the target classifier when the mathematical expectation value of the target parameter of the Nth iterative training converges.

[0145] The voiceprint recognition device provided in the embodiment of the present application can achieve Figure 3 In the method embodiment, the various processes implemented by the voiceprint recognition device are not described here to avoid repetition. In the embodiment of the present application, the voiceprint recognition model and the target classifier connected to the voiceprint recognition model can be used to determine whether the user to be identified and the target user are the same user, thereby improving the accuracy of the recognition result of the voiceprint data of the user to be identified and reducing the loss caused by the inability to accurately identify the voiceprint data of the user to be identified.

[0146] Figure 7 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present application.

[0147] The electronic device 700 includes but is not limited to: a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, a processor 710, and a power supply 711. Those skilled in the art will appreciate that Figure 7 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In the embodiments of the present application, the electronic device includes but is not limited to a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted terminal, a wearable device, and a pedometer.

[0148] When the electronic device is used to execute the steps in the voiceprint recognition model training method, the processor 710 is used to perform the following operations:

[0149] Inputting the labeled first sample data into the encoding network included in the model to be trained, and performing the Nth iteration training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network;

[0150] Inputting the unlabeled second sample data into the decoding network through the encoding network after the Nth iteration training to perform the N+1th iteration training; and inputting the second sample data into the feedforward network to perform the N+1th iteration training;

[0151] When the mean square error between the first vector and the second vector is less than a first threshold, a voiceprint recognition model is obtained; wherein the first vector is output by the decoding network after the N+1th iterative training, the second vector is output by the feedforward network after the N+1th iterative training, and the voiceprint recognition model includes the encoding network after the Nth iterative training, the decoding network after the N+1th iterative training, and the feedforward network after the N+1th iterative training.

[0152] Optionally, the decoding network includes M first convolutional layers, the feedforward network includes M second convolutional layers, the M first convolutional layers and the M second convolutional layers are connected in a one-to-one correspondence, and M is a positive integer;

[0153] The M first convolutional layers output M first vectors, and the M second convolutional layers output M second vectors; the mean square error between the first vectors and the second vectors is less than a first threshold, including: the sum of the M mean square errors is less than the first threshold, and the M mean square errors are obtained by calculating the mean square errors of the M first vectors and the M second vectors.

[0154] Optionally, the processor 710 is further configured to:

[0155] Acquire first initial data and second initial data from a sample pool, wherein the first initial data is labeled, and the second initial data is unlabeled;

[0156] The first initial data is subjected to data amplification processing to obtain the first sample data; and the second initial data is subjected to data amplification processing to obtain the second sample data.

[0157] Optionally, the processor 710 performs data augmentation processing on the first initial data to obtain the first sample data; and performs data augmentation processing on the second initial data to obtain the second sample data, including:

[0158] Performing feature extraction on the first initial data and the first initial data after data amplification processing to obtain a third feature vector; and performing feature extraction on the second initial data and the second initial data after data amplification processing to obtain a fourth feature vector;

[0159] Performing spectrum enhancement on the third eigenvector, and determining the spectrum enhanced third eigenvector as the first sample data; and determining the fourth eigenvector as the second sample data.

[0160] Optionally, the third eigenvector and the fourth eigenvector are both 80-dimensional filter bank features.

[0161] When the electronic device is used to execute the steps in the voiceprint recognition method, the processor 710 is used to perform the following operations:

[0162] Obtaining first voiceprint data of the user to be identified;

[0163] Inputting the first voiceprint data into an encoding network included in a voiceprint recognition model, and outputting a first feature vector corresponding to the first voiceprint data;

[0164] The first feature vector and the pre-stored second feature vector are input into a target classifier, and a likelihood distribution value is output; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is output by inputting the second voiceprint data of the target user into the encoding network included in the voiceprint recognition model;

[0165] When the likelihood distribution value is greater than a second threshold, it is determined that the user to be identified and the target user are the same user.

[0166] Optionally, the processor 710 is further configured to:

[0167] Acquire a target feature vector output by an encoding network included in the voiceprint recognition model, wherein the target feature vector corresponds to the first sample data;

[0168] Input the target feature vector into the classifier to be trained, perform the Nth iteration training, and output the likelihood distribution value after the Nth iteration training; wherein the likelihood distribution value after the Nth iteration training, the target parameter of the Nth iteration training and the target feature vector correspond one to one;

[0169] When the mathematical expectation value of the target parameter of the Nth iteration training converges, the classifier after the Nth iteration training is determined as the target classifier.

[0170] It should be understood that in the embodiment of the present application, the radio frequency unit 701 can be used for receiving and sending signals during information transmission or calls. Specifically, after receiving downlink data from the base station, it is sent to the processor 710 for processing; in addition, uplink data is sent to the base station. Generally, the radio frequency unit 701 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the radio frequency unit 701 can also communicate with the network and other devices through a wireless communication system.

[0171] The electronic device provides users with wireless broadband Internet access through the network module 702, such as helping users to send and receive emails, browse web pages, and access streaming media.

[0172] The audio output unit 703 can convert the audio data received by the RF unit 701 or the network module 702 or stored in the memory 709 into an audio signal and output it as sound. Moreover, the audio output unit 703 can also provide audio output related to a specific function performed by the electronic device 700 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 703 includes a speaker, a buzzer, a receiver, etc.

[0173] The input unit 704 is used to receive audio or video signals. The input unit 704 may include a graphics processor (GPU) 7041 and a microphone 7042, and the graphics processor 7041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 706. The image frame processed by the graphics processor 7041 can be stored in the memory 709 (or other storage medium) or sent via the radio frequency unit 701 or the network module 702. The microphone 7042 can receive sound and can process such sound into audio data. The processed audio data can be converted into a format output that can be sent to a mobile communication base station via the radio frequency unit 701 in the case of a telephone call mode.

[0174] The electronic device 700 also includes at least one sensor 705, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 7061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 7061 and / or the backlight when the electronic device 700 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, which can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 705 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.

[0175] The display unit 706 is used to display information input by the user or information provided to the user. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0176] The user input unit 707 can be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the electronic device. Specifically, the user input unit 707 includes a touch panel 7071 and other input devices 7072. The touch panel 7071, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 7071 or near the touch panel 7071 using any suitable object or accessory such as a finger, stylus, etc.). The touch panel 7071 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the contact point coordinates, and then sends it to the processor 710, receives the command sent by the processor 710 and executes it. In addition, the touch panel 7071 can be implemented using multiple types such as resistive, capacitive, infrared and surface acoustic waves. In addition to the touch panel 7071, the user input unit 707 may also include other input devices 7072. Specifically, other input devices 7072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which are not described in detail here.

[0177] Furthermore, the touch panel 7071 may be overlaid on the display panel 7061. When the touch panel 7071 detects a touch operation on or near it, it transmits the information to the processor 710 to determine the type of the touch event. Then, the processor 710 provides a corresponding visual output on the display panel 7061 according to the type of the touch event. Figure 7 In the figure, the touch panel 7071 and the display panel 7061 are used as two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 7071 and the display panel 7061 can be integrated to realize the input and output functions of the electronic device, which is not limited here.

[0178] The interface unit 708 is an interface for connecting an external device to the electronic device 700. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 708 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the electronic device 700 or may be used to transmit data between the electronic device 700 and an external device.

[0179] The memory 709 can be used to store software programs and various data. The memory 709 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 709 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0180] The processor 710 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 709, and calling data stored in the memory 709, so as to monitor the electronic device as a whole. The processor 710 may include one or more processing units; preferably, the processor 710 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 710.

[0181] The electronic device 700 may also include a power supply 711 (such as a battery) for supplying power to each component. Preferably, the power supply 711 may be logically connected to the processor 710 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system.

[0182] In addition, the electronic device 700 includes some functional modules not shown, which will not be described in detail here.

[0183] Preferably, an embodiment of the present application also provides an electronic device, including a processor 710, a memory 709, and a computer program stored in the memory 709 and executable on the processor 710. When the computer program is executed by the processor 710, the above-mentioned voiceprint recognition model training method or each process of the above-mentioned voiceprint recognition method can be implemented, and the same technical effect can be achieved, which will not be repeated here.

[0184] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 710, each process of the above-mentioned voiceprint recognition model training method or the above-mentioned voiceprint recognition method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0185] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0186] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0187] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A voiceprint recognition model training method, It is characterized in that include: Inputting the labeled first sample data into the encoding network included in the model to be trained, and performing the Nth iteration training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network; Inputting the unlabeled second sample data into the decoding network through the encoding network after the Nth iteration training to perform the N+1th iteration training; and inputting the second sample data into the feedforward network to perform the N+1th iteration training; When the mean square error between the first vector and the second vector is less than a first threshold, a voiceprint recognition model is obtained; wherein the first vector is output by the decoding network after the N+1th iterative training, the second vector is output by the feedforward network after the N+1th iterative training, and the voiceprint recognition model includes the encoding network after the Nth iterative training, the decoding network after the N+1th iterative training, and the feedforward network after the N+1th iterative training.

2. The method according to claim 1, It is characterized in that The decoding network includes M first convolutional layers, the feedforward network includes M second convolutional layers, the M first convolutional layers and the M second convolutional layers are connected in a one-to-one correspondence, and M is a positive integer; The M first convolutional layers output M first vectors, and the M second convolutional layers output M second vectors; the mean square error between the first vectors and the second vectors is less than a first threshold, including: the sum of the M mean square errors is less than the first threshold, and the M mean square errors are obtained by calculating the mean square errors of the M first vectors and the M second vectors.

3. The method according to claim 1 or 2, It is characterized in that The method further comprises: Acquire first initial data and second initial data from a sample pool, wherein the first initial data is labeled, and the second initial data is unlabeled; The first initial data is subjected to data amplification processing to obtain the first sample data; and the second initial data is subjected to data amplification processing to obtain the second sample data.

4. The method according to claim 3, It is characterized in that The performing data amplification processing on the first initial data to obtain the first sample data; and performing data amplification processing on the second initial data to obtain the second sample data, comprises: Performing feature extraction on the first initial data and the first initial data after data amplification processing to obtain a third feature vector; and performing feature extraction on the second initial data and the second initial data after data amplification processing to obtain a fourth feature vector; Performing spectrum enhancement on the third eigenvector, and determining the spectrum enhanced third eigenvector as the first sample data; and determining the fourth eigenvector as the second sample data.

5. The method according to claim 4, It is characterized in that The third eigenvector and the fourth eigenvector are both 80-dimensional filter bank features.

6. A voiceprint recognition method, It is characterized in that include: Obtaining first voiceprint data of the user to be identified; Inputting the first voiceprint data into an encoding network included in a voiceprint recognition model, and outputting a first feature vector corresponding to the first voiceprint data; The first feature vector and the pre-stored second feature vector are input into a target classifier, and a likelihood distribution value is output; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is output by inputting the second voiceprint data of the target user into the encoding network included in the voiceprint recognition model; When the likelihood distribution value is greater than a second threshold, determining that the user to be identified and the target user are the same user; The voiceprint recognition model is trained by the voiceprint recognition model training method according to any one of claims 1 to 5.

7. The method according to claim 6, It is characterized in that The training methods of the target classifier are: Obtaining a target feature vector output by the encoding network included in the voiceprint recognition model, wherein the target feature vector corresponds to the first sample data; Input the target feature vector into the classifier to be trained, perform the Nth iteration training, and output the likelihood distribution value after the Nth iteration training; wherein the likelihood distribution value after the Nth iteration training, the target parameter of the Nth iteration training and the target feature vector correspond one to one; When the mathematical expectation value of the target parameter of the Nth iteration training converges, the classifier after the Nth iteration training is determined as the target classifier.

8. A voiceprint recognition model training device, It is characterized in that include: A first input module, used for inputting the labeled first sample data into the encoding network included in the model to be trained, and performing N-th iterative training; wherein N is a positive integer, the model to be trained also includes a decoding network and a feedforward network, and the encoding network is connected to the feedforward network through the decoding network; A second input module is used to input the unlabeled second sample data into the decoding network through the encoding network after the Nth iteration training, and perform the N+1th iteration training; and input the second sample data into the feedforward network, and perform the N+1th iteration training; The first obtaining module is used to obtain a voiceprint recognition model when the mean square error between the first vector and the second vector is less than a first threshold; wherein the first vector is output by a decoding network after N+1th iterative training, and the second vector is output by a feedforward network after N+1th iterative training, and the voiceprint recognition model includes an encoding network after Nth iterative training, a decoding network after N+1th iterative training, and a feedforward network after N+1th iterative training.

9. A voiceprint recognition device, It is characterized in that include: A first acquisition module, used to acquire first voiceprint data of a user to be identified; A third input module, configured to input the first voiceprint data into an encoding network included in a voiceprint recognition model, and output a first feature vector corresponding to the first voiceprint data; a fourth input module, configured to input the first feature vector and a pre-stored second feature vector into a target classifier, and output a likelihood distribution value; wherein the target classifier is connected to the encoding network included in the voiceprint recognition model, and the second feature vector is output by inputting the second voiceprint data of the target user into the encoding network included in the voiceprint recognition model; A first determination module, configured to determine that the user to be identified and the target user are the same user when the likelihood distribution value is greater than a second threshold; The voiceprint recognition model is trained by the voiceprint recognition model training method according to any one of claims 1 to 5.

10. An electronic device, It is characterized in that include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the voiceprint recognition model training method as described in any one of claims 1 to 5 when executing the computer program, or implements the steps in the voiceprint recognition method as described in claim 6 or 7 when executing the computer program.

11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the voiceprint recognition model training method as described in any one of claims 1 to 5, or, when executed by a processor, implements the steps in the voiceprint recognition method as described in claim 6 or 7.

Citation Information

Patent Citations

  • Speech synthesis method, device and equipment and storage medium

    CN111667812A

  • Customer service telephone voice text transcription method, system and device and storage medium

    CN112217947A