Training of Voiceprint Recognition Model, Voiceprint Recognition Method and Device
By generating support sets and query sets and building specific neural network models, the overfitting problem of voiceprint recognition models when lacking data is solved, and the effect of efficient training under a small amount of data is achieved.
Patent Information
- Application Number
- CN202111214097.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-10-19
AI Technical Summary
Voiceprint recognition models are prone to overfitting when they lack massive labeled audio data, which reduces the recognition accuracy.
By obtaining training data, a support set and a query set are generated, and a neural network model containing feature extraction layer, prototype network layer and full connection layer are built to train the model to avoid overfitting.
With few training data, the voiceprint recognition model is effectively trained to avoid overfitting, which improves the training effect and recognition accuracy of the model.
Smart Images

Figure CN114067805B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to artificial intelligence technology fields such as cloud services, speech processing, and deep learning. Specifically, a method, apparatus, electronic device, and readable storage medium for training a voiceprint recognition model and voiceprint recognition are provided. Background Art
[0002] In related technologies, when training a voiceprint recognition model, a large amount of labeled audio data is required for training. However, audio data may involve personal privacy and often cannot be obtained in large quantities, resulting in overfitting of the voiceprint recognition model in related technologies and reducing the accuracy of the voiceprint recognition model when performing voiceprint recognition. Summary of the Invention
[0003] To solve the technical problem that the voiceprint recognition model in related technologies may cause overfitting when there is no large amount of labeled audio data for training, the present disclosure provides a method for training a voiceprint recognition model and voiceprint recognition, which is used to achieve the purpose of training a voiceprint recognition model with less training data, avoid the problem of overfitting of the trained voiceprint recognition model, and improve the training effect of the voiceprint recognition model.
[0004] According to a first aspect of the present disclosure, a method for training a voiceprint recognition model is provided, including: obtaining training data, where the training data includes multiple sample audio data and category labels of the multiple sample audio data; generating a support set and a query set corresponding to different training tasks according to the training data; constructing a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer, where the feature extraction layer is used to output feature vectors of support audio data included in the support set and feature vectors of query audio data included in the query set, the prototype network layer is used to output prototype vectors of various categories corresponding to the support audio data included in the support set, and the fully connected layer is used to output a probability distribution of the query audio data included in the query set belonging to various categories in the support set; using the support set and the query set corresponding to different training tasks to train the neural network model to obtain a voiceprint recognition model method.
[0005] According to a second aspect of the present disclosure, a voiceprint recognition method is provided, including: obtaining test data, where the test data includes multiple test audio data and category labels of the multiple test audio data; generating a support set according to the test data; obtaining audio data to be recognized as a query set; inputting the support set and the query set into the voiceprint recognition model, and determining a voiceprint recognition result of the audio data to be recognized according to an output result of the voiceprint recognition model.
[0006] According to a third aspect of the present disclosure, there is provided a training device for a voiceprint recognition model, including: a first acquisition unit configured to acquire training data, where the training data includes a plurality of sample audio data and class labels of the plurality of sample audio data; a first generation unit configured to generate a support set and a query set corresponding to different training tasks according to the training data; a construction unit configured to construct a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer, where the feature extraction layer is configured to output feature vectors of support audio data included in the support set and feature vectors of query audio data included in the query set, the prototype network layer is configured to output prototype vectors of various categories corresponding to the support audio data included in the support set, and the fully connected layer is configured to output a probability distribution of the query audio data included in the query set belonging to various categories in the support set; and a training unit configured to train the neural network model using the support set and the query set corresponding to different training tasks to obtain a voiceprint recognition model.
[0007] According to a fourth aspect of the present disclosure, there is provided a voiceprint recognition device, including: a second acquisition unit configured to acquire test data, where the test data includes a plurality of test audio data and class labels of the plurality of test audio data; a second generation unit configured to generate a support set according to the test data; a processing unit configured to acquire audio data to be recognized as a query set; and an identification unit configured to input the support set and the query set into the voiceprint recognition model and determine a voiceprint recognition result of the audio data to be recognized according to an output result of the voiceprint recognition model.
[0008] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0009] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method as described above.
[0010] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program implements the method as described above when executed by a processor.
[0011] It can be seen from the above technical solutions that the present disclosure can achieve the purpose of training a voiceprint recognition model with less training data, and avoid the problem that the trained voiceprint recognition model will have an overfitting phenomenon, thereby improving the training effect of the voiceprint recognition model.
[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0014] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0015] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0016] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;
[0017] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0018] Figure 5 is a block diagram of an electronic device for training a voiceprint recognition model or implementing a voiceprint recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and mechanisms are omitted below for clarity and conciseness.
[0020] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. As Figure 1 shown, the method for training a voiceprint recognition model in this embodiment specifically includes the following steps:
[0021] S101. Obtain training data, where the training data includes a plurality of sample audio data and class labels of the plurality of sample audio data;
[0022] S102. Generate a support set and a query set corresponding to different training tasks according to the training data;
[0023] S103. Construct a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer. The feature extraction layer is used to output the feature vectors of the support audio data included in the support set and the feature vectors of the query audio data included in the query set. The prototype network layer is used to output the prototype vectors of various categories corresponding to the support audio data included in the support set. The fully connected layer is used to output the probability distribution of the query audio data included in the query set belonging to various categories in the support set.
[0024] S104. Use the support set and query set corresponding to different training tasks to train the neural network model to obtain a voiceprint recognition model.
[0025] In the training method of the voiceprint recognition model of this embodiment, after dividing the training data into a support set and a query set corresponding to different training tasks, the support set and query set corresponding to different training tasks are used to train a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer, so as to obtain a voiceprint recognition model. This embodiment can achieve the purpose of training a voiceprint recognition model with less training data, and avoid the problem that the trained voiceprint recognition model will have an overfitting phenomenon, improving the training effect of the voiceprint recognition model.
[0026] The training data obtained by this embodiment executing S101 includes multiple sample audio data. The category label of each sample audio data is used to represent the category to which the sample audio data belongs. The category to which the sample audio data belongs is specifically a certain user or a certain speaker. For example, the category label of sample audio data 1 is user 1, the category label of sample audio data 2 is user 2, and the category label of sample audio data 3 is user 1, etc.
[0027] It can be understood that the sample audio data obtained by this embodiment executing S101 can be the audio data itself or the spectral features extracted from the audio data, such as a spectrogram. That is to say, this embodiment can use the audio data itself to train the neural network model, or use the spectral features extracted from the audio data to train the neural network model.
[0028] If the obtained sample audio data is the audio data itself, after this embodiment executes S101 to obtain the training data, the following content may also be included: preprocess the sample audio data. The preprocessing in this embodiment can be noise reduction processing, voice activity detection (VAD) processing, etc.; use a preset window function to extract spectral features from the preprocessed sample audio data.
[0029] The window function used when this embodiment executes S101 can be a Hamming window, etc. For example, a Hamming window with a preset size (such as 25 ms) is displaced on the audio data at a preset step size (such as 10 ms), so as to obtain the spectral features of the audio data according to the calculation results of the Hamming window at each displacement.
[0030] After this embodiment executes S101 to obtain the training data including multiple sample audio data and the class labels of multiple sample audio data, it executes S102 to generate a support set and a query set corresponding to different training tasks according to the obtained training data.
[0031] In the support sets corresponding to different training tasks generated by this embodiment when executing S102, it includes at least one support audio data and the class labels of at least one support audio data; in the query sets corresponding to different training tasks generated, it includes at least one query audio data and the class labels of at least one query audio data.
[0032] In this embodiment, each training task represents a training of the neural network model; this embodiment executes S102 to generate a support set and a query set corresponding to different training tasks, which means training the neural network model multiple times by dividing the training data into multiple training tasks.
[0033] Specifically, when this embodiment executes S102 to generate a support set and a query set corresponding to different training tasks according to the obtained training data, the optional implementation method that can be adopted is: for each training task, obtain at least one training class corresponding to this training task; according to the class labels of the sample audio data, respectively extract the first preset number of sample audio data from the sample audio data corresponding to the at least one obtained training class as the support audio data to form the support set; extract the second preset number of sample audio data from the remaining sample audio data after extracting the first preset number of sample audio data from the sample audio data corresponding to the at least one training class as the query audio data to form the query set.
[0034] When this embodiment executes S102 to obtain at least one training class corresponding to the training task, it can use the classes corresponding to the randomly selected third preset number of class labels as the at least one training class according to the class labels of the sample audio data.
[0035] For example, if the training data obtained by executing S101 in this embodiment includes sample audio data 1, sample audio data 2, sample audio data 3, sample audio data 4, sample audio data 5, and sample audio data 6, and if the class labels of sample audio data 1 and sample audio data 2 are class 1, the class label of sample audio data 3 is class 2, and the class labels of sample audio data 4, sample audio data 5, and sample audio data 6 are class 3; if at least one training class obtained by executing S102 in this embodiment is class 1 and class 3, and if the first preset quantity is 1, then one sample audio data is respectively extracted from the sample audio data corresponding to class 1 and class 3 to form a support set; if the second preset quantity is 2, then two sample audio data are further extracted from the remaining sample audio data corresponding to class 1 and class 3 to form a query set.
[0036] After this embodiment generates the support set and the query set corresponding to different training tasks by executing S102, it executes S103 to construct a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer.
[0037] Among them, the feature extraction layer in the neural network model constructed by this embodiment executing S103 can be any type of feature extraction network, which is used to encode the input sample audio data into a feature vector with a fixed dimension.
[0038] Specifically, the feature extraction layer in the neural network model constructed by this embodiment executing S103 is composed of a Capsule Network and an Autoencoder; among them, the Capsule Network in the feature extraction layer utilizes the spatial information in the input spectral features (such as the spatial information between high tones and formants) to obtain a feature vector mapped in the space of the Capsule Network; the Autoencoder in the feature extraction layer maps the feature vector output by the Capsule Network from the space of the Capsule Network to a common embedding space, thereby completing the feature extraction of the input sample audio data.
[0039] By constructing a feature extraction layer including a Capsule Network and an Autoencoder, this embodiment can combine the spatial information in the input spectral features to obtain the feature vector corresponding to the sample audio data, thereby improving the accuracy of the feature extraction vector.
[0040] The prototype network layer in the neural network model constructed by this embodiment executing S103 is used to output the prototype vectors of various categories corresponding to the support audio data included in the support set.
[0041] Specifically, when the prototype network layer in the neural network model constructed in S103 of this embodiment outputs the prototype vectors of various categories corresponding to the support audio data included in the support set, an optional implementation method that can be adopted is as follows: for each category in the support set, obtain the feature vectors of the support audio data included in this category; according to the obtained feature vectors, obtain the prototype vector of this category.
[0042] When the prototype network layer in this embodiment obtains the prototype vector of this category according to the obtained feature vectors in S103, it can calculate the mean value of the obtained feature vectors as the prototype vector of this category, or use the accumulated result of the obtained feature vectors as the prototype vector of this category.
[0043] Among them, the prototype network layer in this embodiment can use the following calculation formula to obtain the prototype vectors corresponding to various categories:
[0044]
[0045] In the formula: v c represents the prototype vector of category c; S c represents the support audio data in the support set S that belongs to category c; x i represents S c the spectral feature of the i-th support audio data in; y i represents the category to which the i-th support audio data belongs; f θ (x i ) represents the feature vector of the i-th support audio data.
[0046] The fully connected layer in the neural network model constructed in S103 of this embodiment is used to output the probability distribution of the query audio data included in the query set belonging to various categories in the support set.
[0047] Specifically, when the fully connected layer in the neural network model constructed in S103 of this embodiment outputs the probability distribution of the query audio data included in the query set belonging to various categories in the support set, an optional implementation method that can be adopted is as follows: obtain the feature vectors of each query audio data in the query set; for each query audio data, according to the feature vector of this query audio data and the prototype vectors of various categories in the support set, obtain the probability distribution of this query audio data belonging to various categories in the support set.
[0048] Among them, the fully connected layer in this embodiment can use the following calculation formula to obtain the probability distribution of the query audio data belonging to various categories in the support set:
[0049]
[0050] In the formula: P θ(y = c|x) represents the probability distribution that the query audio data belongs to class c in the support set; x represents the sample audio data belonging to class c in the query set; f θ (x) represents the feature vector of the query audio data; v c represents the prototype vector of class c; v c' represents the prototype vectors of classes other than class c; represents the distance function.
[0051] In this embodiment, after executing S103 to construct a neural network layer model including a feature extraction layer, a prototype network layer, and a fully connected layer, S104 is executed to train the neural network model using the support set and query set corresponding to different training tasks to obtain a voiceprint recognition model.
[0052] Specifically, when this embodiment executes S104 to train the neural network model using the support set and query set corresponding to different training tasks to obtain a voiceprint recognition model, an optional implementation method that can be adopted is: input the support set and query set corresponding to different training tasks into the neural network model; calculate the loss function value according to the probability distribution that the query audio data included in the query set belongs to each class in the support set output by the neural network model for each training task; adjust the parameters of the neural network model according to the calculated loss function value until the neural network model converges to obtain a voiceprint recognition model.
[0053] It can be understood that when the neural network model in this embodiment executes S104 to output the probability that the query audio data included in the corresponding query set belongs to each class in the support set for each training task, it is processed through the feature extraction layer, the prototype network layer, and the fully connected layer in sequence. The specific processing process is as described above and will not be elaborated here.
[0054] Among them, when this embodiment executes S104, the following calculation formula can be used to obtain the loss function value:
[0055] L(θ) = -logP θ (y = c|x)
[0056] In the formula: L θ is the calculated loss function value; P θ (y = c|x) represents the probability distribution that the query audio data belongs to class c in the support set.
[0057] Using the voiceprint recognition model trained by this embodiment, it can output the class corresponding to the audio data in the query set according to the input support set and query set, that is, the voiceprint recognition result of the audio data.
[0058] Through the above method, this embodiment can achieve the purpose of training a voiceprint recognition model with less training data, avoid the problem of overfitting in the trained voiceprint recognition model, and improve the training effect of the voiceprint recognition model.
[0059] Figure 2 It is a schematic diagram according to the second embodiment of the present disclosure. As Figure 2 shown, the voiceprint recognition method of this embodiment specifically includes the following steps:
[0060] S201. Obtain test data, where the test data includes multiple test audio data and class labels of the multiple test audio data;
[0061] S202. Generate a support set according to the test data;
[0062] S203. Obtain the audio data to be recognized as a query set;
[0063] S204. Input the support set and the query set into the voiceprint recognition model, and determine the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model.
[0064] This embodiment generates a support set according to the obtained test data, uses the obtained audio data to be recognized as a query set, and thus uses the pre-trained voiceprint recognition model to process the input support set and query set, and determines the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model. Since the pre-trained voiceprint recognition model can avoid overfitting, this embodiment can improve the accuracy of voiceprint recognition.
[0065] The test data obtained by this embodiment in S201 includes multiple test audio data, and the class label of each test audio data is used to represent the class to which the test audio data belongs. The class to which the test audio data belongs is specifically a certain user or a certain speaker.
[0066] It should be noted that the number of test audio data obtained by this embodiment in S201 is less than the number of sample audio data obtained, and the class labels of the test audio data are different from the class labels of the obtained sample audio data, that is, the test audio data and its class labels in this embodiment do not appear in the training process.
[0067] After this embodiment executes S201 to obtain test data, it executes S202 to generate a support set according to the obtained test data.
[0068] Specifically, when implementing S202 to generate a support set according to the acquired test data in this embodiment, an optional implementation method that can be adopted is as follows: obtain at least one test category; according to the category labels of the test audio data, extract a fourth preset number of test audio data from the test audio data corresponding to the at least one acquired test category respectively to form a support set.
[0069] When implementing S202 to obtain at least one test category in this embodiment, the categories corresponding to the randomly selected fifth preset number of category labels can be used as the at least one test category according to the category labels of the test audio data.
[0070] After implementing S202 to generate a support set according to the test data in this embodiment, S203 is implemented to obtain the audio data to be recognized as a query set.
[0071] That is to say, when implementing S203 in this embodiment, the audio data to be recognized input by the input end is used as a query set, and then the query set composed of the audio data to be recognized is used to recognize the category corresponding to the audio data to be recognized.
[0072] After implementing S203 to obtain the audio data to be recognized as a query set in this embodiment, S204 is implemented to input the acquired support set and query set into the voiceprint recognition model, and determine the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model.
[0073] Specifically, when implementing S204 to determine the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model in this embodiment, an optional implementation method that can be adopted is as follows: obtain the probability distribution of the audio data to be recognized output by the voiceprint recognition model belonging to each category in the support set, and the process of obtaining the probability distribution is the same as the training process of the neural network model; according to the acquired probability distribution, use the category with the largest corresponding probability value as the voiceprint recognition result of the audio data to be recognized.
[0074] That is to say, this embodiment can use the pre-trained voiceprint recognition model through a small amount of labeled test audio data to determine the voiceprint recognition result of the audio data to be recognized (for example, which speaker the audio data to be recognized corresponds to), which can improve the accuracy of voiceprint recognition with fewer resources and has strong generalization ability for audio categories that do not appear in the training process.
[0075] Figure 3 It is a schematic diagram according to the third embodiment of the present disclosure. As Figure 3 shown, the training device 300 of the voiceprint recognition model in this embodiment includes:
[0076] The first acquisition unit 301 is configured to acquire training data, where the training data includes multiple sample audio data and class labels of the multiple sample audio data;
[0077] The first generation unit 302 is configured to generate a support set and a query set corresponding to different training tasks according to the training data;
[0078] The construction unit 303 is configured to construct a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer. The feature extraction layer is configured to output feature vectors of support audio data included in the support set and feature vectors of query audio data included in the query set. The prototype network layer is configured to output prototype vectors of various categories corresponding to the support audio data included in the support set. The fully connected layer is configured to output a probability distribution of the query audio data included in the query set belonging to various categories in the support set;
[0079] The training unit 304 is configured to train the neural network model using the support set and the query set corresponding to different training tasks to obtain a voiceprint recognition model.
[0080] The training data acquired by the first acquisition unit 301 includes multiple sample audio data, and the class label of each sample audio data is used to represent the class to which the sample audio data belongs. The class to which the sample audio data belongs is specifically a certain user or a certain speaker.
[0081] It can be understood that the sample audio data acquired by the first acquisition unit 301 can be the audio data itself or the spectral features extracted from the audio data. That is to say, the first acquisition unit 301 can use the audio data itself to train the neural network model or use the spectral features extracted from the audio data to train the neural network model.
[0082] If the acquired sample audio data is the audio data itself, after the first acquisition unit 301 acquires the training data, it may further include the following: preprocessing the sample audio data; using a preset window function to extract spectral features from the preprocessed sample audio data.
[0083] The window function used by the first acquisition unit 301 can be a Hamming window, etc. For example, a Hamming window with a preset size (such as 25 ms) is used to shift on the audio data with a preset step size (such as 10 ms), so as to obtain the spectral features of the audio data according to the calculation results of the Hamming window at each shift.
[0084] After the training data including a plurality of sample audio data and the class labels of the plurality of sample audio data is acquired by the first acquisition unit 301 in this embodiment, the first generation unit 302 generates a support set and a query set corresponding to different training tasks according to the acquired training data.
[0085] In the support sets corresponding to different training tasks generated by the first generation unit 302, at least one support audio data and the class label of at least one support audio data are included; in the query sets corresponding to different training tasks generated, at least one query audio data and the class label of at least one query audio data are included.
[0086] In this embodiment, each training task represents one training of the neural network model; the first generation unit 302 generating the support set and the query set corresponding to different training tasks means training the neural network model multiple times according to the training data.
[0087] Specifically, when the first generation unit 302 generates the support set and the query set corresponding to different training tasks according to the acquired training data, an optional implementation method that can be adopted is: for each training task, at least one training class corresponding to the training task is acquired; according to the class labels of the sample audio data, the first preset number of sample audio data are respectively extracted from the sample audio data corresponding to the at least one acquired training class as support audio data to form a support set; from the remaining sample audio data after extracting the first preset number of sample audio data from the sample audio data corresponding to the at least one training class, the second preset number of sample audio data are extracted as query audio data to form a query set.
[0088] When the first generation unit 302 acquires at least one training class corresponding to the training task, the classes corresponding to the randomly selected third preset number of class labels can be used as the at least one training class according to the class labels of the sample audio data.
[0089] After the support set and the query set corresponding to different training tasks are generated by the first generation unit 302 in this embodiment, the construction unit 303 constructs a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer.
[0090] Among them, the feature extraction layer in the neural network model constructed by the construction unit 303 can be any type of feature extraction network, which is used to encode the input sample audio data into a feature vector with a fixed dimension.
[0091] Specifically, the feature extraction layer in the neural network model constructed by the construction unit 303 is composed of a Capsule Network and an Autoencoder.
[0092] The building unit 303 can combine the spatial information in the input spectral features to obtain the feature vectors of the corresponding sample audio data by constructing a feature extraction layer including a capsule network and an autoencoder, thereby improving the accuracy of the feature extraction vectors.
[0093] The prototype network layer in the neural network model constructed by the building unit 303 is used to output the prototype vectors of various categories corresponding to the support audio data included in the support set.
[0094] Specifically, when the prototype network layer in the neural network model constructed by the building unit 303 outputs the prototype vectors of various categories corresponding to the support audio data included in the support set, the optional implementation method that can be adopted is: for each category in the support set, obtain the feature vectors of the support audio data included in this category; according to the obtained feature vectors, obtain the prototype vectors of this category.
[0095] When the prototype network layer constructed by the building unit 303 obtains the prototype vectors of this category according to the obtained feature vectors, it can calculate the mean value of the obtained feature vectors as the prototype vectors of this category, or use the accumulation result of the obtained feature vectors as the prototype vectors of this category.
[0096] The fully connected layer in the neural network model constructed by the building unit 303 is used to output the probability distribution of the query audio data included in the query set belonging to various categories in the support set.
[0097] Specifically, when the fully connected layer in the neural network model constructed by the building unit 303 outputs the probability distribution of the query audio data included in the query set belonging to various categories in the support set, the optional implementation method that can be adopted is: obtain the feature vectors of each query audio data in the query set; for each query audio data, according to the feature vector of this query audio data and the prototype vectors of various categories in the support set, obtain the probability distribution of this query audio data belonging to various categories in the support set.
[0098] In this embodiment, after the neural network layer model including a feature extraction layer, a prototype network layer and a fully connected layer is constructed by the building unit 303, the training unit 304 uses the support set and the query set corresponding to different training tasks to train this neural network model to obtain a voiceprint recognition model.
[0099] Specifically, when the training unit 304 trains the neural network model using the support sets and query sets corresponding to different training tasks to obtain a voiceprint recognition model, an optional implementation method that can be adopted is as follows: input the support sets and query sets corresponding to different training tasks into the neural network model; calculate the loss function value according to the probability distribution of the query audio data included in the query set output by the neural network model for each training task belonging to various categories in the support set; adjust the parameters of the neural network model according to the calculated loss function value until the neural network model converges to obtain a voiceprint recognition model.
[0100] It can be understood that when the training unit 304 outputs the probability that the query audio data included in the corresponding query set for each training task belongs to various categories in the support set, it is processed successively through the feature extraction layer, the prototype network layer, and the fully connected layer. The specific processing process is as described above and will not be elaborated here.
[0101] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 4 shown, the voiceprint recognition device 400 of this embodiment includes:
[0102] A second acquisition unit 401, configured to acquire test data, where the test data includes a plurality of test audio data and class labels of the plurality of test audio data;
[0103] A second generation unit 402, configured to generate a support set according to the test data;
[0104] A processing unit 403, configured to acquire the audio data to be recognized as a query set;
[0105] An identification unit 404, configured to input the support set and the query set into the voiceprint recognition model, and determine the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model.
[0106] The test data acquired by the second acquisition unit 401 includes a plurality of test audio data. The class label of each test audio data is used to characterize the class to which the test audio data belongs. The class to which the test audio data belongs is specifically a certain user or a certain speaker.
[0107] It should be noted that the number of test audio data acquired by the second acquisition unit 401 is less than the number of sample audio data acquired, and the class labels of the test audio data are different from the class labels of the acquired sample audio data, that is, the test audio data and its class labels acquired by the second acquisition unit 401 do not appear during the training process.
[0108] After the second acquisition unit 401 acquires the test data in this embodiment, the second generation unit 402 generates a support set according to the acquired test data.
[0109] Specifically, when the second generation unit 402 generates a support set according to the acquired test data, an optional implementation method that can be adopted is: acquiring at least one test category; respectively extracting a fourth preset number of test audio data from the test audio data corresponding to the acquired at least one test category according to the category label of the test audio data to form a support set.
[0110] When the second generation unit 402 acquires at least one test category, it can use the categories corresponding to the randomly selected fifth preset number of category labels as the at least one test category according to the category label of the test audio data.
[0111] After the second generation unit 402 generates a support set according to the test data in this embodiment, the processing unit 403 acquires the audio data to be recognized as a query set.
[0112] That is to say, the processing unit 403 uses the audio data to be recognized input by the input end as a query set, and then uses the query set composed of the audio data to be recognized to recognize the category corresponding to the audio data to be recognized.
[0113] After the processing unit 403 acquires the audio data to be recognized as a query set in this embodiment, the recognition unit 404 inputs the acquired support set and query set into the voiceprint recognition model, and determines the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model.
[0114] Specifically, when the recognition unit 404 determines the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model, an optional implementation method that can be adopted is: acquiring the probability distribution of the audio data to be recognized output by the voiceprint recognition model belonging to various categories in the support set; according to the acquired probability distribution, using the category with the largest corresponding probability value as the voiceprint recognition result of the audio data to be recognized.
[0115] That is to say, the recognition unit 404 can use the pre-trained voiceprint recognition model through a small amount of labeled test audio data to determine the category of the audio data to be recognized, which can improve the accuracy of voiceprint recognition under less resources, and has strong generalization ability for audio categories that do not appear in the training process.
[0116] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0117] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] As Figure 5 shown, it is a block diagram of an electronic device for training a voiceprint recognition model or a voiceprint recognition method according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0119] As Figure 5 shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0120] Multiple components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0121] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as training a voiceprint recognition model or a voiceprint recognition method. For example, in some embodiments, the training of a voiceprint recognition model or a voiceprint recognition method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 508.
[0122] In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the voiceprint recognition model training or the voiceprint recognition method described above may be executed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the voiceprint recognition model training or the voiceprint recognition method by any other suitable means (e.g., by means of firmware).
[0123] The various implementations of the systems and techniques described herein may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable voiceprint recognition model training or voiceprint recognition device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0125] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0126] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0127] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0128] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0129] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0130] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for a voiceprint recognition model, comprising: Obtain training data, where the training data contains multiple sample audio data and class labels of multiple sample audio data; Generate a support set and a query set corresponding to different training tasks according to the training data; Construct a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer. The feature extraction layer is used to output the feature vectors of the support audio data included in the support set and the feature vectors of the query audio data included in the query set. The feature extraction layer is composed of a capsule network and an autoencoder. The prototype network layer is used to output the prototype vectors of each category corresponding to the support audio data included in the support set. The fully connected layer is used to output the probability distribution of the query audio data included in the query set belonging to each category in the support set; Use the support set and the query set corresponding to different training tasks to train the neural network model to obtain a voiceprint recognition model.
2. The method according to claim 1, wherein, The generating a support set and a query set corresponding to different training tasks according to the training data includes: For each training task, obtain at least one training category corresponding to the training task; According to the class labels of the sample audio data, respectively extract a first preset number of sample audio data from the sample audio data corresponding to the at least one training category as support audio data to form the support set; Extract a second preset number of sample audio data from the remaining sample audio data after extracting the first preset number of sample audio data from the sample audio data corresponding to the at least one training category as query audio data to form the query set.
3. The method according to claim 1, wherein, The method further includes: For each category in the support set, obtain the feature vector of the support audio data included in the category; According to the feature vector, obtain the prototype vector of the category.
4. The method according to claim 1 or 3, wherein, The method further includes: Obtain the feature vector of each query audio data in the query set; For each query audio data, according to the feature vector of the query audio data and the prototype vectors of each category in the support set, obtain the probability distribution of the query audio data belonging to each category in the support set.
5. The method according to claim 1, wherein, The using the support set and the query set corresponding to different training tasks to train the neural network model to obtain a voiceprint recognition model includes: Input the support set and the query set corresponding to different training tasks into the neural network model; Calculate the loss function value according to the probability distribution of the query audio data included in the query set output by the neural network model for each training task belonging to each category in the support set; Adjust the parameters of the neural network model according to the loss function value until the neural network model converges to obtain the voiceprint recognition model.
6. A voiceprint recognition method, comprising: Obtain test data, where the test data contains multiple test audio data and class labels of multiple test audio data; Generate a support set according to the test data; Obtain the audio data to be recognized as the query set; Input the support set and the query set into the voiceprint recognition model, and determine the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model; Among them, the voiceprint recognition model is trained according to the method described in any one of claims 1-5.
7. The method according to claim 6, wherein, The generating of the support set according to the test data includes: Obtaining at least one test category; According to the class labels of the test audio data, respectively extracting a fourth preset number of test audio data from the test audio data corresponding to the at least one test category to form the support set.
8. The method according to claim 6, wherein, The determining of the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model includes: Obtaining the probability distribution of the audio data to be recognized output by the voiceprint recognition model belonging to each category in the support set; According to the probability distribution, taking the category with the largest corresponding probability value as the voiceprint recognition result of the audio data to be recognized.
9. A training device for a voiceprint recognition model, comprising: A first obtaining unit, configured to obtain training data, where the training data includes multiple sample audio data and class labels of the multiple sample audio data; A first generating unit, configured to generate a support set and a query set corresponding to different training tasks according to the training data; A constructing unit, configured to construct a neural network model including a feature extraction layer, a prototype network layer, and a fully connected layer, where the feature extraction layer is configured to output feature vectors of support audio data included in the support set and feature vectors of query audio data included in the query set, the feature extraction layer is composed of a capsule network and an autoencoder, the prototype network layer is configured to output prototype vectors of each category corresponding to the support audio data included in the support set, and the fully connected layer is configured to output the probability distribution of the query audio data included in the query set belonging to each category in the support set; A training unit, configured to train the neural network model using the support set and the query set corresponding to different training tasks to obtain a voiceprint recognition model.
10. The device according to claim 9, wherein, When the first generating unit generates a support set and a query set corresponding to different training tasks according to the training data, it performs: For each training task, obtaining at least one training category corresponding to the training task; According to the class labels of the sample audio data, respectively extracting a first preset number of sample audio data from the sample audio data corresponding to the at least one training category as support audio data to form the support set; From the remaining sample audio data after extracting the first preset number of sample audio data from the sample audio data corresponding to the at least one training category, extracting a second preset number of sample audio data as query audio data to form the query set.
11. The device according to claim 9, wherein, When the prototype network layer constructed by the constructing unit outputs prototype vectors of each category corresponding to the support audio data included in the support set, it performs: For each category in the support set, obtaining the feature vector of the support audio data included in the category; According to the feature vector, obtaining the prototype vector of the category.
12. The device according to claim 9 or 11, wherein, When the fully connected layer constructed by the constructing unit outputs the probability distribution of the query audio data included in the query set belonging to each category in the support set, it performs: Obtaining the feature vector of each query audio data in the query set; For each query audio data, based on the feature vector of the query audio data and the prototype vectors of each category in the support set, obtain the probability distribution of the query audio data belonging to each category in the support set.
13. The device according to claim 9, wherein, When the training unit uses the support set and query set corresponding to different training tasks to train the neural network model to obtain a voiceprint recognition model, it performs: Input the support set and query set corresponding to different training tasks into the neural network model; Calculate the loss function value according to the probability distribution of the query audio data included in the query set output by the neural network model for each training task belonging to each category in the support set; Adjust the parameters of the neural network model according to the loss function value until the neural network model converges to obtain the voiceprint recognition model.
14. A voiceprint recognition device, comprising: A second acquisition unit for acquiring test data, where the test data includes multiple test audio data and category labels of the multiple test audio data; A second generation unit for generating a support set according to the test data; A processing unit for acquiring the audio data to be recognized as a query set; An identification unit for inputting the support set and the query set into the voiceprint recognition model, and determining the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model; Wherein, the voiceprint recognition model is trained according to the device described in any one of claims 9-13.
15. The device according to claim 14, wherein, When the second generation unit generates a support set according to the test data, it performs: Acquire at least one test category; According to the category labels of the test audio data, extract a fourth preset number of test audio data from the test audio data corresponding to the at least one test category respectively to form the support set.
16. The device according to claim 14, wherein, When the identification unit determines the voiceprint recognition result of the audio data to be recognized according to the output result of the voiceprint recognition model, it performs: Obtain the probability distribution of the audio data to be recognized output by the voiceprint recognition model belonging to each category in the support set; According to the probability distribution, use the category with the largest corresponding probability value as the voiceprint recognition result of the audio data to be recognized.
17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method described in any one of claims 1-8.
19. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Prototypical network algorithms for few-shot learning
US10963754B1