Training method for voiceprint recognition model, voiceprint recognition method and related devices
By incorporating computational cost into the loss function and using multiple levels of feature extraction in voice recognition models, the method addresses high computational costs while maintaining accuracy, reducing resource usage.
Patent Information
- Application Number
- CN202210173743.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-24
AI Technical Summary
In the existing voiceprint recognition model, the calculation amount is large, resulting in waste of computing resources and inefficient recognition.
A multi-layer cascaded feature extraction network layer is introduced, and the calculation amount of the feature extraction network layer is considered in the loss function. Through layer-by-layer feature extraction and classification, the model training process is optimized to reduce unnecessary calculations.
On the basis of ensuring the accuracy of voiceprint recognition, the calculation amount of feature extraction is reduced, computing resources is saved, and recognition efficiency is improved.
Smart Images

Figure CN114822562B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method for training a voiceprint recognition model, a voiceprint recognition method, and related devices. Background Art
[0002] With the development of artificial intelligence technology, the application of voiceprint recognition is becoming more and more extensive. In related technologies, voiceprint recognition is generally performed through a voiceprint recognition model. Specifically, the voiceprint recognition model includes a multi-layer feature extraction network layer for feature extraction and a classification layer for classification. In this voiceprint recognition model, only the feature information extracted by the last layer of the feature extraction network layer is input into the classification layer for classification, resulting in a relatively large computational load for voiceprint recognition. Summary of the Invention
[0003] In view of the above problems, the present application provides a method for training a voiceprint recognition model, a voiceprint recognition method, and related devices to improve the above problems.
[0004] According to one aspect of an embodiment of the present application, there is provided a method for training a voiceprint recognition model. The voiceprint recognition model includes a classification layer and a multi-layer cascaded feature extraction network layer. The method includes: obtaining a sample set, where the sample set includes a plurality of sample audio and the voiceprint labels corresponding to each sample audio; for each sample audio, each feature extraction network layer in the voiceprint recognition model performs feature extraction layer by layer based on the sample audio to obtain the feature information output by each feature extraction network layer; the classification layer respectively performs voiceprint classification according to the feature information output by each feature extraction network layer to obtain the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer; according to the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer, the voiceprint labels corresponding to the sample audio, the target computational load corresponding to each feature extraction network layer, and a preset loss function, calculate the loss values of each feature extraction network layer for the sample audio respectively; where the target computational load corresponding to a feature extraction network layer is equal to the sum of the computational load of the feature extraction network layer and the computational load of the feature extraction network layer before the feature extraction network layer; according to the loss values of each feature extraction network layer for the sample audio, calculate the target loss value; and adjust the parameters of the voiceprint recognition model in the reverse direction according to the target loss value until the model training end condition is reached.
[0005] According to one aspect of the embodiments of the present application, a voiceprint recognition method is provided, including: obtaining a target audio to be recognized; using the first-layer feature extraction network layer in the voiceprint recognition model as the target feature extraction network layer, and using the target audio as target information; the voiceprint recognition model is trained according to the training method of the voiceprint recognition model as described above; the target feature extraction network layer extracts features from the target information to obtain feature information output by the target feature extraction network layer; the classification layer in the voiceprint recognition model performs voiceprint classification according to the feature information output by the target feature extraction network layer to obtain the probabilities of the target audio corresponding to each voiceprint category; according to the probabilities of the target audio corresponding to each voiceprint category, determine the maximum probability; if the maximum probability is greater than a set probability threshold, then use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio.
[0006] According to one aspect of the embodiments of the present application, a training device for a voiceprint recognition model is provided, including: a sample acquisition module, configured to obtain a sample set, the sample set including a plurality of sample audios and the voiceprint labels corresponding to each sample audio; a feature extraction module, configured to, for each sample audio, the feature extraction network layers in the voiceprint recognition model sequentially extract features based on the sample audio to obtain the feature information output by each feature extraction network layer; a voiceprint classification module, configured to the classification layer respectively perform voiceprint classification according to the feature information output by each feature extraction network layer to obtain the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer; a loss value calculation module, configured to respectively calculate the loss values of each feature extraction network layer for the sample audio according to the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer, the voiceprint labels corresponding to the sample audio, the target calculation amounts corresponding to each feature extraction network layer, and a preset loss function; wherein, the target calculation amount corresponding to a feature extraction network layer is equal to the sum of the calculation amount of the feature extraction network layer and the calculation amount of the feature extraction network layer before the feature extraction network layer; a target loss value calculation module, configured to calculate a target loss value according to the loss values of each feature extraction network layer for the sample audio; a model adjustment module, configured to reversely adjust the parameters of the voiceprint recognition model according to the target loss value until the model training end condition is reached.
[0007] According to one aspect of the embodiments of the present application, a voiceprint recognition device is provided, including: a target audio acquisition module, configured to acquire a target audio to be recognized; a feature extraction module, configured to use the first-layer feature extraction network layer in the voiceprint recognition model as the target feature extraction network layer and the target audio as target information; the voiceprint recognition model is trained according to the above-mentioned training method of the voiceprint recognition model; a feature output module, configured to perform feature extraction on the target information by the target feature extraction network layer to obtain the feature information output by the target feature extraction network layer; a voiceprint classification module, configured to perform voiceprint classification on the feature information output by the target feature extraction network layer by the classification layer in the voiceprint recognition model to obtain the probabilities of the target audio corresponding to each voiceprint category; a maximum probability determination module, configured to determine the maximum probability according to the probabilities of the target audio corresponding to each voiceprint category; a voiceprint recognition module, configured to, if the maximum probability is greater than a set probability threshold, use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio.
[0008] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the above-mentioned training method of the voiceprint recognition model and the voiceprint recognition method are implemented.
[0009] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the above-mentioned training method of the voiceprint recognition model and the voiceprint recognition method are implemented.
[0010] In the solution of the present application, the target calculation amount corresponding to the feature extraction network layer is introduced into the loss function, and voiceprint classification is performed based on the feature information output by each feature extraction network layer to obtain the sample voiceprint recognition results corresponding to the feature information output by each feature extraction network layer for the sample audio, rather than only performing voiceprint classification on the feature information output by the last-layer feature extraction network layer. On this basis, according to the sample voiceprint recognition results corresponding to the feature information output by each feature extraction network layer for the sample audio, the voiceprint label corresponding to the sample audio, and the target calculation amount corresponding to each feature extraction network layer, the loss value of each feature extraction network layer for the sample audio is calculated. The target calculation amount corresponding to the feature extraction network layer reflects the amount of computing resources required to obtain the feature information output by the feature extraction network layer.
[0011] For voiceprint recognition, the more voiceprint information is reflected in the extracted feature information, the higher the accuracy of the voiceprint classification result obtained by using this feature information for voiceprint classification. Therefore, according to the sample voiceprint recognition result corresponding to the feature information output by the feature extraction network layer for the sample audio and the voiceprint label corresponding to the sample audio, it is possible to calculate how much voiceprint information is reflected in the feature information extracted by this feature extraction network layer, and further reflect the accuracy of the feature information extracted by this feature extraction network layer.
[0012] That is to say, according to the sample voiceprint recognition result corresponding to the feature information output by each feature extraction network layer for the sample audio, the voiceprint label corresponding to the sample audio, and the target calculation amount corresponding to each feature extraction network layer, the loss value of each feature extraction network layer for the sample audio is calculated, which is calculated by taking into account the influence of both the accuracy of voiceprint recognition and the calculation amount required for feature extraction. Thus, by using the target loss value calculated according to the loss value of each feature extraction network layer for the sample audio to inversely adjust the parameters of the voiceprint recognition model, it is possible to make each feature extraction network layer in the trained voiceprint recognition model compromise between ensuring the accuracy of voiceprint recognition and reducing the calculation amount used for feature extraction. Thus, on the basis of ensuring the accuracy of voiceprint recognition, the calculation amount required for feature extraction is reduced, and further, the calculation resources used for voiceprint recognition are reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0014] Figure 1 The structure diagram of the computer system of the electronic device suitable for implementing the embodiments of the present application is shown.
[0015] Figure 2 The flowchart of the training method of the voiceprint recognition model according to an embodiment of the present application is shown.
[0016] Figure 3 It is the flowchart of the steps before step 220 shown according to an embodiment of the present application.
[0017] Figure 4 The flowchart of the voiceprint recognition method according to an embodiment of the present application is shown.
[0018] Figure 5The flowchart of voiceprint recognition performed by the last layer feature extraction network layer according to an embodiment of the present application is shown.
[0019] Figure 6 The block diagram of a training device for a voiceprint recognition model shown according to an embodiment of the present application.
[0020] Figure 7 The block diagram of a voiceprint recognition device shown according to an embodiment of the present application. Detailed implementation manners
[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0022] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0023] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0024] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0025] It should be noted that: "a plurality" mentioned herein means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0026] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0027] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0028] Figure 1 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown, as Figure 1 shown, in the memory 1005 as a storage medium, an operating system, a data storage module, a network communication module, a user interface module, and an electronic program may be included.
[0029] In Figure 1 the shown electronic device, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; in the electronic device of the present invention, the processor 1001 and the memory 1005 may be arranged in the electronic device, and the electronic device calls the training device of the voiceprint recognition model and the voiceprint recognition device stored in the memory 1005 through the processor 1001, and respectively executes the training method and the voiceprint recognition method provided by the embodiments of the present application.
[0030] The processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire electronic device through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling the data stored in the memory 1005, it performs various functions of the electronic device and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and may be implemented separately through a communication chip.
[0031] The memory 1005 may include random access memory (RAM) and may also include read-only memory. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, alarm function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device (such as fake response commands, obtained process states), etc.
[0032] Those skilled in the art can understand that Figure 1 the structure shown in
[0033] Figure 2It is a flowchart of a method for training a voiceprint recognition model shown according to an embodiment of the present application. This method can be executed by a computer device with processing capabilities, such as a server, a cloud server, or other terminal devices with processing capabilities, which are not specifically limited here. The voiceprint recognition model includes a classification layer and a multi-level cascaded feature extraction network layer. Referring to Figure 2 As shown, this method at least includes steps 210 to 260, which are introduced in detail as follows:
[0034] Step 210, obtain a sample set, where the sample set includes multiple sample audios and the voiceprint labels corresponding to each of the sample audios.
[0035] The sample audio can be obtained by collecting audio of a sound-producing body. The voiceprint label corresponding to the sample audio is used to indicate the sound-producing body from which the sample audio is sourced. For example, if sample audio I is obtained by collecting audio of user A1, then the voiceprint label corresponding to sample audio I is used to indicate that the sound-producing body of this sample audio I is user A1.
[0036] Step 220, for each sample audio, each feature extraction network layer in the voiceprint recognition model performs feature extraction layer by layer based on the sample audio to obtain the feature information output by each feature extraction network layer.
[0037] The feature extraction network layer refers to a neural network layer used for feature extraction. This neural network layer can be a convolutional neural network layer, a recurrent neural network layer, a fully connected neural network layer, a long short-term memory neural network layer, a feedforward neural network layer, a pooling neural network layer, etc., which are not specifically limited here.
[0038] In the present application, the voiceprint recognition model includes a multi-level cascaded feature extraction network layer. Among them, different feature extraction network layers can be neural network layers of the same type or different types. For example, each feature extraction network layer in the voiceprint recognition model is a convolutional neural network layer. Another example is that some feature extraction network layers of the voiceprint recognition model are convolutional neural network layers, and some feature extraction network layers are fully connected neural network layers.
[0039] After the sample audio enters the voiceprint recognition model, the first-layer feature extraction network layer extracts features from the sample audio to obtain the feature information of the sample audio. After that, the feature information output by the first-layer feature extraction network layer is used as the input of the second-layer feature extraction network layer. The second-layer feature extraction network layer extracts features from the feature information output by the first-layer feature extraction network layer again to obtain the feature information output by the second-layer feature extraction network layer for the sample audio. After that, the feature information output by the second-layer feature extraction network layer for the sample audio is input into the third-layer feature extraction network layer to continue feature extraction, and so on, to realize the layer-by-layer feature extraction of the sample audio.
[0040] The feature information output by each feature extraction network layer for the sample audio is used to characterize the audio features of the sample audio. Further, this feature information is used to reflect the voiceprint of the sample audio.
[0041] It can be understood that in the voiceprint recognition model, since the feature extraction network layer further extracts features based on the feature information output by the previous feature extraction network layer, the deeper the corresponding layer number of the feature extraction network layer, the more voiceprint information of the sample audio is contained in the feature information output by this feature extraction network layer, and the more accurate the voiceprint information of the sample audio is. However, the feature information output by each feature extraction network layer is obtained by extracting features again on the basis of the feature extraction by the previous feature extraction network layer. Therefore, the deeper the corresponding layer number of the feature extraction network layer, the more computing power is required to obtain the corresponding feature information.
[0042] Step 230: The classification layer respectively performs voiceprint classification according to the feature information output by each feature extraction network layer to obtain the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer.
[0043] That is to say, in this application, the feature information output by each feature extraction network layer for the sample audio needs to be subjected to voiceprint classification, rather than only classifying the feature information output by the last feature extraction network layer in the voiceprint recognition model.
[0044] The classification layer can perform voiceprint classification through a classification function. Among them, the classification function can be a softmax function, a sigmoid function, etc. The classification function, such as the softmax function, classifies the feature information output by each feature extraction network layer to obtain the probabilities corresponding to each voiceprint category of the sample audio, and takes the voiceprint category corresponding to the maximum probability as the sample voiceprint classification result of the sample audio. Among them, one voiceprint category is used to identify a sound-emitting body. For example, if the sound-emitting body is a person, different voiceprint categories are used to indicate different users.
[0045] Step 240: Calculate the loss value of each feature extraction network layer for the sample audio respectively according to the sample voiceprint classification result corresponding to the feature information output by each feature extraction network layer, the voiceprint label corresponding to the sample audio, the target calculation amount corresponding to each feature extraction network layer, and a preset loss function; wherein, the target calculation amount corresponding to a feature extraction network layer is equal to the sum of the calculation amount of the feature extraction network layer and the calculation amounts of the feature extraction network layers before the feature extraction network layer.
[0046] The sample voiceprint classification result refers to the voiceprint classification result obtained by the classification layer for the feature information of the sample audio, and this sample voiceprint classification result indicates the speaking user corresponding to the predicted sample audio.
[0047] Specifically, the target calculation amount corresponding to each feature extraction network layer can be calculated according to the Figure 3 process shown. As Figure 3 shown, it includes:
[0048] Step 102: Obtain the calculation amount information respectively set for each feature extraction network layer from the 1st layer to the i-th layer feature extraction network layers; the calculation amount information is used to indicate the calculation amount required for the corresponding feature extraction network layer to perform feature extraction; wherein, 1 ≤ i ≤ N, i is a positive integer, and N is the total number of feature extraction network layers in the voiceprint recognition model.
[0049] Step 104: Add up the calculation amounts respectively corresponding to all the feature extraction network layers from the 1st layer to the i-th layer feature extraction network layers to obtain the target calculation amount corresponding to the i-th layer feature extraction network layer.
[0050] The calculations performed by each feature extraction network layer in the voiceprint recognition model for feature extraction are generally matrix operations, such as dot product operations, etc., which can be further decomposed into addition, subtraction, multiplication, division, exponentiation, exponential operations, etc. In a specific embodiment, the number of operations of the feature extraction network layer can be used as the calculation amount of the feature extraction network layer. Specifically, the number of operations can be measured by FLOPs (Floating-point Operations), and one floating-point operation can be defined as one multiplication operation and one addition operation. On this basis, according to the type of neurons included in each feature extraction network layer (such as fully connected neurons, convolutional neurons, pooling neurons, etc.), such as the number of neurons included and the calculations performed by the feature extraction network layer, the number of floating-point operations required for the feature extraction network layer to perform feature extraction can be determined.
[0051] It can be understood that different neurons are included in the feature extraction network layer, and the calculations performed by the feature extraction network layer for feature extraction are also different. Correspondingly, the amount of calculations required for the feature extraction network layer to perform feature extraction is also correspondingly different. Further, there are also differences in the number of neurons included in different feature extraction network layers, and the number of neurons determines the dimension of the output of the feature extraction network layer. Therefore, the number of neurons included in the feature extraction layer will also correspondingly affect the amount of calculations required for the feature extraction network layer to perform feature extraction.
[0052] Therefore, in a specific embodiment, the amount of calculations required for the feature extraction network layer to perform feature extraction can be set according to the type of neurons included in the feature extraction network layer and the number of neurons included.
[0053] For example, if a feature extraction network layer is a fully connected neural network layer, the calculation performed by the fully connected neural network layer is:
[0054] y = matmaul(x, W) + b; (Formula 1)
[0055] Where x is the input information of the fully connected neural network layer (i.e., the feature information output by the previous feature extraction network layer); W is the I×J weight matrix of the fully connected neural network layer (where the dimension of the weight matrix is related to the number of fully connected neurons included in the fully connected neural network layer); b is the bias matrix; y is the output information of the fully connected neural network layer (i.e., the feature information output by the fully connected neural network layer). It can be understood that the dimension of x is I, and the dimension of y is J. It can be determined that the number of floating-point operations required for the fully connected neural network layer to perform feature extraction is (2I - 1)×J.
[0056] On the basis of setting the amount of calculations required for each feature extraction network layer to perform feature extraction, for each feature extraction network layer, the corresponding target amount of calculations for each feature extraction network layer can be calculated according to the process of step 102 - step 104 above.
[0057] The target amount of calculations corresponding to a feature extraction network layer indicates the amount of calculations required to obtain the feature information output by the feature extraction network layer.
[0058] In this application, the loss function set for the voiceprint recognition model includes a first parameter term and a second parameter term. Among them, the first parameter term is used to characterize the degree of difference between the sample voiceprint recognition result of the sample audio and the voiceprint label corresponding to the sample audio, and the second parameter term is used to characterize the amount of calculations required for the feature information used to obtain the sample voiceprint recognition result.
[0059] In a specific embodiment, the first parameter item may be a cross-entropy loss function, a squared loss function, an average absolute value loss function, a Huber loss function, a logarithmic loss function, etc., and no specific limitation is made here.
[0060] The second parameter item may be equal to the target calculation amount corresponding to the feature extraction network layer, or may be equal to the value obtained by processing the target calculation amount corresponding to the feature extraction network layer. The processing performed on the target calculation amount corresponding to the feature extraction network layer may be a normalization process, and no specific limitation is made here.
[0061] In summary, if Z represents the function value of the loss function, K1 represents the first parameter vector, and K2 represents the second parameter vector, then it can be expressed as:
[0062] Z = t1 * K1 + t2 * K2; (Formula 2)
[0063] Among them, t1 is the weight coefficient corresponding to the first parameter item, t2 is the weight coefficient corresponding to the second parameter item, and t1 and t2 can be set according to actual needs. If K1 is a cross-entropy loss function, then
[0064] where y i is the voiceprint label corresponding to the i-th sample audio, is the sample voiceprint classification result corresponding to the i-th sample audio.
[0065] If K1 is a squared loss function, then
[0066] If K1 is an average absolute value loss function, then
[0067] In some embodiments, K2 may be equal to the target calculation amount corresponding to the feature extraction network layer, and may also be equal to the ratio of the target calculation amount corresponding to a feature extraction network layer to the total calculation amount, where the total calculation amount is equal to the target calculation amount corresponding to the last feature extraction network layer in the voiceprint recognition model.
[0068] It can be understood that there may be differences in the value ranges of the first parameter item and the second parameter item. In some embodiments, if the difference in the value ranges of the first parameter item and the second parameter item is relatively large. For example, if the value range of the first parameter item is 0 to 1 and the value range of the second parameter item is 50 to 1000. In this case, if the value of the first parameter item obtained according to the voiceprint label corresponding to the sample audio and the sample voiceprint recognition result corresponding to the sample audio and the value of the second parameter item calculated according to the target calculation amount corresponding to the feature extraction network layer are directly added, it can be seen that at this time, the added result is greatly affected by the value of the second parameter item and less affected by the value of the first parameter item, or it can be approximately considered that it is not affected by the value of the first parameter item. Therefore, in order to avoid this situation, the value ranges of the first parameter item and the second parameter item can be further transformed into the same value range or a similar value range. Specifically, the parameter item with a larger value (the first parameter item or the second parameter item) can be divided by the first specified number to transform the value range of this parameter item. For example, if the parameter item with a larger value is the second parameter item, the first specified number can be the total calculation amount mentioned above; or the parameter item with a smaller value can be multiplied by the second specified number to transform the value range of this parameter item.
[0069] The loss value of a feature extraction network layer for the sample audio is equal to substituting the sample voiceprint classification result corresponding to the feature information output by the feature extraction network layer, the voiceprint label corresponding to the sample audio, and the target calculation amount corresponding to the feature extraction network layer into a loss function (such as formula 2 in the above text), and calculating the function value of the loss function.
[0070] Please continue to refer to Figure 2 , step 250, calculate the target loss value according to the loss values of each feature extraction network layer for the sample audio.
[0071] Specifically, step 250 includes: adding up the loss values of all feature extraction network layers in the voiceprint recognition model for the sample audio to obtain the target loss value.
[0072] Step 260: Adjust the parameters of the voiceprint recognition model in the reverse direction according to the target loss value until the model training end condition is reached. Specifically, a loss value range can be set. If the target loss value calculated for the sample audio exceeds this loss value range, then adjust the parameters of the voiceprint recognition model in the reverse direction. Then, based on the voiceprint recognition model with adjusted parameters, perform the above steps 220 - 250 again until the newly calculated target loss value for the sample audio is within the loss value range; conversely, if the target loss value calculated for the sample audio is within this loss value range, then continue to use the next sample audio to continue training this voiceprint recognition model.
[0073] The adjusted parameters of the voiceprint recognition model include at least one of the weight parameters of each feature extraction network layer and the weight parameter of the classification layer in the voiceprint recognition model.
[0074] The model training end condition can be that the number of iterations of the voiceprint recognition model reaches the set number threshold, or that the accuracy of voiceprint recognition of the voiceprint recognition model reaches the accuracy threshold. Of course, it can also be other conditions, which are not specifically limited here.
[0075] As described above, in the voiceprint recognition model, the deeper the depth of the feature extraction network layer, the more voiceprint information the output feature information represents, and the higher the accuracy of using the output feature information for voiceprint feature classification. However, the more computational resources are required to obtain the corresponding output feature information. In the related art, in order to ensure the accuracy of voiceprint recognition, generally only the feature information output by the last feature extraction network layer in the voiceprint recognition model is input to the classification layer for voiceprint classification. In this way, since all feature extraction network layers are required to perform feature extraction each time, the computational cost of voiceprint recognition using this method is relatively large.
[0076] In the solution of the present application, the target computational amount corresponding to the feature extraction network layer is introduced into the loss function, and voiceprint classification is performed based on the feature information output by each feature extraction network layer to obtain the sample voiceprint recognition results corresponding to the feature information output by each feature extraction network layer for the sample audio, rather than only performing voiceprint classification on the feature information output by the last feature extraction network layer. On this basis, according to the sample voiceprint recognition results corresponding to the feature information output by each feature extraction network layer for the sample audio, the voiceprint label corresponding to the sample audio, and the target computational amount corresponding to each feature extraction network layer, calculate the loss value of each feature extraction network layer for the sample audio. The target computational amount corresponding to the feature extraction network layer reflects the amount of computational resources required to obtain the feature information output by this feature extraction network layer.
[0077] For voiceprint recognition, the more voiceprint information reflected by the extracted feature information, the higher the accuracy of the voiceprint classification result obtained by using this feature information for voiceprint classification. Therefore, according to the sample voiceprint recognition result corresponding to the feature information output by the feature extraction network layer for the sample audio and the voiceprint label corresponding to the sample audio, it is possible to calculate the amount of voiceprint information reflected by the feature information extracted by this feature extraction network layer, and further reflect the accuracy of the feature information extracted by this feature extraction network layer.
[0078] That is to say, according to the sample voiceprint recognition result corresponding to the feature information output by each feature extraction network layer for the sample audio, the voiceprint label corresponding to the sample audio, and the target calculation amount corresponding to each feature extraction network layer, the loss value of each feature extraction network layer for the sample audio is calculated, which is calculated by taking into account the influence of both the accuracy of voiceprint recognition and the calculation amount required for feature extraction. Thus, by using the target loss value calculated based on the loss value of each feature extraction network layer for the sample audio to inversely adjust the parameters of the voiceprint recognition model, it is possible to make each feature extraction network layer in the trained voiceprint recognition model compromise between ensuring the accuracy of voiceprint recognition and reducing the calculation amount used for feature extraction. Thus, on the basis of ensuring the accuracy of voiceprint recognition, the calculation amount required for feature extraction is reduced, and further, the calculation resources used for voiceprint recognition are reduced.
[0079] Figure 4 It is a flowchart of a voiceprint recognition method shown in an embodiment of the present application. This method can be executed by a computer device with processing capabilities, such as a server, a cloud server, etc., which is not specifically limited here. Refer to Figure 4 As shown, this method at least includes steps 310 to 360, which are introduced in detail as follows:
[0080] Step 310, obtain the target audio to be recognized.
[0081] The target audio generally refers to any audio to be used for voiceprint recognition.
[0082] Step 320, use the first-layer feature extraction network layer in the voiceprint recognition model as the target feature extraction network layer, and use the target audio as the target information; the voiceprint recognition model is trained according to any embodiment of the above-mentioned voiceprint recognition model training method.
[0083] Step 330, the target feature extraction network layer extracts features from the target information to obtain the feature information output by the target feature extraction network layer.
[0084] Step 340: The classification layer in the voiceprint recognition model performs voiceprint classification based on the feature information output by the target feature extraction network layer to obtain the probabilities of the target audio corresponding to each voiceprint category.
[0085] Step 350: Determine the maximum probability according to the probabilities of the target audio corresponding to each voiceprint category.
[0086] Step 360: If the maximum probability is greater than a set probability threshold, then use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio.
[0087] As described above, each feature extraction network layer in the voiceprint recognition model trained according to the voiceprint recognition model training method provided in this application can make a trade-off between accurately extracting voiceprint features and reducing the amount of calculation. On this basis, during the specific application process of the voiceprint recognition model, starting from the first-layer feature extraction network layer of the voiceprint recognition model, the feature information output by this first-layer feature extraction network layer for the target audio is input into the classification layer for voiceprint classification. If the maximum probability among the probabilities of the target audio corresponding to each voiceprint category determined in the classification layer is greater than the set probability threshold, it indicates that the probability that the target audio at this time corresponds to the voiceprint category corresponding to the maximum probability is relatively high. On this basis, use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio, and there is no need for the feature extraction network layer after the first-layer feature extraction network layer in the voiceprint recognition model to continue feature extraction. Therefore, compared with the prior art in which only the feature information output by the last-layer feature extraction network layer in the voiceprint recognition model is used for voiceprint classification, the voiceprint recognition method of this application can greatly reduce the amount of calculation for feature extraction, thereby reducing the computing resources used for voiceprint recognition and saving computing resources.
[0088] In some embodiments of this application, after step 350, the method further includes: if the maximum probability is not greater than the probability threshold, then use the next-layer feature extraction network layer of the target feature extraction network layer as the new target feature extraction network layer, use the feature information output by the target feature extraction network layer as the new target information, and return to execute the step of performing feature extraction on the new target information by the target feature extraction network layer to obtain the feature information output by the target feature extraction network layer until the newly obtained maximum probability is greater than the probability threshold, or the new target feature extraction network layer is the last-layer feature extraction network layer in the voiceprint recognition model.
[0089] That is to say, on the basis of performing voiceprint classification on the feature information output by the first-layer feature extraction network layer, if the maximum probability is not greater than the probability threshold, the feature information output by the first-layer feature network layer is continuously input into the second-layer feature extraction network layer for further feature extraction, and then the feature information output by the second-layer feature extraction network layer is subjected to voiceprint classification to obtain the probabilities of the target audio corresponding to each voiceprint category, and the maximum probability is determined. If the maximum probability is greater than the probability threshold, the voiceprint category corresponding to the maximum probability is used as the voiceprint recognition result of the target audio; otherwise, if the maximum probability is not greater than the probability threshold, the feature information output by the second-layer feature extraction network layer is input into the third feature extraction network layer for further feature extraction, and so on, until the maximum probability determined by performing voiceprint classification on the feature information output by a certain feature extraction network layer is greater than the probability threshold, or the target feature extraction network layer outputting the current feature information is the penultimate layer feature extraction network in the voiceprint recognition model (that is, the next feature extraction network layer of the target feature extraction network layer is the last layer feature extraction network layer in the voiceprint recognition model). In some embodiments of the present application, when the new target feature extraction network layer is the last layer feature extraction network layer in the voiceprint recognition model, refer to Figure 5 As shown, the voiceprint recognition method further includes:
[0090] Step 372, obtaining the feature information output by the last layer feature extraction network layer in the voiceprint recognition model.
[0091] Step 374, performing voiceprint classification on the feature information output by the last layer feature extraction network layer by the classification layer to obtain the candidate probabilities of the target audio corresponding to each voiceprint category.
[0092] In the present application, for the sake of distinction, the probabilities of the target audio corresponding to each voiceprint category obtained by performing voiceprint classification on the feature information output by the last layer feature extraction network layer by the classification layer are called candidate probabilities.
[0093] Step 376, determining the maximum candidate probability according to the candidate probabilities of the target audio corresponding to each voiceprint category.
[0094] The maximum candidate probability is the maximum value among the candidate probabilities of the target audio corresponding to each voiceprint category.
[0095] Step 378, using the voiceprint category corresponding to the maximum candidate probability as the voiceprint recognition result of the target audio.
[0096] In this embodiment, if the maximum probabilities corresponding to the feature information output by the feature extraction network layer before the last feature extraction network layer of the voiceprint recognition model for voiceprint classification are all less than the probability threshold, after obtaining the feature information output by the last feature extraction network layer and performing voiceprint classification based on the feature information output by the last feature extraction network layer, the voiceprint category corresponding to the maximum candidate probability is directly used as the voiceprint recognition result of the target audio.
[0097] The following introduces the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.
[0098] Figure 6 is a block diagram of a training device for a voiceprint recognition model shown according to an embodiment of the present application, as Figure 6 shown. The training device for the voiceprint recognition model includes:
[0099] A sample acquisition module 410; configured to acquire a sample set, where the sample set includes a plurality of sample audio and voiceprint labels corresponding to each of the sample audio;
[0100] A feature extraction module 420, configured to, for each sample audio, perform feature extraction layer by layer on the sample audio by each feature extraction network layer in the voiceprint recognition model to obtain the feature information output by each feature extraction network layer;
[0101] A voiceprint classification module 430, configured to perform voiceprint classification by the classification layer respectively according to the feature information output by each feature extraction network layer to obtain the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer;
[0102] A loss value calculation module 440, configured to calculate the loss value of each feature extraction network layer for the sample audio respectively according to the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer, the voiceprint label corresponding to the sample audio, the target calculation amount corresponding to each feature extraction network layer, and a preset loss function; wherein, the target calculation amount corresponding to a feature extraction network layer is equal to the sum of the calculation amount of the feature extraction network layer and the calculation amount of the feature extraction network layer before the feature extraction network layer;
[0103] A target loss value calculation module 450, configured to calculate a target loss value according to the loss value of each feature extraction network layer for the sample audio;
[0104] A model adjustment module 460, configured to reversely adjust the parameters of the voiceprint recognition model according to the target loss value until the model training end condition is reached.
[0105] In some embodiments of the present application, the target loss value calculation module 450 includes a target loss value acquisition module, which is configured to add up the loss values of all the feature extraction network layers in the voiceprint recognition model for the sample audio to obtain the target loss value.
[0106] In some embodiments of the present application, the training device of the voiceprint recognition model further includes a computation amount acquisition module, which is configured to acquire the computation amount information respectively set for each feature extraction network layer from the first layer to the i-th layer of the feature extraction network layers; the computation amount information is used to indicate the computation amount required for the corresponding feature extraction network layer to perform feature extraction; where 1 ≤ i ≤ N, i is a positive integer, and N is the total number of feature extraction network layers in the voiceprint recognition model; a target computation amount calculation module, which is configured to add up the computation amounts respectively corresponding to all the feature extraction network layers from the first layer to the i-th layer of the feature extraction network layers to obtain the target computation amount corresponding to the i-th layer of the feature extraction network layer.
[0107] It should be noted that each module in the training device of the voiceprint recognition model in this embodiment corresponds one-to-one to each step in the training method of the voiceprint recognition model in the foregoing embodiment. Therefore, the specific implementation manners of this embodiment can refer to the implementation manners of the foregoing training method of the voiceprint recognition model, and will not be elaborated here.
[0108] It should be understood that the above is only for illustration and does not constitute any limitation to the technical solution of the present application. Those skilled in the art can set it based on needs in actual applications, and no limitation is made here.
[0109] Figure 7 is a block diagram of a voiceprint recognition device shown according to an embodiment of the present application, as Figure 7 shown, the voiceprint recognition device includes:
[0110] A target audio acquisition module 510; configured to acquire a target audio to be recognized;
[0111] A feature extraction module 520, which is configured to use the first-layer feature extraction network layer in the voiceprint recognition model as the target feature extraction network layer and the target audio as the target information; the voiceprint recognition model is trained according to the training method of the voiceprint recognition model in any of the foregoing embodiments;
[0112] A feature output module 530, which is configured to perform feature extraction on the target information by the target feature extraction network layer to obtain the feature information output by the target feature extraction network layer;
[0113] A voiceprint classification module 540, configured to perform voiceprint classification on the basis of the feature information output by the target feature extraction network layer by a classification layer in the voiceprint recognition model, so as to obtain probabilities of the target audio corresponding to each voiceprint category;
[0114] A maximum probability determination module 550, configured to determine a maximum probability according to the probabilities of the target audio corresponding to each voiceprint category;
[0115] A voiceprint recognition module 560, configured to, if the maximum probability is greater than a set probability threshold, use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio.
[0116] It should be noted that in this embodiment, each module in the voiceprint recognition device corresponds one by one to each step in the voiceprint recognition method in the foregoing embodiment. Therefore, the specific implementation manners of this embodiment may refer to the implementation manners of the foregoing voiceprint recognition method, and will not be elaborated here.
[0117] In some embodiments of the present application, the voiceprint recognition device further includes a maximum probability judgment module, configured to, if the maximum probability is not greater than the probability threshold, use the next feature extraction network layer of the target feature extraction network layer as a new target feature extraction network layer, use the feature information output by the target feature extraction network layer as new target information, and return to execute the step of performing feature extraction on the new target information by the target feature extraction network layer to obtain the feature information output by the target feature extraction network layer, until the newly obtained maximum probability is greater than the probability threshold, or the new target feature extraction network layer is the last feature extraction network layer in the voiceprint recognition model.
[0118] In some embodiments of the present application, the new target feature extraction network layer is the last feature extraction network layer in the voiceprint recognition model, and the voiceprint recognition device further includes a last layer feature information acquisition module, configured to acquire the feature information output by the last feature extraction network layer in the voiceprint recognition model; a candidate probability acquisition module, configured to perform voiceprint classification on the feature information output by the last feature extraction network layer by the classification layer to obtain candidate probabilities of the target audio corresponding to each voiceprint category; a maximum candidate probability determination module, configured to determine a maximum candidate probability according to the candidate probabilities of the target audio corresponding to each voiceprint category; and a voiceprint recognition result output module, configured to use the voiceprint category corresponding to the maximum candidate probability as the voiceprint recognition result of the target audio.
[0119] It should be understood that the above is only an example for illustration, and does not impose any limitation on the technical solutions of the present application. Those skilled in the art can make settings based on needs in actual applications, and no limitation is imposed here.
[0120] The present application also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the methods in any of the above method embodiments are implemented.
[0121] The computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for computer-readable instructions for performing any method steps in the above methods. These computer-readable instructions may be read out from or written into one or more computer program products. The computer-readable instructions may be compressed in a suitable form, for example.
[0122] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods in any of the above embodiments.
[0123] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.
[0124] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of the present application.
[0125] Other embodiments of the present application will be readily contemplated by those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0126] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A training method for a voiceprint recognition model, characterized in that, The voiceprint recognition model includes a classification layer and a multi-level cascaded feature extraction network layer. The method includes: Obtain a sample set, where the sample set includes multiple sample audios and voiceprint labels corresponding to each of the sample audios; For each sample audio, each feature extraction network layer in the voiceprint recognition model performs feature extraction layer by layer based on the sample audio to obtain the feature information output by each feature extraction network layer; The classification layer performs voiceprint classification respectively according to the feature information output by each feature extraction network layer to obtain the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer; According to the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer, the voiceprint label corresponding to the sample audio, the target calculation amount corresponding to each feature extraction network layer, and a preset loss function, calculate the loss value of each feature extraction network layer for the sample audio respectively; wherein, the target calculation amount corresponding to a feature extraction network layer is equal to the sum of the calculation amount of the feature extraction network layer and the calculation amount of the feature extraction network layer before the feature extraction network layer; Calculate the target loss value according to the loss values of each feature extraction network layer for the sample audio; According to the target loss value, adjust the parameters of the voiceprint recognition model in reverse until the model training end condition is reached.
2. The method according to claim 1, wherein The calculating the target loss value according to the loss values of each feature extraction network layer for the sample audio includes: Adding up the loss values of all feature extraction network layers in the voiceprint recognition model for the sample audio to obtain the target loss value.
3. The method according to claim 1, characterized in that Before calculating the loss values of each feature extraction network layer for the sample audio respectively according to the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer, the voiceprint label corresponding to the sample audio, the target calculation amount corresponding to each feature extraction network layer, and a preset loss function, the method further includes: Obtain the calculation amount information set for each feature extraction network layer in the feature extraction network layers from the first layer to the i-th layer; the calculation amount information is used to indicate the calculation amount required for the corresponding feature extraction network layer to perform feature extraction; wherein, 1≤i≤N, i is a positive integer, and N is the total number of feature extraction network layers in the voiceprint recognition model; Add up the calculation amounts corresponding to all feature extraction network layers in the feature extraction network layers from the first layer to the i-th layer to obtain the target calculation amount corresponding to the i-th layer feature extraction network layer.
4. A voiceprint recognition method, characterized in that, Includes: Obtain the target audio to be recognized; Take the first-layer feature extraction network layer in the voiceprint recognition model as the target feature extraction network layer, and take the target audio as the target information; The voiceprint recognition model is trained according to the method described in any one of claims 1-3; The target feature extraction network layer performs feature extraction on the target information to obtain the feature information output by the target feature extraction network layer; The classification layer in the voiceprint recognition model performs voiceprint classification according to the feature information output by the target feature extraction network layer to obtain the probabilities of the target audio corresponding to each voiceprint category; Determine the maximum probability according to the probabilities of the target audio corresponding to each voiceprint category. If the maximum probability is greater than a set probability threshold, then use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio.
5. The method according to claim 4, characterized in that, After determining the maximum probability according to the probabilities of the target audio corresponding to each voiceprint category, the method further includes: If the maximum probability is not greater than the probability threshold, then use the next feature extraction network layer of the target feature extraction network layer as the new target feature extraction network layer, use the feature information output by the target feature extraction network layer as the new target information, and return to execute the step of performing feature extraction on the new target information by the target feature extraction network layer to obtain the feature information output by the target feature extraction network layer, until the newly obtained maximum probability is greater than the probability threshold, or the new target feature extraction network layer is the last feature extraction network layer in the voiceprint recognition model.
6. The method according to claim 5, wherein If the new target feature extraction network layer is the last feature extraction network layer in the voiceprint recognition model, the method further includes: Obtain the feature information output by the last feature extraction network layer in the voiceprint recognition model. Perform voiceprint classification on the feature information output by the last feature extraction network layer by the classification layer to obtain the candidate probabilities of the target audio corresponding to each voiceprint category. Determine the maximum candidate probability according to the candidate probabilities of the target audio corresponding to each voiceprint category. Use the voiceprint category corresponding to the maximum candidate probability as the voiceprint recognition result of the target audio.
7. A training device for a voiceprint recognition model, characterized in that, The voiceprint recognition model includes a classification layer and multiple cascaded feature extraction network layers, and includes: A sample acquisition module, configured to acquire a sample set, where the sample set includes multiple sample audios and voiceprint labels corresponding to each sample audio. A feature extraction module, configured to, for each sample audio, perform layer-by-layer feature extraction on the sample audio by each feature extraction network layer in the voiceprint recognition model to obtain the feature information output by each feature extraction network layer. A voiceprint classification module, configured to perform voiceprint classification on the feature information output by each feature extraction network layer by the classification layer respectively to obtain the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer. A loss value calculation module, configured to calculate the loss value of each feature extraction network layer for the sample audio respectively according to the sample voiceprint classification results corresponding to the feature information output by each feature extraction network layer, the voiceprint label corresponding to the sample audio, the target calculation amount corresponding to each feature extraction network layer, and a preset loss function; wherein, the target calculation amount corresponding to a feature extraction network layer is equal to the sum of the calculation amount of the feature extraction network layer and the calculation amount of the feature extraction network layer before the feature extraction network layer. A target loss value calculation module, configured to calculate the target loss value according to the loss values of each feature extraction network layer for the sample audio. A model adjustment module, configured to adjust the parameters of the voiceprint recognition model in reverse according to the target loss value until the model training end condition is reached.
8. A voiceprint recognition device, characterized in that, Includes: A target audio acquisition module, configured to acquire a target audio to be recognized; A feature extraction module, configured to use the first-layer feature extraction network layer in the voiceprint recognition model as the target feature extraction network layer, and use the target audio as target information; The voiceprint recognition model is trained by the method according to any one of claims 1-3; A feature output module, configured to extract features from the target information by the target feature extraction network layer to obtain feature information output by the target feature extraction network layer; A voiceprint classification module, configured to perform voiceprint classification on the basis of the feature information output by the target feature extraction network layer by a classification layer in the voiceprint recognition model to obtain probabilities of the target audio corresponding to each voiceprint category; A maximum probability determination module, configured to determine a maximum probability according to the probabilities of the target audio corresponding to each voiceprint category; A voiceprint recognition module, configured to, if the maximum probability is greater than a set probability threshold, use the voiceprint category corresponding to the maximum probability as the voiceprint recognition result of the target audio.
9. An electronic device, characterized in that, Including: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.
10. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Acoustic model training method and device, computer equipment and storage medium
CN111128137A
Training calculation amount calculation method and device of neural network model and medium
CN111814978A