Method for training timbre recognition model, related components, and timbre recognition method
Through adversarial training technology, the generator network and classifier network in the tone recognition model are trained, which solves the problem that the timbre of the same subject in the existing technology cannot be recognized in different scenarios, and achieves a high-accurate tone recognition effect.
Patent Information
- Application Number
- CN202211667038.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-12-22
AI Technical Summary
The existing tone recognition model cannot effectively identify the tone of the same subject in different scenarios, resulting in deviations in subject identity confirmation.
Through adversarial training, the generator network in the trained tone model is trained with the introduction of the discriminator model, so that it can extract robust tone embedding features. The classifier network is also trained to ensure that the target tone recognition model can accurately identify the identity of the same subject in different scenarios.
It realizes accurate recognition of the same subject audio in different scenarios, and improves the recognition accuracy.
Smart Images

Figure CN116013267B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a method, device, equipment, and storage medium for training a timbre recognition model, and a timbre recognition method. Background Art
[0002] Existing timbre recognition models generally recognize the timbres corresponding to the audio of the same subject in different scenarios as the timbres of different subjects, and cannot perform cross-scenario recognition, which will lead to deviations in the confirmation of the subject's identity. For example, in current music and karaoke software, the timbre recognition function is widely used in scenarios such as song recommendation and singer identity confirmation. However, for entertainment stars in speaking scenarios such as interviews and acting and singing scenarios, although the timbres are generally the same, the recognized timbres will be different. The main reason is that in the singing scenario, the pitch changes relatively more, and the rhythm, tone, etc. are also different.
[0003] Therefore, how to provide a timbre recognition solution that is not affected by scenarios is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment, and storage medium for training a timbre recognition model, and a timbre recognition model method, so that the trained target timbre recognition model can recognize the subject identities corresponding to the audio of the same subject in different scenarios as the subject, and the recognition accuracy is relatively high. The specific solutions are as follows:
[0005] The first aspect of the present application provides a method for training a timbre recognition model, including:
[0006] Input audio sample one and audio sample two into the timbre recognition model to be trained, so as to use the generator network of the timbre recognition model to be trained to extract features from the input audio sample one and audio sample two, and obtain timbre embedding feature one and timbre embedding feature two; the audio sample one and the audio sample two belong to different scenarios respectively;
[0007] Input the timbre embedding feature one and the timbre embedding feature two into the discriminator model, so as to use the discriminator model to perform scenario judgment on the timbre embedding feature one and the timbre embedding feature two, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model judges the timbre embedding feature one and the timbre embedding feature two as the same scenario;
[0008] Perform backpropagation according to the loss value of the discriminator loss function, and use the generator loss function to perform adversarial training on the generator network until the generator network converges;
[0009] Training the classifier network in the to-be-trained timbre recognition model using the timbre embedding feature one and the timbre embedding feature two until the classifier network converges, to obtain a target timbre recognition model including at least the trained generator network and the trained classifier network.
[0010] Optionally, before inputting the audio sample one and the audio sample two into the to-be-trained timbre recognition model, it further includes:
[0011] Performing scene annotation on the audio sample one and the audio sample two to obtain the audio sample one and the audio sample two carrying scene labels;
[0012] Correspondingly, before performing adversarial training using the loss function, it further includes:
[0013] Determining the discriminator loss function and the generator loss function corresponding to the timbre embedding feature one and the timbre embedding feature two according to the scene labels; wherein, the scene label of the timbre embedding feature is consistent with the scene label of the corresponding audio sample.
[0014] Optionally, the using the generator network of the to-be-trained timbre recognition model to extract features from the input audio sample one and audio sample two includes:
[0015] Using the generator network to extract features from the input audio sample one and audio sample two in a weight-sharing manner; wherein, the generator network is a siamese network.
[0016] Optionally, the generator network includes a TDNN delay layer, an SE residual layer, an attention statistical pooling layer, and a fully connected layer;
[0017] Correspondingly, the using the generator network of the to-be-trained timbre recognition model to extract features from the input audio sample one and audio sample two includes:
[0018] Using the TDNN delay layer to initialize the number of channels of the audio sample one and the audio sample two to a dimension of a fixed size;
[0019] Using the SE residual layer to add multi-scale features to the output of the TDNN delay layer;
[0020] Using the attention statistical pooling layer to probabilize the output of the SE residual layer;
[0021] Using the fully connected layer to perform a full connection on the output of the attention statistical pooling layer to obtain the timbre embedding feature one and the timbre embedding feature two.
[0022] Optionally, before inputting the first audio sample and the second audio sample into the to-be-trained timbre recognition model, it further includes:
[0023] Performing subject identity annotation on the first audio sample and the second audio sample to obtain the first audio sample and the second audio sample carrying subject identity labels.
[0024] Optionally, training the classifier network in the to-be-trained timbre recognition model by using the first timbre embedding feature and the second timbre embedding feature includes:
[0025] Training the classifier network by using the classifier loss function of the classifier network until the subject identity of the audio sample recognized by the classifier network is consistent with the subject identity label corresponding to the audio sample.
[0026] The second aspect of the present application provides a timbre recognition method, including:
[0027] Obtaining an audio to be recognized;
[0028] Inputting the audio to be recognized into a target timbre recognition model, so that the target timbre recognition model uses a generator network to perform feature extraction on the input audio to be recognized to obtain a to-be-recognized timbre embedding feature, and uses a classifier network to perform timbre recognition on the to-be-recognized timbre embedding feature and then outputs the subject identity corresponding to the audio to be recognized; wherein, the target timbre recognition model is obtained based on the foregoing timbre recognition model training method.
[0029] The third aspect of the present application provides a timbre recognition model training device, including:
[0030] A feature extraction module, configured to input a first audio sample and a second audio sample into a to-be-trained timbre recognition model, so as to use the generator network of the to-be-trained timbre recognition model to perform feature extraction on the input first audio sample and second audio sample to obtain a first timbre embedding feature and a second timbre embedding feature; the first audio sample and the second audio sample belong to different scenarios respectively;
[0031] A discriminator model training module, configured to input the first timbre embedding feature and the second timbre embedding feature into a discriminator model, so as to use the discriminator model to perform scene judgment on the first timbre embedding feature and the second timbre embedding feature, and perform adversarial training on the discriminator model by using a discriminator loss function until the discriminator model determines that the first timbre embedding feature and the second timbre embedding feature are in the same scene;
[0032] A generator network training module, configured to perform backpropagation according to the loss value of the discriminator loss function, and perform adversarial training on the generator network by using the generator loss function until the generator network converges;
[0033] A classifier network training module, configured to train the classifier network in the to-be-trained voiceprint recognition model by using the first voiceprint embedding feature and the second voiceprint embedding feature until the classifier network converges, so as to obtain a target voiceprint recognition model including at least the trained generator network and the trained classifier network.
[0034] A fourth aspect of the present application provides an electronic device, which includes a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the foregoing voiceprint recognition model training method.
[0035] A fifth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, the foregoing voiceprint recognition model training method is implemented.
[0036] In this application, first, audio sample one and audio sample two are input into the to-be-trained timbre recognition model, so as to extract features from the input audio sample one and audio sample two by using the generator network of the to-be-trained timbre recognition model, obtaining timbre embedding feature one and timbre embedding feature two; the audio sample one and the audio sample two belong to different scenarios respectively; then, the timbre embedding feature one and the timbre embedding feature two are input into the discriminator model, so as to use the discriminator model to perform scenario judgment on the timbre embedding feature one and the timbre embedding feature two, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model judges the timbre embedding feature one and the timbre embedding feature two as the same scenario; then, backpropagation is performed according to the loss value of the discriminator loss function, and the generator network is subjected to adversarial training by using the generator loss function until the generator network converges; finally, the classifier network in the to-be-trained timbre recognition model is trained by using the timbre embedding feature one and the timbre embedding feature two until the classifier network converges, obtaining a target timbre recognition model including at least the trained generator network and the trained classifier network. It can be seen that in this application, the generator network in the to-be-trained timbre model is trained by means of adversarial training with the introduction of the discriminator model, so that the trained generator network can extract robust timbre embedding features for audio of the same subject in different scenarios. At the same time, the classifier network in the to-be-trained timbre recognition model is trained, so that the trained target timbre recognition model can identify the subject identities corresponding to the audio of the same subject in different scenarios as the subject, and the recognition accuracy is relatively high. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0038] Figure 1 It is a hardware composition framework diagram of a timbre recognition model training method and / or a timbre recognition method provided by this application;
[0039] Figure 2 It is a hardware composition framework diagram of a specific timbre recognition model training method and / or a timbre recognition method provided by this application;
[0040] Figure 3 It is a flowchart of a timbre recognition model training provided by this application;
[0041] Figure 4A flowchart of a specific method for training a timbre recognition model provided by this application;
[0042] Figure 5 A schematic diagram of a specific method for training a timbre recognition model provided by this application;
[0043] Figure 6 A flowchart of a specific method for training a timbre recognition model provided by this application;
[0044] Figure 7 A structural diagram of a specific generator network provided by this application;
[0045] Figure 8 A flowchart of a specific method for extracting features using a generator network provided by this application;
[0046] Figure 9 A flowchart of a timbre recognition method provided by this application;
[0047] Figure 10 A schematic diagram of the timbre recognition process in a specific application scenario provided by this application;
[0048] Figure 11 A structural diagram of a timbre recognition model training device provided by this application. Specific implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] Existing voice recognition models generally identify the voices corresponding to the audio of the same subject in different scenarios as the voices of different subjects, and cannot perform cross-scenario recognition, which may lead to deviations in the confirmation of the subject's identity. For example, in current music and karaoke software, the voice recognition function is widely used in scenarios such as song recommendation and singer identity confirmation. However, for entertainment stars in speaking scenarios such as interviews and acting and singing scenarios, although the voices are generally the same, the recognized voices may be different. The main reason is that in the singing scenario, the pitch changes relatively more, and the rhythm, tone, etc. are also different. To address the above technical deficiencies, this application provides a voice recognition model training solution. By means of adversarial training and introducing a discriminator model, the generator network in the voice model to be trained is trained, so that the trained generator network can extract robust voice embedding features for the audio of the same subject in different scenarios. At the same time, the classifier network in the voice recognition model to be trained is trained, so that the trained target voice recognition model can recognize the subject identities corresponding to the audio of the same subject in different scenarios as the subject, and the recognition accuracy is relatively high. Based on the above voice recognition model training solution, this application also correspondingly provides a voice recognition solution. Please refer to the specific embodiments and details are not described here.
[0051] For ease of understanding, first, the hardware composition framework applicable to the voice recognition model training method and / or voice recognition method corresponding to this application is introduced. It can be seen Figure 1 , where Figure 1 It shows a schematic diagram of the hardware composition framework applicable to a voice recognition model training method and / or voice recognition method of this application.
[0052] From Figure 1 it can be seen that the hardware composition framework may include: an electronic device 10, where the electronic device 10 may include: a processor 11, a memory 12, a communication interface 13, a multimedia component 14, an input / output interface 15, and a communication bus 16. The processor 11, the memory 12, the communication interface 13, the multimedia component 14, and the input / output interface 15 all complete communication with each other through the communication bus 16.
[0053] In the embodiments of this application, the processor 11 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field-programmable gate array, or other programmable logic devices, etc. The processor may call the program stored in the memory 12. Specifically, the processor may execute the operations performed on the computer device side in the following embodiments of the voice recognition model training method and / or voice recognition method.
[0054] The memory 12 is used to store one or more programs. The program may include program code, and the program code includes computer operation instructions, which can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. In the embodiments of the present application, the memory stores at least a program for implementing the following functions:
[0055] Input audio sample one and audio sample two into the to-be-trained timbre recognition model, and use the generator network of the to-be-trained timbre recognition model to extract features from the input audio sample one and audio sample two, obtaining timbre embedding feature one and timbre embedding feature two; the audio sample one and the audio sample two belong to different scenarios respectively;
[0056] Input the timbre embedding feature one and the timbre embedding feature two into the discriminator model, and use the discriminator model to perform scenario judgment on the timbre embedding feature one and the timbre embedding feature two, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model determines that the timbre embedding feature one and the timbre embedding feature two are in the same scenario;
[0057] Perform backpropagation according to the loss value of the discriminator loss function, and use the generator loss function to perform adversarial training on the generator network until the generator network converges;
[0058] Use the timbre embedding feature one and the timbre embedding feature two to train the classifier network in the to-be-trained timbre recognition model until the classifier network converges, obtaining a target timbre recognition model that at least includes the trained generator network and the trained classifier network.
[0059] In a possible implementation, the memory 12 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area may store data created during the use of the electronic device, such as audio sample data, etc.
[0060] The communication interface 13 may be an interface of a communication module, such as an interface of a GSM module.
[0061] The multimedia component 14 may include a screen and an audio component. The screen may be a touch screen, for example. The audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 12 or sent through the communication interface 13. The audio component further includes at least one speaker for outputting audio signals.
[0062] The input / output interface 15 provides an interface between the processor 11 and other interface modules. The other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication interface 13 is used for the electronic device 10 to communicate with other devices in a wired or wireless manner. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them. Accordingly, the communication component 105 may include: a Wi-Fi component, a Bluetooth component, an NFC component.
[0063] The electronic device 10 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the tone recognition model training method and / or the tone recognition method.
[0064] Of course, Figure 1 The structure of the shown electronic device 10 does not limit the computer device in the embodiments of the present application. In practical applications, the electronic device may include more components than Figure 1More or fewer components as shown, or combining certain components.
[0065] Among them, Figure 1 the electronic device 10 in can be a terminal (such as a mobile terminal like a mobile phone, a tablet computer, or a fixed terminal like a PC), a server, or a smart electronic device.
[0066] It can be understood that in the embodiments of the present application, the number of electronic devices is not limited. It can be multiple electronic devices that cooperate to complete the functions of the tone recognition model training method and / or the tone recognition method. In a possible case, please refer to Figure 2 . From Figure 2 it can be seen that the hardware composition framework may include: a first electronic device 101 and a second electronic device 102. The first electronic device 101 and the second electronic device 102 are communicatively connected through a network 103.
[0067] In the embodiments of the present application, the hardware structures of the first electronic device 101 and the second electronic device 102 may refer to Figure 1 the electronic device 10 in. It can be understood that there are two electronic devices 10 in this embodiment, and the two perform data interaction to implement the functions of tone recognition model training and / or tone recognition. Further, in the embodiments of the present application, the form of the network 103 is not limited. For example, the network 103 can be a wireless network (such as WIFI, Bluetooth, etc.) or a wired network.
[0068] Among them, the first electronic device 101 and the second electronic device 102 can be the same type of electronic device. For example, both the first electronic device 101 and the second electronic device 102 are servers; they can also be different types of electronic devices. For example, the first electronic device 101 can be a terminal or a smart electronic device, and the second electronic device 102 can be a server. In another possible case, a server with strong computing power can be used as the second electronic device 102 to improve data processing efficiency and reliability, thereby improving the tone recognition model training efficiency and / or tone recognition efficiency. At the same time, a terminal or a smart electronic device with low cost and wide application range is used as the first electronic device 101 to realize the interaction between the second electronic device 102 and the user.
[0069] It can be understood that the interaction process can be: the terminal collects the audio to be recognized and transmits it to the server, and the server inputs the audio to be recognized into the target tone recognition model, so that the target tone recognition model uses the generator network to extract features from the input audio to be recognized to obtain the embedded features of the tone to be recognized, and uses the classifier network to perform tone recognition on the embedded features of the tone to be recognized and then outputs the main identity corresponding to the audio to be recognized, and outputs the main identity to the terminal.
[0070] Further, to facilitate the user in obtaining the subject identity, the first electronic device 101 may also output the subject identity when receiving the audio to be recognized. The embodiments of the present application do not limit the output form of the first electronic device 101. For example, the subject identity may be output using the display in the multimedia component, or may be output through the voice device in the multimedia component.
[0071] Figure 3 FIG. is a flowchart of a method for training a voiceprint recognition model provided by an embodiment of the present application. Refer to Figure 3 As shown, the method for training the voiceprint recognition model includes:
[0072] S11: Input the first audio sample and the second audio sample into the voiceprint recognition model to be trained, so as to extract features of the input first audio sample and the second audio sample using the generator network of the voiceprint recognition model to be trained, and obtain a first voiceprint embedding feature and a second voiceprint embedding feature; the first audio sample and the second audio sample belong to different scenarios respectively.
[0073] Before model training, training samples need to be obtained. The training samples for training the voiceprint recognition model are audio samples. Here, mainly two audio samples in different scenarios are used for training. For example, in order to recognize voices in a singing scenario and a speaking scenario, the first audio sample and the second audio sample may be samples in the singing scenario and samples in the speaking scenario respectively. In this embodiment, the training samples do not necessarily need to include data of the speaker in different scenarios. The speaker is also the subject. For a speaker, only a sample of the speaker in a certain scenario is required, thereby reducing the difficulty of obtaining samples. That is, it is relatively difficult to obtain data of the same speaker in different scenarios at the same time. For example, if you want the trained voiceprint recognition model to recognize a certain singer in an audio, generally only a sample of the singer when singing or speaking is required.
[0074] In this embodiment, the audio samples are audio of multiple speakers, and for each speaker, only the audio in one scenario is required, and the scenarios corresponding to the audio of multiple speakers are different. After obtaining the audio samples (including the first audio sample and the second audio sample), input the first audio sample and the second audio sample into the voiceprint recognition model to be trained. The voiceprint recognition model to be trained is the voiceprint recognition model before training, including the generator network before training and the classifier network before training. The generator network is denoted as Generator. After the audio samples are input, the generator network of the voiceprint recognition model to be trained extracts features of the input first audio sample and the second audio sample, and obtains a first voiceprint embedding feature and a second voiceprint embedding feature. The voiceprint embedding feature is voiceprint embedding.
[0075] S12: Input the first timbre embedding feature and the second timbre embedding feature into the discriminator model to use the discriminator model to perform scene judgment on the first timbre embedding feature and the second timbre embedding feature, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model determines that the first timbre embedding feature and the second timbre embedding feature are in the same scene.
[0076] In this embodiment, after using the generator network to generate the timbre embeddings of the first audio sample and the second audio sample, the timbre embedding features in different scenes are then input into the discriminator model. The discriminator model is denoted as Discriminator. In this embodiment, the network structure of the discriminator model Discriminator is not limited and can be any discriminator network. The discriminator model performs scene judgment on the first timbre embedding feature and the second timbre embedding feature.
[0077] It should be noted that the discriminator model uses different discriminator loss functions for the first timbre embedding feature and the second timbre embedding feature. Therefore, it is necessary to use the discriminator loss function corresponding to the timbre embedding feature to perform adversarial training on the discriminator model until the discriminator model determines that the first timbre embedding feature and the second timbre embedding feature are in the same scene. For example, for the first audio sample in the singing scene and the second audio sample in the speaking scene input above, the training objective of the above discriminator model is to recognize the timbre embeddings of the singing scene and the speaking scene as the same scene.
[0078] S13: Perform backpropagation according to the loss value of the discriminator loss function, and use the generator loss function to perform adversarial training on the generator network until the generator network converges.
[0079] In this embodiment, the training result of the discriminator model will affect the training of the generator network. After the discriminator model is trained, further perform backpropagation according to the loss value of the discriminator loss function, and use the generator loss function to perform adversarial training on the generator network until the generator network converges. The training objective of the above generator network is to make the discriminator model unable to identify the timbre embedding generated by itself as which scene.
[0080] S14: Use the first timbre embedding feature and the second timbre embedding feature to train the classifier network in the timbre recognition model to be trained until the classifier network converges, and obtain a target timbre recognition model including at least the trained generator network and the trained classifier network.
[0081] In this embodiment, in order to enable the to-be-trained voiceprint recognition model to distinguish the voiceprints of different speakers, that is, identities, the to-be-trained voiceprint recognition model further includes a classifier network, and the classifier network is used to recognize the identity of the speaker. During training, the classifier network in the to-be-trained voiceprint recognition model is trained using the first voiceprint embedding feature and the second voiceprint embedding feature until the classifier network converges. The trained model corresponding to the to-be-trained voiceprint recognition model is the target voiceprint recognition model, and the target voiceprint recognition model at least includes the generator network and the trained classifier network in the target voiceprint recognition model. It can be understood that the generator network in the target voiceprint recognition model is used to extract voiceprint features, and the extracted voiceprint features are not affected by the scene, and the classifier network in the target voiceprint recognition model is used to recognize the identity of the speaker according to the extracted voiceprint features.
[0082] It can be seen that in the embodiment of the present application, the first audio sample and the second audio sample are first input into the to-be-trained voiceprint recognition model, so as to use the generator network of the to-be-trained voiceprint recognition model to extract features from the input first audio sample and the second audio sample, obtaining the first voiceprint embedding feature and the second voiceprint embedding feature; the first audio sample and the second audio sample belong to different scenes respectively; then the first voiceprint embedding feature and the second voiceprint embedding feature are input into the discriminator model, so as to use the discriminator model to perform scene judgment on the first voiceprint embedding feature and the second voiceprint embedding feature, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model determines that the first voiceprint embedding feature and the second voiceprint embedding feature are in the same scene; then, backpropagation is performed according to the loss value of the discriminator loss function, and the generator network is subjected to adversarial training using the generator loss function until the generator network converges; finally, the classifier network in the to-be-trained voiceprint recognition model is trained using the first voiceprint embedding feature and the second voiceprint embedding feature until the classifier network converges, obtaining a target voiceprint recognition model that at least includes the trained generator network and the trained classifier network. In the embodiment of the present application, the generator network in the to-be-trained voice model is trained by means of adversarial training in the case of introducing a discriminator model, so that the trained generator network can extract robust voiceprint embedding features for audio of the same subject in different scenes. At the same time, the classifier network in the to-be-trained voiceprint recognition model is trained, so that the trained target voiceprint recognition model can recognize the subject identities corresponding to the audio of the same subject in different scenes as the subject, and the recognition accuracy is relatively high.
[0083] Figure 4 It is a flowchart of a specific voiceprint recognition model training method provided by the embodiment of the present application. Refer to Figure 4As shown, the timbre recognition model training method includes:
[0084] S21: Perform scene labeling on the audio sample 1 and the audio sample 2 to obtain the audio sample 1 and the audio sample 2 carrying scene labels.
[0085] In this embodiment, after obtaining audio sample 1 and audio sample 2, it is necessary to perform scene labeling on the audio samples in different scenes to obtain the audio sample 1 and the audio sample 2 carrying scene labels. For example, the scene label of audio sample 1 in the singing scene can be set to 0, and the scene label of audio sample 2 in the speaking scene can be set to 1.
[0086] S22: Input the audio sample 1 and the audio sample 2 into the timbre recognition model to be trained, so as to utilize the generator network to extract features of the input audio sample 1 and the audio sample 2 by weight sharing, and obtain the timbre embedding feature 1 and the timbre embedding feature 2.
[0087] In this embodiment, for the specific process of the above step S21, reference can be made to the corresponding contents disclosed in the above embodiments, which will not be repeated here. It should be noted that the generator network of the timbre recognition model to be trained is a twin network. The twin network is also called a Siamese network, and the Siamese in the network is achieved by sharing weights.
[0088] S23: Input the timbre embedding feature one and the timbre embedding feature two into the discriminator model, and determine the discriminator loss function corresponding to the timbre embedding feature one and the timbre embedding feature two according to the scene label, so as to perform adversarial training on the discriminator model using the discriminator loss function corresponding to the timbre embedding feature one and the timbre embedding feature two, until the discriminator model judges the timbre embedding feature one and the timbre embedding feature two as the same scene.
[0089] S24: Back-propagation is performed according to the loss value of the discriminator loss function, and the generator loss function corresponding to the timbre embedding feature one and the timbre embedding feature two is determined according to the scene label, so as to perform adversarial training on the generator network using the generator loss function corresponding to the timbre embedding feature until the generator network converges.
[0090] In this embodiment, different scenarios correspond to different discriminator loss functions. After inputting the first timbre embedding feature and the second timbre embedding feature into the discriminator network, it is necessary to first determine the discriminator loss functions corresponding to the first timbre embedding feature and the second timbre embedding feature according to the scenario label, and then use the discriminator loss functions corresponding to the first timbre embedding feature and the second timbre embedding feature to perform adversarial training on the discriminator model until the discriminator model determines that the first timbre embedding feature and the second timbre embedding feature are in the same scenario.
[0091] Similarly, different scenarios correspond to different generator loss functions. When performing backpropagation according to the loss value of the discriminator loss function, it is necessary to first determine the generator loss functions corresponding to the first timbre embedding feature and the second timbre embedding feature according to the scenario label, and then use the generator loss functions corresponding to the first timbre embedding feature and the second timbre embedding feature to perform adversarial training on the generator network until the generator network converges.
[0092] Suppose the audio samples in this embodiment include a singing scenario and a speaking scenario. The scenario label of the first audio sample in the singing scenario is 0, and the scenario label of the second audio sample in the speaking scenario is 1. The discriminative loss function (Discriminative Loss) is:
[0093]
[0094] When the scenario label is 0, the corresponding discriminator loss function is D(x) 2 , when the scenario label is 1, the corresponding discriminator loss function is (1 - D(x)) 2 .
[0095] The generative loss function (Generative Loss) is:
[0096]
[0097] When the scenario label is 0, the corresponding discriminator loss function is (1 - D(x)) 2 , when the scenario label is 1, the corresponding discriminator loss function is D(x) 2 .
[0098] L D and L G can refer to the Maximum Mean Discrepancy (MMD) loss function, which will not be elaborated in this embodiment. The entire above training process is as Figure 5 shown.
[0099] S25: Train the classifier network in the to-be-trained timbre recognition model by using the first timbre embedding feature and the second timbre embedding feature until the classifier network converges, so as to obtain a target timbre recognition model including at least the trained generator network and the trained classifier network.
[0100] In this embodiment, for the specific process of the above step S26, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0101] Figure 6 It is a flowchart of a specific timbre recognition model training method provided by an embodiment of the present application. Refer to Figure 6 As shown, the timbre recognition model training method includes:
[0102] S31: Perform scene annotation on the first audio sample and the second audio sample and perform main body identity annotation on the first audio sample and the second audio sample to obtain the first audio sample and the second audio sample carrying scene labels and main body identity labels.
[0103] In this embodiment, after obtaining the audio samples, it is necessary to perform scene annotation and main body identity annotation on the audio samples in different scenes at the same time to obtain audio samples carrying scene labels and main body identity labels. For example, the scene label of the first audio sample in the singing scene can be set to 0, the scene label of the second audio sample in the speaking scene can be set to 1, and the main body identity label of the audio sample can be set to A or B or C, etc. There are as many main body identity labels as there are speakers in the audio sample.
[0104] S32: Input the first audio sample and the second audio sample into the to-be-trained timbre recognition model, so as to use the generator network of the to-be-trained timbre recognition model to extract features from the input first audio sample and second audio sample, and obtain a first timbre embedding feature and a second timbre embedding feature.
[0105] In this embodiment, the generator network includes a TDNN time delay layer, an SE residual layer, an attention statistical pooling layer, and a fully connected layer, as Figure 7 shown. Correspondingly, the process of using the generator network of the to-be-trained timbre recognition model to extract features from the input first audio sample and second audio sample specifically includes the following steps (as Figure 8 shown):
[0106] S321: Use the TDNN time delay layer to initialize the number of channels of the first audio sample and the second audio sample to a dimension of a fixed size.
[0107] S322: Use the SE residual layer to add multi-scale features to the output of the TDNN time delay layer.
[0108] S323: Probabilize the output of the SE residual layer using the attention statistical pooling layer.
[0109] S324: Obtain the timbre embedding features by performing a fully connected operation on the output of the attention statistical pooling layer using the fully connected layer.
[0110] In this embodiment, the input of the generator network is 80-dimensional Fbank features with a length of T. The first part is processed by a TDNN delay layer, and the TDNN delay layer initializes the number of channels of the audio sample one and the audio sample two to a dimension of a fixed size. The specific structure of the TDNN delay layer is Conv1D + ReLU + BN. Since it is a one-dimensional convolution, it is equivalent to a TDNN module. The second part is processed by an SE residual layer, and the structure of the SE residual layer is SE-Res2Net. Generally, there are N layers of SE-Res2Net. Res2Net is used to increase multi-scale features, and SENET is a compression and excitation network. The third part is processed by an attention statistical pooling layer, and the attention statistical pooling layer is an Attentive Stat Pooling + BN layer, which is used to perform probabilistic pooling on the output. It should be noted that before the third part, that is, after the second part, a TDNN module can also be connected. Its function is multi-layer feature fusion, so the input here is the concatenation of the outputs of each previous SE-Res2Net module. The last part is processed by a fully connected layer, and the specific structure is FC + BN. The fully connected layer performs a fully connected linear transformation on the output of the attention statistical pooling layer to obtain the timbre embedding feature one and the timbre embedding feature two, and the output dimension is 192.
[0111] S33: Input the timbre embedding feature one and the timbre embedding feature two into the discriminator model to use the discriminator model to perform a scene judgment on the timbre embedding feature one and the timbre embedding feature two, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model determines that the timbre embedding feature one and the timbre embedding feature two are in the same scene.
[0112] S34: Perform backpropagation according to the loss value of the discriminator loss function, and use the generator loss function to perform adversarial training on the generator network until the generator network converges.
[0113] In this embodiment, for the specific processes of the above steps S33 and S34, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0114] S35: Train the classifier network using the classifier loss function of the classifier network until the main identity of the audio sample recognized by the classifier network is consistent with the main identity label corresponding to the audio sample.
[0115] In this embodiment, the classifier network is mainly trained using the classifier loss function of the classifier network until the main identity of the audio sample recognized by the classifier network is consistent with the main identity label corresponding to the audio sample. The classifier network can be AAM-Softmax. The AAM-Softmax layer is used to classify the output, and the number of classifications is the number of speakers. For example, if there are three speakers A, B, and C in the training data, the possible probability output of the trained classifier network may be (A 0.4, B 0.3, C 0.3), indicating that the probability that the speaker in the audio is A is 0.4, the probability that it is B is 0.3, and the probability that it is C is 0.3. Take A with the highest probability as the main identity corresponding to the speaker.
[0116] It can be understood that in the testing and using stage of the model, AAM-Softmax requires that the speaker of the audio appears in the training data. If the speaker of the audio does not appear in the training data, AAM-Softmax cannot be used at this time. In this case, the user can be required to record a registration audio and calculate the timbre embedding. When the user needs to perform identity verification each time, record a verification audio, calculate the timbre embedding, and compare it with the timbre embedding of the registration data of the corresponding user in the database, such as calculating the cosine score CDS between the two timbre embeddings.
[0117] Figure 9 It is a flowchart of a specific timbre recognition method provided by an embodiment of the present application. Refer to Figure 9 As shown, the timbre recognition method includes:
[0118] S41: Obtain the audio to be recognized.
[0119] S42: Input the audio to be recognized into the target timbre recognition model so that the target timbre recognition model uses the generator network to extract features from the input audio to be recognized to obtain the timbre embedding feature to be recognized, and uses the classifier network to perform timbre recognition on the timbre embedding feature to be recognized and then output the main identity corresponding to the audio to be recognized.
[0120] In this embodiment, the audio to be recognized is first obtained, and then the audio to be recognized is input into the target voiceprint recognition model, so that the target voiceprint recognition model uses the generator network to extract features from the input audio to be recognized to obtain the voiceprint embedding features to be recognized, and uses the classifier network to perform voiceprint recognition on the voiceprint embedding features to be recognized and then output the subject identity corresponding to the audio to be recognized; wherein, the target voiceprint recognition model is trained based on the voiceprint recognition model training method as described above.
[0121] It can be seen that in the embodiment of the present application, the audio to be recognized is first obtained, and then the audio to be recognized is input into the target voiceprint recognition model, so that the target voiceprint recognition model uses the generator network to extract features from the input audio to be recognized to obtain the voiceprint embedding features to be recognized, and uses the classifier network to perform voiceprint recognition on the voiceprint embedding features to be recognized and then output the subject identity corresponding to the audio to be recognized; wherein, the target voiceprint recognition model is trained based on the foregoing voiceprint recognition model training method. The present application can recognize the same speaker as the same speaker in different scenarios, and the recognition accuracy is greatly improved.
[0122] For ease of understanding, please refer to Figure 10 , and introduce it in combination with an application scenario of the present solution. Hereinafter, the process of voiceprint recognition is described by taking a terminal and a server as an application scenario.
[0123] After the user turns on the terminal, the terminal establishes a wireless connection with the server. The user inputs audio data in the terminal interface, and after obtaining the audio data, the terminal sends the audio data to the server through the wireless network. After receiving the audio data, the server calls the generator network in the target voiceprint recognition model to extract features from the input audio data to obtain the voiceprint embedding features corresponding to the input audio data, and calls the classifier network in the target voiceprint recognition model to perform voiceprint recognition on the voiceprint embedding features corresponding to the input audio data and then output the subject identity corresponding to the input audio data. After obtaining the subject identity corresponding to the input audio data, the server sends the subject identity to the terminal. When the terminal obtains the subject identity, it outputs the subject identity through the screen.
[0124] Furthermore, it should be noted that the way for the user to input audio data in the terminal interface can be input by using the voice device of the terminal through voice collection, or can be input by file upload in the terminal. Correspondingly, when the terminal outputs the subject identity, it can be output by voice or through the screen of the terminal.
[0125] On the other hand, the present application also provides a voiceprint recognition model training device. For example, refer to Figure 11, which shows a schematic structural diagram of a composition of an embodiment of a tone recognition model training device according to the present application. The device of this embodiment can be applied to the electronic device in the above embodiment. The device includes:
[0126] A feature extraction module 21, configured to input an audio sample one and an audio sample two into a tone recognition model to be trained, so as to use a generator network of the tone recognition model to be trained to extract features from the input audio sample one and the audio sample two, and obtain a tone embedding feature one and a tone embedding feature two; the audio sample one and the audio sample two belong to different scenarios respectively;
[0127] A discriminator model training module 22, configured to input the tone embedding feature one and the tone embedding feature two into a discriminator model, so as to use the discriminator model to perform scene judgment on the tone embedding feature one and the tone embedding feature two, and use a discriminator loss function to perform adversarial training on the discriminator model until the discriminator model determines that the tone embedding feature one and the tone embedding feature two are in the same scene;
[0128] A generator network training module 23, configured to perform backpropagation according to the loss value of the discriminator loss function, and use a generator loss function to perform adversarial training on the generator network until the generator network converges;
[0129] A classifier network training module 24, configured to use the tone embedding feature one and the tone embedding feature two to train a classifier network in the tone recognition model to be trained until the classifier network converges, and obtain a target tone recognition model including at least the trained generator network and the trained classifier network.
[0130] It can be seen that in the embodiment of the present application, the audio sample one and the audio sample two are first input into the to-be-trained timbre recognition model, so as to extract features from the input audio sample one and audio sample two by using the generator network of the to-be-trained timbre recognition model, and obtain timbre embedding feature one and timbre embedding feature two; the audio sample one and the audio sample two belong to different scenarios respectively; then the timbre embedding feature one and the timbre embedding feature two are input into the discriminator model, so as to use the discriminator model to perform scenario judgment on the timbre embedding feature one and the timbre embedding feature two, and use the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model judges the timbre embedding feature one and the timbre embedding feature two as the same scenario; then, backpropagation is performed according to the loss value of the discriminator loss function, and the generator network is subjected to adversarial training by using the generator loss function until the generator network converges; finally, the classifier network in the to-be-trained timbre recognition model is trained by using the timbre embedding feature one and the timbre embedding feature two until the classifier network converges, and a target timbre recognition model including at least the trained generator network and the trained classifier network is obtained. In the embodiment of the present application, the generator network in the to-be-trained timbre model is trained by means of adversarial training in the case of introducing a discriminator model, so that the trained generator network can extract robust timbre embedding features for audio of the same subject in different scenarios. At the same time, the classifier network in the to-be-trained timbre recognition model is trained, so that the trained target timbre recognition model can identify the subject identities corresponding to the audio of the same subject in different scenarios as the subject, and the recognition accuracy is relatively high.
[0131] In some specific embodiments, the feature extraction module 21 is specifically configured to extract features from the input audio sample one and audio sample two by using the generator network in a weight-sharing manner; wherein, the generator network is a siamese network.
[0132] In some specific embodiments, the timbre recognition model training device further includes:
[0133] The first annotation module is configured to perform scenario annotation on the audio sample one and the audio sample two to obtain the audio sample one and the audio sample two carrying scenario labels;
[0134] The second annotation module is configured to perform subject identity annotation on the audio sample one and the audio sample two to obtain the audio sample one and the audio sample two carrying subject identity labels;
[0135] A loss function selection module, configured to determine discriminator loss functions and generator loss functions corresponding to the first timbre embedding feature and the second timbre embedding feature according to the scenario label; wherein, the scenario label of the timbre embedding feature is consistent with the scenario label of the corresponding audio sample.
[0136] In some specific embodiments, the generator network includes a TDNN delay layer, an SE residual layer, an attention statistical pooling layer, and a fully connected layer. The feature extraction module 21 specifically includes:
[0137] A dimension fixing unit, configured to use the TDNN delay layer to initialize the number of channels of the first audio sample and the second audio sample to a dimension of a fixed size;
[0138] A feature scale increasing unit, configured to use the SE residual layer to increase multi-scale features for the output of the TDNN delay layer;
[0139] A probability unit, configured to use the attention statistical pooling layer to probabilize the output of the SE residual layer;
[0140] A fully connected unit, configured to use the fully connected layer to perform a full connection on the output of the attention statistical pooling layer to obtain the timbre embedding feature.
[0141] In some specific embodiments, the classifier network training module 24 is specifically configured to train the classifier network by using the classifier loss function of the classifier network until the main identity of the audio sample recognized by the classifier network is consistent with the main identity label corresponding to the audio sample.
[0142] On the other hand, the present application also provides a timbre recognition device. The device in this embodiment can be applied to the electronic device in the above embodiment. The device includes:
[0143] An acquisition module 31, configured to acquire an audio to be recognized;
[0144] A recognition module 32, configured to input the audio to be recognized into a target timbre recognition model, so that the target timbre recognition model uses the generator network to extract features from the input audio to be recognized to obtain a timbre embedding feature to be recognized, and uses the classifier network to perform timbre recognition on the timbre embedding feature to be recognized and then output the main identity corresponding to the audio to be recognized; wherein, the target timbre recognition model is trained based on the foregoing timbre recognition model training method.
[0145] It can be seen that in the embodiments of the present application, the audio to be recognized is first obtained, and then the audio to be recognized is input into the target voice recognition model, so that the target voice recognition model uses the generator network to extract features from the input audio to be recognized to obtain the embedded features of the voice to be recognized, and uses the classifier network to perform voice recognition on the embedded features of the voice to be recognized and then output the subject identity corresponding to the audio to be recognized; wherein, the target voice recognition model is trained based on the foregoing voice recognition model training method. The present application can recognize the same speaker as the same speaker in different scenarios, and the recognition accuracy is greatly improved.
[0146] On the other hand, the present application also provides an electronic device, which may include a processor and a memory. The relationship between the processor and the memory in the electronic device can refer to Figure 1 .
[0147] Among them, the processor of the electronic device is used to execute the program stored in the memory;
[0148] The memory of the electronic device is used to store a program, and the program is at least used for:
[0149] Obtain the audio to be recognized;
[0150] Input the audio to be recognized into the target voice recognition model, so that the target voice recognition model uses the generator network to extract features from the input audio to be recognized to obtain the embedded features of the voice to be recognized, and uses the classifier network to perform voice recognition on the embedded features of the voice to be recognized and then output the subject identity corresponding to the audio to be recognized; wherein, the target voice recognition model is trained based on the foregoing voice recognition model training method.
[0151] Of course, the electronic device may further include a communication interface, a multimedia component, an input / output interface, etc., which are not specifically limited herein.
[0152] On the other hand, the embodiments of the present application also disclose a storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the steps of the voice recognition model training method and / or the voice recognition method disclosed in any of the foregoing embodiments are implemented.
[0153] It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method part.
[0154] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0155] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0156] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for training a timbre recognition model, characterized in that, it includes: Inputting audio sample one and audio sample two into the timbre recognition model to be trained, and using the generator network of the timbre recognition model to be trained to extract features from the input audio sample one and audio sample two, obtaining timbre embedding feature one and timbre embedding feature two; the audio sample one and the audio sample two belong to different scenarios respectively; Inputting the timbre embedding feature one and the timbre embedding feature two into the discriminator model, and using the discriminator model to perform scenario judgment on the timbre embedding feature one and the timbre embedding feature two, and using the discriminator loss function to perform adversarial training on the discriminator model until the discriminator model judges the timbre embedding feature one and the timbre embedding feature two as the same scenario; Performing backpropagation according to the loss value of the discriminator loss function, and using the generator loss function to perform adversarial training on the generator network until the generator network converges; Using the timbre embedding feature one and the timbre embedding feature two to train the classifier network in the timbre recognition model to be trained until the classifier network converges, obtaining a target timbre recognition model including at least the trained generator network and the trained classifier network.
2. The timbre recognition model training method according to claim 1, characterized in that, before inputting the audio sample one and the audio sample two into the timbre recognition model to be trained, it further includes: Performing scenario annotation on the audio sample one and the audio sample two to obtain the audio sample one and the audio sample two carrying scenario labels; Correspondingly, before performing adversarial training using the loss function, it further includes: Determining the discriminator loss function and the generator loss function corresponding to the timbre embedding feature one and the timbre embedding feature two according to the scenario labels; wherein, the scenario label of the timbre embedding feature is consistent with the scenario label of the corresponding audio sample.
3. The timbre recognition model training method according to claim 1, characterized in that, the using the generator network of the timbre recognition model to be trained to extract features from the input audio sample one and audio sample two includes: Using the generator network to extract features from the input audio sample one and audio sample two in a weight-sharing manner; wherein, the generator network is a siamese network.
4. The timbre recognition model training method according to claim 1, characterized in that, the generator network includes a TDNN time delay layer, an SE residual layer, an attention statistical pooling layer and a fully connected layer; Correspondingly, the using the generator network of the timbre recognition model to be trained to extract features from the input audio sample one and audio sample two includes: Using the TDNN time delay layer to initialize the number of channels of the audio sample one and the audio sample two to a dimension of a fixed size; Using the SE residual layer to add multi-scale features to the output of the TDNN time delay layer; Using the attention statistical pooling layer to probabilize the output of the SE residual layer; The output of the attention statistical pooling layer is fully connected by using the fully connected layer to obtain the first timbre embedding feature and the second timbre embedding feature.
5. The method for training a timbre recognition model according to any one of claims 1 to 4, wherein, before inputting the first audio sample and the second audio sample into the timbre recognition model to be trained, it further includes: performing main body identity annotation on the first audio sample and the second audio sample to obtain the first audio sample and the second audio sample carrying main body identity labels.
6. The method for training a timbre recognition model according to claim 5, wherein, the training of the classifier network in the timbre recognition model to be trained by using the first timbre embedding feature and the second timbre embedding feature includes: training the classifier network by using the classifier loss function of the classifier network until the main body identity of the audio sample recognized by the classifier network is consistent with the main body identity label corresponding to the audio sample.
7. A timbre recognition method, wherein, it includes: obtaining an audio to be recognized; inputting the audio to be recognized into a target timbre recognition model, so that the target timbre recognition model uses a generator network to extract features of the input audio to be recognized to obtain a timbre embedding feature to be recognized, and uses a classifier network to perform timbre recognition on the timbre embedding feature to be recognized and then output the main body identity corresponding to the audio to be recognized; wherein, the target timbre recognition model is trained according to the method for training a timbre recognition model according to any one of claims 1 to 6.
8. A device for training a timbre recognition model, wherein, it includes: a feature extraction module, configured to input a first audio sample and a second audio sample into a timbre recognition model to be trained, so as to use the generator network of the timbre recognition model to be trained to extract features of the input first audio sample and the second audio sample, and obtain a first timbre embedding feature and a second timbre embedding feature; the first audio sample and the second audio sample belong to different scenarios respectively; a discriminator model training module, configured to input the first timbre embedding feature and the second timbre embedding feature into a discriminator model, so as to use the discriminator model to perform scene judgment on the first timbre embedding feature and the second timbre embedding feature, and use a discriminator loss function to perform adversarial training on the discriminator model until the discriminator model determines that the first timbre embedding feature and the second timbre embedding feature are in the same scene; a generator network training module, configured to perform backpropagation according to the loss value of the discriminator loss function, and use a generator loss function to perform adversarial training on the generator network until the generator network converges; a classifier network training module, configured to use the first timbre embedding feature and the second timbre embedding feature to train the classifier network in the timbre recognition model to be trained until the classifier network converges, and obtain a target timbre recognition model including at least the trained generator network and the trained classifier network.
9. An electronic device, wherein, The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the timbre recognition model training method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that it is used to store computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the timbre recognition model training method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice scene recognition method and device, voice control method and equipment and air conditioner
CN109741747A
Scenarized intelligent voice recognition method and system
CN114120972A