Training Method for Human Voice Character Recognition Model, Human Voice Character Recognition Method and Device
By training the first network and the second network, the human voice symbol sequence is directly recognized from the accompaniment audio, which solves the problem of high computational complexity in the prior art and realizes efficient human voice symbol recognition.
Patent Information
- Application Number
- CN202280004816.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-16
AI Technical Summary
In the prior art, the recognition of song people's voice signs needs to be processed based on the vocal accompaniment separation algorithm, and the calculation complexity is high.
By obtaining the annotated human voice audio, pure human voice audio and accompaniment audio, the first and second networks are trained to obtain the human voice symbol recognition model, and the human voice symbol sequence is directly identified from the target audio with accompaniment, avoiding the call of the human voice accompaniment separation algorithm.
The computational complexity of human voice sign recognition is reduced, and a small number of labeled samples are used to train models with strong generalization performance through semi-supervised training methods, which reduces the acquisition cost of training samples.
Smart Images

Figure CN116034425B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a training method for a human voice note recognition model, a human voice note recognition method, and a human voice note recognition device. Background Art
[0002] The vocal note recognition of a song refers to obtaining the vocal note sequence of the song based on the song with accompaniment.
[0003] In addition to vocals, songs often include accompaniment from various instruments. Some live performances also contain background noise or reverberation, which poses significant challenges for recognizing vocal notes. Related technologies use a vocal accompaniment separation algorithm to separate the vocal audio from the song. This is then processed using a vocal note recognition model to obtain the song's vocal note sequence.
[0004] However, the above method requires vocal note recognition based on the vocal accompaniment separation algorithm, which has high computational complexity. Summary of the Invention
[0005] The present invention provides a method for training a human voice note recognition model, a method for human voice note recognition, and a device for human voice note recognition. The technical solution is as follows:
[0006] According to one aspect of an embodiment of the present application, a method for training a human voice note recognition model is provided, the method comprising:
[0007] Obtain at least one annotated vocal audio, vocal note annotation results corresponding to each of the annotated vocal audio, at least one pure vocal audio, and at least one accompaniment audio;
[0008] Training a first network based on the labeled vocal audio, the accompaniment audio, and the vocal note labeling results corresponding to the labeled vocal audio to obtain a trained first network; the first network is configured to output a vocal note recognition result corresponding to the labeled vocal audio based on a synthesized audio of the labeled vocal audio and the accompaniment audio;
[0009] Based on the trained first network, the pure human voice audio and the accompaniment audio, a second network is trained to obtain a human voice note recognition model; the second network is used to output a human voice note recognition result corresponding to the pure human voice audio based on the synthesized audio of the pure human voice audio and the accompaniment audio.
[0010] According to one aspect of an embodiment of the present application, a method for recognizing human voice notes is provided, the method comprising:
[0011] Acquire target audio with accompaniment, where the target audio includes human voice and accompaniment;
[0012] Obtain the audio features of the target audio, where the audio features include features related to the target audio in the time-frequency domain;
[0013] Process the audio features through a human voice note recognition model to obtain the note features of the target audio, where the note features include features related to the human voice notes of the target audio;
[0014] Process the note features through the human voice note recognition model to obtain the human voice note sequence of the target audio;
[0015] Wherein, the human voice note recognition model is obtained by training a second network based on a trained first network, pure human voice audio, and accompaniment audio; the first network is used to output the human voice note recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio; the second network is used to output the human voice note recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
[0016] According to one aspect of the embodiments of the present application, there is provided a training device for a human voice note recognition model, the device includes:
[0017] A sample acquisition module, configured to acquire a first training sample set, a second training sample set, and a third training sample set, where the first training sample set includes at least one labeled human voice audio and the human voice note annotation result corresponding to the labeled human voice audio, the second training sample set includes at least one pure human voice audio, and the third training sample set includes at least one accompaniment audio;
[0018] A first network training module, configured to train a first network based on the labeled human voice audio, the accompaniment audio, and the human voice note annotation result corresponding to the labeled human voice audio to obtain a trained first network; the first network is used to output the human voice note recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio;
[0019] A second network training module, configured to train a second network based on the trained first network, the pure human voice audio, and the accompaniment audio to obtain a human voice note recognition model; the second network is used to output the human voice note recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
[0020] According to one aspect of the embodiments of the present application, there is provided a human voice note recognition device, the device includes:
[0021] An audio acquisition module for acquiring a target audio with accompaniment, where the target audio includes human voices and accompaniment;
[0022] A feature acquisition module for acquiring audio features of the target audio, where the audio features include features related to the target audio in the time-frequency domain;
[0023] A feature extraction module for processing the audio features through a human voice note recognition model to obtain note features of the target audio, where the note features include features related to the human voice notes of the target audio;
[0024] A result obtaining module for processing the note features through the human voice note recognition model to obtain a human voice note sequence of the target audio;
[0025] Wherein, the human voice note recognition model is obtained by training a second network based on a trained first network, a pure human voice audio, and an accompaniment audio; the first network is used to output a human voice note recognition result corresponding to the annotated human voice audio according to a synthesized audio of the annotated human voice audio and the accompaniment audio; the second network is used to output a human voice note recognition result corresponding to the pure human voice audio according to a synthesized audio of the pure human voice audio and the accompaniment audio.
[0026] According to one aspect of an embodiment of the present application, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory, and the processor executes the computer program to implement the training method of the above human voice note recognition model, or to implement the above human voice note recognition method.
[0027] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided. A computer program is stored in the storage medium, and the computer program is used to be executed by a processor to implement the training method of the above human voice note recognition model, or to implement the above human voice note recognition method.
[0028] According to one aspect of an embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor reads and executes the computer instructions from the computer-readable storage medium to implement the training method of the above human voice note recognition model, or to implement the above human voice note recognition method.
[0029] The technical solution provided by the embodiment of the present application may include the following beneficial effects:
[0030] The human voice symbol recognition model obtained through the above training method can directly recognize the corresponding human voice symbol sequence from the target audio with accompaniment. Therefore, in the model usage stage, there is no need to call a human voice and accompaniment separation algorithm to extract the human voice audio from the target audio, reducing the computational complexity of human voice symbol recognition. In addition, this application adopts a semi-supervised training method, training the first network with a small number of labeled samples, and then training the second network with the first network and a large number of unlabeled samples. In this way, only a small number of labeled samples are required to train a model with strong generalization performance, reducing the acquisition cost of training samples. Brief Description of the Drawings
[0031] Figure 1 is a schematic diagram of the implementation environment of the solution provided by an embodiment of this application;
[0032] Figure 2 is a flowchart of the training method of the human voice symbol recognition model provided by an embodiment of this application;
[0033] Figure 3 is a flowchart of the training method of the human voice symbol recognition model provided by another embodiment of this application;
[0034] Figure 4 is a flowchart of the training method of the human voice symbol recognition model provided by another embodiment of this application;
[0035] Figure 5 is a schematic diagram of the training method of the human voice symbol recognition model provided by an embodiment of this application;
[0036] Figure 6 is a flowchart of the human voice symbol recognition method provided by an embodiment of this application;
[0037] Figure 7 is a schematic diagram of the human voice symbol recognition model provided by an embodiment of this application;
[0038] Figure 8 is a block diagram of the training device of the human voice symbol recognition model provided by an embodiment of this application;
[0039] Figure 9 is a block diagram of the training device of the human voice symbol recognition model provided by another embodiment of this application;
[0040] Figure 10 is a block diagram of the human voice symbol recognition device provided by an embodiment of this application;
[0041] Figure 11 is a schematic diagram of the structure of the computer device provided by an embodiment of this application. Detailed Description of the Embodiments
[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0043] Please refer to Figure 1 , which shows a schematic diagram of the solution implementation environment provided by an embodiment of this application. The solution implementation environment may include: a model usage device 10 and a model training device 20.
[0044] The model usage device 10 is used to execute the human voice symbol recognition method in the embodiments of this application. The model usage device 10 may be a terminal device 11 or a server 12. The terminal device 11 may be an electronic device such as a mobile phone, a tablet computer, a game console, an e-book reader, a multimedia playback device, a wearable device, a PC (Personal Computer), or a vehicle-mounted terminal. A target application or a client of the target application may be run in the terminal device 11. In the embodiments of this application, the above-mentioned target application refers to an application that provides the human voice symbol recognition function. Optionally, the target application may be a system-level application, such as an operating system or a native application provided by the operating system; or it may be a third-party application, such as a third-party application downloaded and installed by the user himself / herself. The embodiments of this application do not make any limitations in this regard.
[0045] The server 12 may be the background server of the above-mentioned target application, and is used to provide background services for the target application in the terminal device 11. The server 12 may be a single server, a server cluster composed of multiple servers, or a cloud computing service center. Optionally, the server 12 provides background services for the target applications in multiple terminal devices 11 at the same time.
[0046] The terminal device 11 and the server 12 may communicate with each other through the network 13. The network 13 may be a wired network or a wireless network.
[0047] For the human voice symbol recognition method provided in the embodiments of this application, the execution subject of each step may be a computer device, and the computer device refers to an electronic device with data calculation, processing, and storage capabilities. For example, the human voice symbol recognition method may be executed by the terminal device 11 (such as the client of the target application installed and running in the terminal device 11 executes the human voice symbol recognition method), or may be executed by the server 12, or may be executed by the interaction and cooperation of the terminal device 11 and the server 12. This application does not make any limitations in this regard. For example, the terminal device 11 obtains a target audio and sends the target audio to the server 12, and the server 12 executes the human voice symbol recognition method to obtain a human voice symbol sequence.
[0048] The model training device 20 is used to execute the training method for the vocal note recognition model in the embodiments of the present application. The model training device 20 can be a server or a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. The model training device 20 trains the vocal note recognition model, and the trained vocal note recognition model is deployed in the model utilization device 10.
[0049] Please refer to Figure 2 , which shows a flow chart of a method for training a human voice note recognition model provided by an embodiment of the present application. The method may include at least one of the following steps 210 to 230.
[0050] Step 210 , obtaining at least one annotated vocal audio, vocal note annotating results corresponding to each annotated vocal audio, at least one pure vocal audio, and at least one accompaniment audio.
[0051] In some embodiments, a first training sample set, a second training sample set, and a third training sample set can be obtained. The first training sample set includes at least one labeled vocal audio and the vocal note labeling results corresponding to the labeled vocal audio. The second training sample set includes at least one pure vocal audio. The third training sample set includes at least one accompaniment audio.
[0052] Vocals refer to the parts of a song that are sung by the human voice, such as lyrics and harmony. Non-vocals refer to the parts of a song other than the human voice, such as accompaniment, reverb, noise, etc.
[0053] The labeled vocal audio refers to a cappella singing audio, in which the vocal notes corresponding to each audio frame contained in the audio are labeled. The vocal note labeling result corresponding to the labeled vocal audio refers to a vocal note sequence consisting of the vocal notes corresponding to each audio frame contained in the labeled vocal audio.
[0054] Pure vocal audio refers to the audio containing only vocals separated from song audio with accompaniment.
[0055] Accompaniment audio refers to the audio containing only the accompaniment obtained by separating the audio of the song with accompaniment.
[0056] In some embodiments, a vocal accompaniment separation algorithm can be used to separate pure vocal audio and accompaniment audio from songs with accompaniment. By performing the above separation operation on multiple songs, multiple pure vocal audios can be obtained to construct the second training sample set, and multiple accompaniment audios can be obtained to construct the third training sample set.
[0057] In some embodiments, the number of labeled human voice audio in the first training sample set is much less than the number of pure human voice audio in the second training sample set. Exemplarily, the first training sample set contains 100 pieces of labeled human voice audio, and the second training sample set contains 10,000 pieces of pure human voice audio.
[0058] The present application does not limit the number of accompaniment audio in the third training sample set. For example, the number of accompaniment audio in the third training sample set may be the same as or different from the number of pure human voice audio in the second training sample set.
[0059] Step 220: Train the first network based on the labeled human voice audio, the accompaniment audio, and the human voice symbol annotation result corresponding to the labeled human voice audio to obtain the trained first network; the first network is used to output the human voice symbol recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio.
[0060] The first network refers to an initialized human voice symbol recognition model. In some embodiments, the first network may also be referred to as a teacher network, and the second network may also be referred to as a student network.
[0061] In some embodiments, synthesize the accompaniment audio and the labeled human voice audio to obtain the synthesized audio corresponding to the labeled human voice audio; train the first network based on the synthesized audio corresponding to the labeled human voice audio and the human voice symbol annotation result corresponding to the labeled human voice audio to obtain the trained first network.
[0062] In some embodiments, the synthesized audio corresponding to the labeled human voice audio includes the accompaniment audio and the labeled human voice audio.
[0063] In some embodiments, process the synthesized audio corresponding to the labeled human voice audio through the first network to obtain the human voice symbol recognition result corresponding to the labeled human voice audio as the first human voice symbol recognition result; train the first network according to the first human voice symbol recognition result and the human voice symbol annotation result to obtain the trained first network.
[0064] The first human voice symbol recognition result refers to the human voice symbol sequence of the pure human voice audio obtained through the first network. By inputting the synthesized audio corresponding to the labeled human voice audio into the first network, the first network processes the synthesized audio corresponding to the labeled human voice audio and outputs the first human voice symbol recognition result corresponding to the labeled human voice audio. In some embodiments, train the first network according to the loss function to obtain the trained first network. The present application does not limit the specific loss function. Exemplarily, a cross-entropy loss function, an exponential loss function, a log logarithmic loss function, an absolute value loss function, a Focal-Loss loss function, etc. can be adopted.
[0065] In some embodiments, the parameters of the first network are adjusted by calculating the loss function value between the first voice note recognition result and the voice note labeling result to obtain the trained first network.
[0066] In some embodiments, the first network is trained by adjusting parameters of the first network by calculating a loss function value between the first voice note recognition result and the voice note labeling result.
[0067] In some embodiments, the first network includes an input layer, an intermediate layer, and an output layer. The input layer is configured to input audio features of synthesized audio corresponding to the labeled vocal audio; the intermediate layer is configured to extract note features of the synthesized audio corresponding to the labeled vocal audio based on the audio features; and the output layer is configured to obtain a vocal note sequence of the synthesized audio corresponding to the labeled vocal audio based on the note features.
[0068] In some embodiments, the input layer obtains the audio features of the synthesized audio corresponding to the labeled human voice audio based on the synthesized audio corresponding to the labeled human voice audio, and transmits the audio features to the middle layer.
[0069] In some embodiments, the input layer directly obtains the audio features of the synthesized audio corresponding to the labeled human voice audio and transmits them to the middle layer.
[0070] In some embodiments, the output layer is also used to identify the vocal and non-vocal parts of the note features.
[0071] In some embodiments, the first network is trained based on the vocal part of the note feature, the first vocal note recognition result, and the vocal note labeling result to obtain a trained first network.
[0072] In some embodiments, the first network is a neural network, and this application does not limit the specific network structure.
[0073] Step 230 : Based on the trained first network, the pure vocal audio, and the accompaniment audio, a second network is trained to obtain a vocal note recognition model; the second network is configured to output a vocal note recognition result corresponding to the pure vocal audio based on the synthesized audio of the pure vocal audio and the accompaniment audio.
[0074] In some embodiments, the second network is trained based on the trained first network, pure human voice audio, and accompaniment audio.
[0075] The second network refers to an initialized human voice note recognition model. In some embodiments, the second network is a neural network, and this application does not limit the specific network structure.
[0076] In some embodiments, the second network and the first network are two networks with the same structure and the same initialization parameters.
[0077] In some embodiments, the trained first network processes the pure human voice audio to obtain the human voice symbol recognition result corresponding to the pure human voice audio as the second human voice symbol recognition result; the second human voice symbol recognition result is determined as the pseudo-label information corresponding to the pure human voice audio; the second network is trained according to the pure human voice audio, the accompaniment audio, and the pseudo-label information corresponding to the pure human voice audio.
[0078] In some embodiments, the second human voice symbol recognition result can be directly determined as the pseudo-label information. The solution is simple and easy to implement, and the calculation cost is low.
[0079] In some embodiments, the second human voice symbol recognition result is corrected, and the corrected human voice symbol sequence is determined as the pseudo-label information. Correcting the second human voice symbol recognition result improves the accuracy of the pseudo-label information and further improves the accuracy of the human voice symbol recognition model obtained after training.
[0080] In some embodiments, the accompaniment audio and the pure human voice audio are synthesized to obtain the synthesized audio corresponding to the pure human voice audio; the second network is trained according to the synthesized audio corresponding to the pure human voice audio and the pseudo-label information.
[0081] In some embodiments, the synthesized audio corresponding to the pure human voice audio includes the accompaniment audio and the pure human voice audio.
[0082] In some embodiments, the second network processes the synthesized audio corresponding to the pure human voice audio to obtain the human voice symbol recognition result corresponding to the pure human voice audio as the third human voice symbol recognition result; the second network is trained according to the third human voice symbol recognition result and the pseudo-label information. The third human voice symbol recognition result refers to the human voice symbol sequence of the pure human voice audio obtained through the second network. The synthesized audio corresponding to the pure human voice audio is input into the second network, and the second network processes the synthesized audio corresponding to the pure human voice audio and outputs the third human voice symbol recognition result.
[0083] In some embodiments, the second network is trained according to the loss function. The specific loss function is not limited in this application. Exemplarily, a cross-entropy loss function, an exponential loss function, a log loss function, an absolute value loss function, a Focal-Loss loss function, etc. can be adopted.
[0084] In some embodiments, by calculating the loss function value between the third human voice symbol recognition result and the pseudo-label information, the parameters of the second network are adjusted to obtain the human voice symbol recognition model.
[0085] In some embodiments, the parameters of the second network are adjusted by calculating the loss function value between the third recognition result of human voice notes and the pseudo-label information, and the second network is trained.
[0086] In some embodiments, the second network includes an input layer, an intermediate layer, and an output layer. The input layer is used to input the audio features of the synthesized audio corresponding to the pure human voice audio; the intermediate layer is used to extract the note features of the synthesized audio corresponding to the pure human voice audio according to the audio features; the output layer is used to obtain the human voice note sequence of the synthesized audio corresponding to the pure human voice audio according to the note features.
[0087] In some embodiments, the output layer is further used to identify the human voice part and the non-human voice part of the note features.
[0088] In some embodiments, the input layer is used to obtain the audio features of the synthesized audio corresponding to the pure human voice audio according to the synthesized audio corresponding to the pure human voice audio, and transmit them to the intermediate layer.
[0089] In some embodiments, the input layer is used to directly obtain the audio features of the synthesized audio corresponding to the pure human voice audio, and transmit them to the intermediate layer.
[0090] In some embodiments, the second network is trained according to the human voice part of the note features, the second recognition result of human voice notes, and the pseudo-label information.
[0091] In some embodiments, the loss function for training the first network and the loss function for training the second network may be the same or different, and the present application does not limit this. Exemplarily, the loss function for training the first network and the loss function for training the second network are both cross-entropy loss functions. Exemplarily, the loss function for training the first network is a cross-entropy loss function, and the loss function for training the second network is an absolute value loss function.
[0092] The human voice note sequence refers to a note sequence representing the pitch range of the human voice, which includes the starting point, offset point, and pitch value of different pitch ranges. The offset point refers to the end point of the pitch range, which can be represented by its offset relative to the starting point, so it is called the offset point. Pitch refers to various sounds with different pitch levels, that is, the height of the sound, which is one of the basic characteristics of the sound. The pitch range refers to a section of audio range with the same pitch.
[0093] In some embodiments, the human voice note sequence is a MIDI (Musical Instrument Digital Interface) sequence.
[0094] In some embodiments, the training stopping condition is that the second network converges, that is, the second recognition result of the human voice note corresponding to the pure human voice audio obtained by the second network is infinitely close to the pseudo label information corresponding to the pure human voice audio.
[0095] In some embodiments, whether the second network meets the stop training condition is determined based on the loss function. For example, the stop training condition of the second network is that the loss function value reaches a minimum value.
[0096] In some embodiments, the training stop condition can be set as the number of iterations, and the training stop condition is satisfied when the set number of iterations is reached. The number of iterations can be calculated based on the number of executions of step 230.
[0097] In some embodiments, as Figure 3 As shown, the method further includes step 232, determining whether the second network meets the stop training condition; if so, determining the trained second network as the vocal note recognition model; if not, determining the trained second network as the trained first network, and executing the above step 230 again. That is, if the second network does not meet the stop training condition, the trained second network is determined as the trained first network, and the step of training the second network based on the trained first network, the pure vocal audio, and the accompaniment audio (step 230) is executed again.
[0098] Exemplarily, the second network meets the training stop condition after the nth training. For the i-th training among the n trainings, the second network after the i-1th training is determined as the first network for the i-th training, and the step of training the second network based on the trained first network, pure human voice audio and accompaniment audio (step 230) is started again, where n is an integer greater than 2 and i is an integer greater than 1.
[0099] The technical solution provided by the embodiments of the present application, using the vocal note recognition model obtained through the above-mentioned training method, can directly identify the corresponding vocal note sequence from the target audio with accompaniment. Therefore, during the model usage phase, there is no need to call the vocal accompaniment separation algorithm to extract the vocal audio from the target audio, thereby reducing the computational complexity of vocal note recognition. In addition, the present application adopts a semi-supervised training method, training the first network with a small number of labeled samples, and then training the second network with the first network and a large number of unlabeled samples. In this way, only a small number of labeled samples are required to train a model with strong generalization performance, reducing the cost of obtaining training samples.
[0100] Please refer to Figure 4 , which shows a flow chart of a method for training a human voice note recognition model provided by another embodiment of the present application. The method may include at least one of the following steps 410 to 440.
[0101] Step 410: Obtain at least one annotated vocal audio, vocal note annotation results corresponding to each of the annotated vocal audio, at least one pure vocal audio, and at least one accompaniment audio.
[0102] In some embodiments, a cappella data set and a song data set are obtained, wherein the cappella data set includes at least one a cappella audio and a vocal note labeling result corresponding to the a cappella audio, and the song data set includes at least one song audio with accompaniment.
[0103] A cappella audio refers to vocal audio performed without accompaniment. The vocal note annotation result corresponding to the a cappella audio is a vocal note sequence consisting of the vocal notes corresponding to each audio frame contained in the a cappella audio.
[0104] Song audio refers to the audio combined with lyrics and accompaniment, which includes accompaniment and vocals. In certain embodiments, song audio also includes noise and reverberation.
[0105] In some embodiments, based on the a cappella audio and the vocal note labeling results corresponding to the a cappella audio, labeled vocal audio and the vocal note labeling results corresponding to the labeled vocal audio are generated to construct a first training sample set.
[0106] In some embodiments, a cappella audio is detected to obtain silent parts and unvoiced parts in the a cappella audio; the a cappella audio is determined as labeled vocal audio; and from the vocal note labeling results corresponding to the a cappella audio, the vocal note labeling results corresponding to the silent parts and the vocal note labeling results corresponding to the unvoiced parts are deleted to generate vocal note labeling results corresponding to the labeled vocal audio, thereby constructing a first training sample set.
[0107] In some embodiments, the a cappella audio is detected using a human voice detection algorithm to obtain silent parts and unvoiced parts in the a cappella audio.
[0108] By adopting the above method, it is ensured that the vocal note annotation result corresponding to the a cappella audio only has pitch in the vocal part, and the silent part and the unvoiced part have no pitch, thereby ensuring the accuracy of the vocal note annotation result corresponding to the a cappella audio.
[0109] In some embodiments, a vocal separation operation is performed on the song audio to obtain vocal audio and accompaniment audio; based on the vocal audio, pure vocal audio is generated to construct a second training sample set; based on the accompaniment audio, a third training sample set is constructed.
[0110] This application does not limit the specific method of performing vocal separation on song audio. For example, a vocal separation operation is performed on the song audio using a vocal accompaniment separation algorithm to obtain vocal audio and accompaniment audio.
[0111] In some embodiments, the human voice audio is detected to obtain the non-human voice part in the human voice audio; the non-human voice part in the human voice audio is deleted to generate a pure human voice audio; and a second training sample set is constructed based on the pure human voice audio.
[0112] In some embodiments, the human voice audio is detected by a human voice detection algorithm to obtain the non-human voice part in the human voice audio, and the non-human voice part in the human voice audio is deleted to generate a pure human voice audio. Exemplarily, the human voice audio is detected by a human voice detection algorithm to obtain the non-human voice part in the human voice audio, and the non-human voice part in the human voice audio that exceeds 3 seconds is deleted to generate a pure human voice audio. In a general song, the human voice only occupies a part of it, and the number of training samples in the second training sample set required for training is large. Deleting the non-human voice part in the human voice audio can improve the training efficiency and save the storage space required for the second training sample set.
[0113] In some embodiments, all the obtained pure human voice audios are used to construct a second training sample set.
[0114] Since the human voice and accompaniment separation algorithm cannot ensure perfect separation of the human voice and accompaniment of each song, it is necessary to clean the pure human voice audio and remove the pure human voice audio with remaining accompaniment.
[0115] In some embodiments, for each audio frame in the pure human voice audio, it is detected whether the audio frame is a human voice audio frame, and the energy of the audio frame is calculated; if the audio frame is not a human voice audio frame and the energy of the audio frame is less than a second threshold, the audio frame is determined to be an invalid frame; if the proportion of the number of invalid frames in the pure human voice audio in the total number of audio frames included in the pure human voice audio is greater than a third threshold, the pure human voice audio is determined to be an invalid pure human voice audio; and a pure human voice audio is generated based on the pure human voice audio other than the invalid pure human voice audio.
[0116] In some embodiments, the specific values of the second threshold and the third threshold can be set according to actual needs, and this application does not make a limitation. Exemplarily, for songs of different styles, the value of the second threshold can be different. For example, the second threshold for rock songs is higher than the second threshold for ancient style songs.
[0117] Exemplarily, the value of the third threshold is set to 30%. If the proportion of the number of invalid frames in the pure human voice audio in the total number of audio frames included in the pure human voice audio is greater than 30%, the pure human voice audio is determined to be an invalid pure human voice audio.
[0118] In some embodiments, all the obtained pure human voice audios other than the invalid pure human voice audio are used to generate a pure human voice audio.
[0119] Step 420: Synthesize the accompaniment audio and the labeled human voice audio to obtain a synthesized audio corresponding to the labeled human voice audio.
[0120] In some embodiments, an accompaniment audio is randomly selected from at least one accompaniment audio as a target accompaniment audio; the labeled human voice audio is subjected to data augmentation processing to obtain the processed labeled human voice audio; wherein, the data augmentation processing includes at least one of the following: adding reverberation and changing the fundamental frequency; the target accompaniment audio and the processed labeled human voice audio are synthesized to obtain a synthesized audio corresponding to the labeled human voice audio.
[0121] In some embodiments, an accompaniment audio is randomly selected from the third training sample set as a target accompaniment audio.
[0122] When sound waves encounter obstacles during propagation, they will be reflected by the obstacles, and each reflection will absorb some by the obstacles. In this way, after the sound source stops emitting sound, the sound waves will still be reflected and absorbed multiple times before finally disappearing. We then feel that there are several sound waves mixed and lasting for a period of time after the sound source stops emitting sound. This phenomenon is called reverberation. Adding reverberation to the labeled human voice audio can change the sound quality of the labeled human voice audio.
[0123] Changing the fundamental frequency means changing the fundamental frequency of the labeled human voice audio within a certain range, as well as the human voice symbol annotation result corresponding to the labeled human voice audio. The application does not limit the range of changing the fundamental frequency. Exemplarily, the fundamental frequency of the labeled human voice audio is changed within the range of -200 to +300 cents, and the human voice symbol annotation result corresponding to the labeled human voice audio is adjusted to the corresponding pitch. For example, the fundamental frequency of the labeled human voice audio is raised by 200 cents, and the pitch of the human voice symbol annotation result corresponding to the labeled human voice audio is also raised by 200 cents.
[0124] In some embodiments, the fundamental frequency of any one or more audio frames included in the labeled human voice audio can be changed, as well as the pitch of the human voice symbol annotation result corresponding to the one or more audio frames.
[0125] Step 430: Based on the synthesized audio corresponding to the labeled human voice audio and the human voice symbol annotation result corresponding to the labeled human voice audio, the first network is trained to obtain the trained first network.
[0126] In some embodiments, the synthesized audio corresponding to the labeled human voice audio is processed by the first network to obtain a human voice symbol recognition result corresponding to the labeled human voice audio as the first human voice symbol recognition result; according to the first human voice symbol recognition result and the human voice symbol annotation result, the loss function value of the first network is determined; according to the loss function value of the first network, the parameters of the first network are adjusted to obtain the trained first network.
[0127] In some embodiments, the first network is trained using a cross-entropy loss function.
[0128] In some embodiments, the first network is trained based on the synthesized audio corresponding to the labeled human voice audio and the human voice symbol annotation result until convergence, and the trained first network is obtained.
[0129] Step 440, based on the trained first network, the pure human voice audio, and the accompaniment audio, the second network is trained to obtain a human voice symbol recognition model.
[0130] In some embodiments, the trained first network is used to process the pure human voice audio to obtain the human voice symbol recognition result corresponding to the pure human voice audio, which is used as the second human voice symbol recognition result; the second human voice symbol recognition result is determined as the pseudo-label information corresponding to the pure human voice audio; and the second network is trained according to the pure human voice audio, the accompaniment audio, and the pseudo-label information.
[0131] In some embodiments, the fundamental frequency of the pure human voice audio is extracted; according to the fundamental frequency of the pure human voice audio, the second human voice symbol recognition result is corrected to obtain the pseudo-label information corresponding to the pure human voice audio.
[0132] In some embodiments, the fundamental frequency of the pure human voice audio is extracted by a fundamental frequency extraction algorithm.
[0133] In some embodiments, for each note included in the second human voice symbol recognition result, the pitch difference between the note and the fundamental frequency of the corresponding pronunciation position of the note is calculated; if the pitch difference is greater than the first threshold, the pitch of the note is corrected to the pitch of the fundamental frequency of the corresponding pronunciation position of the note; if the pitch difference is less than or equal to the first threshold, the pitch of the note remains unchanged.
[0134] In some embodiments, the present application does not limit the value of the first threshold.
[0135] Exemplarily, the value of the first threshold is 3 MIDI values. Then, if the pitch difference between the note and the fundamental frequency of the corresponding pronunciation position of the note is greater than 3 MIDI values, the pitch of the note is corrected to the pitch of the fundamental frequency of the corresponding pronunciation position of the note; if the pitch difference is less than or equal to 3 MIDI values, the pitch of the note remains unchanged.
[0136] For example, if the fundamental frequency of the corresponding pronunciation position of the note is 5 MIDI values, and if the pitch of the note is less than 2 MIDI values, or the pitch of the note is greater than 8 MIDI values, the pitch of the note is corrected to 5 MIDI values; if the pitch of the note is between 2 MIDI values and 8 MIDI values, the pitch of the note remains unchanged.
[0137] By correcting the second human voice symbol recognition result in the above manner, the accuracy of the pseudo-label information corresponding to the pure human voice audio is ensured, making the semi-supervised training method more efficient and stable.
[0138] In some embodiments, an accompaniment audio and a pure human voice audio are synthesized to obtain a synthesized audio corresponding to the pure human voice audio; the synthesized audio corresponding to the pure human voice audio is processed by a second network to obtain a human voice symbol recognition result corresponding to the pure human voice audio as the third human voice symbol recognition result; the second network is trained according to the third human voice symbol recognition result and the pseudo-label information.
[0139] In some embodiments, a loss function value of the second network is determined according to the third human voice symbol recognition result and the pseudo-label information; the parameters of the second network are adjusted according to the loss function value of the second network to obtain a human voice symbol recognition model.
[0140] In some embodiments, the second network is trained using a cross-entropy loss function.
[0141] In some embodiments, the second network can also perform human voice recognition on the synthesized audio corresponding to the pure human voice audio to obtain the human voice part of the synthesized audio corresponding to the pure human voice audio and the non-human voice part of the synthesized audio corresponding to the pure human voice audio, and then train the second network according to the human voice part of the synthesized audio corresponding to the pure human voice audio, the non-human voice part of the synthesized audio corresponding to the pure human voice audio, and the pure human voice audio.
[0142] In some embodiments, a fully connected layer can be used to perform human voice recognition on the synthesized audio corresponding to the pure human voice audio to obtain the human voice part of the synthesized audio corresponding to the pure human voice audio and the non-human voice part of the synthesized audio corresponding to the pure human voice audio. Exemplarily, Softmax can be used as a classifier to classify the human voice part of the synthesized audio corresponding to the pure human voice audio and the non-human voice part of the synthesized audio corresponding to the pure human voice audio.
[0143] In some embodiments, the method further includes step 442 of determining whether the second network meets the stop training condition; if so, the trained second network is determined as the human voice symbol recognition model; if not, the trained second network is determined as the trained first network, and the above step 440 is executed again.
[0144] Exemplarily, please refer to Figure 5 which shows a schematic diagram of a method for training a human voice symbol recognition model provided in an embodiment of the present application.
[0145] Step 1: Randomly select an accompaniment audio from a third training sample set (which can also be referred to as dataset 3) 511 as a target accompaniment audio; perform data augmentation processing on the labeled human voice audio in a first training sample set (which can also be referred to as dataset 1) 512 to obtain the processed labeled human voice audio; synthesize the target accompaniment audio and the processed labeled human voice audio to obtain a synthesized audio corresponding to the labeled human voice audio.
[0146] Process the synthesized audio corresponding to the annotator's voice audio through the teacher network 513 to obtain the human voice character recognition result corresponding to the annotator's voice audio as the first human voice character recognition result; determine the loss function value 514 (cross-entropy loss function) of the teacher network according to the first human voice character recognition result and the human voice character annotation result corresponding to the annotator's voice audio; train the teacher network 513 according to the loss function value 514 (cross-entropy loss function) of the teacher network to obtain the trained teacher network 521.
[0147] Step 2: Process the pure human voice audio in the second training sample set (which can also be called dataset 2) 522 through the trained teacher network 521 to obtain the human voice character recognition result corresponding to the pure human voice audio as the second human voice character recognition result (which can also be called the pseudo-label corresponding to the pure human voice audio) 523; determine the pseudo-label information corresponding to the pure human voice audio (which can also be called the pseudo-label correction corresponding to the pure human voice audio) 524 based on the second human voice character recognition result 523.
[0148] Step 3: Randomly select an accompaniment audio from the third training sample set 511 as the target accompaniment audio; perform data augmentation processing on the pure human voice audio in at least one pure human voice audio 522 to obtain the processed pure human voice audio; synthesize the target accompaniment audio with the processed pure human voice audio to obtain the synthesized audio corresponding to the pure human voice audio.
[0149] Process the synthesized audio corresponding to the pure human voice audio through the student network 525 to obtain the human voice character student recognition result corresponding to the pure human voice audio as the third human voice character recognition result (which can also be called the prediction corresponding to the pure human voice audio) 526.
[0150] Step 4: Determine the loss function value 527 (cross-entropy loss function) of the student network according to the human voice character student recognition result 526 corresponding to the pure human voice audio and the pseudo-label information 524 corresponding to the pure human voice audio; train the student network 525 according to the loss function value 527 (cross-entropy loss function) of the student network to obtain the trained student network 531.
[0151] Inference: When the trained student network 531 does not meet the stop training condition, determine the trained student network 531 as the trained teacher network and start executing from step 2 again. That is, replace the trained teacher network 521 in step 2 with the trained student network 531 and start executing from step 2 again.
[0152] When the trained student network 531 meets the stop training condition, the trained student network 531 is determined as a human voice note recognition model. Inputting a song with accompaniment, the human voice note recognition model processes the song with accompaniment, and a human voice note sequence 533 corresponding to the song with accompaniment can be obtained.
[0153] The technical solution provided by the embodiments of the present application, through the strategy of random data augmentation, further expands the number of training samples on the basis of the existing training samples to train the human voice note recognition model, and further improves the robustness of the human voice note recognition model.
[0154] Please refer to Figure 6 , which shows a flowchart of a human voice note recognition method provided by an embodiment of the present application. The method may include at least one of the following steps 610 to 640.
[0155] Step 610, obtaining a target audio with accompaniment, where the target audio includes human voice and accompaniment.
[0156] In some embodiments, the target audio further includes noise and reverberation.
[0157] In some embodiments, the present application does not limit the type of the target audio with accompaniment. Exemplarily, the target audio may be a song with accompaniment or a live song recording.
[0158] Step 620, obtaining audio features of the target audio, where the audio features include features related to the target audio in the time-frequency domain.
[0159] In some embodiments, performing a time-frequency transform on the target audio to obtain frequency-domain features of the target audio; performing a filtering process on the frequency-domain features to obtain audio features of the target audio.
[0160] The present application does not limit the specific method for performing a time-frequency transform on the target audio. Exemplarily, the CWT-ESS (Continuous Wavelet Transform) algorithm, the STFT-ESS (Short-Time Fourier Transform) algorithm, the OpenGAN algorithm, etc. can be used.
[0161] The present application does not limit the method for performing a filtering process on the frequency-domain features. Exemplarily, low-pass filtering, high-pass filtering, band-pass filtering, band-stop filtering, etc. can be used.
[0162] Step 630, processing the audio features through a human voice note recognition model to obtain note features of the target audio, where the note features include features related to the human voice notes of the target audio.
[0163] The human voice symbol recognition model is obtained by training a second network based on a trained first network, pure human voice audio, and accompaniment audio; the first network is used to output the human voice symbol recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio; the second network is used to output the human voice symbol recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
[0164] In some embodiments, for each audio frame included in the target audio, the human voice symbol recognition model processes the audio feature of the audio frame and the context information of the audio feature of the audio frame to obtain a first intermediate feature corresponding to the audio frame; according to the first intermediate feature corresponding to the audio frame, a second intermediate feature corresponding to the audio frame is extracted; according to the second intermediate feature corresponding to the audio frame and the context information of the second intermediate feature corresponding to the audio frame, a note feature corresponding to the audio frame is obtained; wherein, the note feature of the target audio includes the note features corresponding to each audio frame included in the target audio.
[0165] The first intermediate feature corresponding to the audio frame includes the audio feature corresponding to the audio frame and the context information of the audio feature corresponding to the audio frame.
[0166] The second intermediate feature corresponding to the audio frame is used to characterize the pitch feature of the audio frame.
[0167] The note feature corresponding to the audio frame includes the second intermediate feature corresponding to the audio frame and the context information of the second intermediate feature corresponding to the audio frame.
[0168] The context information refers to the association information between the target audio frame and adjacent audio frames. The adjacent audio frames refer to the adjacent audio frames and / or similar audio frames of the target audio frame. The adjacent audio frame refers to the audio frame that does not contain other audio frames between it and the target audio frame. The similar audio frames refer to the audio frames within a certain range of the target audio frame. For example, the five audio frames before and after the target audio frame can be called adjacent audio frames. The application does not limit the method for determining the range of the similar audio frames.
[0169] The application does not limit the method for obtaining the first intermediate feature corresponding to the audio frame according to the audio feature of the audio frame and the context information of the audio feature of the audio frame. Exemplarily, a recurrent neural network can be used. For example, it can be implemented through an LSTM (Long Short Term Memory Network) model, or it can be implemented through a GRU (Gate Recurrent Unit) model.
[0170] This application does not limit the method for extracting the second intermediate feature corresponding to an audio frame based on the first intermediate feature corresponding to the audio frame. Exemplarily, it can be implemented through a convolutional neural network. For example, it can be implemented through a CNN (Convolutional Neural Network), or it can also be implemented through a residual convolutional neural network (ResNet).
[0171] This application does not limit the method for obtaining the note feature corresponding to an audio frame based on the second intermediate feature corresponding to the audio frame and the context information of the second intermediate feature corresponding to the audio frame. Exemplarily, it can be implemented using a recurrent neural network. For example, it can be implemented through an LSTM (Long Short Term Memory Network) model, or it can also be implemented through a GRU (Gate Recurrent Unit) model.
[0172] Step 640, process the note feature through a human voice note recognition model to obtain the human voice note sequence of the target audio.
[0173] In some embodiments, classify the note feature of the target audio through a human voice note recognition model to obtain the human voice note sequence of the target audio.
[0174] In some embodiments, classify the note feature of the target audio according to the pitch of the note feature of the target note to obtain the human voice note sequence of the target audio.
[0175] Exemplarily, the human voice note sequence of the target audio is a MIDI sequence. According to the pitch of the note feature of the target note, classify the note feature of the target audio into different MIDI values to obtain the MIDI sequence of the target audio.
[0176] In some embodiments, the human voice note recognition model includes: an input layer, an intermediate layer, and an output layer.
[0177] The input layer is used to input the audio feature of the target audio.
[0178] The intermediate layer is used to extract the note feature of the target audio according to the audio feature.
[0179] The intermediate layer includes a first intermediate feature extraction layer, a second intermediate feature extraction layer, and a note feature extraction layer.
[0180] For each audio frame included in the target audio, the first intermediate feature extraction layer is used to obtain the first intermediate feature corresponding to the audio frame according to the audio feature of the audio frame and the context information of the audio feature of the audio frame. The second intermediate feature extraction layer is used to extract the second intermediate feature corresponding to the audio frame according to the first intermediate feature corresponding to the audio frame. The note feature extraction layer is used to obtain the note feature corresponding to the audio frame according to the second intermediate feature corresponding to the audio frame and the context information of the second intermediate feature corresponding to the audio frame.
[0181] In some embodiments, the first feature extraction layer is a bidirectional LSTM model, the second feature extraction layer is a CNN model, and the note feature extraction layer is a bidirectional LSTM model. In some embodiments, the second feature extraction layer can be configured with one or more CNN networks according to actual needs to form a CNN model, and this application does not limit this. For example, a CNN model is composed of 5 layers of CNN networks.
[0182] The output layer is used to obtain the human voice note sequence of the target audio according to the note feature.
[0183] In some embodiments, the output layer is a fully connected layer. In some embodiments, the output layer uses Softmax as a classifier.
[0184] Exemplarily, as Figure 7 shown, the human voice note recognition model 700 includes an input layer 710, an intermediate layer 720, and an output layer 730. The intermediate layer 720 includes a first intermediate feature extraction layer 721, a second intermediate feature extraction layer 722, and a note feature extraction layer 730.
[0185] It should be noted that the above embodiments of the human voice note recognition method and the above embodiments of the training method of the human voice note recognition model belong to the same concept. Please refer to the above embodiments of the training method of the human voice note recognition model, and details will not be repeated here.
[0186] The technical solution provided by the embodiments of this application can identify the human voice note sequence of the target note with accompaniment through the human voice note recognition model, without invoking the human voice and accompaniment separation algorithm, reducing the computational complexity, thereby reducing the production cost. At the same time, the accuracy is not affected by the human voice and accompaniment separation algorithm, ensuring the accuracy of the human voice note sequence.
[0187] The following are the device embodiments of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the device embodiments of this application, please refer to the method embodiments of this application.
[0188] Please refer to Figure 8, which shows a block diagram of a training device for a human voice symbol recognition model provided by an embodiment of the present application. This device has the functions of implementing the above method examples, and these functions can be implemented by hardware or by hardware executing corresponding software. This device can be the terminal device introduced above or can be set in the terminal device. As Figure 8 shown, the device 800 may include: a sample acquisition module 810, a first network training module 820, and a second network training module 830.
[0189] The sample acquisition module 810 is configured to acquire at least one labeled human voice audio, the human voice symbol annotation results respectively corresponding to each of the labeled human voice audios, at least one pure human voice audio, and at least one accompaniment audio.
[0190] The first network training module 820 is configured to train a first network based on the labeled human voice audio, the accompaniment audio, and the human voice symbol annotation result corresponding to the labeled human voice audio, to obtain a trained first network; the first network is configured to output a human voice symbol recognition result corresponding to the labeled human voice audio according to a synthesized audio of the labeled human voice audio and the accompaniment audio.
[0191] The second network training module 830 is configured to train a second network based on the trained first network, the pure human voice audio, and the accompaniment audio, to obtain a human voice symbol recognition model; the second network is configured to output a human voice symbol recognition result corresponding to the pure human voice audio according to a synthesized audio of the pure human voice audio and the accompaniment audio.
[0192] In some embodiments, as Figure 9 shown, the first network training module 820 includes a first synthesis unit 821 and a first training unit 822.
[0193] The first synthesis unit 821 is configured to synthesize the accompaniment audio and the labeled human voice audio to obtain a synthesized audio corresponding to the labeled human voice audio;
[0194] The first training unit 822 is configured to train the first network based on the synthesized audio corresponding to the labeled human voice audio and the human voice symbol annotation result corresponding to the labeled human voice audio, to obtain the trained first network.
[0195] In some embodiments, the first synthesis unit 821 is configured to randomly select an accompaniment audio from the at least one accompaniment audio as a target accompaniment audio; perform data augmentation processing on the labeled human voice audio to obtain processed labeled human voice audio; wherein the data augmentation processing includes at least one of the following: adding reverberation and changing the fundamental frequency; synthesize the target accompaniment audio and the processed labeled human voice audio to obtain a synthesized audio corresponding to the labeled human voice audio.
[0196] In some embodiments, the first training unit 822 is configured to process the synthesized audio corresponding to the labeled human voice audio through the first network to obtain a human voice symbol recognition result corresponding to the labeled human voice audio as a first human voice symbol recognition result; determine a loss function value of the first network according to the first human voice symbol recognition result and the human voice symbol annotation result; adjust parameters of the first network according to the loss function value of the first network to obtain the trained first network.
[0197] In some embodiments, as Figure 9 shown, the second network training module 830 includes a first processing unit 831, a determination unit 832, a second synthesis unit 833, a second processing unit 834, and a second training unit 835.
[0198] The first processing unit 831 is configured to process the pure human voice audio through the trained first network to obtain a human voice symbol recognition result corresponding to the pure human voice audio as a second human voice symbol recognition result.
[0199] The determination unit 832 is configured to determine the second human voice symbol recognition result as pseudo-label information corresponding to the pure human voice audio.
[0200] The second synthesis unit 833 is configured to synthesize the accompaniment audio and the pure human voice audio to obtain a synthesized audio corresponding to the pure human voice audio.
[0201] The second processing unit 834 is configured to process the synthesized audio corresponding to the pure human voice audio through the second network to obtain a human voice symbol recognition result corresponding to the pure human voice audio as a third human voice symbol recognition result.
[0202] The second training unit 835 is configured to train the second network according to the third human voice symbol recognition result and the pseudo-label information corresponding to the pure human voice audio to obtain a human voice symbol recognition model.
[0203] In some embodiments, the determining unit 832 is configured to extract the fundamental frequency of the pure human voice audio; and correct the second recognition result of the human voice characters according to the fundamental frequency of the pure human voice audio to obtain the pseudo-label information corresponding to the pure human voice audio.
[0204] In some embodiments, the determining unit 832 is configured to, for each note included in the second recognition result of the human voice characters, calculate the pitch difference between the note and the fundamental frequency of the pronunciation position corresponding to the note; if the pitch difference is greater than a first threshold, correct the pitch of the note to the pitch of the fundamental frequency of the pronunciation position corresponding to the note; if the pitch difference is less than or equal to the first threshold, keep the pitch of the note unchanged; and determine the second recognition result of the human voice characters with adjusted pitch as the pseudo-label information corresponding to the pure human voice audio.
[0205] In some embodiments, the second training unit 835 is configured to determine the loss function value of the second network according to the third recognition result of the human voice characters and the pseudo-label information; and adjust the parameters of the second network according to the loss function value of the second network to obtain the human voice character recognition model.
[0206] In some embodiments, the second network training module 830 is further configured to, when the second network does not meet the stop training condition, determine the trained second network as the trained first network, and start executing again from the step of training the second network based on the trained first network, the pure human voice audio, and the accompaniment audio.
[0207] In some embodiments, the sample acquisition module 810 is configured to acquire at least one a cappella audio, the human voice character annotation results respectively corresponding to the a cappella audios, and at least one accompanied song audio; generate the annotated human voice audio and the human voice character annotation result corresponding to the annotated human voice audio according to the a cappella audio and the human voice character annotation result corresponding to the a cappella audio; perform a human voice separation operation on the song audio to obtain a human voice audio and an accompaniment audio; and generate the pure human voice audio according to the human voice audio.
[0208] In some embodiments, the sample acquisition module 810 is configured to detect the a cappella audio to obtain the silent part and the clear voice part in the a cappella audio; determine the a cappella audio as the annotated human voice audio; and delete the human voice character annotation results corresponding to the silent part and the clear voice part from the human voice character annotation result corresponding to the a cappella audio to generate the human voice character annotation result corresponding to the annotated human voice audio.
[0209] In some embodiments, the sample acquisition module 810 is configured to detect the human voice audio to obtain the non-human voice part in the human voice audio; delete the non-human voice part in the human voice audio to generate a pure human voice audio; for each audio frame in the pure human voice audio, detect whether the audio frame is a human voice audio frame and calculate the energy of the audio frame; if the audio frame is not a human voice audio frame and the energy of the audio frame is less than a second threshold, determine the audio frame as an invalid frame; if the proportion of the number of invalid frames in the pure human voice audio in the total number of audio frames included in the pure human voice audio is greater than a third threshold, determine the pure human voice audio as an invalid pure human voice audio; generate the pure human voice audio according to the pure human voice audio other than the invalid pure human voice audio.
[0210] The technical solution provided by the embodiments of the present application, through the human voice character recognition model obtained by the above training method, can directly recognize the corresponding human voice character sequence from the target audio with accompaniment. Therefore, in the model usage stage, there is no need to call the human voice-accompaniment separation algorithm to extract the human voice audio from the target audio, reducing the computational complexity of human voice character recognition. In addition, the present application adopts a semi-supervised training method, training the first network with a small number of labeled samples, and then training the second network with the first network and a large number of unlabeled samples. In this way, only a small number of labeled samples are required to train a model with strong generalization performance, reducing the acquisition cost of training samples.
[0211] Please refer to Figure 10 , which shows a block diagram of a human voice character recognition device provided by an embodiment of the present application. The device has the functions of implementing the above method examples, and the functions can be implemented by hardware or by hardware executing corresponding software. The device can be the terminal device introduced above or can be set in the terminal device. As Figure 10 shown, the device 1000 may include: an audio acquisition module 1010, a feature acquisition module 1020, a feature extraction module 1030, and a result obtaining module 1040.
[0212] The audio acquisition module 1010 is configured to acquire a target audio with accompaniment, where the target audio includes a human voice and accompaniment.
[0213] The feature acquisition module 1020 is configured to acquire audio features of the target audio, where the audio features include features related to the target audio in the time-frequency domain.
[0214] The feature extraction module 1030 is configured to process the audio features through a human voice character recognition model to obtain note features of the target audio, where the note features include features related to the human voice characters of the target audio.
[0215] As a result, a result obtaining module 1040 is obtained, which is used to process the note features through the human voice note recognition model to obtain the human voice note sequence of the target audio; wherein, the human voice note recognition model is obtained by training a second network based on a trained first network, a pure human voice audio, and an accompaniment audio; the first network is used to output the human voice note recognition result corresponding to the annotated human voice audio according to the synthesized audio of the annotated human voice audio and the accompaniment audio; the second network is used to output the human voice note recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
[0216] In some embodiments, the feature extraction module 1030 is configured to, for each audio frame included in the target audio, obtain a first intermediate feature corresponding to the audio frame through the human voice note recognition model according to the audio feature of the audio frame and the context information of the audio feature of the audio frame; extract a second intermediate feature corresponding to the audio frame according to the first intermediate feature corresponding to the audio frame; and obtain a note feature corresponding to the audio frame according to the second intermediate feature corresponding to the audio frame and the context information of the second intermediate feature corresponding to the audio frame; wherein, the note features of the target audio include the note features corresponding to each audio frame included in the target audio.
[0217] In some embodiments, the feature acquisition module 1020 is configured to perform time-frequency transformation on the target audio to obtain the frequency domain feature of the target audio; and perform filtering processing on the frequency domain feature to obtain the audio feature of the target audio.
[0218] In some embodiments, the result obtaining module 1040 is configured to perform classification processing on the note features of the target audio through the human voice note recognition model to obtain the human voice note sequence of the target audio.
[0219] In some embodiments, the human voice note sequence is obtained by a human voice note recognition model, and the human voice note recognition model includes: an input layer, an intermediate layer, and an output layer; the input layer is used to input the audio feature of the target audio; the intermediate layer is used to extract the note feature of the target audio according to the audio feature; and the output layer is used to obtain the human voice note sequence of the target audio according to the note feature.
[0220] The technical solution provided by the embodiments of the present application can identify the human voice note sequence of the target note with accompaniment through the human voice note recognition model, without invoking a human voice-accompaniment separation algorithm, reducing the computational complexity, and at the same time, the accuracy is not affected by the human voice-accompaniment separation algorithm, ensuring the accuracy of the human voice note sequence.
[0221] It should be noted that when the device provided in the above embodiments realizes its functions, only the division of the above-mentioned functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to actual needs, that is, the content structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0222] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0223] Please refer to Figure 11 , which shows a schematic structural diagram of a computer device provided in an embodiment of the present application. The computer device can be any electronic device with data calculation, processing, and storage functions. The computer device can be used to implement the training method of the human voice character recognition model provided in the above embodiments, or to implement the human voice character recognition method provided in the above embodiments. Specifically:
[0224] The computer device 1100 includes a central processing unit (such as a CPU (Central Processing Unit, central processor), a GPU (Graphics Processing Unit, graphics processor), and an FPGA (Field Programmable Gate Array, field programmable logic gate array), etc.) 1101, a system memory 1104 including a RAM (Random-Access Memory, random access memory) 1102 and a ROM (Read-Only Memory, read-only memory) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 further includes a basic input / output system (Input Output System, I / O system) 1106 for facilitating the transmission of information between various components within the server, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1111.
[0225] In some embodiments, the basic input / output system 1106 includes a display 1108 for displaying information and input devices 1109 such as a mouse, keyboard, etc. for user input of information. Among them, both the display 1108 and the input devices 1109 are connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include an input / output controller 1110 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides outputs to a display screen, printer, or other types of output devices.
[0226] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the computer device 1100. That is to say, the mass storage device 1107 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0227] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art know that the computer storage media is not limited to the above several types. The above system memory 1104 and mass storage device 1107 may be collectively referred to as memory.
[0228] According to an embodiment of the present application, the computer device 1100 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 1100 can be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105, or in other words, the network interface unit 1111 can also be used to connect to other types of networks or remote computer systems (not shown).
[0229] A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned training method of the human voice character recognition model, or to implement the above-mentioned human voice character recognition method.
[0230] In an exemplary embodiment, there is also provided a computer-readable storage medium, in which a computer program is stored, and the computer program is loaded and executed by a processor to implement the above-mentioned training method of the human voice character recognition model, or to implement the above-mentioned human voice character recognition method.
[0231] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0232] In an exemplary embodiment, there is also provided a computer program product, which includes a computer program stored in a computer-readable storage medium, and the processor reads and executes the computer program from the computer-readable storage medium to implement the above-mentioned training method of the human voice character recognition model, or to implement the above-mentioned human voice character recognition method.
[0233] In the description of the embodiments of the present application, the term "corresponding" may indicate a direct or indirect corresponding relationship between two parties, may also indicate an associated relationship between two parties, or may be a relationship such as indication and being indicated, configuration and being configured, etc.
[0234] As used herein, "a plurality of" means two or more. "And / or" describes the associated relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0235] In addition, the step numbers described in this document only exemplarily show one possible execution sequence among the steps. In some other embodiments, the above steps may not be executed in the numbered order. For example, two steps with different numbers may be executed simultaneously, or two steps with different numbers may be executed in the reverse order of the illustration. The embodiments of the present application do not limit this.
[0236] In addition, the embodiments provided in this document can be combined arbitrarily to form new embodiments, which are all within the protection scope of the present application.
[0237] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0238] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A training method for a human voice symbol recognition model, characterized in that The method includes: obtaining at least one labeled person's voice audio, the corresponding person's voice symbol annotation results for each of the labeled person's voice audios, at least one pure person's voice audio, and at least one accompaniment audio; training a first network based on the labeled person's voice audio, the accompaniment audio, and the corresponding person's voice symbol annotation results of the labeled person's voice audio to obtain a trained first network; the first network is used to output the person's voice symbol recognition result corresponding to the labeled person's voice audio according to the synthesized audio of the labeled person's voice audio and the accompaniment audio; training a second network based on the trained first network, the pure person's voice audio, and the accompaniment audio to obtain a person's voice symbol recognition model; the second network is used to output the person's voice symbol recognition result corresponding to the pure person's voice audio according to the synthesized audio of the pure person's voice audio and the accompaniment audio.
2. The method according to claim 1, wherein The training of the first network based on the labeled person's voice audio, the accompaniment audio, and the corresponding person's voice symbol annotation results of the labeled person's voice audio to obtain a trained first network includes: synthesizing the accompaniment audio and the labeled person's voice audio to obtain the synthesized audio corresponding to the labeled person's voice audio; training the first network based on the synthesized audio corresponding to the labeled person's voice audio and the corresponding person's voice symbol annotation results of the labeled person's voice audio to obtain the trained first network.
3. The method according to claim 2, wherein The synthesizing the accompaniment audio and the labeled person's voice audio to obtain the synthesized audio corresponding to the labeled person's voice audio includes: randomly selecting an accompaniment audio from the at least one accompaniment audio as the target accompaniment audio; performing data augmentation processing on the labeled person's voice audio to obtain a processed labeled person's voice audio; wherein, the data augmentation processing includes at least one of the following: adding reverberation, changing the fundamental frequency; synthesizing the target accompaniment audio and the processed labeled person's voice audio to obtain the synthesized audio corresponding to the labeled person's voice audio.
4. The method according to claim 2, wherein The training of the first network based on the synthesized audio corresponding to the labeled person's voice audio and the corresponding person's voice symbol annotation results of the labeled person's voice audio to obtain the trained first network includes: processing the synthesized audio corresponding to the labeled person's voice audio through the first network to obtain the person's voice symbol recognition result corresponding to the labeled person's voice audio as the first person's voice symbol recognition result; determining the loss function value of the first network according to the first person's voice symbol recognition result and the person's voice symbol annotation result; adjusting the parameters of the first network according to the loss function value of the first network to obtain the trained first network.
5. The method according to claim 1, characterized in that The training of the second network based on the trained first network, the pure person's voice audio, and the accompaniment audio to obtain a person's voice symbol recognition model includes: processing the pure person's voice audio through the trained first network to obtain the person's voice symbol recognition result corresponding to the pure person's voice audio as the second person's voice symbol recognition result; determining the second person's voice symbol recognition result as the pseudo-label information corresponding to the pure person's voice audio; Synthesize the described accompaniment audio and the described pure human voice audio to obtain the synthesized audio corresponding to the pure human voice audio; Process the synthesized audio corresponding to the pure human voice audio through the second network to obtain the human voice symbol recognition result corresponding to the pure human voice audio, as the third human voice symbol recognition result; Train the second network according to the third human voice symbol recognition result and the pseudo-label information to obtain a human voice symbol recognition model.
6. The method according to claim 5, wherein The step of determining the second human voice symbol recognition result as the pseudo-label information corresponding to the pure human voice audio includes: Extract the fundamental frequency of the pure human voice audio; According to the fundamental frequency of the pure human voice audio, correct the second human voice symbol recognition result to obtain the pseudo-label information corresponding to the pure human voice audio.
7. The method according to claim 6, wherein The step of correcting the second human voice symbol recognition result according to the fundamental frequency of the pure human voice audio to obtain the pseudo-label information corresponding to the pure human voice audio includes: For each note included in the second human voice symbol recognition result, calculate the pitch difference between the note and the fundamental frequency of the pronunciation position corresponding to the note; If the pitch difference is greater than the first threshold, correct the pitch of the note to the pitch of the fundamental frequency of the pronunciation position corresponding to the note; If the pitch difference is less than or equal to the first threshold, keep the pitch of the note unchanged; Determine the second human voice symbol recognition result with adjusted pitch as the pseudo-label information corresponding to the pure human voice audio.
8. The method according to claim 5, wherein The step of training the second network according to the third human voice symbol recognition result and the pseudo-label information to obtain a human voice symbol recognition model includes: Determine the loss function value of the second network according to the third human voice symbol recognition result and the pseudo-label information; Adjust the parameters of the second network according to the loss function value of the second network to obtain the human voice symbol recognition model.
9. The method according to claim 1, wherein The method further includes: When the second network does not meet the stop training condition, determine the trained second network as the trained first network, and start executing again from the step of training the second network based on the trained first network, the pure human voice audio, and the accompaniment audio.
10. The method according to claim 1, characterized in that, The step of obtaining at least one annotated human voice audio, the human voice symbol annotation results respectively corresponding to each of the annotated human voice audios, at least one pure human voice audio, and at least one accompaniment audio includes: Obtain at least one a cappella audio, the human voice symbol annotation results respectively corresponding to each of the a cappella audios, and at least one accompanied song audio; Generate the annotated human voice audio and the human voice symbol annotation result corresponding to the annotated human voice audio according to the a cappella audio and the human voice symbol annotation result corresponding to the a cappella audio; Perform a vocal separation operation on the song audio to obtain a human voice audio and the accompaniment audio; Generate the pure human voice audio according to the human voice audio.
11. The method according to claim 10, wherein The step of generating the annotated human voice audio and the human voice symbol annotation result corresponding to the annotated human voice audio according to the a cappella audio and the human voice symbol annotation result corresponding to the a cappella audio includes: Detect the a cappella audio to obtain the silent part and the voiceless part in the a cappella audio; Determine the a cappella audio as the labeled human voice audio; Delete the human voice symbol annotation results corresponding to the silent part and the voiceless part from the human voice symbol annotation results corresponding to the a cappella audio, and generate the human voice symbol annotation results corresponding to the labeled human voice audio.
12. The method according to claim 10, wherein The generating the pure human voice audio according to the human voice audio includes: Detect the human voice audio to obtain the non-human voice part in the human voice audio; Delete the non-human voice part in the human voice audio to generate a pure human voice audio; For each audio frame in the pure human voice audio, detect whether the audio frame is a human voice audio frame and calculate the energy of the audio frame; If the audio frame is not a human voice audio frame and the energy of the audio frame is less than a second threshold, determine the audio frame as an invalid frame; If the proportion of the number of invalid frames in the pure human voice audio in the total number of audio frames included in the pure human voice audio is greater than a third threshold, determine the pure human voice audio as an invalid pure human voice audio; Generate the pure human voice audio according to the pure human voice audio other than the invalid pure human voice audio.
13. A method for identifying human voice notes, characterized in that, The method includes: Obtain a target audio with accompaniment, where the target audio includes human voices and accompaniment; Obtain the audio features of the target audio, where the audio features include features related to the target audio in the time-frequency domain; Process the audio features through a human voice symbol recognition model to obtain the note features of the target audio, where the note features include features related to the human voice symbols of the target audio; Process the note features through the human voice symbol recognition model to obtain the human voice symbol sequence of the target audio; Among them, the human voice symbol recognition model is trained by training a second network based on a trained first network, a pure human voice audio, and an accompaniment audio; the first network is used to output the human voice symbol recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio; the second network is used to output the human voice symbol recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
14. The method according to claim 13, wherein The extracting the note features of the target audio by the human voice symbol recognition model according to the audio features includes: For each audio frame included in the target audio, process the audio features of the audio frame and the context information of the audio features of the audio frame through the human voice symbol recognition model to obtain a first intermediate feature corresponding to the audio frame; Extract a second intermediate feature corresponding to the audio frame according to the first intermediate feature corresponding to the audio frame; Obtain the note feature corresponding to the audio frame according to the second intermediate feature corresponding to the audio frame and the context information of the second intermediate feature corresponding to the audio frame; Among them, the note features of the target audio include the note features corresponding to each audio frame included in the target audio.
15. The method according to claim 13, wherein The obtaining the audio features of the target audio includes: Perform time-frequency transformation on the target audio to obtain the frequency-domain features of the target audio; Perform filtering processing on the frequency-domain features to obtain the audio features of the target audio.
16. The method according to claim 13, characterized in that The obtaining, by the human voice symbol recognition model according to the symbol features, of the human voice symbol sequence of the target audio includes: Classify the symbol features of the target audio through the human voice symbol recognition model to obtain the human voice symbol sequence of the target audio.
17. The method according to claim 13, wherein The human voice symbol recognition model includes: an input layer, an intermediate layer, and an output layer; The input layer is used to input the audio features of the target audio; The intermediate layer is used to extract the symbol features of the target audio according to the audio features; The output layer is used to obtain the human voice symbol sequence of the target audio according to the symbol features.
18. A training device for a human voice symbol recognition model, characterized in that, The device includes: A sample acquisition module, configured to acquire at least one labeled human voice audio, the human voice symbol annotation results respectively corresponding to each of the labeled human voice audios, at least one pure human voice audio, and at least one accompaniment audio; A first network training module, configured to train a first network based on the labeled human voice audio, the accompaniment audio, and the human voice symbol annotation result corresponding to the labeled human voice audio, to obtain a trained first network; the first network is used to output the human voice symbol recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio; A second network training module, configured to train a second network based on the trained first network, the pure human voice audio, and the accompaniment audio, to obtain a human voice symbol recognition model; the second network is used to output the human voice symbol recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
19. A human voice symbol recognition device, characterized in that, The device includes: An audio acquisition module, configured to acquire a target audio with accompaniment, where the target audio includes human voice and accompaniment; A feature acquisition module, configured to acquire the audio features of the target audio, where the audio features include features related to the target audio in the time-frequency domain; A feature extraction module, configured to process the audio features through a human voice symbol recognition model to obtain the symbol features of the target audio, where the symbol features include features related to the human voice symbols of the target audio; A result obtaining module, configured to process the symbol features through the human voice symbol recognition model to obtain the human voice symbol sequence of the target audio; Wherein, the human voice symbol recognition model is obtained by training a second network based on the trained first network, the pure human voice audio, and the accompaniment audio; the first network is used to output the human voice symbol recognition result corresponding to the labeled human voice audio according to the synthesized audio of the labeled human voice audio and the accompaniment audio; the second network is used to output the human voice symbol recognition result corresponding to the pure human voice audio according to the synthesized audio of the pure human voice audio and the accompaniment audio.
20. A computer device, characterized in that, The computer device includes a processor and a memory. A computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 12, or to implement the method according to any one of claims 13 to 17.
21. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is used to be executed by a processor to implement the method according to any one of claims 1 to 12, or to implement the method according to any one of claims 13 to 17.
22. A computer program product, characterized in that, The computer program product includes a computer program. The computer program is stored in a computer-readable storage medium, and the processor reads and executes the computer program from the computer-readable storage medium to implement the method according to any one of claims 1 to 12, or to implement the method according to any one of claims 13 to 17.
Citation Information
Patent Citations
Audio separation network training method and device, audio separation method and device and medium
CN111341341A
Lyric acoustic model training method, lyric recognition method, device and product
CN115083397A