Audio processing method and related device
By training the fundamental frequency prediction model and acoustic model, the music score file and dry sound audio are used to generate synthetic songs, which solves the problem that it is difficult to generate real and natural songs in singing synthesis, and improves the sound quality of the synthetic songs.
Patent Information
- Application Number
- CN202211471824.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-11-22
AI Technical Summary
It is difficult to generate real and natural vocal vocal synthesis, mainly because the vocal vocal contains more tone, rhythm and melody information, which increases the difficulty of synthesis.
By obtaining the score file and dry sound audio of the song training sample, determine the syllable sequence, note sequence, fundamental frequency sequence and pronunciation sequence, input the initial fundamental frequency prediction model and initial acoustic model for training, obtaining the target fundamental frequency prediction model and target acoustic model. These models are used to generate predicted base frequency sequences and predicted acoustic features based on the score files of the song to be synthesized, thereby performing audio synthesis.
The sound quality of the synthetic songs is improved, making the synthetic songs more realistic and natural.
Smart Images

Figure CN115862592B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio processing method and related devices. Background Art
[0002] Singing synthesis refers to the use of speech synthesis technology to enable computers to sing songs like humans. As a new application field of speech synthesis technology, singing synthesis has great application value and prospects in the fields of virtual singers, record production, digital music creation, etc. Since singing carries more information such as pitch, rhythm and melody than speech, it inevitably increases the difficulty of singing synthesis. How to synthesize real and natural singing is a technical problem that researchers urgently need to solve. Summary of the invention
[0003] The embodiments of the present application provide an audio processing method and related devices, which can improve the sound quality of synthesized songs and make the synthesized songs more realistic and natural.
[0004] On the one hand, an embodiment of the present application provides an audio processing method, the method comprising:
[0005] Obtain the music score file and dry audio of the song training sample;
[0006] Determine the syllable sequence and note sequence of the song training sample according to the music score file, and determine the first fundamental frequency sequence and pronunciation sequence of the song training sample according to the dry sound audio;
[0007] Inputting the syllable sequence and the note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, and training the initial fundamental frequency prediction model according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain a target fundamental frequency prediction model, wherein the target fundamental frequency prediction model is used to generate the predicted fundamental frequency sequence of the song to be synthesized according to the score file of the song to be synthesized;
[0008] The first fundamental frequency sequence and pronunciation sequence of the song training sample are input into the initial acoustic model to obtain the first acoustic feature, and the initial acoustic model is trained according to the first acoustic feature and the second acoustic feature of the dry sound audio to obtain a target acoustic model, wherein the target acoustic model is used to generate the predicted acoustic features of the song to be synthesized according to the predicted fundamental frequency sequence of the song to be synthesized, and the predicted acoustic features are used to generate the synthesized audio of the song to be synthesized.
[0009] On the one hand, an embodiment of the present application provides an audio processing method, the method comprising:
[0010] Acquire a music score file of a song to be synthesized, and determine a syllable sequence and a note sequence of the song to be synthesized according to the music score file;
[0011] Inputting the syllable sequence and the note sequence of the song to be synthesized into the target fundamental frequency prediction model to obtain the predicted fundamental frequency sequence of the song to be synthesized;
[0012] Determining a target pronunciation sequence according to the syllable sequence of the song to be synthesized, and inputting the target pronunciation sequence and the predicted fundamental frequency sequence into a target acoustic model to obtain predicted acoustic features of the song to be synthesized;
[0013] The vocoder is called to perform audio synthesis processing on the predicted acoustic features to obtain the synthesized audio of the song to be synthesized.
[0014] On the one hand, an embodiment of the present application provides an audio processing device, the device comprising:
[0015] An acquisition unit, used to acquire the music score file and dry audio of the song training sample;
[0016] A processing unit, used to determine the syllable sequence and note sequence of the song training sample according to the music score file, and to determine the first fundamental frequency sequence and pronunciation sequence of the song training sample according to the dry sound audio;
[0017] The processing unit is further used to input the syllable sequence and the note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, and train the initial fundamental frequency prediction model according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain a target fundamental frequency prediction model, wherein the target fundamental frequency prediction model is used to generate the predicted fundamental frequency sequence of the song to be synthesized according to the score file of the song to be synthesized;
[0018] The processing unit is further used to input the first fundamental frequency sequence and pronunciation sequence of the song training sample into the initial acoustic model to obtain a first acoustic feature, and train the initial acoustic model according to the first acoustic feature and the second acoustic feature of the dry sound audio to obtain a target acoustic model, wherein the target acoustic model is used to generate the predicted acoustic features of the song to be synthesized according to the predicted fundamental frequency sequence of the song to be synthesized, and the predicted acoustic features are used to generate the synthesized audio of the song to be synthesized.
[0019] On the one hand, an embodiment of the present application provides an audio processing device, the device comprising:
[0020] An acquisition unit, used for acquiring a music score file of a song to be synthesized;
[0021] A processing unit, used for determining a syllable sequence and a note sequence of the song to be synthesized according to the music score file;
[0022] The processing unit is further used to input the syllable sequence and note sequence of the song to be synthesized into the target fundamental frequency prediction model to obtain the predicted fundamental frequency sequence of the song to be synthesized;
[0023] The processing unit is further used to determine a target pronunciation sequence according to the syllable sequence of the song to be synthesized, and input the target pronunciation sequence and the predicted fundamental frequency sequence into a target acoustic model to obtain predicted acoustic features of the song to be synthesized;
[0024] The processing unit is also used to call a vocoder to perform audio synthesis processing on the predicted acoustic features to obtain the synthesized audio of the song to be synthesized.
[0025] On the one hand, an embodiment of the present application provides a computer device, which includes a processor, a communication interface and a memory, wherein the processor, the communication interface and the memory are interconnected, wherein the memory stores a computer program, and the processor is used to call the computer program to execute the audio processing method of any possible implementation method described above.
[0026] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the audio processing method of any possible implementation manner is implemented.
[0027] On the one hand, an embodiment of the present application further provides a computer program product, which includes a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the steps of the audio processing method provided in the embodiment of the present application.
[0028] On the one hand, an embodiment of the present application also provides a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing method provided by the embodiment of the present application.
[0029] In an embodiment of the present application, by cultivating the initial fundamental frequency prediction model's deep modeling capability for the syllable sequence and note sequence of the song training samples, the target fundamental frequency prediction model can accurately predict the fundamental frequency sequence of the song to be synthesized; at the same time, the initial acoustic model can enhance the initial acoustic model's ability to characterize acoustic features by taking into account the learning of the fundamental frequency sequence and pronunciation sequence of the song training samples, so that the target acoustic model can generate acoustic features with higher accuracy for the song to be synthesized, thereby improving the sound quality of the synthesized song and making the synthesized singing more realistic and natural. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical methods of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 A schematic diagram of the system architecture of an audio processing system provided in an embodiment of the present application;
[0032] Figure 2 A flowchart of an audio processing method provided in an embodiment of the present application;
[0033] Figure 3 An example schematic diagram of a music score file provided in an embodiment of the present application;
[0034] Figure 4 A flowchart of another audio processing method provided in an embodiment of the present application;
[0035] Figure 5 A flowchart of another audio processing method provided in an embodiment of the present application;
[0036] Figure 6 A flowchart of another audio processing method provided in an embodiment of the present application;
[0037] Figure 7 A flowchart of another audio processing method provided in an embodiment of the present application;
[0038] Figure 8 A flowchart of another audio processing method provided in an embodiment of the present application;
[0039] Fig. 9 A flowchart of another audio processing method provided in an embodiment of the present application;
[0040] Fig.10 A flowchart of another audio processing method provided in an embodiment of the present application;
[0041] Fig.11 A schematic diagram of the structure of an audio processing device provided in an embodiment of the present application;
[0042] Fig.12 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical methods in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0044] This application proposes an audio processing method that can be applied to various fields or scenarios such as cloud technology, artificial intelligence, blockchain, Internet of Vehicles, smart transportation, smart home, etc. Specifically, the audio processing method proposed in this application can be implemented based on the speech processing technology in artificial intelligence technology. Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level technology and software-level technology. Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, large video processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and smart transportation. The key technologies of speech processing technology include automatic speech recognition technology (Automatic Speech Recognition, ASR), speech synthesis technology (Text To Speech, TTS), and voiceprint recognition technology. Allowing computers to listen, see, speak, and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future.
[0045] See also Figure 1 , Figure 1 A schematic diagram of the system architecture of an audio processing system provided in an embodiment of the present application. Figure 1 As shown, the system architecture includes a terminal device 11 and a server 12. The terminal device 11 and the server 12 can communicate with each other through a network. There can be one or more terminal devices 11.
[0046] The terminal device 11 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The server 12 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0047] Figure 1The system architecture shown can implement the audio processing method provided in the embodiment of the present application. Taking the terminal device 11 and the server 12 jointly executing the method as an example, the implementation process generally includes:
[0048] In the first stage, the server 12 needs to obtain the score file and dry audio of the song training sample, determine the syllable sequence and note sequence of the song training sample according to the score file of the song training sample, and determine the first fundamental frequency sequence and pronunciation sequence of the song training sample according to the dry audio of the song training sample. Then, the syllable sequence and note sequence of the song training sample are input into the initial fundamental frequency prediction model to obtain the second fundamental frequency sequence, and the initial fundamental frequency prediction model is trained according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain the target fundamental frequency prediction model. And the first fundamental frequency sequence and pronunciation sequence of the song training sample are input into the initial acoustic model to obtain the first acoustic feature, and the initial acoustic model is trained according to the first acoustic feature and the second acoustic feature of the dry audio to obtain the target acoustic model.
[0049] In the second stage, the server 12 will be equipped with the trained target fundamental frequency prediction model and target acoustic model. When the server 12 receives the audio synthesis request for the song to be synthesized sent by the terminal device 11, the score file of the song to be synthesized is obtained, and the syllable sequence and note sequence of the song to be synthesized are determined according to the score file of the song to be synthesized. Then the syllable sequence and note sequence of the song to be synthesized are input into the carried target fundamental frequency prediction model to obtain the predicted fundamental frequency sequence of the song to be synthesized. And the target pronunciation sequence is determined according to the syllable sequence of the song to be synthesized, and the target pronunciation sequence and the predicted fundamental frequency sequence are input into the carried target acoustic model to obtain the predicted acoustic features of the song to be synthesized. Finally, the vocoder is called to perform audio synthesis processing on the predicted acoustic features to obtain the synthesized audio of the song to be synthesized. The method of the present application can improve the sound quality of the synthesized song and make the synthesized singing more real and natural.
[0050] The specific implementation of the audio processing method provided by this application is described in detail below. Figure 2 , Figure 2 A flowchart of an audio processing method provided in an embodiment of the present application is provided. The method can be applied to the above Figure 1 The server 12 in the method comprises:
[0051] S201, obtaining the music score file and dry audio of the song training sample.
[0052] The music score in the music score file refers to a regular combination of various written symbols that record the pitch or rhythm of music, such as common simple musical notation, five-line musical notation, guitar musical notation, guqin musical notation, etc. In one embodiment, the server can construct a first data set, the first data set includes music score files of multiple songs, one of the multiple music score files of the multiple songs is a music score file of a song training sample, and the server can obtain the music score file of the song training sample from the first data set.
[0053] Dry audio refers to the vocal audio of a song, excluding the accompaniment audio of the song; it can be all or part of the vocal audio of a song. In one embodiment, the server can construct a second data set, the second data set includes the dry audio of at least one song, one of the dry audio of the at least one song is the dry audio of the song training sample, and the server can obtain the dry audio of the song training sample from the second data set.
[0054] It should be noted that the above-mentioned first data set and the above-mentioned second data set may correspond to different song sets, that is, the songs included in the above-mentioned multiple songs and the above-mentioned at least one song may be different. Therefore, the present application does not require paired music score texts and dry audio to complete the training of subsequent models (including the initial fundamental frequency prediction model and the initial acoustic model).
[0055] S202: Determine a syllable sequence and a note sequence of a song training sample according to the music score file, and determine a first fundamental frequency sequence and a pronunciation sequence of the song training sample according to the dry sound audio.
[0056] The music score file of the song training sample contains the lyrics of the song training sample and multiple notes (music symbols), for example Figure 3 is an example of a music score file. In one embodiment, the time information of each note in the song training sample can be determined through the music score file of the song training sample, and the time information of each note includes the note start time (reflecting the start time of the note pronunciation) and the note end time (reflecting the end time of the note pronunciation), as well as the note duration (reflecting the duration of the note pronunciation) determined according to the difference between the note start time and the note end time. Understandably, in the music score, the number of beats in each measure and the duration of each beat are usually represented by the time signature. For example, 4 / 4 beats means that a quarter note is one beat and each measure has four beats, and 3 / 8 beats means that an eighth note is one beat and each measure has three beats. When the music score stipulates 60 beats for 1 minute, the length of 1 beat is 1 second; taking 4 / 4 beats as an example, each measure has four beats, that is, 4 seconds, and so on. Therefore, the time information of each note can be known by the number of beats the note lasts or the duration of the note expressed in proportion.
[0057] The music score file of the song training sample records the correspondence between syllables and notes. Each syllable can correspond to one or more notes, and the one or more notes can be determined as the auxiliary notes corresponding to each syllable. Then, the time information of each syllable is determined according to the time information of each auxiliary note. The time information of each syllable includes the syllable start time (reflecting the start time of the syllable pronunciation), the syllable end time (reflecting the end time of the syllable pronunciation) and the syllable duration (reflecting the duration of the syllable pronunciation). Specifically, the sum of the note durations of the auxiliary notes is determined as the syllable duration of each syllable. The earliest note start time among the note start times of the auxiliary notes is determined as the syllable start time of each syllable. The latest note end time among the note end times of the auxiliary notes is determined as the syllable end time of each syllable.
[0058] Among them, a syllable corresponds to a complete pronunciation unit (such as a Chinese character, an English word). In addition, a syllable usually contains one or more phonemes (the smallest pronunciation unit), such as "b (wave), p (slope), m (mo)" in Chinese and " / i: / , / I / , / e / " in English.
[0059] In one embodiment, after obtaining the time information of each syllable in the music score file of the song training sample, the server can determine the syllable sequence of the song training sample based on a preset unit frame length (which can be set manually, such as 10-20 milliseconds) and the time information of each syllable.
[0060] Specifically, for any syllable in each syllable, the number of any syllable is determined according to the preset unit frame length and the syllable duration included in the time information of any syllable. Among them, the number of any syllable is obtained by dividing the syllable duration of the syllable by the preset unit frame length and performing rounding operation, and the rounding operation can be rounding up or rounding down. The number of any syllable represents the number of preset unit frame lengths that the syllable needs to last, which can reflect the duration that the syllable needs to last when singing. For example, the syllable duration of the syllable "I" is 63 milliseconds, the preset unit frame length is 15 milliseconds, and the number of syllables "I" can be 4 (rounding down) or 5 (rounding up), indicating that the syllable "I" needs to last 4 or 5 preset unit frame lengths. Further, according to the preset unit frame length and the syllable start time included in the time information of any syllable, the sequence start position of any syllable is determined. Among them, the sequence start position of any syllable is obtained by dividing the syllable start time of any syllable by the preset unit frame length and performing rounding operation. For example, the syllable start time of the syllable "I" is 15 milliseconds, the preset unit frame length is 15 milliseconds, and the sequence start position of the syllable "I" is 1. Finally, according to the number of any syllable and the starting position of the sequence, each syllable is sorted and processed to obtain the syllable sequence of the song training sample. For example, the number of syllables "I", "love", and "you" in the lyrics "I love you" are 2, 4, and 3 respectively, and the sequence start positions are 0, 3, and 8 respectively, then the corresponding syllable sequence is [I, I, placeholder, love, love, love, love, placeholder, you, you]. It can be seen that the syllable sequence can not only reflect the syllable's own information in the song training sample, but also reflect the pronunciation position, pronunciation duration and other information of the syllable in the song training sample, which can effectively characterize the pronunciation of the song training sample.
[0061] Furthermore, after obtaining the time information of each note in the score file of the song training sample, the server can determine the note sequence of the song training sample according to a preset unit frame length (which can be set manually, such as 10-20 milliseconds) and the time information of each note.
[0062] Specifically, for any note among the notes, the number of any note is determined according to the preset unit frame length and the note duration of any note. The number of any note is obtained by dividing the note duration of the note by the preset unit frame length and performing rounding operation, and the rounding operation can be rounding up or rounding down. The number of any note represents the number of preset unit frame lengths that the any note needs to last, which can reflect the duration that the any note needs to last when playing. Then, according to the preset unit frame length and the note start time of any note, the sequence start position of any note is determined. The sequence start position of any note is obtained by dividing the note start time of any note by the preset unit frame length and performing rounding operation. Finally, according to the number of any note and the sequence start position, each note is sorted and processed to obtain the note sequence of the song training sample. It can be understood that the processing logic for obtaining the note sequence of the song training sample is consistent with the aforementioned processing method for obtaining the syllable sequence of the song training sample, which will not be repeated here. Similarly, the note sequence can not only reflect the information of the notes in the song training sample, but also reflect the pronunciation position, pronunciation duration and other information of the notes in the song training sample, which can effectively represent the melody of the song training sample. It should be noted that the sequence length of the syllable sequence and the note sequence obtained through the score file of the song training sample is the same, so the preset unit frame length used should be the same.
[0063] See also Figure 4 For the dry audio of the song training sample, on the one hand, the lyrics text of the song training sample can be obtained by performing lyrics recognition processing on the dry audio of the song training sample. The lyrics text includes each pronunciation element in the dry audio and the time information of each pronunciation element. The time information of each pronunciation element includes the pronunciation start time (reflecting the start time of the pronunciation of the pronunciation element), the pronunciation end time (reflecting the end time of the pronunciation of the pronunciation element), and the pronunciation duration determined according to the pronunciation start time and the pronunciation end time (reflecting the duration of the pronunciation of the pronunciation element). Further, the pronunciation sequence of the song training sample can be determined according to the time information of each pronunciation element in the dry audio and the preset unit frame length (which can be set artificially, such as 10-20 milliseconds). It should be noted that the pronunciation element can refer to any one of a syllable or a phoneme. On the other hand, the dry audio can be subjected to fundamental frequency extraction processing according to the preset unit frame length to obtain the fundamental frequency sequence of the song training sample. In order to facilitate the distinction from the fundamental frequency sequence obtained by the model below, the fundamental frequency sequence here can be referred to as the first fundamental frequency sequence.
[0064] In a feasible embodiment, the dry audio of the song training sample can be first segmented according to the breath positions in the dry audio, and the multiple audios obtained after the segmentation can be further processed for lyrics recognition to obtain the lyrics text of the song training sample, wherein the breath position refers to the gap given to the singer to breathe and inhale in the song.
[0065] In one embodiment, the pronunciation sequence of the song training sample is determined according to the time information of each pronunciation element in the dry audio and the preset unit frame length, including: for any pronunciation element among the pronunciation elements, according to the preset unit frame length and the pronunciation duration of any pronunciation element, the number of any pronunciation element is determined. Wherein, the number of any pronunciation element is obtained by dividing the pronunciation duration of the pronunciation element by the preset unit frame length and performing rounding operation, and the rounding operation can be rounding up or rounding down. The number of any pronunciation element represents the number of preset unit frame lengths that the any pronunciation element needs to last. According to the preset unit frame length and the pronunciation start time of any pronunciation element, the sequence start position of any pronunciation element is determined. Wherein, the sequence start position of any pronunciation element is obtained by dividing the pronunciation start time of any pronunciation element by the preset unit frame length and performing rounding operation. According to the number of any pronunciation element and the sequence start position, each pronunciation element is sorted to obtain the pronunciation sequence of the song training sample. It can be understood that the processing logic for obtaining the pronunciation sequence of the song training sample is consistent with the aforementioned processing method for obtaining the syllable sequence of the song training sample, which will not be repeated here. The pronunciation sequence can not only reflect the information of the pronunciation elements (syllables or notes) in the song training samples, but also reflect the pronunciation position and pronunciation duration of the pronunciation elements in the song training samples, which can effectively characterize the pronunciation of the song training samples.
[0066] Further, the dry audio of the song training sample is subjected to fundamental frequency extraction processing according to the preset unit frame length to obtain the first fundamental frequency sequence of the song training sample, including: the dry audio of the song training sample is subjected to frame processing according to the preset unit frame length to obtain multiple frames of sub-audio. The fundamental frequency value of each frame of sub-audio in the multiple frames of sub-audio is obtained by a fundamental frequency extraction algorithm. Commonly used fundamental frequency extraction algorithms include autocorrelation algorithm, parallel processing method, cepstrum method and simplified inverse filtering method. Then, according to the time sequence of the multiple frames of sub-audio, the multiple fundamental frequency values of the multiple frames of sub-audio are sorted and processed to obtain the first fundamental frequency sequence of the song training sample. Among them, the sine wave with the lowest frequency in the sound signal is the fundamental tone, and the frequency of the fundamental tone is the fundamental frequency value. The first fundamental frequency sequence can reflect the change of the tone of the dry audio over time. It should be noted that the sequence length of the pronunciation sequence obtained by the dry audio of the song training sample and the first fundamental frequency sequence are the same, so the preset unit frame length used should be the same.
[0067] In an embodiment of the present application, the target fundamental frequency prediction model can use the music score file of the song to be synthesized to predict the fundamental frequency sequence of the song to be synthesized, so there is no need to extract the fundamental frequency information from the template audio, and the method of extracting the fundamental frequency information is simpler.
[0068] S203. Input the syllable sequence and note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, and train the initial fundamental frequency prediction model according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain a target fundamental frequency prediction model, wherein the target fundamental frequency prediction model is used to generate a predicted fundamental frequency sequence of the song to be synthesized according to the score file of the song to be synthesized.
[0069] See also Figure 5 , inputting the syllable sequence and note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, including: encoding each syllable in the syllable sequence of the song training sample to obtain a syllable coding sequence, the encoding process can be implemented by using N serial numbers to encode N states, for example, the three syllables "I", "love", and "you" in the syllable sequence can be encoded with serial numbers 1, 2, and 3 respectively. Encoding each note in the note sequence of the song training sample to obtain a note coding sequence, the note can express different characteristics of the sound, such as pitch, volume, length, etc. The encoding process can be encoding each note by the corresponding pitch (or frequency), the pitch can reflect the frequency of the sound object when vibrating, for example, the pitch of the note C4 is 72 (the frequency is 523 Hz), and 72 (or 523) can be used to encode C4 in the note sequence.
[0070] Furthermore, by respectively inputting the syllable coding sequence and the note coding sequence into the embedding layer (embedding layer) included in the initial fundamental frequency prediction model, feature extraction processing of the syllable coding sequence and the note coding sequence is implemented to obtain syllable sequence features corresponding to the syllable coding sequence and note sequence features corresponding to the note coding sequence (both can be vectors of fixed size). The syllable sequence features are then input into the semantic coding module included in the initial fundamental frequency prediction model to obtain semantic coding features. The language coding features can indicate the semantic content representation of the song training sample. The semantic coding module can be an encoder constructed based on a model such as a Long Short-Term Memory (LSTM) model and a BERT (Bidirectional Encoder Representations from Transformers) model, and the present application does not impose any restrictions on this.
[0071] Furthermore, the semantic coding features and the note sequence features are fused and input into the decoding module included in the initial fundamental frequency prediction model to obtain the note residual features. The fusion can be the addition or concatenation of the semantic coding features and the note sequence features. The decoding module can be a decoder built based on a deep neural network, for example, the deep neural network can be a bidirectional RNN (Recurrent Neural Network). Finally, the fundamental frequency sequence is determined based on the note residual features and the note coding sequence. In order to distinguish it from the first fundamental frequency sequence, the fundamental frequency sequence obtained here can be called the second fundamental frequency sequence, as shown in the following formula (1), which can be the second fundamental frequency sequence F0_hat obtained by adding the note residual feature F0_res and the note coding sequence Note. When singing, the pitch within the same syllable usually changes and jitters. By training the note residual features, the changes and jitters of real singing can be simulated. The dimensionality of the note residual features can be the same as the sequence length of the note coding sequence.
[0072] F0_hat=F0_res+Note (1)
[0073] In one embodiment, the server may determine the second base frequency sequence F0_hat and the first base frequency sequence y according to the following formula (2): true The difference data loss1 between the second fundamental frequency sequence and the first fundamental frequency sequence is used to train the initial fundamental frequency prediction model. One training refers to using the difference data loss1 to reversely adjust the model parameters of the initial fundamental frequency prediction model. It is understandable that the initial fundamental frequency prediction model can be trained multiple times using the score files of multiple songs in the first training set. When the difference data loss1 obtained in a certain training is less than a first preset threshold (which can be set manually) or the number of training times reaches a first preset number, the trained initial fundamental frequency prediction model can be determined as the target fundamental frequency prediction model. The target fundamental frequency prediction model can be used to generate a predicted fundamental frequency sequence of a song to be synthesized based on the score file of the song to be synthesized. The detailed process can be found in the following Figure 8 The illustrated embodiment.
[0074] loss1=(y true -F0_hat) 2 (2)
[0075] It should be explained that the present application will train the initial fundamental frequency prediction model in the direction of decreasing the difference data loss1. During this training process, the initial fundamental frequency prediction model will cultivate the ability to deeply model syllable sequences and note sequences by making the second fundamental frequency sequence more and more similar to the first fundamental frequency sequence. By using multiple information such as syllable information and note information, it can ensure that the target fundamental frequency prediction model can accurately predict the fundamental frequency sequence for the song to be synthesized.
[0076] S204, inputting the first fundamental frequency sequence and pronunciation sequence of the song training sample into the initial acoustic model to obtain the first acoustic feature, and training the initial acoustic model according to the first acoustic feature and the second acoustic feature of the dry audio to obtain a target acoustic model, wherein the target acoustic model is used to generate the predicted acoustic features of the song to be synthesized according to the predicted fundamental frequency sequence of the song to be synthesized, and the predicted acoustic features are used to generate the synthesized audio of the song to be synthesized. In order to distinguish from the acoustic features obtained by the model, the acoustic features of the dry audio can be called the second acoustic features, and the acoustic features obtained by the model can be called the first acoustic features.
[0077] See also Figure 6 , input the first fundamental frequency sequence and pronunciation sequence of the song training sample into the initial acoustic model to obtain the first acoustic feature, including: encoding each pronunciation element in the pronunciation sequence of the song training sample to obtain a pronunciation coding sequence. The encoding process can be implemented by using N serial numbers to encode N states. For example, the three phonemes "b", "p", and "m" in the pronunciation sequence can be encoded with serial numbers 1, 2, and 3 respectively; one-hot encoding can also be used. For example, the three phonemes "b", "p", and "m" in the pronunciation sequence can be encoded with
[001] ,
[010] , and
[100] respectively; of course, other encoding methods can also be used, and this application does not limit this.
[0078] Further, by respectively inputting the first base frequency sequence and the pronunciation coding sequence into the embedding layer (embedding layer) included in the initial acoustic feature, the feature extraction processing of the first base frequency sequence and the pronunciation coding sequence is realized, and the base frequency sequence features corresponding to the first base frequency sequence and the pronunciation sequence features corresponding to the pronunciation coding sequence are obtained (both can be vectors of fixed size). The base frequency sequence features and the pronunciation sequence features are then fused to obtain fused features, and the fusion processing can refer to splicing or adding the base frequency sequence features and the pronunciation sequence features. Then the fused features are input into the encoding and decoding module included in the initial acoustic model to obtain conversion features, and the encoding and decoding module includes an encoding module and a decoding module. The present application does not specifically limit the structure of the encoding module and the decoding module. For example, the encoding module can be implemented using a bidirectional recurrent neural network, such as a bidirectional LSTM (Long Short-Term Memory, long short-term memory network), and the decoder can be implemented using a unidirectional recurrent neural network, such as a unidirectional LSTM. Finally, the conversion features are input into the linear module included in the initial acoustic model to obtain the first acoustic feature, and the linear module can refer to the linear layer in the neural network.
[0079] Each singer has a unique identifier (i.e. singer identifier), see Figure 7 In a feasible embodiment, the singer identification of the singer of the dry audio of the song training sample can be obtained, and the singer identification is input into the embedding layer included in the initial acoustic model to obtain the singer feature, and then the singer feature, the fundamental frequency sequence feature and the pronunciation sequence feature are fused to obtain the above fusion feature and continue to perform subsequent steps. By integrating the singer feature, the initial acoustic feature can learn the personal characteristics of the singer when singing.
[0080] In one embodiment, the server may obtain the second acoustic feature of the dry audio of the song training sample. The second acoustic feature may be an acoustic feature such as a Mel-cephalometric coefficient obtained based on the dry audio. Furthermore, the initial acoustic model is trained once based on the difference data loss2 between the first acoustic feature and the second acoustic feature of the dry audio (which may be obtained based on the square of the difference between the first acoustic feature and the second acoustic feature of the dry audio). One training refers to reversely adjusting the model parameters of the initial acoustic model once using the difference data loss2. It is understandable that the initial acoustic model may be trained multiple times using the dry audio of at least one song in the second training set. When the difference data loss2 obtained during a certain training is less than a second preset threshold (which may be set manually) or the number of training times reaches a second preset number, the trained initial acoustic model is determined as the target acoustic model. The target acoustic model may be used to generate the predicted acoustic features of the song to be synthesized based on the predicted fundamental frequency sequence of the song to be synthesized. The predicted acoustic features are used to generate the synthesized audio of the song to be synthesized. The detailed process may be referred to as follows: Figure 8 It can be seen that the training of the initial acoustic model does not require paired dry audio and music score files, which can reduce the difficulty of obtaining training samples.
[0081] It should be explained that the present application will train the initial acoustic model in the direction of reducing the difference data loss2. During this training process, the initial acoustic model will make the first acoustic feature and the second acoustic feature more and more similar, thereby taking into account the learning of the fundamental frequency sequence and the pronunciation sequence, and can enhance the ability to characterize the acoustic features. Ultimately, by using multiple information such as pronunciation information and fundamental frequency information, it can ensure that the target generation model can accurately predict the acoustic features of the song to be synthesized.
[0082] In an embodiment of the present application, by cultivating the initial fundamental frequency prediction model's deep modeling capability for the syllable sequence and note sequence of the song training samples, the target fundamental frequency prediction model can accurately predict the fundamental frequency sequence of the song to be synthesized; at the same time, the initial acoustic model can enhance the initial acoustic model's ability to characterize acoustic features by taking into account the learning of the fundamental frequency sequence and pronunciation sequence of the song training samples, so that the target acoustic model can generate acoustic features with higher accuracy for the song to be synthesized, thereby improving the sound quality of the synthesized song and making the synthesized singing more realistic and natural.
[0083] See also Figure 8 , Figure 8 A flowchart of another audio processing method provided in an embodiment of the present application. The method can be applied to the above Figure 1 The server 12 in the method comprises:
[0084] S801, obtaining a music score file of a song to be synthesized, and determining a syllable sequence and a note sequence of the song to be synthesized according to the music score file.
[0085] In one embodiment, the terminal device may send an audio synthesis request to the server, and the server may obtain the score file of the song to be synthesized carried by the audio synthesis request; or the server may obtain the score file of the song to be synthesized stored from a local or cloud database. Then, the time information of each note in the score file of the song to be synthesized may be obtained. And the auxiliary notes corresponding to each syllable in the score file are determined, and the time information of each syllable is determined according to the time information of each auxiliary note. Specifically, the sum of the note durations of each auxiliary note is determined as the syllable duration of each syllable. The earliest note start time among the note start times of each auxiliary note is determined as the syllable start time of each syllable. The latest note end time among the note end times of each auxiliary note is determined as the syllable end time of each syllable.
[0086] Furthermore, for any syllable in each syllable, the number of any syllable is determined according to the preset unit frame length (which can be set manually, for example, 10-20 milliseconds) and the syllable duration included in the time information of any syllable. The number of any syllable is obtained by rounding the syllable duration of the syllable divided by the preset unit frame length, and the rounding operation can be rounding up or rounding down. The number of any syllable represents the number of preset unit frame lengths that the any syllable needs to last. Then, according to the preset unit frame length and the syllable start time included in the time information of any syllable, the sequence start position of any syllable is determined. The sequence start position of any syllable is obtained by rounding the syllable start time of any syllable divided by the preset unit frame length. Finally, according to the number of any syllables and the sequence start position, each syllable is sorted to obtain the syllable sequence of the song to be synthesized. And for any note among the notes, the number of any note is determined according to the preset unit frame length (which can be set manually, for example, 10-20 milliseconds) and the note duration included in the time information of any note. The number of any note is obtained by rounding the note duration of the note divided by the preset unit frame length, and the rounding operation can be rounding up or rounding down. The number of any note represents the number of preset unit frame lengths that the any note needs to last. Then, according to the preset unit frame length and the note start time included in the time information of any note, the sequence start position of any note is determined. The sequence start position of any note is obtained by rounding the note start time of any note divided by the preset unit frame length. Finally, according to the number of any note and the sequence start position, each note is sorted to obtain the note sequence of the song to be synthesized. This syllable (or note) sequence can not only reflect the syllable (or note) information itself in the song to be synthesized, but also reflect the pronunciation position, pronunciation duration and other information of the syllable (or note) in the song to be synthesized, which can effectively characterize the pronunciation condition (or melody condition) of the song to be synthesized.
[0087] S802: Input the syllable sequence and note sequence of the song to be synthesized into a target fundamental frequency prediction model to obtain a predicted fundamental frequency sequence of the song to be synthesized.
[0088] In one embodiment, the server may encode each syllable in the syllable sequence of the song to be synthesized to obtain a syllable coding sequence. The coding process can be achieved by using N serial numbers to encode N states. And each note in the note sequence of the song to be synthesized is encoded to obtain a note coding sequence. The encoding process can be encoding each note by the corresponding pitch (or frequency). Further, feature extraction is performed on the syllable coding sequence and the note coding sequence respectively to obtain syllable sequence features and note sequence features. The syllable sequence features and note sequence features can be obtained by inputting the syllable coding sequence and the note coding sequence into the embedding layer (embedding layer) included in the target fundamental frequency prediction model. The syllable sequence features are then input into the semantic coding module included in the target fundamental frequency prediction model to obtain semantic coding features. Then, the semantic coding features and the note sequence features are fused and input into the decoding module included in the target fundamental frequency prediction model to obtain note residual features. The fusion can be the addition or concatenation of the semantic coding features and the note sequence features. Finally, the predicted fundamental frequency sequence of the song to be synthesized is determined according to the note residual features and the note coding sequence. Specifically, the predicted fundamental frequency sequence of the song to be synthesized is obtained by adding the note residual features and the note coding sequence.
[0089] On the one hand, the note coding sequence (corresponding to pitch or frequency) fused in the predicted fundamental frequency sequence can indicate the frequency of vibration of the sound-producing object in each frame of sub-audio. On the other hand, the note residual features fused in the predicted fundamental frequency sequence can simulate the real sound (specifically pitch) changes and jitters. Therefore, the predicted fundamental frequency sequence can effectively simulate the real fundamental frequency sequence of the song to be synthesized.
[0090] S803. Determine a target pronunciation sequence according to the syllable sequence of the song to be synthesized, and input the target pronunciation sequence and the predicted fundamental frequency sequence into a target acoustic model to obtain predicted acoustic features of the song to be synthesized.
[0091] In one embodiment, if the pronunciation element is a syllable, the syllable sequence of the song to be synthesized is determined as the target pronunciation sequence.
[0092] In another embodiment, if the pronunciation element is a phoneme, the syllable sequence of the song to be synthesized is input into the phoneme duration prediction model to obtain the predicted phoneme duration of each phoneme in the song to be synthesized. Then, a predicted phoneme sequence is constructed according to a preset unit frame length (which can be set manually, such as 10-20 milliseconds), the predicted phoneme duration of each phoneme in the song to be synthesized, and the time information of the syllable corresponding to each phoneme. For example, the preset unit frame length is 15 milliseconds, and the syllable start times of the syllables "I", "love", and "you" in the lyrics "I love you" are 45 milliseconds, 120 milliseconds, and 145 milliseconds respectively. The phoneme durations of the phonemes "w" and "o" corresponding to the syllable "I" are 30 milliseconds and 30 milliseconds respectively, the phoneme duration of the phoneme "ai" corresponding to the syllable "love" is 45 milliseconds, and the phoneme durations of the phonemes "n" and "i" corresponding to the syllable "you" are 15 and 45 milliseconds respectively. Then the predicted phoneme sequence is [placeholder, placeholder, placeholder, w, w, o, o, placeholder, ai, ai, ai, n, i, i, i]. Finally, the predicted phoneme sequence can be determined as the target pronunciation sequence. It can be understood that the number of phonemes represents the number of preset unit frame lengths that the phoneme needs to last, which can reflect the duration that the phoneme needs to last when singing.
[0093] In one embodiment, a syllable sequence of a song training sample can be obtained, and each syllable in the syllable sequence of the song training sample is encoded to obtain a syllable encoding sequence. The encoding process can be achieved by using N serial numbers to encode N states, and can also be achieved by encoding methods such as one-hot encoding. The syllable encoding sequence is input into the embedding layer (embedding layer) included in the initial duration prediction model to obtain syllable sequence features, and then the syllable sequence features are input into the semantic encoding module included in the initial duration prediction model to obtain semantic encoding features. Then the semantic encoding features are input into the decoding module included in the initial duration prediction model to obtain decoding features, and then the decoding features are input into the linear module included in the initial duration prediction model to obtain the training phoneme duration of each phoneme in the song training sample. The initial duration prediction model is further trained according to the real phoneme duration and the training phoneme duration of each phoneme in the song training sample to obtain the phoneme duration prediction model. Specifically, the initial duration prediction model can be trained once according to the difference data loss3 between the real phoneme duration and the training phoneme duration of each phoneme in the song training sample (which can be obtained according to the square of the difference between the real phoneme duration and the training phoneme duration). One training refers to reversely adjusting the model parameters of the initial duration prediction model once using the difference data loss3. It is understandable that the initial duration prediction model can be trained multiple times using syllable training of multiple songs. When the difference data loss3 obtained in a certain training is less than a third preset threshold (which can be set manually) or the number of training times reaches a third preset number, the trained initial duration prediction model is determined as the phoneme duration prediction model. The above processing logic of inputting the syllable sequence of the song to be synthesized into the phoneme duration prediction model to obtain the predicted phoneme duration of each phoneme in the song to be synthesized is consistent with the processing logic of inputting the syllable sequence of the song training sample into the initial duration prediction model during the training process to obtain the training phoneme duration of each phoneme in the song training sample, and will not be repeated here.
[0094] In one embodiment, each pronunciation element in the target pronunciation sequence of the song to be synthesized is encoded to obtain a pronunciation coding sequence. The encoding process can be implemented by using N serial numbers to encode N states. For example, the four phonemes "b", "ao", "n", and "i" in the target pronunciation sequence can be encoded by 1, 2, 3, and 4, respectively. It can also be implemented by other encoding methods such as one-hot encoding. Then, the predicted base frequency sequence and pronunciation coding sequence of the song to be synthesized are subjected to feature extraction processing to obtain base frequency sequence features and pronunciation sequence features. The base frequency sequence features and pronunciation sequence features can be obtained by inputting the first base frequency sequence and the pronunciation coding sequence into the embedding layer (embedding layer) included in the target acoustic feature. The base frequency sequence features and the pronunciation sequence features are further fused to obtain fused features. The fusion process can refer to splicing or adding the base frequency sequence features and the pronunciation sequence features. Then the fused features are input into the encoding and decoding module included in the target acoustic model to obtain conversion features. Finally, the conversion features are input into the linear module included in the target acoustic model to obtain predicted acoustic features. The target acoustic model can generate a highly accurate predicted acoustic feature by fusing a target pronunciation sequence indicating the pronunciation of the song to be synthesized and a predicted fundamental frequency sequence indicating the pitch change of the song to be synthesized.
[0095] In a feasible embodiment, if one wants to generate audio with personal characteristics, one can determine the singer identification, input the singer identification into the embedding layer included in the target acoustic model, obtain the singer features, fuse the singer features, fundamental frequency sequence features and pronunciation sequence features to obtain the above-mentioned fused features, and finally use the predicted acoustic features generated by the above-mentioned fused features to generate synthetic audio with personal characteristics.
[0096] In summary, see Fig. 9 When the pronunciation element is a syllable, the score file of the song to be synthesized can be used to obtain each syllable in the song to be synthesized and the corresponding time information, thereby obtaining a syllable sequence. And the score file of the song to be synthesized can be used to obtain each note in the song to be synthesized and the corresponding time information, thereby obtaining a note sequence. Further, the syllable sequence and the note sequence are input into the target fundamental frequency prediction model to obtain a predicted fundamental frequency sequence. The predicted fundamental frequency sequence and the syllable sequence are input into the target acoustic model to obtain predicted acoustic features. See Fig.10When the pronunciation element is a phoneme, the score file of the song to be synthesized can be used to obtain each syllable and the corresponding time information in the song to be synthesized, thereby obtaining a syllable sequence, and the syllable sequence is input into the phoneme duration prediction model to obtain the predicted phoneme duration of each phoneme in the song to be synthesized, and the predicted phoneme sequence is determined according to the predicted phoneme duration of each phoneme in the song to be synthesized. And the score file of the song to be synthesized can be used to obtain each note in the song to be synthesized and the corresponding time information, thereby obtaining a note sequence, and further, the syllable sequence and the note sequence are input into the target fundamental frequency prediction model to obtain a predicted fundamental frequency sequence. Finally, the predicted fundamental frequency sequence and the predicted phoneme sequence are input into the target acoustic model to obtain the predicted acoustic features.
[0097] S804: Calling a vocoder to perform audio synthesis processing on the predicted acoustic features to obtain synthesized audio of the song to be synthesized.
[0098] The vocoder can encode the acoustic features to generate a sound waveform. Therefore, the present application uses the vocoder to perform audio synthesis processing on the predicted acoustic features to reconstruct the speech waveform and obtain the synthesized audio of the song to be synthesized, which is a dry audio. The accompaniment audio of the song to be synthesized can also be obtained, and the accompaniment audio of the song to be synthesized and the synthesized audio are used to generate the playback audio of the song to be synthesized.
[0099] In an embodiment of the present application, the target fundamental frequency prediction model can be used to accurately predict the fundamental frequency sequence of the song to be synthesized, so that the target acoustic model can generate highly accurate acoustic features for the song to be synthesized through the obtained fundamental frequency sequence and pronunciation sequence, thereby improving the sound quality of the synthesized song and making the synthesized singing more realistic and natural.
[0100] It can be understood that in the specific implementation of the present application, relevant data such as music score files and dry audio are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0101] The above detailed description of the method of the embodiment of the present application, in order to facilitate better implementation of the above method of the embodiment of the present application, the following provides a device of the embodiment of the present application. Fig.11 , Fig.11 1 is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present application. In one embodiment, the audio processing device 110 may include:
[0102] An acquisition unit 1101 is used to acquire a music score file and dry audio of a song training sample;
[0103] The processing unit 1102 is used to determine the syllable sequence and the note sequence of the song training sample according to the music score file, and determine the first fundamental frequency sequence and the pronunciation sequence of the song training sample according to the dry sound audio;
[0104] The processing unit 1102 is further configured to input the syllable sequence and the note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, and train the initial fundamental frequency prediction model according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain a target fundamental frequency prediction model, wherein the target fundamental frequency prediction model is used to generate the predicted fundamental frequency sequence of the song to be synthesized according to the score file of the song to be synthesized;
[0105] The processing unit 1102 is further used to input the first fundamental frequency sequence and pronunciation sequence of the song training sample into the initial acoustic model to obtain a first acoustic feature, and train the initial acoustic model according to the first acoustic feature and the second acoustic feature of the dry audio to obtain a target acoustic model, wherein the target acoustic model is used to generate the predicted acoustic features of the song to be synthesized according to the predicted fundamental frequency sequence of the song to be synthesized, and the predicted acoustic features are used to generate the synthesized audio of the song to be synthesized.
[0106] In one embodiment, the acquisition unit 1101 is specifically used to: acquire the time information of each syllable in the music score file, and determine the syllable sequence of the song training sample according to the preset unit frame length and the time information of each syllable; acquire the time information of each note in the music score file, and determine the note sequence of the song training sample according to the preset unit frame length and the time information of each note; acquire the time information of each pronunciation element in the dry sound audio, and determine the pronunciation sequence of the song training sample according to the preset unit frame length and the time information of each pronunciation element, the pronunciation element includes any one of a syllable and a phoneme;
[0107] The processing unit 1102 is specifically used to: perform fundamental frequency extraction processing on the dry sound audio according to the preset unit frame length to obtain a first fundamental frequency sequence of the song training sample.
[0108] In one embodiment, the time information of the pronunciation elements includes the pronunciation duration and the pronunciation start time; the processing unit 1102 is specifically used to: for each of the pronunciation elements, determine the number of the pronunciation elements according to the preset unit frame length and the pronunciation duration of the pronunciation element, and determine the sequence start position of the pronunciation element according to the preset unit frame length and the pronunciation start time of the pronunciation element; sort the pronunciation elements according to the number of each of the pronunciation elements and the sequence start position to obtain the pronunciation sequence of the song training sample.
[0109] In one embodiment, the processing unit 1102 is specifically configured to: perform frame processing on the dry audio according to the preset unit frame length to obtain multiple frames of sub-audio;
[0110] The acquisition unit 1101 is specifically used to: acquire the fundamental frequency value of each frame of sub-audio in the multiple frames of sub-audio;
[0111] The processing unit 1102 is specifically used to sort the multiple fundamental frequency values of the multiple frames of sub-audio to obtain the first fundamental frequency sequence of the song training sample.
[0112] In one embodiment, the processing unit 1102 is specifically used to: encode each pronunciation element in the pronunciation sequence of the song training sample to obtain a pronunciation coding sequence; extract features from the first fundamental frequency sequence and the pronunciation coding sequence of the song training sample to obtain fundamental frequency sequence features and pronunciation sequence features; fuse the fundamental frequency sequence features and the pronunciation sequence features to obtain fused features; input the fused features into the encoding and decoding module included in the initial acoustic model to obtain conversion features; input the conversion features into the linear module included in the initial acoustic model to obtain the first acoustic features.
[0113] In one embodiment, the processing unit 1102 is specifically used to: encode each syllable in the syllable sequence of the song training sample to obtain a syllable code sequence, and encode each note in the note sequence of the song training sample to obtain a note code sequence; perform feature extraction on the syllable code sequence and the note code sequence to obtain syllable sequence features and note sequence features; input the syllable sequence features into a semantic coding module included in an initial fundamental frequency prediction model to obtain semantic coding features, and fuse the semantic coding features with the note sequence features and input them into a decoding module included in the initial fundamental frequency prediction model to obtain note residual features; determine a second fundamental frequency sequence based on the note residual features and the note code sequence.
[0114] In another embodiment, the audio processing device 110 may include:
[0115] The acquisition unit 1101 is used to acquire the music score file of the song to be synthesized;
[0116] The processing unit 1102 is used to determine the syllable sequence and the note sequence of the song to be synthesized according to the music score file;
[0117] The processing unit 1102 is further configured to input the syllable sequence and the note sequence of the song to be synthesized into a target fundamental frequency prediction model to obtain a predicted fundamental frequency sequence of the song to be synthesized;
[0118] The processing unit 1102 is further used to determine a target pronunciation sequence according to the syllable sequence of the song to be synthesized, and input the target pronunciation sequence and the predicted fundamental frequency sequence into a target acoustic model to obtain predicted acoustic features of the song to be synthesized;
[0119] The processing unit 1102 is further used to call a vocoder to perform audio synthesis processing on the predicted acoustic features to obtain the synthesized audio of the song to be synthesized.
[0120] In one embodiment, the processing unit 1102 is specifically used for: if the pronunciation element is a syllable, then determining the syllable sequence of the song to be synthesized as the target pronunciation sequence; if the pronunciation element is a phoneme, then inputting the syllable sequence of the song to be synthesized into a phoneme duration prediction model to obtain the predicted phoneme duration of each phoneme in the song to be synthesized, and determining the target pronunciation sequence according to a preset unit frame length and the predicted phoneme duration of each phoneme in the song to be synthesized; wherein the phoneme duration prediction model is obtained by training an initial duration prediction model according to the actual phoneme duration and the training phoneme duration of each phoneme in the song training sample, and the training phoneme duration is obtained by inputting the syllable sequence of the song training sample into the initial duration prediction model.
[0121] It can be understood that the functions of each functional unit of the audio processing device described in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment, and its specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.
[0122] In an embodiment of the present application, by cultivating the initial fundamental frequency prediction model's deep modeling capability for the syllable sequence and note sequence of the song training samples, the target fundamental frequency prediction model can accurately predict the fundamental frequency sequence of the song to be synthesized; at the same time, the initial acoustic model can enhance the initial acoustic model's ability to characterize acoustic features by taking into account the learning of the fundamental frequency sequence and pronunciation sequence of the song training samples, so that the target acoustic model can generate acoustic features with higher accuracy for the song to be synthesized, thereby improving the sound quality of the synthesized song and making the synthesized singing more realistic and natural.
[0123] like Fig.12 As shown, Fig.12 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The internal structure of the computer device 120 is as follows Fig.12 As shown, it includes: one or more processors 1201, memory 1202, and communication interface 1203. The processor 1201, memory 1202, and communication interface 1203 may be connected via bus 1204 or other methods. The embodiment of the present application takes the connection via bus 1204 as an example.
[0124] Among them, the processor 1201 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device 120, which can parse various instructions in the computer device 120 and process various data of the computer device 120. For example, the CPU can be used to parse the power on and off instructions sent by the user to the computer device 120, and control the computer device 120 to perform power on and off operations; for another example, the CPU can transmit various interactive data between the internal structures of the computer device 120, and so on. The communication interface 1203 can optionally include a standard wired interface, a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), which is controlled by the processor 1201 to send and receive data. The memory 1202 (Memory) is a memory device in the computer device 120, which is used to store computer programs and data. It can be understood that the memory 1202 here can include both the built-in memory of the computer device 120 and the extended memory supported by the computer device 120. The memory 1202 provides a storage space, which stores the operating system of the computer device 120, which may include but is not limited to: Windows system, Linux system, Android system, iOS system, etc., and this application does not limit this. In one embodiment, the processor 1201 performs the following operations by running the computer program stored in the memory 1202:
[0125] Obtain the music score file and dry audio of the song training sample;
[0126] Determine the syllable sequence and note sequence of the song training sample according to the music score file, and determine the first fundamental frequency sequence and pronunciation sequence of the song training sample according to the dry sound audio;
[0127] Inputting the syllable sequence and the note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, and training the initial fundamental frequency prediction model according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain a target fundamental frequency prediction model, wherein the target fundamental frequency prediction model is used to generate the predicted fundamental frequency sequence of the song to be synthesized according to the score file of the song to be synthesized;
[0128] The first fundamental frequency sequence and pronunciation sequence of the song training sample are input into the initial acoustic model to obtain the first acoustic feature, and the initial acoustic model is trained according to the first acoustic feature and the second acoustic feature of the dry sound audio to obtain a target acoustic model, wherein the target acoustic model is used to generate the predicted acoustic features of the song to be synthesized according to the predicted fundamental frequency sequence of the song to be synthesized, and the predicted acoustic features are used to generate the synthesized audio of the song to be synthesized.
[0129] In one embodiment, the processor 1201 is specifically used to: obtain time information of each syllable in the music score file, and determine the syllable sequence of the song training sample according to a preset unit frame length and the time information of each syllable; obtain time information of each note in the music score file, and determine the note sequence of the song training sample according to the preset unit frame length and the time information of each note; obtain time information of each pronunciation element in the dry sound audio, and determine the pronunciation sequence of the song training sample according to the preset unit frame length and the time information of each pronunciation element, the pronunciation element includes any one of a syllable and a phoneme; perform fundamental frequency extraction processing on the dry sound audio according to the preset unit frame length to obtain a first fundamental frequency sequence of the song training sample.
[0130] In one embodiment, the time information of the pronunciation elements includes the pronunciation duration and the pronunciation start time; the processor 1201 is specifically used to: for each of the pronunciation elements, determine the number of the pronunciation elements according to the preset unit frame length and the pronunciation duration of the pronunciation element, and determine the sequence start position of the pronunciation element according to the preset unit frame length and the pronunciation start time of the pronunciation element; sort the pronunciation elements according to the number of each pronunciation element and the sequence start position to obtain the pronunciation sequence of the song training sample.
[0131] In one embodiment, the processor 1201 is specifically used to: perform frame processing on the dry sound audio according to the preset unit frame length to obtain multiple frames of sub-audio; obtain the fundamental frequency value of each frame of sub-audio in the multiple frames of sub-audio; sort the multiple fundamental frequency values of the multiple frames of sub-audio to obtain the first fundamental frequency sequence of the song training sample.
[0132] In one embodiment, the processor 1201 is specifically used to: encode each pronunciation element in the pronunciation sequence of the song training sample to obtain a pronunciation coding sequence; extract features from the first fundamental frequency sequence and the pronunciation coding sequence of the song training sample to obtain fundamental frequency sequence features and pronunciation sequence features; fuse the fundamental frequency sequence features and the pronunciation sequence features to obtain fused features; input the fused features into the encoding and decoding module included in the initial acoustic model to obtain conversion features; input the conversion features into the linear module included in the initial acoustic model to obtain the first acoustic features.
[0133] In one embodiment, the processor 1201 is specifically used to: encode each syllable in the syllable sequence of the song training sample to obtain a syllable code sequence, and encode each note in the note sequence of the song training sample to obtain a note code sequence; perform feature extraction on the syllable code sequence and the note code sequence to obtain syllable sequence features and note sequence features; input the syllable sequence features into a semantic coding module included in an initial fundamental frequency prediction model to obtain semantic coding features, and fuse the semantic coding features and the note sequence features and input them into a decoding module included in the initial fundamental frequency prediction model to obtain note residual features; determine a second fundamental frequency sequence based on the note residual features and the note code sequence.
[0134] In another embodiment, the processor 1201 performs the following operations by running the computer program stored in the memory 1202:
[0135] Acquire a music score file of a song to be synthesized, and determine a syllable sequence and a note sequence of the song to be synthesized according to the music score file;
[0136] Inputting the syllable sequence and the note sequence of the song to be synthesized into the target fundamental frequency prediction model to obtain the predicted fundamental frequency sequence of the song to be synthesized;
[0137] Determining a target pronunciation sequence according to the syllable sequence of the song to be synthesized, and inputting the target pronunciation sequence and the predicted fundamental frequency sequence into a target acoustic model to obtain predicted acoustic features of the song to be synthesized;
[0138] The vocoder is called to perform audio synthesis processing on the predicted acoustic features to obtain the synthesized audio of the song to be synthesized.
[0139] In one embodiment, the processor 1201 is specifically used for: if the pronunciation element is a syllable, then determining the syllable sequence of the song to be synthesized as the target pronunciation sequence; if the pronunciation element is a phoneme, then inputting the syllable sequence of the song to be synthesized into a phoneme duration prediction model to obtain the predicted phoneme duration of each phoneme in the song to be synthesized, and determining the target pronunciation sequence according to a preset unit frame length and the predicted phoneme duration of each phoneme in the song to be synthesized; wherein the phoneme duration prediction model is obtained by training an initial duration prediction model according to the actual phoneme duration and the training phoneme duration of each phoneme in the song training sample, and the training phoneme duration is obtained by inputting the syllable sequence of the song training sample into the initial duration prediction model.
[0140] In a specific implementation, the processor 1201, memory 1202 and communication interface 1203 described in the embodiment of the present application can execute the implementation method described in an audio processing method provided in an embodiment of the present application, and can also execute the implementation method described in an audio processing device provided in an embodiment of the present application, which will not be repeated here.
[0141] In an embodiment of the present application, by cultivating the initial fundamental frequency prediction model's deep modeling capability for the syllable sequence and note sequence of the song training samples, the target fundamental frequency prediction model can accurately predict the fundamental frequency sequence of the song to be synthesized; at the same time, the initial acoustic model can enhance the initial acoustic model's ability to characterize acoustic features by taking into account the learning of the fundamental frequency sequence and pronunciation sequence of the song training samples, so that the target acoustic model can generate acoustic features with higher accuracy for the song to be synthesized, thereby improving the sound quality of the synthesized song and making the synthesized singing more realistic and natural.
[0142] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is run on a computer device, the computer device executes the audio processing method of any possible implementation method described above. The specific implementation method can be referred to the above description, and will not be repeated here.
[0143] The embodiment of the present application also provides a computer program product, which includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the steps of the audio processing method provided in the embodiment of the present application are implemented. The specific implementation method can be referred to the above description, which will not be repeated here.
[0144] The embodiment of the present application also provides a computer program, the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, the processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing method provided in the embodiment of the present application. The specific implementation method can be referred to the above description, which will not be repeated here.
[0145] It should be noted that, for the above-mentioned various method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0146] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0147] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. An audio processing method, characterized in that: The method comprises: Obtain the music score file and dry audio of the song training sample; Determine the syllable sequence and note sequence of the song training sample according to the music score file, and determine the first fundamental frequency sequence and pronunciation sequence of the song training sample according to the dry sound audio; Inputting the syllable sequence and the note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence, and training the initial fundamental frequency prediction model according to the second fundamental frequency sequence and the first fundamental frequency sequence to obtain a target fundamental frequency prediction model, wherein the target fundamental frequency prediction model is used to generate the predicted fundamental frequency sequence of the song to be synthesized according to the score file of the song to be synthesized; The first fundamental frequency sequence and pronunciation sequence of the song training sample are input into the initial acoustic model to obtain the first acoustic feature, and the initial acoustic model is trained according to the first acoustic feature and the second acoustic feature of the dry sound audio to obtain a target acoustic model, wherein the target acoustic model is used to generate the predicted acoustic features of the song to be synthesized according to the predicted fundamental frequency sequence of the song to be synthesized, and the predicted acoustic features are used to generate the synthesized audio of the song to be synthesized.
2. The method according to claim 1, characterized in that The step of determining the syllable sequence and the note sequence of the song training sample according to the music score file, and determining the first fundamental frequency sequence and the pronunciation sequence of the song training sample according to the dry sound audio, comprises: Acquire the time information of each syllable in the music score file, and determine the syllable sequence of the song training sample according to the preset unit frame length and the time information of each syllable; Acquire the time information of each note in the music score file, and determine the note sequence of the song training sample according to the preset unit frame length and the time information of each note; Acquire time information of each pronunciation element in the dry audio, and determine the pronunciation sequence of the song training sample according to the preset unit frame length and the time information of each pronunciation element, wherein the pronunciation element includes any one of a syllable and a phoneme; The baseband frequency of the dry audio is extracted according to the preset unit frame length to obtain a first baseband frequency sequence of the song training sample.
3. The method according to claim 2, characterized in that The time information of the pronunciation elements includes the pronunciation duration and the pronunciation start time; and determining the pronunciation sequence of the song training sample according to the preset unit frame length and the time information of each pronunciation element includes: For each of the pronunciation elements, the number of the pronunciation elements is determined according to the preset unit frame length and the pronunciation duration of the pronunciation element, and the sequence start position of the pronunciation element is determined according to the preset unit frame length and the pronunciation start time of the pronunciation element; According to the number of each pronunciation element and the starting position of the sequence, the pronunciation elements are sorted to obtain the pronunciation sequence of the song training sample.
4. The method according to claim 2, characterized in that: The step of performing fundamental frequency extraction processing on the dry sound audio according to the preset unit frame length to obtain a first fundamental frequency sequence of the song training sample includes: Performing frame processing on the dry audio according to the preset unit frame length to obtain multiple frames of sub-audio; Obtaining a fundamental frequency value of each frame of sub-audio in the multiple frames of sub-audio; The multiple fundamental frequency values of the multiple frames of sub-audio are sorted to obtain a first fundamental frequency sequence of the song training sample.
5. The method according to any one of claims 1 to 4, characterized in that The step of inputting the first fundamental frequency sequence and pronunciation sequence of the song training sample into the initial acoustic model to obtain the first acoustic feature comprises: Encoding each pronunciation element in the pronunciation sequence of the song training sample to obtain a pronunciation coding sequence; Performing feature extraction processing on the first fundamental frequency sequence and the pronunciation coding sequence of the song training sample respectively to obtain fundamental frequency sequence features and pronunciation sequence features, and performing fusion processing on the fundamental frequency sequence features and the pronunciation sequence features to obtain fusion features; The fusion feature is input into a codec module included in the initial acoustic model to obtain a conversion feature, and the conversion feature is input into a linear module included in the initial acoustic model to obtain a first acoustic feature.
6. The method according to any one of claims 1 to 4, characterized in that The step of inputting the syllable sequence and the note sequence of the song training sample into the initial fundamental frequency prediction model to obtain a second fundamental frequency sequence comprises: Encoding each syllable in the syllable sequence of the song training sample to obtain a syllable encoding sequence, and encoding each note in the note sequence of the song training sample to obtain a note encoding sequence; Performing feature extraction processing on the syllable code sequence and the note code sequence respectively to obtain syllable sequence features and note sequence features; Inputting the syllable sequence feature into a semantic coding module included in an initial fundamental frequency prediction model to obtain a semantic coding feature, and fusing the semantic coding feature with the note sequence feature and inputting it into a decoding module included in the initial fundamental frequency prediction model to obtain a note residual feature; A second fundamental frequency sequence is determined according to the note residual feature and the note encoding sequence.
7. An audio processing method, characterized in that: The method comprises: Acquire a music score file of a song to be synthesized, and determine a syllable sequence and a note sequence of the song to be synthesized according to the music score file; Inputting the syllable sequence and note sequence of the song to be synthesized into the target fundamental frequency prediction model of the audio processing method according to any one of claims 1 to 6 to obtain the predicted fundamental frequency sequence of the song to be synthesized; Determine a target pronunciation sequence according to the syllable sequence of the song to be synthesized, and input the target pronunciation sequence and the predicted fundamental frequency sequence into a target acoustic model of an audio processing method according to any one of claims 1 to 6 to obtain predicted acoustic features of the song to be synthesized; The vocoder is called to perform audio synthesis processing on the predicted acoustic features to obtain the synthesized audio of the song to be synthesized.
8. The method according to claim 7, characterized in that The step of determining the target pronunciation sequence according to the syllable sequence of the song to be synthesized comprises: If the pronunciation element is a syllable, the syllable sequence of the song to be synthesized is determined as the target pronunciation sequence; If the pronunciation element is a phoneme, the syllable sequence of the song to be synthesized is input into a pre-trained phoneme duration prediction model to obtain the predicted phoneme duration of each phoneme in the song to be synthesized, and the target pronunciation sequence is determined based on the preset unit frame length and the predicted phoneme duration of each phoneme in the song to be synthesized; wherein the phoneme duration prediction model is trained by the phoneme duration of each phoneme in the song training sample.
9. A computer device, characterized in that: The computer device includes a memory, a communication interface and a processor, wherein the memory, the communication interface and the processor are interconnected; the memory stores a computer program, and the processor calls the computer program stored in the memory to implement the audio processing method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the audio processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Hybrid speech synthesizer, method and use
CN101156196A
Audio synthesis method and device, computer equipment and storage medium
CN114360492A