Singing voice synthesis method, device, computer device and storage medium
By determining the syllable and phoneme length information based on the music score and lyrics in the singing synthesis technology, and obtaining the acoustic characteristic parameters of the singing voice, the problem of difficulty in improving the naturalness of singing synthesis in the existing technology is solved, and a more natural and expressive singing synthesis is achieved.
Patent Information
- Application Number
- CN202110942245.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-08-17
AI Technical Summary
The existing singing vocal synthesis technology is difficult to improve the naturalness of singing vocal synthesis, and the accuracy of the acoustic model is limited, which makes it difficult to improve the naturalness of the synthetic singing after reaching a bottleneck.
By determining the syllable duration information and phoneme duration information based on the music score and lyrics, acoustic characteristic parameters of multiple audio frames to be generated are obtained, and the singing voice is generated using these characteristic parameters.
The naturalness of singing vocal synthesis is improved, and through precise control at the phoneme level, it exceeds the accuracy limitations of traditional acoustic models, and significantly improves the naturalness and expressiveness of synthetic vocals.
Smart Images

Figure CN114283789B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular, to a singing synthesis method, apparatus, computer device, and storage medium. Background Art
[0002] With the development and progress of audio technology, singing synthesis technology has always been highly concerned. In singing synthesis technology, a digital score is provided to a computer device so that the computer device synthesizes the singing corresponding to the score, and then with accompaniment, a song similar to human singing can be obtained.
[0003] Currently, when synthesizing singing, a given score can be converted into corresponding labels, and then a trained acoustic model is used to predict the acoustic features required for synthesizing singing. Finally, a vocoder is used to synthesize the singing. Since the accuracy of the acoustic model has an upper limit, it is difficult to improve the naturalness of the synthesized singing after reaching a bottleneck. Therefore, a method that can improve the naturalness of singing synthesis is needed. Summary of the Invention
[0004] Embodiments of this application provide a singing synthesis method, apparatus, computer device, and storage medium, which can improve the naturalness of singing synthesis. The technical solution is as follows:
[0005] On the one hand, a singing generation method is provided, and the method includes:
[0006] Based on the score and the lyrics corresponding to the score, determine syllable duration information, where the syllable duration information is used to represent the number of audio frames occupied by each of the multiple syllables in the lyrics;
[0007] Based on the score, the lyrics, and the syllable duration information, determine phoneme duration information, where the phoneme duration information is used to represent the number of audio frames occupied by each of the multiple phonemes included in the multiple syllables;
[0008] Based on the score, the lyrics, and the phoneme duration information, obtain acoustic feature parameters of multiple audio frames in the singing to be generated, where the acoustic feature parameters are used to represent the acoustic features of the corresponding audio frames;
[0009] Generate the singing to be generated based on the acoustic feature parameters of the multiple audio frames.
[0010] On the other hand, a singing generation apparatus is provided, and the apparatus includes:
[0011] A first determination module, configured to determine syllable duration information based on a score and the lyrics corresponding to the score, where the syllable duration information is used to represent the number of audio frames occupied by each of the multiple syllables in the lyrics;
[0012] A second determination module, configured to determine phoneme duration information based on the musical score, the lyrics, and the syllable duration information, where the phoneme duration information is used to represent the number of audio frames occupied by each of the multiple phonemes included in the multiple syllables;
[0013] An acquisition module, configured to acquire acoustic feature parameters of multiple audio frames in the to-be-generated singing voice based on the musical score, the lyrics, and the phoneme duration information, where the acoustic feature parameters are used to represent the acoustic features of the corresponding audio frames;
[0014] A generation module, configured to generate the to-be-generated singing voice based on the acoustic feature parameters of the multiple audio frames.
[0015] In a possible implementation manner, the second determination module includes:
[0016] A first acquisition sub-module, configured to acquire the multiple syllables and the multiple phonemes in the lyrics;
[0017] The first acquisition sub-module is further configured to acquire multiple pitches corresponding to multiple notes in the musical score, where the multiple notes correspond to the multiple syllables;
[0018] A second acquisition sub-module, configured to acquire a first semantic feature of the multiple syllables based on the multiple syllables, the multiple phonemes, the multiple pitches, and the singer identifier, where the first semantic feature is used to represent the semantics when the singer sings the corresponding syllable at the corresponding pitch;
[0019] A determination sub-module, configured to determine the phoneme duration information based on the syllable duration information and the first semantic feature of the multiple syllables.
[0020] In a possible implementation manner, the determination sub-module includes:
[0021] A first acquisition unit, configured to acquire initial features of the multiple audio frames based on the syllable duration information and the first semantic feature of the multiple syllables, where the initial features are the first semantic features of the syllables to which the corresponding audio frames belong;
[0022] A first determination unit, configured to determine multiple phonemes corresponding to each of the multiple audio frames based on the initial features of the multiple audio frames and the position features of the multiple audio frames;
[0023] A second determination unit, configured to determine the phoneme duration information based on the correspondence between the multiple audio frames and the multiple phonemes.
[0024] In a possible implementation manner, the first determination unit includes:
[0025] A splicing subunit, configured to splice the initial features of the multiple audio frames and the position features of the multiple audio frames respectively to obtain multiple first spliced features;
[0026] A convolution subunit, configured to perform convolution processing on the multiple first spliced features to obtain multiple first target features;
[0027] A weighting subunit, configured to perform weighting processing on the multiple first target features to obtain multiple second target features;
[0028] A prediction subunit, configured to predict multiple phonemes corresponding to the multiple audio frames respectively based on the multiple second target features.
[0029] In a possible implementation manner, the prediction subunit is configured to:
[0030] Perform a fully connected process on any one of the multiple second target features to obtain multiple prediction probabilities of the audio frame corresponding to the second target feature, where the prediction probability is used to represent the possibility that the audio frame corresponds to a phoneme;
[0031] Determine the phoneme with the highest prediction probability as the phoneme corresponding to the audio frame.
[0032] In a possible implementation manner, the first obtaining unit is configured to:
[0033] For any one of the multiple syllables, copy the first semantic feature of the syllable for a first target number of times to obtain the initial features of the multiple audio frames included in the syllable, where the first target number of times is a value obtained by subtracting one from the number of audio frames occupied by the syllable.
[0034] In a possible implementation manner, the obtaining module includes:
[0035] A third obtaining sub-module, configured to obtain second semantic features of the multiple audio frames based on the lyrics, the musical score, and the phoneme duration information, where the second semantic features are used to represent the semantics when the singer sings the corresponding phoneme at the corresponding pitch;
[0036] An encoding sub-module, configured to encode the second semantic features of the multiple audio frames to obtain intermediate features of the multiple audio frames;
[0037] A decoding sub-module, configured to decode the intermediate features of the multiple audio frames to obtain third semantic features of the multiple audio frames;
[0038] A processing sub-module, configured to process the third semantic features of the multiple audio frames to obtain acoustic feature parameters of the multiple audio frames.
[0039] In a possible implementation manner, the third acquisition sub-module includes:
[0040] A second acquisition unit, configured to acquire the phoneme features of the multiple phonemes in the lyrics and the pitch features of the multiple pitches corresponding to the multiple notes in the music score, where the multiple notes correspond to the multiple syllables;
[0041] A third acquisition unit, configured to acquire the frame-level phoneme features of the multiple audio frames based on the phoneme features of the multiple phonemes and the phoneme duration information, where the frame-level phoneme features are the phoneme features of the phoneme to which the corresponding audio frame belongs;
[0042] A fourth acquisition unit, configured to acquire the frame-level pitch features of the multiple audio frames based on the pitch features of the multiple pitches and the phoneme duration information, where the frame-level pitch features are the pitch features of the pitch corresponding to the phoneme to which the corresponding audio frame belongs;
[0043] A splicing unit, configured to splice the frame-level phoneme features of the multiple audio frames with the frame-level pitch features of the multiple audio frames and the singer features of the singer identifier respectively, to obtain the second semantic features of the multiple audio frames.
[0044] In a possible implementation manner, the third acquisition unit is configured to:
[0045] For any phoneme in the multiple phonemes, duplicate the phoneme features of the phoneme a second target number of times to obtain the frame-level phoneme features of the multiple audio frames included in the phoneme, where the second target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme.
[0046] In a possible implementation manner, the fourth acquisition unit is configured to:
[0047] For any pitch in the multiple pitches, duplicate the pitch features of the pitch a third target number of times to obtain the frame-level pitch features of the multiple audio frames included in the phoneme corresponding to the pitch, where the third target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme corresponding to the pitch.
[0048] In a possible implementation manner, the apparatus further includes:
[0049] A linear processing module, configured to perform linear processing on the third semantic features of the multiple audio frames to obtain multiple third target features;
[0050] The linear processing module is further configured to perform linear processing on the frame-level pitch features of the multiple audio frames to obtain multiple fourth target features;
[0051] A splicing module, configured to splice the multiple third target features and the multiple fourth target features respectively to obtain multiple second spliced features;
[0052] A convolution module, configured to perform convolution processing on the multiple second spliced features to obtain at least one of the silence parameters or the fundamental frequency parameters of each of the multiple audio frames, where the silence parameter is used to characterize whether the corresponding audio frame is a silent segment, and the fundamental frequency parameter is used to characterize the logarithm of the fundamental frequency of the corresponding audio frame.
[0053] In a possible implementation manner, the generation module is configured to:
[0054] Splice the acoustic feature parameters of the multiple audio frames and at least one of the silence parameters or the fundamental frequency parameters of each of the multiple audio frames to obtain the target acoustic features of the multiple audio frames;
[0055] Input the target acoustic features of the multiple audio frames into a vocoder, and synthesize the audio signal of the to-be-generated singing voice through the vocoder.
[0056] On the one hand, a computer device is provided, which includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the singing voice synthesis method as described above.
[0057] On the one hand, a storage medium is provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the singing voice synthesis method as described above.
[0058] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes one or more program codes, and the one or more program codes are stored in a computer-readable storage medium. One or more processors of a computer device can read the one or more program codes from the computer-readable storage medium, and the one or more processors execute the one or more program codes, so that the computer device can execute the singing voice synthesis method as described above.
[0059] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:
[0060] Based on the syllable duration information given in the musical score, the phoneme duration information is further predicted. Since the phoneme duration information can characterize the number of audio frames occupied by each phoneme, the accuracy of the acoustic model is no longer limited to the rough syllable level, but can reach the precise control at the phoneme level, greatly improving the naturalness of the singing voice synthesis. Description of the Drawings
[0061] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0062] Figure 1 It is a schematic diagram of the principle of a singing synthesis system provided by an embodiment of the present application;
[0063] Figure 2 It is a schematic diagram of the implementation environment of a singing generation method provided by an embodiment of the present application;
[0064] Figure 3 It is a schematic diagram of the principle of a singing generation system provided by an embodiment of the present application;
[0065] Figure 4 It is a flowchart of a singing generation method provided by an embodiment of the present application;
[0066] Figure 5 It is a flowchart of a singing generation method provided by an embodiment of the present application;
[0067] Figure 6 It is a flowchart of obtaining a first semantic feature provided by an embodiment of the present application;
[0068] Figure 7 It is a flowchart of determining phoneme duration information provided by an embodiment of the present application;
[0069] Figure 8 It is a schematic diagram of the principle of a duration model provided by an embodiment of the present application;
[0070] Figure 9 It is a flowchart of obtaining a second semantic feature provided by an embodiment of the present application;
[0071] Figure 10 It is a schematic diagram of the principle of a singing generation system provided by an embodiment of the present application;
[0072] Figure 11 It is a schematic diagram of the principle of a pitch prediction model provided by an embodiment of the present application;
[0073] Figure 12 It is a schematic diagram of the structure of a singing generation device provided by an embodiment of the present application;
[0074] Figure 13 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application;
[0075] Figure 14It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0076] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe in detail the embodiments of the present application with reference to the accompanying drawings.
[0077] In the present application, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions. It should be understood that there is no logical or temporal dependence between "first", "second", and "nth", nor are the quantity and execution order limited.
[0078] In the present application, the term "at least one" means one or more, and the meaning of "multiple" means two or more. For example, multiple first positions mean two or more first positions.
[0079] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making.
[0080] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as audio processing technology, computer vision technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0081] Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction. Among them, audio processing technology (Speech Technology, also known as speech processing technology) has become one of the most promising human-computer interaction methods in the future, specifically including text-to-speech technology (TTS, also known as text-to-speech conversion technology), automatic speech recognition technology (ASR), speech separation technology, and voiceprint recognition technology, etc.
[0082] With the development of AI technology, audio processing technology has been studied and applied in many fields, such as common smart speakers, intelligent voice assistants, voice shopping systems, voice recognition products, voiceprint recognition products, smart homes, smart wearable devices, virtual assistants, intelligent marketing, driverless, autonomous driving, drones, robots, intelligent healthcare, intelligent customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of AI technology, audio processing technology will be applied in more fields and play an increasingly important role.
[0083] The embodiments of this application relate to singing synthesis technology in the field of audio processing. Singing synthesis technology is a branch of speech synthesis technology. Music is not only pleasant to listen to, but also can convey the mood of the composer and singer. Among them, various music genres can bring various auditory experiences to the listeners. While enriching the spiritual life of the listeners, music is also an art form that can reflect the emotions of people's real life. And singing is a musical form expressed through the human voice and is also the most expressive way of human language. Singing synthesis technology refers to combining relevant speech synthesis technologies, inputting sheet music and lyrics into a computer, and enabling the computer to emit beautiful and pleasant singing like a human.
[0084] Hereinafter, the terms related to the embodiments of this application will be explained.
[0085] Singing Synthesis: That is, singing generation technology, which means that a computer converts sheet music information into standard and smooth singing.
[0086] Silence parameter: An index used to measure whether an audio segment is a silent segment. For example, the silence parameter can be vuv (voice / unvoice), where vuv = 1 represents a voiced segment with pitch, and vuv = 0 represents an unvoiced segment without pitch.
[0087] MGC (Mel Generalized Cepstral) feature: It refers to the Mel-Frequency Cepstral Coefficients (MFCC) feature after dimensionality reduction.
[0088] BAP (Band Aperiodicity): Represents an aperiodic excitation signal.
[0089] Fundamental frequency parameter: Denoted as lf0 (log f0), where f0 is the fundamental frequency / pitch. Therefore, lf0 represents the value of the fundamental frequency / pitch of the audio after taking the logarithm.
[0090] Phoneme: The smallest unit of pronunciation in human language that can distinguish meanings.
[0091] In a singing voice synthesis system, two concepts are involved: musical score and lyrics. The musical score is used to represent the melody (i.e., tune) of a song, while the lyrics are used to represent the text content of a song being sung. A musical score may include a series of notes, each note having its corresponding pitch (i.e., tone), and the lyrics may include a series of characters, each character having its corresponding syllable. For example, in Chinese lyrics, a character refers to a Chinese character, and a syllable refers to the pinyin of the Chinese character. In English lyrics, a character refers to a word, and a syllable refers to the complete pronunciation unit of the word (rather than the smallest pronunciation unit). It can be understood that usually multiple phonemes are included in a syllable, and each phoneme can correspond to multiple audio frames in the speech signal.
[0092] Figure 1 is a schematic diagram of the principle of a singing voice synthesis system provided by an embodiment of the present application. As Figure 1 shown, digital musical score information 101 is provided to a computer. For example, this musical score information 101 is generally in the musicxml format. Usually, the musical score information 101 refers to: the musical score, the lyrics, and the correspondence between the notes in the musical score and the syllables in the lyrics. After performing data preprocessing on the musical score information 101, the duration information 1021, pitch information 1022, and text information 1023 included in the musical score information 101 can be obtained. The duration information 1021 refers to the duration occupied by each note in the musical score (i.e., each syllable in the lyrics), the pitch information 1022 refers to the pitch occupied by each note in the musical score, and the text information 1023 refers to each note in the musical score corresponding to each character in the lyrics. In other words, a note in the musical score corresponds to a syllable in the singing voice and a character in the lyrics. It should be noted that if it is a Chinese song, usually one note corresponds to the pinyin of one Chinese character. If it is an English song, usually one note corresponds to one English word. The duration information 1021, pitch information 1022, and text information 1023 are input into the singing voice synthesis model 103. Through the singing voice synthesis model 103, the singing voice corresponding to the musical score and the lyrics can be synthesized. Finally, combined with the accompaniment 104, an audio 105 similar to human singing can be obtained.
[0093] In the above singing synthesis system, on the one hand, although the duration information 1021 of each note is provided in the musical score information 101, due to the different lyrics context and the characteristics of the singer himself, the pronunciation duration of each phoneme within the same syllable is not directly obtained by evenly dividing the duration of the note corresponding to this syllable. For example, different singers have different durations of each phoneme within the same syllable when singing the same syllable. Another example is that even when the same singer sings the same syllable, the durations of each phoneme within this syllable are not the same according to the different lyrics context. Another example is that even when the singer and the lyrics context are fixed, the durations of each phoneme within this syllable are not necessarily evenly divided. Therefore, the above singing synthesis model 103 lacks detailed modeling of the duration, resulting in low naturalness and insufficient expressiveness of the synthesized audio 105. On the other hand, compared with speech, the pitch in singing has more drastic changes. Although the overall change trend of the pitch follows the musical score, even if the pitch markings of two characters in the musical score are the same, their actual pitches will be different according to the content of the characters. In addition, the pitch within the same character usually also has changes and jitters, that is, the above singing synthesis model 103 lacks detailed and direct control of the pitch.
[0094] In view of this, the embodiments of the present application provide a singing generation method, which can consider and model the duration in combination with the lyrics context and the characteristics of the singer himself to improve the naturalness and expressiveness of the synthesized singing. Moreover, it can also achieve detailed and direct frame-level control of the pitch, which has important value for the human-computer interaction industry and can be applied to various application scenarios such as virtual divas, digital music creation, and music education. Due to the enhanced modeling of the duration and pitch, the synthesized songs are more standard and smooth.
[0095] Figure 2 It is a schematic diagram of the implementation environment of a singing generation method provided by the embodiments of the present application. Refer to Figure 2 In this implementation environment, the terminal 201 and the server 202 are involved.
[0096] The terminal 201 is used to provide the musical score and lyrics. An application program supporting audio processing, such as an audio-video application, an audio synthesis application, a singing synthesis application, an audio post-production application, etc., can be installed and run on the terminal 201. The user can log in to this application program on the terminal 201, input the digital musical score and lyrics into this application program, and click the singing generation option to trigger the terminal 201 to send a singing generation request to the server 202. The singing generation request at least carries the user's account identifier and the digital musical score and lyrics. Among them, the digital musical score and lyrics can be data files in the musicxml format.
[0097] The terminal 201 and the server 202 can be directly or indirectly connected through wired or wireless communication means, which are not limited in this application.
[0098] The server 202 is used to provide a singing generation service. Optionally, the server 202 includes at least one of a single server, multiple servers, a cloud computing platform, or a virtualization center. Optionally, the server 202 can undertake the main computing work, and the terminal 201 can undertake the secondary computing work; or, the server 202 undertakes the secondary computing work, and the terminal 201 undertakes the main computing work; or, the terminal 201 and the server 202 adopt a distributed computing architecture for collaborative computing.
[0099] Schematically, the server 202 receives a singing generation request sent by the terminal 201, parses the singing synthesis request to obtain the user's account identifier, the digital music score and lyrics, performs authentication verification based on the account identifier, and when the authentication passes, synthesizes a segment of singing audio based on the digital music score and lyrics. Optionally, directly return the singing audio to the terminal 201, or, on the basis of the singing audio, an accompaniment audio can also be added to obtain a synthesized song audio, and the song audio is returned to the terminal 201. Optionally, the terminal 201 can also independently complete the singing generation work without sending a singing generation request to the server 202, or, the user can directly trigger and implement the singing generation work on the server 202 side without communicating with the terminal 201, which can save communication overhead.
[0100] In some embodiments, the server 202 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0101] In some embodiments, the terminal 201 can generally refer to one of multiple terminals. The device types of the terminal 201 include but are not limited to: smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, vehicle-mounted terminals, televisions, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, etc., but are not limited thereto.
[0102] Those skilled in the art will understand that the number of the above-mentioned terminals 201 may be more or less. For example, there may be only one of the above-mentioned terminals 201, or there may be dozens or hundreds of the above-mentioned terminals 201, or even more. The embodiments of the present application do not limit the number and device type of the terminals 201.
[0103] Figure 3 It is a schematic diagram of the principle of a singing voice generation system provided by an embodiment of the present application. As Figure 3 shown, when singing voice generation is performed, the user input includes: musical score 321 and lyrics 322. Since there is a corresponding relationship between the notes in the musical score 321 and the characters in the lyrics 322, the musical score 321 and the lyrics 322 can also be collectively referred to as musical score information 320. The system output includes: computer-synthesized singing voice audio 340.
[0104] In an exemplary scenario, the user inputs digitized and segmented musical score information 320. The computer device (such as a server) performs data preprocessing on the musical score information 320, and the musical score 321 and the lyrics 322 can be obtained. Among them, the musical score 321 can be a numbered musical notation or a staff notation, and the lyrics 322 can be lyrics in Chinese characters, lyrics in pinyin form, lyrics in English form, or lyrics in other language forms, etc. The embodiments of the present application do not specifically limit this. Then, the musical score 321 and the lyrics 322 are input into a singing voice synthesis system 330 based on a neural network acoustic parameter prediction model. Among them, the singing voice synthesis system 330 may include: a duration model 331, a pitch prediction model 332, an acoustic parameter prediction model 333, and a vocoder 334. Optionally, what is input into the singing voice synthesis system 330 may be the musical score 321 and the lyrics 322. Due to further processing by the singing voice synthesis system 330, a pinyin sequence corresponding to the lyrics, and the duration and pitch of each note shown in the musical score are obtained. Or, the pinyin sequence corresponding to the lyrics, and the duration and pitch of each note shown in the musical score can also be directly extracted from the further processing of the musical score 321, and the information obtained from the above processing is input into the singing voice synthesis system 330.
[0105] Optionally, the musical score 321 is input into the duration model 331 to predict the pronunciation duration of each phoneme based on the total duration occupied by each syllable given in the musical score 321. Optionally, the musical score 321 is also input into the pitch prediction model 332 to predict the voice-unvoice parameter vuv and the fundamental frequency parameter lf0 of each frame in the audio in combination with the intermediate output of the acoustic parameter prediction model 333 and the pitch of each syllable given in the musical score 321. Optionally, the lyrics 322, the output of the duration model 331 (i.e., the pronunciation duration of each phoneme), and the output of the pitch prediction model 332 (i.e., the voice-unvoice parameter vuv and the fundamental frequency parameter lf0 of each frame in the audio) are input into the acoustic parameter prediction model 333 to predict the acoustic feature parameters of the audio. Finally, the predicted acoustic feature parameters, the voice-unvoice parameter vuv, and the fundamental frequency parameter lf0 are concatenated and input into the vocoder 334 together to synthesize the final singing audio 340. Schematically, the acoustic parameter prediction model 333 can be a neural network acoustic parameter prediction model such as the FastSpeech model, and the vocoder 334 can be a WORLD vocoder, etc. The embodiments of the present application do not specifically limit this. It should be noted that the duration model 331 can be trained based on the number of frames occupied by each phoneme in the musical score and the real singing voice. Since the pitch prediction model 332 needs to use the intermediate output of the acoustic parameter prediction model 333 and needs to use the output of the duration model 331 as input, the pitch prediction model 332 and the acoustic parameter prediction model 333 can be trained together after obtaining the output of the trained duration model 331.
[0106] The singing voice synthesis system provided by the embodiments of the present application, that is, a neural network singing voice synthesis scheme based on accurate duration and pitch modeling and control. For Chinese singing voice synthesis, the user input includes digitized and segmented musical score information and lyrics pinyin information, and the system output is the computer-synthesized singing audio. Since the acoustic parameter prediction model, the vocoder, the duration model, and the pitch prediction model cooperate well, the finally synthesized singing audio has clear pronunciation and high singing voice synthesis quality. The duration model fully considers various information such as the lyrics context, the current phoneme, and the singer, ensuring that it can make a phoneme duration prediction that conforms to the characteristics of the singer himself, increasing the naturalness of the synthesized singing voice. Since the results of the pitch prediction model (voice-unvoice parameter vuv and fundamental frequency parameter lf0) are directly concatenated with the acoustic feature parameters predicted by the acoustic parameter prediction model and sent into the vocoder, the f0 of the audio generated by the vocoder highly depends on the voice-unvoice parameter vuv and the fundamental frequency parameter lf0, achieving the purpose of accurately controlling the pitch of each frame of the synthesized audio using the pitch prediction model and increasing the expressiveness of the synthesized singing voice.
[0107] Figure 4 is a flowchart of a singing voice generation method provided by the embodiments of the present application. See Figure 4, this embodiment is applied to a computer device. Taking this computer device as a server as an example for illustration, this embodiment includes the following steps:
[0108] 401. The server determines syllable duration information based on the musical score and the lyrics corresponding to the musical score. The syllable duration information is used to represent the number of audio frames occupied by each of the multiple syllables in the lyrics.
[0109] In some embodiments, the server can parse the digital musical score information to obtain the musical score and the lyrics corresponding to the musical score. Optionally, the musical score information can be a data file in the musicxml format; the musical score can be a staff notation or a numbered musical notation; the lyrics can be Chinese lyrics, English lyrics, or lyrics in other languages. The embodiments of the present application do not make specific limitations in this regard.
[0110] Schematically, the user directly inputs the digital musical score information to the server, or the user uploads the digital musical score information to the server through an application on the terminal. Alternatively, the server loads the digital musical score information from another cloud database. The embodiments of the present application do not make specific limitations on the acquisition method of the musical score information. For example, the user inputs the digital musical score information to an audio application on the terminal, clicks the singing generation option, triggers the terminal to send a singing generation request carrying the digital musical score information to the server, and the server receives the singing generation request and parses the singing generation request to obtain the digital musical score information.
[0111] In some embodiments, after the server obtains the musical score and the corresponding lyrics, it can parse the musical score to obtain the multiple pitches corresponding to the multiple notes in the musical score. The multiple pitches can be regarded as constituting a pitch sequence, which is also called a tone sequence. Optionally, the server can parse the lyrics to obtain the multiple characters in the lyrics. The multiple characters can be regarded as constituting a character sequence. For example, for Chinese lyrics, the character sequence is the Chinese character sequence, and for English lyrics, the character sequence is the word sequence. Schematically, since each Chinese character has its own pinyin, usually, the pinyin of a Chinese character corresponds to a syllable uttered during singing. Therefore, for Chinese lyrics, in addition to obtaining the Chinese character sequence, the Chinese character sequence can also be converted into a corresponding pinyin sequence, and the pinyin sequence represents the syllable sequence of the singer during singing. Optionally, since in addition to recording the notes, the musical score also implies the duration of each note in the musical score, and when the sampling rate is fixed, the duration occupied by each audio frame in the audio signal is usually also fixed (this duration can be called the frame length of the audio frame). Therefore, the server can determine the frame length of a single audio frame in the audio signal based on the sampling rate of the audio signal. Further, for each note in the musical score, dividing the duration of the note in the musical score by the frame length can obtain the number of audio frames occupied by the syllable corresponding to the note. Repeating the above steps can obtain a series of audio frame numbers, and this series of audio frame numbers can constitute a duration sequence, which is also the syllable duration information and represents the number of audio frames occupied by each syllable corresponding to each note in the lyrics. It should be noted that since a note in the musical score corresponds to both a character in the lyrics and a syllable during singing, the sequence lengths of the above pitch sequence, character sequence, syllable sequence (such as pinyin sequence), and duration sequence are equal, that is, the number of elements contained in the above four sequences is equal.
[0112] 402. The server determines phoneme duration information based on the musical score, the lyrics, and the syllable duration information, and the phoneme duration information is used to characterize the number of audio frames occupied by each of the multiple phonemes included in the multiple syllables.
[0113] In some embodiments, when determining the phoneme duration information, the server can obtain the multiple syllables and the multiple phonemes in the lyrics; obtain the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables; obtain a first semantic feature of the multiple syllables based on the multiple syllables, the multiple phonemes, the multiple pitches, and the singer identifier, and the first semantic feature is used to characterize the semantics when the singer sings the corresponding syllable at the corresponding pitch; and determine the phoneme duration information based on the syllable duration information and the first semantic feature of the multiple syllables.
[0114] Optionally, since a plurality of syllables in the lyrics, i.e., a syllable sequence, have been obtained in step 401 above, each of the plurality of syllables can be divided according to the smallest pronunciation unit capable of distinguishing meanings, so as to obtain one or more phonemes included in each syllable. Thus, by combining all the phonemes included in all the syllables, a plurality of phonemes in the lyrics, i.e., a phoneme sequence, can be obtained.
[0115] Optionally, since a plurality of pitches corresponding to the respective plurality of notes in the musical score, i.e., a pitch sequence, have been obtained in step 401 above, there is no need to repeat the obtaining step, which can save the processing resources of the server.
[0116] Optionally, the singer identifier can be used to uniquely identify the singer (i.e., the speaker). The singer identifier (Identification, ID) can be specified by the user. For example, the user selects the singer to which the voice to be synthesized this time belongs, and the server determines the singer identifier of the singer selected by the user. Alternatively, the user may not specify the singer, and the server randomly selects a singer identifier from the singer database. Alternatively, the server obtains the singer identifier with the highest selection frequency, etc. The embodiments of the present application do not specifically limit the source of the singer identifier.
[0117] Optionally, the process of the server obtaining the first semantic feature, that is, the process of the server obtaining the first semantic feature of each syllable based on the syllable sequence, the phoneme sequence, the pitch sequence, and the singer identifier. In some embodiments, the server obtains the syllable feature of each syllable based on the syllable sequence; obtains the pooling feature map of each phoneme based on the phoneme sequence; obtains the pitch feature of each note based on the pitch sequence; obtains the singer feature based on the singer identifier; then, taking the syllable as a unit, fuses the pooling feature map of each phoneme belonging to this syllable, the syllable feature of this syllable, the pitch feature of the pitch corresponding to this syllable, and the singer feature to obtain the initial semantic feature of each syllable; performs convolutional processing on the initial semantic feature of each syllable to obtain the semantic feature map of each syllable; and encodes the semantic feature map of each syllable to obtain the first semantic feature of each syllable.
[0118] In some embodiments, the server performs embedding encoding on each syllable in the syllable sequence to obtain the embedding feature of each syllable, and determines the embedding feature of each syllable as the syllable feature of each syllable to enhance the expression ability of the syllable feature. Optionally, other encoding methods such as one-hot encoding can also be used for each syllable. The embodiments of the present application do not specifically limit the encoding method of the syllable feature.
[0119] In some embodiments, the server performs embedding encoding on each phoneme in the phoneme sequence to obtain the embedding feature of each phoneme. Then, based on the sorting information of each phoneme in the phoneme sequence, the position feature of each phoneme is obtained. The embedding feature and the position feature of each phoneme are concatenated to obtain the target concatenated feature of each phoneme. The target concatenated feature of each phoneme is input into the first feature extraction model, and the first feature extraction model performs convolutional processing on the target concatenated feature of each phoneme to obtain the feature map of each phoneme. The feature map of each phoneme is input into a pooling layer, and the pooling layer performs pooling processing on the feature map of each phoneme to obtain the pooled feature map of each phoneme, where the first feature extraction model is used to extract the feature map of the input phoneme.
[0120] Schematically, the first feature extraction model can be a CNN (Convolutional Neural Network). The CNN includes one or more convolutional blocks. When the number of CNN blocks is multiple, the multiple CNN blocks are connected in series, that is, the output of the previous CNN block is used as the input of the current CNN block. Similarly, the output of the current CNN block is used as the input of the next CNN block. Among them, the input of the first CNN block is the target concatenated feature of each phoneme. At this time, the server inputs the target concatenated feature of each phoneme into the first CNN block, and the first CNN block performs convolutional processing on the target concatenated feature of each phoneme to output the first feature map of each phoneme, and inputs the first feature map of each phoneme into the second CNN block, and so on, until the last CNN block finishes processing, and obtains the last feature map of each phoneme output by the last CNN block, which is the final feature map of each phoneme, and inputs the feature map of each phoneme into the pooling layer.
[0121] Optionally, each CNN block includes a two-layer structure: a one-dimensional convolutional layer and a normalization layer. Here, the one-dimensional convolutional layer refers to a convolutional layer with a kernel size of 1×1, which can be used to deepen the depth of the CNN without increasing the receptive field to introduce more non-linear factors. Moreover, by flexibly controlling the number of kernels in the one-dimensional convolutional layer, it is possible to increase or decrease the dimension of the input feature map (for the first CNN block, it is the target stitching feature; for the i-th CNN block, it is the output of the (i - 1)-th CNN block, i ≥ 2). In other words, the number of kernels in the one-dimensional convolutional layer of each CNN block is equal to the dimension of the feature map output by the current CNN block. The normalization layer is used to pull the feature map output by the one-dimensional convolutional layer from a non-standard normal distribution back to a standard normal distribution with a mean of 0 and a variance of 1, thereby avoiding the gradient saturation of the CNN and accelerating the convergence of the CNN. For example, the (i - 1)-th standard feature map of each phoneme is input into the i-th CNN block. Through the one-dimensional convolutional layer of the i-th CNN block, using a 1×1 convolutional kernel, the (i - 1)-th standard feature map of each phoneme is convolved to obtain the i-th feature map of each phoneme, and the i-th feature map of each phoneme is input into the normalization layer for normalization processing to obtain the i-th standard feature map of each phoneme, and the i-th standard feature map of each phoneme is input into the (i + 1)-th CNN block, where the standard feature map refers to the feature map obtained after being normalized by the normalization layer.
[0122] Optionally, when performing pooling on the feature map of each phoneme (i.e., the standard feature map output by the last CNN block) through this pooling layer, either mean pooling or max pooling can be performed to obtain the pooled feature map of each phoneme. The present application embodiment does not specifically limit the pooling method within the pooling layer. For example, the mean pooling method is adopted in this pooling layer.
[0123] In some embodiments, the server performs embedding encoding on each pitch in the pitch sequence to obtain the embedding feature of each pitch, and determines the embedding feature of each pitch as the pitch feature of each pitch to enhance the expression ability of the pitch feature. Optionally, other encoding methods such as one-hot encoding can also be used for each pitch. The present application embodiment does not specifically limit the encoding method of the pitch feature.
[0124] In some embodiments, the server performs embedding encoding on the singer identifier to obtain the embedding features of the singer, and determines the embedding features of the singer as the singer features, so as to enhance the expression ability of the singer features. Optionally, other encoding methods such as one-hot encoding can also be used for the singer identifier, and the encoding method of the singer features in the embodiments of the present application is not specifically limited.
[0125] In some embodiments, after the server obtains the pooled feature maps of each phoneme, for each syllable, since it is known which one or more phonemes are included in each syllable, that is, the correspondence between the syllable and the phoneme is known, the pooled feature maps of the phonemes belonging to the syllable, the syllable features of the syllable, the pitch features of the pitch corresponding to the syllable, and the singer features can be concatenated to obtain the initial semantic features of the syllable, which can save the computational amount of the fusion process. Optionally, in addition to the concatenation method, other methods such as element-wise addition, element-wise multiplication, and bilinear pooling can also be used for fusion to improve the expression ability of the initial semantic features. The embodiments of the present application do not specifically limit the fusion method.
[0126] In some embodiments, after obtaining the initial semantic features of each syllable, the initial semantic features of each syllable can be input into a second feature extraction model, and the second feature extraction model performs convolutional processing on the initial semantic features of each syllable to obtain the semantic feature maps of each syllable, where the second feature extraction model is used to extract the semantic feature maps of the input syllables. Schematically, the second feature extraction model can also be a CNN. For example, it has a similar structure to the first feature extraction model, that is, it also includes one or more CNN blocks, and the CNN block also includes two layers: a 1D convolutional layer and a normalization layer, which will not be elaborated here. It should be noted that even if both the first feature extraction model and the second feature extraction model are CNN models including multiple CNN blocks, the two can have different numbers of CNN blocks. For example, the first feature extraction model includes 6 CNN blocks, and the second feature extraction model includes 4 CNN blocks, etc. The embodiments of the present application do not specifically limit this.
[0127] In some embodiments, after obtaining the semantic feature maps of each syllable, the semantic feature maps of each syllable can be input into the first encoder, and the first encoder encodes the semantic feature maps of each syllable to output the first semantic feature of each syllable. Schematically, the first encoder can be a Long Short-Term Memory (LSTM) artificial neural network. One or more hidden layers can be included in the LSTM, and multiple memory units are included in each hidden layer. Each memory unit corresponds to a syllable. When the number of hidden layers in the LSTM is multiple, the multiple hidden layers can be connected in series, that is, the output of the previous hidden layer serves as the input of the current hidden layer. Similarly, the output of the current hidden layer serves as the input of the next hidden layer.
[0128] Taking the j-th (j≥2) memory unit in the i-th (i≥2) hidden layer of the LSTM as an example, the inputs of this memory unit include: the output feature of the (j - 1)-th memory unit in the i-th hidden layer and the output feature of the j-th memory unit in the (i - 1)-th hidden layer. This memory unit performs a weighted transformation on the above two output features to obtain the output feature of this memory unit. Then, the output feature of the memory unit is respectively input into the (j + 1)-th memory unit in the i-th hidden layer and the j-th memory unit in the (i + 1)-th hidden layer. Among them, the inputs of each memory unit in the first hidden layer include the output feature of the previous memory unit in the first hidden layer and the semantic feature map of the corresponding syllable. The output feature of each memory unit in the last hidden layer is the first semantic feature of each syllable. Performing the above operations on each memory unit within each hidden layer is equivalent to performing recursive weighted processing in the entire LSTM, that is, forward encoding of the semantic feature map of each syllable is performed.
[0129] In some embodiments, in addition to the LSTM, the first encoder can also be a Bidirectional Long Short-Term Memory (BLSTM) network, or a sequence-to-sequence model such as a Recurrent Neural Network (RNN). The embodiments of the present application do not specifically limit the structure of the first encoder.
[0130] In some embodiments, after obtaining the first semantic feature of each syllable, in combination with the syllable duration information obtained in step 401 above, the phoneme duration information can be determined through the following operations: Based on the syllable duration information and the first semantic features of the multiple syllables, obtain the initial features of the multiple audio frames, where the initial features are the first semantic features of the syllables to which the corresponding audio frames belong; Based on the initial features of the multiple audio frames and the position features of the multiple audio frames, determine the multiple phonemes corresponding to each of the multiple audio frames; Based on the correspondence between the multiple audio frames and the multiple phonemes, determine the phoneme duration information.
[0131] Optionally, for any one of the multiple syllables, when obtaining the initial features of each audio frame, the server can copy the first semantic feature of the syllable the first target number of times to obtain the initial features of the multiple audio frames included in the syllable, where the first target number is the value obtained by subtracting one from the number of audio frames occupied by the syllable.
[0132] In some embodiments, the server determines, from the syllable duration information, the number of audio frames occupied by the current syllable, subtracts one from the number of audio frames occupied by the current syllable to obtain the first target number, and then copies the first semantic feature of the syllable the first target number of times. The first target number of copied features and the original first semantic feature of the current syllable are together determined as the initial features of the multiple audio frames. For example, if the first syllable occupies 10 frames, then the first semantic feature of the first syllable is copied 9 times to obtain 9 features, and the 9 copied features and the original first semantic feature of the first syllable are determined as the initial features of a total of 10 audio frames.
[0133] In some embodiments, when determining the phoneme corresponding to each audio frame, the server can splice the initial features of the multiple audio frames and the position features of the multiple audio frames respectively to obtain multiple first spliced features; perform convolutional processing on the multiple first spliced features to obtain multiple first target features; perform weighted processing on the multiple first target features to obtain multiple second target features; and predict the multiple phonemes corresponding to each of the multiple audio frames based on the multiple second target features.
[0134] Optionally, for each audio frame, the server obtains the position feature (e.g., position encoding) of the audio frame based on the position information of the audio frame in all audio frames, and splices the initial feature of the audio frame and the position feature of the audio frame to obtain a first spliced feature of the audio frame.
[0135] Optionally, when the server obtains the first target feature, it may input the multiple first concatenated features into a third feature extraction model, and perform convolution processing on each first concatenated feature through the third feature extraction model to obtain the corresponding first target feature, where the third feature extraction model is used to extract the first target feature of the input audio frame. Schematically, the third feature extraction model may also be a CNN. For example, it has a similar structure to the first feature extraction model, that is, it also includes one or more CNN blocks, and the CNN block also includes two layers: a one-dimensional convolutional layer and a normalization layer, which will not be elaborated here. It should be noted that even if both the first feature extraction model and the third feature extraction model are CNN models including multiple CNN blocks, they may have different numbers of CNN blocks. For example, the first feature extraction model includes 6 CNN blocks, and the third feature extraction model includes 4 CNN blocks, etc. The embodiments of the present application do not specifically limit this.
[0136] Optionally, when the server obtains the second target feature, it may input the multiple first target features into a second encoder, and perform weighting processing on each first target feature through the second encoder to obtain the corresponding second target feature. The structure of the second encoder includes but is not limited to: LSTM, BLSTM, RNN, etc. For example, the second encoder is also an LSTM network, but it may have a different number of hidden layers or different parameters from the first encoder, which will not be elaborated here.
[0137] Optionally, when the server predicts the phoneme corresponding to each audio frame, it may perform a fully connected process on any one of the multiple second target features to obtain multiple prediction probabilities of the audio frame corresponding to the second target feature, and the prediction probability is used to represent the possibility that the audio frame corresponds to a phoneme; determine the phoneme with the highest prediction probability as the phoneme corresponding to the audio frame.
[0138] In some embodiments, the server inputs the second target feature of each audio frame output by the second encoder into a fully connected layer, performs a fully connected process on each second target feature through the fully connected layer, and outputs multiple prediction probabilities for each audio frame corresponding to each second target feature. Each prediction probability is used to represent the possibility that the audio frame corresponds to a phoneme, obtains the maximum prediction probability among the multiple prediction probabilities, and determines the phoneme corresponding to the maximum prediction probability as the phoneme corresponding to the audio frame. In other words, the above prediction process can be regarded as a classification process, and the phoneme ID of each phoneme represents a classification label.
[0139] In some embodiments, precisely because the server determines the phonemes corresponding to each audio frame, it can, conversely, determine the phoneme duration information based on the correspondence between the multiple audio frames and the multiple phonemes. For example, among the 10 audio frames of the first syllable, there are 3 phonemes. The 1st - 3rd audio frames all correspond to phoneme A, the 4th audio frame corresponds to phoneme B, and the 5th - 10th audio frames correspond to phoneme C. Thus, it can be determined that in the first syllable, the number of audio frames occupied by phoneme A is 3, the number of audio frames occupied by phoneme B is 1, and the number of audio frames occupied by phoneme C is 6.
[0140] 403. The server obtains the acoustic feature parameters of multiple audio frames in the to - be - generated singing voice based on the musical score, the lyrics, and the phoneme duration information, and the acoustic feature parameters are used to characterize the acoustic features of the corresponding audio frames.
[0141] In some embodiments, when obtaining the acoustic feature parameters, the server can obtain the second semantic features of the multiple audio frames based on the lyrics, the musical score, and the phoneme duration information. The second semantic features are used to characterize the semantics when the singer sings the corresponding phoneme at the corresponding pitch; encode the second semantic features of the multiple audio frames to obtain the intermediate features of the multiple audio frames; decode the intermediate features of the multiple audio frames to obtain the third semantic features of the multiple audio frames; and process the third semantic features of the multiple audio frames to obtain the acoustic feature parameters of the multiple audio frames.
[0142] In some embodiments, when obtaining the second semantic features, the server can obtain the phoneme features of the multiple phonemes in the lyrics and the pitch features of the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables; obtain the frame - level phoneme features of the multiple audio frames based on the phoneme features of the multiple phonemes and the phoneme duration information, and the frame - level phoneme features are the phoneme features of the phoneme to which the corresponding audio frame belongs; obtain the frame - level pitch features of the multiple audio frames based on the pitch features of the multiple pitches and the phoneme duration information, and the frame - level pitch features are the pitch features of the pitch corresponding to the phoneme to which the corresponding audio frame belongs; splice the frame - level phoneme features of the multiple audio frames with the frame - level pitch features of the multiple audio frames and the singer features of the singer identifier respectively to obtain the second semantic features of the multiple audio frames.
[0143] Optionally, when the server obtains the phoneme features, it can perform embedding encoding on each phoneme in the phoneme sequence to obtain the embedding features of each phoneme. Then, based on the sorting information of each phoneme in the phoneme sequence, the position features of each phoneme are obtained. The embedding features and position features of each phoneme are concatenated to obtain the target concatenated features of each phoneme. The target concatenated features of each phoneme are input into the fourth feature extraction model, and the fourth feature extraction model encodes and decodes the target concatenated features of each phoneme to obtain the phoneme features of each phoneme, where the fourth feature extraction model is used to extract the phoneme features of the input phonemes.
[0144] Schematically, the fourth feature extraction model can be an FFT (Feed-Forward Transformer). The FFT may include an encoding module and a decoding module. The encoding module is used to encode the features of the phonemes, and the decoding module is used to decode the features of the phonemes. The encoding module includes one or more feed-forward transformation blocks (FFT blocks), and the decoding module also includes one or more FFT blocks. Usually, the number of FFT blocks in the encoding module and the decoding module is the same. For the encoding module or the decoding module, when the number of FFT blocks is multiple, the multiple FFT blocks are connected in series (i.e., cascaded), that is, the output of the previous FFT block is used as the input of the current FFT block. Similarly, the output of the current FFT block is used as the input of the next FFT block. Among them, the input of the first FFT block in the encoding module is the target concatenated feature of each phoneme.
[0145] For example, the server inputs the target splicing features of each phoneme into the first FFT block of the encoding module, weights the target splicing features of each phoneme through the first FFT block, outputs the first hidden vector of each phoneme, and inputs the first hidden vector of each phoneme into the second FFT block, and so on, until the last FFT block is processed, obtaining the last hidden vector of each phoneme output by the last FFT block, that is, the final hidden state features of each phoneme, and based on adjusting the length of the hidden state features of each phoneme through a length regulator. Then, the adjusted hidden state features of each phoneme are input into the first FFT block of the decoding module, the hidden state features of each phoneme are weighted through the first FFT block, the first hidden vector of each phoneme is output, and the first hidden vector of each phoneme is input into the second FFT block, and so on, until the last FFT block is processed, obtaining the last hidden vector of each phoneme output by the last FFT block, which is the final phoneme feature of each phoneme.
[0146] Optionally, each FFT block includes a four-layer structure: a multi-head attention layer, a first residual normalization layer, a one-dimensional convolutional layer, and a second residual normalization layer. Among them, multiple attention layers are used to comprehensively extract the correlation relationships between phonemes in the pinyin sequence from multiple expression subspaces. The one-dimensional convolutional layer is similar to the above-mentioned CNN block and will not be elaborated here. The input of the first residual normalization layer is the first residual feature obtained by concatenating the input of the multi-head attention layer and the output of the multi-head attention layer. The input of the second residual normalization layer is the second residual feature obtained by concatenating the input of the one-dimensional convolutional layer and the output of the one-dimensional convolutional layer. The above two residual normalization layers actually still belong to the normalization layer, but there is a residual connection between the input and output of the normalization layer and the previous layer, so they are called residual normalization layers. For example, the target concatenated feature of each phoneme is input into the multi-head attention layer of the first FFT block of the FFT encoding module. The multi-head attention layer weights the target concatenated feature of each phoneme to obtain the attention feature of each phoneme. The target concatenated feature and the attention feature of each phoneme are concatenated to obtain the first residual feature of each phoneme. The first residual feature of each phoneme is input into the first residual normalization layer for normalization processing to obtain the first standard feature of each phoneme. Then, the first standard feature of each phoneme is input into the one-dimensional convolutional layer, and a 1×1 convolutional kernel is used to perform convolutional processing on the first standard feature of each phoneme to obtain the convolutional feature map of each phoneme. The first standard feature and the convolutional feature map of each phoneme are concatenated to obtain the second residual feature of each phoneme. The second residual feature of each phoneme is input into the second residual normalization layer for normalization processing to obtain the second standard feature of each phoneme. The second standard feature of each phoneme above will be used as the input of the second FFT block in the encoding module, and so on. The processing process of the FFT block in the decoding module is the same and will not be elaborated here. Finally, the second standard feature of each phoneme output by the last FFT block in the decoding module is obtained as the phoneme feature of each phoneme.
[0147] In some embodiments, the server performs embedding encoding on each pitch in the pitch sequence to obtain the embedding feature of each pitch, and determines the embedding feature of each pitch as the pitch feature of each pitch to enhance the expression ability of the pitch feature. Optionally, other encoding methods such as one-hot encoding can also be used for each pitch. The embodiments of the present application do not specifically limit the encoding method of the pitch feature.
[0148] In some embodiments, the server performs embedding encoding on the singer identifier to obtain the embedding feature of the singer, and determines the embedding feature of the singer as the singer feature, so as to improve the expression ability of the singer feature. Optionally, other encoding methods such as one-hot encoding can also be used for the singer identifier, and the embodiments of the present application do not specifically limit the encoding method of the singer feature.
[0149] In some embodiments, for any one of the multiple phonemes, when the server obtains the corresponding frame-level phoneme feature, it can copy the phoneme feature of the phoneme for a second target number of times to obtain the frame-level phoneme features of multiple audio frames included in the phoneme, where the second target number of times is a value obtained by subtracting one from the number of audio frames occupied by the phoneme.
[0150] In some embodiments, the server determines the number of audio frames occupied by the current phoneme from the phoneme duration information, subtracts one from the number of audio frames occupied by the current phoneme to obtain the second target number of times, and then copies the phoneme feature of the phoneme for the second target number of times. The second target number of times of copied features and the original phoneme feature of the current phoneme are together determined as the frame-level phoneme features of the multiple audio frames. For example, if the first phoneme occupies 3 frames, the phoneme feature of the first phoneme is copied 2 times to obtain 2 features, and the 2 copied features and the original phoneme feature of the first phoneme are determined as the frame-level phoneme features of 3 audio frames in total.
[0151] In some embodiments, for any one of the multiple pitches, when the server obtains the corresponding frame-level pitch feature, it can copy the pitch feature of the pitch for a third target number of times to obtain the frame-level pitch features of multiple audio frames included in the phoneme corresponding to the pitch, where the third target number of times is a value obtained by subtracting one from the number of audio frames occupied by the phoneme corresponding to the pitch.
[0152] In some embodiments, the server first determines the phoneme corresponding to the current pitch, then determines the number of audio frames occupied by the phoneme from the phoneme duration information, subtracts one from the number of audio frames occupied by the phoneme to obtain the third target number of times, and then copies the pitch feature of the current pitch for the third target number of times. The third target number of times of copied features and the original pitch feature of the current pitch are together determined as the frame-level pitch features of the multiple audio frames. For example, the m-th (m≥1) pitch corresponds to the m-th phoneme, and the m-th phoneme occupies 5 frames, then the pitch feature of the m-th pitch is copied 4 times to obtain 4 features, and the 4 copied features and the original pitch feature of the m-th pitch are determined as the frame-level pitch features of 5 audio frames in total.
[0153] In some embodiments, for each audio frame, when obtaining the corresponding second semantic feature, the server may splice the frame-level phoneme feature of the audio frame, the frame-level pitch feature of the audio frame, and the singer feature to obtain the second semantic feature of the audio frame.
[0154] In some embodiments, after the server obtains the second semantic feature of each audio frame, it may input the second semantic feature of each audio frame into an acoustic parameter prediction model. Optionally, the acoustic parameter prediction model includes a first linear layer, an encoding module, a decoding module, a second linear layer, and a post-processing layer. For example, input the second semantic feature of each audio frame into the first linear layer, and perform a linear transformation on the second semantic feature of each audio frame through the first linear layer to obtain the second semantic feature of each audio frame after the linear transformation. Then, input the second semantic feature of each audio frame after the linear transformation into the encoding module, and encode the second semantic feature of each audio frame after the linear transformation through the encoding module to obtain the intermediate feature of each audio frame. Input the intermediate feature of each audio frame into the decoding module, and decode the intermediate feature of each audio frame through the decoding module to obtain the third semantic feature of each audio frame. Then, input the third semantic feature of each audio frame into the second linear layer, and perform a linear transformation on the third semantic feature of each audio frame through the second linear layer to obtain the third semantic feature of each audio frame after the linear transformation. Input the third semantic feature of each audio frame after the linear transformation into the post-processing layer, and process the third semantic feature of each audio frame after the linear transformation through the post-processing layer to obtain the standard semantic feature of each audio frame. Splice the standard semantic feature of each audio frame and the third semantic feature after the linear transformation to obtain the acoustic feature parameters of each audio frame.
[0155] Schematically, the acoustic parameter prediction model may have a structure similar to that of the FFT. That is, the encoding module of the acoustic parameter prediction model includes one or more FFT blocks, and the decoding module also includes one or more FFT blocks. Usually, the number of FFT blocks in the encoding module and the decoding module is the same. The number of FFT blocks in the acoustic parameter prediction model may be the same as or different from the number of FFT blocks in the fourth feature extraction model. This application embodiment does not specifically limit this. The connection method, internal structure, and processing process of the FFT blocks in the acoustic parameter prediction model are similar to those of the FFT blocks in the fourth feature extraction model, and will not be elaborated here.
[0156] Schematically, in the post-processing layer of the acoustic parameter prediction model, one or more 1D convolutional layers and a normalization layer may be included. For example, the post-processing layer includes 5 1D convolutional layers and 1 normalization layer. The 1D convolutional layer and the normalization layer have both been introduced in step 402 above and will not be elaborated here.
[0157] 404. The server generates the to-be-generated singing voice based on the acoustic feature parameters of the multiple audio frames.
[0158] Optionally, the acoustic feature parameter may be at least one of the MGC parameter or the BAP parameter of each audio frame. Of course, the acoustic feature parameter may also be the MFCC parameter of each audio frame, etc. The embodiments of the present application do not specifically limit the type of the acoustic feature parameter.
[0159] In some embodiments, the server may directly input the acoustic feature parameters of the multiple audio frames into a vocoder, and synthesize a singing voice matching the acoustic feature parameters through the vocoder, which can simplify the singing voice generation process.
[0160] Optionally, the vocoder may be a WORLD vocoder or a vocoder with other structures. The embodiments of the present application do not specifically limit the structure of the vocoder.
[0161] In some embodiments, the server may also use the intermediate output of the acoustic parameter prediction model, that is, the third semantic features of the multiple audio frames output by the decoding module, to perform precise pitch control modeling on each audio frame. Optionally, the server performs linear processing on the third semantic features of the multiple audio frames to obtain multiple third target features; performs linear processing on the frame-level pitch features of the multiple audio frames to obtain multiple fourth target features; respectively splices the multiple third target features and the multiple fourth target features to obtain multiple second splicing features; performs convolutional processing on the multiple second splicing features to obtain at least one of the silence parameters or fundamental frequency parameters of each of the multiple audio frames. The silence parameter is used to represent whether the corresponding audio frame is a silent segment, and the fundamental frequency parameter is used to represent the logarithm of the fundamental frequency of the corresponding audio frame.
[0162] In some embodiments, the server may input the third semantic features of the plurality of audio frames and the frame-level pitch features of the plurality of audio frames into a pitch prediction model, which includes a third linear layer, a fourth linear layer, and a one-dimensional convolutional layer. Optionally, the third semantic features of the plurality of audio frames are input into the third linear layer, and the third linear layer performs linear processing on the third semantic features of the plurality of audio frames to obtain a plurality of third target features; the frame-level pitch features of the plurality of audio frames are input into the fourth linear layer, and the fourth linear layer performs linear processing on the frame-level pitch features of the plurality of audio frames to obtain a plurality of fourth target features. Then, the plurality of third target features and the plurality of fourth target features are respectively concatenated to obtain a plurality of second concatenated features, and the plurality of second concatenated features are input into the one-dimensional convolutional layer, and the one-dimensional convolutional layer performs convolutional processing on the plurality of second concatenated features to obtain at least one of the silence parameters or fundamental frequency parameters of each of the plurality of audio frames.
[0163] In the above process, through the pitch prediction model, at least one of the silence parameters or fundamental frequency parameters is additionally obtained for each audio frame, so that at least one of the silence parameters or fundamental frequency parameters can be input into the vocoder together with the acoustic feature parameters after concatenation to perform precise pitch control on the generated singing voice.
[0164] In some embodiments, after obtaining at least one of the silence parameters or fundamental frequency parameters for each audio frame, the server may also concatenate the acoustic feature parameters of the plurality of audio frames and at least one of the silence parameters or fundamental frequency parameters of each of the plurality of audio frames to obtain the target acoustic features of the plurality of audio frames; the target acoustic features of the plurality of audio frames are input into the vocoder, and the vocoder synthesizes the audio signal of the singing voice to be generated.
[0165] In the above process, the vocoder not only passes the acoustic feature parameters of each audio frame, but also provides at least one of the silence parameters or fundamental frequency parameters of each audio frame, so that not only can the number of audio frames that each phoneme lasts be precisely controlled by introducing phoneme duration information, but also the pitch of each phoneme can be precisely controlled by introducing at least one of the silence parameters or fundamental frequency parameters, greatly improving the naturalness and expressiveness of the output singing voice.
[0166] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated here one by one.
[0167] The method provided by the embodiments of the present application further predicts phoneme duration information on the basis of the syllable duration information given in the music score. Since the phoneme duration information can represent the number of audio frames occupied by each phoneme, the accuracy of the acoustic model is no longer limited to the rough syllable level, but can reach the precise control at the phoneme level, greatly improving the naturalness of singing synthesis.
[0168] In the above embodiment, it shows how to generate a piece of singing audio similar to human singing by a computer. In the embodiments of the present application, the process of singing generation will be introduced in detail in combination with a duration model, a pitch prediction model, an acoustic model, and a vocoder, and the structural descriptions of the respective models will be given schematically. The following will be described in detail.
[0169] Figure 5 It is a flowchart of a singing generation method provided by the embodiments of the present application. As Figure 5 shown, this embodiment is applied to a computer device. Taking the computer device as a server as an example for illustration, this embodiment includes the following steps:
[0170] 501. The server obtains the digitized music score and the lyrics corresponding to the music score.
[0171] In some embodiments, the server can parse the digitized music score information to obtain the music score and the lyrics corresponding to the music score. Optionally, the music score information can be a data file in the musicxml format; the music score can be a staff notation or a numbered musical notation; the lyrics can be Chinese lyrics, English lyrics, or lyrics in other languages. The embodiments of the present application do not make specific limitations in this regard.
[0172] Schematically, the user directly inputs the digitized music score information to the server, or the user uploads the digitized music score information to the server through an application on the terminal. Alternatively, the server loads the digitized music score information from another cloud database. The embodiments of the present application do not make specific limitations on the acquisition method of the music score information. For example, the user inputs the digitized music score information to an audio application on the terminal, clicks the singing generation option, triggers the terminal to send a singing generation request carrying the digitized music score information to the server, and the server receives the singing generation request and parses the singing generation request to obtain the digitized music score information.
[0173] 502. The server obtains a pitch sequence based on the music score, and the pitch sequence includes multiple pitches corresponding to multiple notes in the music score.
[0174] In some embodiments, after the server obtains the music score, it can parse the music score to obtain multiple pitches corresponding to multiple notes in the music score. The multiple pitches can be regarded as constituting a pitch sequence, and the pitch sequence is also called a tone sequence.
[0175] 503. The server obtains a syllable sequence, a phoneme sequence, and syllable duration information based on the lyrics. The syllable sequence includes multiple syllables in the lyrics. The phoneme sequence includes multiple phonemes included in the multiple syllables in the lyrics. The syllable duration information is used to represent the number of audio frames occupied by each of the multiple syllables in the lyrics.
[0176] Among them, the multiple syllables correspond to the multiple notes.
[0177] In some embodiments, after the server obtains the lyrics, it can parse the lyrics to obtain multiple characters in the lyrics. The multiple characters can be regarded as constituting a character sequence. For example, for Chinese lyrics, the character sequence is also a Chinese character sequence. For English lyrics, the character sequence is also a word sequence. Schematically, since each Chinese character has its own pinyin, usually, the pinyin of a Chinese character corresponds to a syllable uttered during singing. Therefore, for Chinese lyrics, in addition to obtaining the Chinese character sequence, the Chinese character sequence can also be converted into a corresponding pinyin sequence, and the pinyin sequence represents the syllable sequence of the singer during singing.
[0178] In some embodiments, for each syllable in the syllable sequence, dividing it according to the smallest pronunciation unit that can distinguish meanings can obtain one or more phonemes included in each syllable. Thus, by combining all the phonemes included in all the syllables, the multiple phonemes in the lyrics, that is, the phoneme sequence, can be obtained.
[0179] In some embodiments, since in addition to recording notes, the music score also implies the duration of each note in the music score, and when the sampling rate is fixed, the duration occupied by each audio frame in the audio signal is usually also fixed (this duration can be called the frame length of the audio frame). Therefore, the server can determine the frame length of a single audio frame in the audio signal based on the sampling rate of the audio signal. Further, for each note in the music score, dividing the duration of the note in the music score by the frame length can obtain the number of audio frames occupied by the syllable corresponding to the note. Repeating the above steps can obtain a series of audio frames, and this series of audio frames can constitute a duration sequence. The duration sequence is also the syllable duration information, representing the number of audio frames occupied by each of the syllables corresponding to the respective notes in the music score in the lyrics.
[0180] In the above steps 501-503, the server determines the syllable duration information based on the music score and the lyrics corresponding to the music score. Based on this syllable duration information, the phoneme duration information for the singer can be determined by combining the phoneme sequence, the syllable sequence, the pitch sequence, and the specified singer identifier, so as to achieve precise control of the number of audio frames occupied by each phoneme.
[0181] 504. The server obtains a first semantic feature of the multiple syllables based on the syllable sequence, the phoneme sequence, the pitch sequence, and the singer identifier, where the first semantic feature is used to characterize the semantics when the singer sings the corresponding syllable at the corresponding pitch.
[0182] Optionally, the singer identifier can be used to uniquely identify the singer (i.e., the speaker). The singer identifier can be specified by the user. For example, the user selects the singer to which the song to be synthesized this time belongs, and the server determines the singer identifier of the singer selected by the user. Or, the user can also not specify the singer, and the server randomly selects a singer identifier from the singer database, or the server obtains the singer identifier with the highest selection frequency, etc. The embodiments of the present application do not specifically limit the source of the singer identifier.
[0183] Figure 6 It is a flowchart of obtaining a first semantic feature provided by an embodiment of the present application. As Figure 6 shown, when the server obtains the first semantic feature of each syllable, it can execute the following steps 5041-5047, which are described in detail below:
[0184] 5041. The server obtains the syllable feature of each syllable based on the syllable sequence.
[0185] In some embodiments, the server performs embedding encoding on each syllable in the syllable sequence to obtain the embedding feature of each syllable, and determines the embedding feature of each syllable as the syllable feature of each syllable to enhance the expression ability of the syllable feature.
[0186] Optionally, other encoding methods such as one-hot encoding can also be used for each syllable. The embodiments of the present application do not specifically limit the encoding method of the syllable feature.
[0187] 5042. The server obtains the pooling feature map of each phoneme based on the phoneme sequence.
[0188] In some embodiments, the server performs embedding encoding on each phoneme in the phoneme sequence to obtain the embedding feature of each phoneme. Then, based on the sorting information of each phoneme in the phoneme sequence, the position feature of each phoneme is obtained. The embedding feature and the position feature of each phoneme are concatenated to obtain the target concatenated feature of each phoneme. The target concatenated feature of each phoneme is input into the first feature extraction model, and the first feature extraction model performs convolution processing on the target concatenated feature of each phoneme to obtain the feature map of each phoneme. The feature map of each phoneme is input into a pooling layer, and the pooling layer performs pooling processing on the feature map of each phoneme to obtain the pooled feature map of each phoneme, where the first feature extraction model is used to extract the feature map of the input phoneme.
[0189] In the above process, by further using the first feature extraction model to extract the feature maps of each phoneme from the embedding features of each phoneme and performing pooling on the feature maps of each phoneme to obtain the pooled feature maps of each phoneme, it can make the pooled feature maps of each phoneme contain more deep and hidden feature information, and improve the expression ability of the features compared with directly using the embedding features of each phoneme.
[0190] Schematically, the first feature extraction model can be a CNN. The CNN includes one or more CNN blocks. When the number of CNN blocks is multiple, the multiple CNN blocks are connected in series, that is, the output of the previous CNN block is used as the input of the current CNN block. Similarly, the output of the current CNN block is used as the input of the next CNN block, where the input of the first CNN block is the target concatenated feature of each phoneme. At this time, the server inputs the target concatenated feature of each phoneme into the first CNN block, and the first CNN block performs convolution processing on the target concatenated feature of each phoneme to output the first feature map of each phoneme, and inputs the first feature map of each phoneme into the second CNN block, and so on, until the last CNN block finishes processing, and obtains the last feature map of each phoneme output by the last CNN block, which is the final feature map of each phoneme, and inputs the feature map of each phoneme into the pooling layer.
[0191] Optionally, each CNN block includes two-layer structure: a one-dimensional convolutional layer and a normalization layer. Herein, the one-dimensional convolutional layer refers to a convolutional layer with a convolution kernel size of 1×1, which can be used to deepen the depth of the CNN without increasing the receptive field to introduce more non-linear factors. Moreover, by flexibly controlling the number of convolution kernels in the one-dimensional convolutional layer, it is possible to increase or decrease the dimension of the input feature map (for the first CNN block, it is the target stitching feature; for the i-th CNN block, it is the output of the (i - 1)-th CNN block, i≥2). In other words, the number of convolution kernels in the one-dimensional convolutional layer of each CNN block is equal to the dimension of the feature map output by the current CNN block. The normalization layer is used to pull the feature map output by the one-dimensional convolutional layer from a non-standard normal distribution back to a standard normal distribution with a mean of 0 and a variance of 1, so as to avoid gradient saturation of the CNN and accelerate the convergence of the CNN. For example, the (i - 1)-th standard feature map of each phoneme is input into the i-th CNN block. Through the one-dimensional convolutional layer of the i-th CNN block, using a 1×1 convolution kernel, the (i - 1)-th standard feature map of each phoneme is convolved to obtain the i-th feature map of each phoneme, and the i-th feature map of each phoneme is input into the normalization layer for normalization processing to obtain the i-th standard feature map of each phoneme, and the i-th standard feature map of each phoneme is input into the (i + 1)-th CNN block, where the standard feature map refers to the feature map obtained after being normalized by the normalization layer.
[0192] Optionally, when performing pooling processing on the feature map of each phoneme (i.e., the standard feature map output by the last CNN block) through this pooling layer, average pooling or max pooling can be performed to obtain the pooled feature map of each phoneme. The embodiments of the present application do not specifically limit the pooling method in the pooling layer. For example, average pooling is used in this pooling layer.
[0193] In the above process, by using CNN as the first feature extraction model, since a one-dimensional convolutional layer is adopted in each CNN block, it is possible to deepen the depth of the CNN without increasing the receptive field to introduce more non-linear factors, and it is also possible to flexibly increase or decrease the dimension of the feature maps of each phoneme output, which is convenient for extracting deeper feature information. Of course, in addition to CNN, a first feature extraction model with a structure such as DNN (Deep Neural Network) or VGG (Visual Geometry Group) can also be used. The embodiments of the present application do not specifically limit the structure of the first feature extraction model.
[0194] 5043. The server obtains the pitch feature of each note based on the pitch sequence.
[0195] In some embodiments, the server performs embedding encoding on each pitch in the pitch sequence to obtain the embedding feature of each pitch, and determines the embedding feature of each pitch as the pitch feature of each pitch to enhance the expression ability of the pitch feature.
[0196] Optionally, other encoding methods such as one-hot encoding can also be used for each pitch. The embodiments of the present application do not specifically limit the encoding method of the pitch feature.
[0197] 5044. Based on the singer identifier, the server obtains the singer feature.
[0198] In some embodiments, the server performs embedding encoding on the singer identifier to obtain the embedding feature of the singer, and determines the embedding feature of the singer as the singer feature to enhance the expression ability of the singer feature.
[0199] Optionally, other encoding methods such as one-hot encoding can also be used for the singer identifier. The embodiments of the present application do not specifically limit the encoding method of the singer feature.
[0200] 5045. For each syllable in the syllable sequence, the server fuses the pooled feature maps of the phonemes belonging to the syllable, the syllable feature of the syllable, the pitch feature of the pitch corresponding to the syllable, and the singer feature to obtain the initial semantic feature of each syllable.
[0201] In some embodiments, after the server obtains the pooled feature map of each phoneme, for each syllable, since the one or more phonemes included in each syllable are known, that is, the correspondence between the syllable and the phoneme is known, the pooled feature maps of the phonemes belonging to the present syllable, the syllable feature of the present syllable, the pitch feature of the pitch corresponding to the present syllable, and the singer feature can be concatenated to obtain the initial semantic feature of the present syllable, which can save the computational amount in the fusion process.
[0202] In some embodiments, for each syllable, the pooled feature maps of the phonemes belonging to the present syllable, the syllable feature of the present syllable, and the pitch feature of the pitch corresponding to the present syllable can be concatenated to obtain the initial semantic feature of the present syllable, and after concatenating the initial semantic features of all syllables, they are concatenated with the singer feature as a whole, which can reduce the computational complexity in the fusion process.
[0203] Optionally, in addition to the splicing method, fusion can also be performed by element-wise addition, element-wise multiplication, bilinear pooling, etc. to improve the expression ability of the initial semantic features. The embodiments of the present application do not specifically limit the fusion method.
[0204] 5046. The server performs convolutional processing on the initial semantic features of each syllable to obtain a semantic feature map for each syllable.
[0205] In some embodiments, after obtaining the initial semantic features of each syllable, the initial semantic features of each syllable can be input into a second feature extraction model, and the second feature extraction model performs convolutional processing on the initial semantic features of each syllable to obtain a semantic feature map for each syllable, where the second feature extraction model is used to extract the semantic feature map of the input syllable. Schematically, the second feature extraction model can also be a CNN. For example, it has a similar structure to the first feature extraction model, that is, it also includes one or more CNN blocks, and the CNN block also includes two layers: a 1D convolutional layer and a normalization layer, which will not be elaborated here. It should be noted that even if both the first feature extraction model and the second feature extraction model are CNN models including multiple CNN blocks, they can have different numbers of CNN blocks. For example, the first feature extraction model includes 6 CNN blocks, and the second feature extraction model includes 4 CNN blocks, etc. The embodiments of the present application do not specifically limit this.
[0206] In the above process, by using a CNN as the second feature extraction model, since a 1D convolutional layer is used in the CNN block, the depth of the CNN can be deepened without increasing the receptive field to introduce more non-linear factors, and the semantic feature maps of each output syllable can be flexibly upsampled or downsampled to facilitate the extraction of deeper feature information. Of course, in addition to the CNN, a second feature extraction model with a structure such as DNN or VGG can also be used. The embodiments of the present application do not specifically limit the structure of the second feature extraction model.
[0207] 5047. The server encodes the semantic feature map of each syllable to obtain the first semantic feature of each syllable.
[0208] In some embodiments, after obtaining the semantic feature map of each syllable, the semantic feature map of each syllable can be input into a first encoder, and the first encoder encodes the semantic feature map of each syllable to output the first semantic feature of each syllable.
[0209] Schematically, the first encoder may be an LSTM. One or more hidden layers may be included in the LSTM. A plurality of memory units are included in each hidden layer. Each memory unit corresponds to a syllable. When the number of hidden layers in the LSTM is plural, the plurality of hidden layers may be connected in series, that is, the output of the previous hidden layer serves as the input of the current hidden layer. Similarly, the output of the current hidden layer serves as the input of the next hidden layer.
[0210] Taking the j-th (j≥2) memory unit in the i-th (i≥2) hidden layer of the LSTM as an example, the input of this memory unit includes: the output feature of the (j - 1)-th memory unit in the i-th hidden layer and the output feature of the j-th memory unit in the (i - 1)-th hidden layer. This memory unit performs a weighted transformation on the above two output features to obtain the output feature of this memory unit. Then, the output feature of the memory unit is respectively input into the (j + 1)-th memory unit in the i-th hidden layer and the j-th memory unit in the (i + 1)-th hidden layer. Among them, the input of each memory unit in the first hidden layer includes the output feature of the previous memory unit in the first hidden layer and the semantic feature map of the corresponding syllable. The output feature of each memory unit in the last hidden layer is the first semantic feature of each syllable. Performing the above operations on each memory unit in each hidden layer is equivalent to performing recursive weighted processing in the entire LSTM, that is, performing forward encoding on the semantic feature map of each syllable.
[0211] In the above process, the LSTM model can perform forward encoding on the semantic feature map of each syllable. That is, when encoding the semantic feature map of each syllable, historical information of each syllable before this syllable needs to be introduced, which can improve the expression ability of the first semantic feature of each syllable. In some embodiments, in addition to the LSTM, the first encoder may also be a BLSTM. The BLSTM model can perform bidirectional encoding on the semantic feature map of each syllable, including both forward encoding and backward encoding. That is, when encoding the semantic feature map of each syllable, historical information of each syllable before this syllable and future information of each syllable after this syllable need to be introduced, which can further improve the expression ability of the first semantic feature of each syllable, but consumes more computing resources. Or, the first encoder may also be an RNN to simplify the complexity of the model, or other sequence-to-sequence models. The embodiments of the present application do not specifically limit the structure of the first encoder.
[0212] 505. The server determines phoneme duration information based on the syllable duration information and the first semantic features of the plurality of syllables. The phoneme duration information is used to represent the number of audio frames occupied by each of the plurality of phonemes included in the plurality of syllables.
[0213] In some embodiments, the server can predict the number of audio frames occupied by each phoneme in this syllable by combining the first semantic feature of each syllable and the number of audio frames occupied by this syllable indicated in the syllable duration information. The number of audio frames occupied by all phonemes can constitute the phoneme duration information. This predicted phoneme duration information is not a simple and mechanical equal division of each phoneme within the syllable, but is a comprehensive and accurate prediction of the audio frames by combining the context of this syllable in the lyrics, the first semantic feature of this syllable itself, and multiple aspects of information such as the singer, etc., which can achieve precise control of the continuous number of frames, i.e., the continuous duration, of each phoneme in the computer-synthesized singing audio.
[0214] Figure 7 FIG. 4 is a flowchart of determining phoneme duration information provided by an embodiment of the present application. As Figure 7 shown, when determining the phoneme duration information, the server can perform the following steps 5051-5053:
[0215] 5051. The server obtains the initial features of the multiple audio frames based on the syllable duration information and the first semantic features of the multiple syllables. The initial features are the first semantic features of the syllables to which the corresponding audio frames belong.
[0216] Optionally, for any one of the multiple syllables, when obtaining the initial features of each audio frame, the server can copy the first semantic feature of the syllable the first target number of times to obtain the initial features of the multiple audio frames included in the syllable, where the first target number is the value obtained by subtracting one from the number of audio frames occupied by the syllable.
[0217] In some embodiments, the server determines the number of audio frames occupied by this syllable from the syllable duration information, subtracts one from the number of audio frames occupied by this syllable to obtain the first target number, and then copies the first semantic feature of the syllable the first target number of times. The first target number of copied features and the original first semantic feature of this syllable are together determined as the initial features of the multiple audio frames. For example, if the first syllable occupies 10 frames, then the first semantic feature of the first syllable is copied 9 times to obtain 9 features, and the 9 copied features and the original first semantic feature of the first syllable are determined as the initial features of a total of 10 audio frames.
[0218] In the above step 5051, by using the syllable duration information to copy the first semantic features of each syllable at the original syllable level, the initial features of each audio frame at the frame level are obtained. This is equivalent to using a length regulator to limit the total number of audio frames occupied by each syllable, and based on this, the phoneme corresponding to each frame is predicted, that is, the phoneme duration information is intelligently predicted.
[0219] 5052. The server determines multiple phonemes corresponding to each of the multiple audio frames based on the initial features of the multiple audio frames and the position features of the multiple audio frames.
[0220] In some embodiments, for each audio frame, the server can predict the phoneme corresponding to this audio frame by using the corresponding initial feature and position feature. Optionally, in the prediction process, the following steps 5052A - 5052D are performed:
[0221] 5052A. The server splices the initial features of the multiple audio frames and the position features of the multiple audio frames respectively to obtain multiple first spliced features.
[0222] Optionally, for each audio frame, the server obtains the position feature (e.g., position encoding) of this audio frame based on the position information of this audio frame in all audio frames, and splices the initial feature of this audio frame and the position feature of this audio frame to obtain a first spliced feature of this audio frame.
[0223] In the above step 5052A, the splicing method can save the fusion complexity of the initial feature and the position feature. Alternatively, fusion methods such as element - wise addition, element - wise multiplication, and bilinear pooling can also be used to improve the expression ability of the first spliced feature. The embodiments of the present application do not specifically limit the fusion method of the initial feature and the position feature.
[0224] 5052B. The server performs convolution processing on the multiple first spliced features to obtain multiple first target features.
[0225] Optionally, when the server obtains the first target feature, it can input the multiple first spliced features into a third feature extraction model, and the third feature extraction model performs convolution processing on each first spliced feature to obtain the corresponding first target feature, where the third feature extraction model is used to extract the first target feature of the input audio frame. Schematically, the third feature extraction model can also be a CNN. For example, it has a similar structure to the first feature extraction model, that is, it also includes one or more CNN blocks, and the CNN block also includes two layers: a 1 - D convolutional layer and a normalization layer, which will not be elaborated here. It should be noted that even if both the first feature extraction model and the third feature extraction model are CNN models including multiple CNN blocks, they can have different numbers of CNN blocks. For example, the first feature extraction model includes 6 CNN blocks, and the third feature extraction model includes 4 CNN blocks, etc. The embodiments of the present application do not specifically limit this.
[0226] In the above process, by using a CNN as the third feature extraction model, since a 1D convolutional layer is adopted in each CNN block, the depth of the CNN can be deepened without increasing the receptive field to introduce more non-linear factors, and the dimensionality of each first target feature in the output can be flexibly increased or decreased, facilitating the extraction of deeper feature information. Of course, in addition to the CNN, third feature extraction models with structures such as DNN and VGG can also be used. The embodiment of the present application does not specifically limit the structure of the third feature extraction model.
[0227] 5052C. The server performs a weighting process on the multiple first target features to obtain multiple second target features.
[0228] Optionally, when obtaining the second target features, the server can input the multiple first target features into a second encoder, and the second encoder performs a weighting process on each first target feature to obtain the corresponding second target feature. The structure of the second encoder includes but is not limited to: LSTM, BLSTM, RNN, etc. For example, the second encoder is also an LSTM network. Through the LSTM network, forward encoding can be performed on each first target feature. That is, when encoding each first target feature, the historical information of each audio frame before this audio frame needs to be introduced, which can improve the expression ability of the second target feature of each audio frame. However, the second encoder may have a different number of hidden layers or different parameters from the first encoder, which will not be elaborated here.
[0229] 5052D. The server predicts multiple phonemes corresponding to each of the multiple audio frames based on the multiple second target features.
[0230] Optionally, when the server predicts the phoneme corresponding to each audio frame, it can perform a fully connected process on any one of the multiple second target features to obtain multiple prediction probabilities for the audio frame corresponding to the second target feature. The prediction probability is used to represent the possibility that the audio frame corresponds to a phoneme; the phoneme with the highest prediction probability is determined as the phoneme corresponding to the audio frame.
[0231] In some embodiments, the server inputs the second target feature of each audio frame output by the second encoder into a fully connected layer. The fully connected layer performs a fully connected process on each second target feature, and for each audio frame corresponding to each second target feature, multiple prediction probabilities are output. Each prediction probability is used to represent the possibility that the audio frame corresponds to a phoneme. The maximum prediction probability among the multiple prediction probabilities is obtained, and the phoneme corresponding to the maximum prediction probability is determined as the phoneme corresponding to the audio frame. In other words, the above prediction process can be regarded as a classification process, and the phoneme ID of each phoneme represents a classification label.
[0232] Figure 8 This is a schematic diagram of the principle of a duration model provided by an embodiment of the present application. As Figure 8 shown, taking the first feature extraction model, the second feature extraction model, and the third feature extraction model as CNNs, and the first encoder and the second encoder as LSTMs as examples for illustration. Among them, the CNN includes N CNN blocks (N≥1), and each CNN block includes a one-dimensional convolutional layer and a normalization layer.
[0233] The input of the duration model includes: the syllable sequence 801, the phoneme sequence 802, and the syllable duration information 803 obtained in the above step 503, the pitch sequence 804 obtained in the above step 502, and the specified singer identifier 805. The syllable sequence 801, the phoneme sequence 802, the pitch sequence 804, and the singer identifier 805 are respectively subjected to embedding processing to obtain the syllable feature 811, the embedding feature 8121 of each phoneme, the pitch feature 814, and the singer feature 815. Then, the position feature 8122 of each phoneme is obtained, and the embedding feature 8121 and the position feature 8122 of each phoneme are concatenated to obtain the target concatenated feature of each phoneme. The target concatenated feature of each phoneme is input into the first feature extraction model 8123, and the feature map of each phoneme is output. The feature map of each phoneme is input into the pooling layer 8124, and the pooled feature map of each phoneme is output. Then, for each syllable, the pooled feature maps of the phonemes belonging to the syllable, the syllable feature 811 of the syllable, the pitch feature 814 corresponding to the pitch of the syllable, and the singer feature 815 are concatenated to obtain the initial semantic feature of each syllable. The initial semantic feature of each syllable is input into the second feature extraction model 821, and the semantic feature map of each syllable is output. The semantic feature map of each syllable is input into the first encoder 822, and the first semantic feature of each syllable is output. The first semantic feature of each syllable and the syllable duration information 803 are input into the length regulator 823, and the initial feature of each audio frame is output. The initial feature of each audio frame is concatenated with the position feature of each audio frame to obtain the first concatenated feature of each audio frame. The first concatenated feature of each audio frame is input into the third feature extraction model 824, and the first target feature of each audio frame is output. The first target feature of each audio frame is input into the second encoder 825, and the second target feature of each audio frame is output. Finally, the second target feature of each audio frame is input into the fully connected layer 826, and the prediction probability corresponding to all phoneme IDs of each audio frame is output. Based on the multiple prediction probabilities corresponding to each audio frame, the phoneme corresponding to the phoneme ID with the highest prediction probability is selected as the phoneme corresponding to the current audio frame.
[0234] As Figure 8The duration model shown in [description] can, by setting the length regulator 823, use the syllable duration information 803 given in the musical score as a constraint to first ensure the total number of audio frames occupied by the entire syllable, and then, based on the first semantic features obtained by combining a series of encodings of the input phoneme sequence, pitch sequence, singer identifier, and syllable sequence, and in combination with the subsequent CNN block (i.e., the third feature extraction model) and LSTM (i.e., the second encoder), perform probability prediction for each frame relative to each phoneme ID. Subsequently, the phoneme with the highest prediction probability for each frame is selected as the phoneme corresponding to each frame, and thus the duration number of frames (i.e., phoneme duration information) of each phoneme can be converted, achieving the effect of accurately predicting and controlling the duration and number of frames of each phoneme.
[0235] 5053. The server determines the phoneme duration information based on the correspondence between the multiple audio frames and the multiple phonemes.
[0236] In some embodiments, precisely because the server determines the phoneme corresponding to each audio frame, it can, conversely, determine the phoneme duration information based on the correspondence between the multiple audio frames and the multiple phonemes. For example, among the 10 audios of the first syllable, there are 3 phonemes. The first to third audio frames all correspond to phoneme A, the fourth audio frame corresponds to phoneme B, and the fifth to tenth audio frames correspond to phoneme C. Thus, it can be determined that in the first syllable, the number of audio frames occupied by phoneme A is 3, the number of audio frames occupied by phoneme B is 1, and the number of audio frames occupied by phoneme C is 6.
[0237] In the above steps 5052A - 5052D, since the number of audio frames occupied by each phoneme is relatively diverse and constantly changing for the overall duration model, it is relatively complex to directly learn the number of audio frames occupied by each phoneme. However, given an audio, the type of phoneme corresponding to each audio frame is fixed. Therefore, after changing the thinking, first use the syllable duration information of the musical score to limit the total number of audio frames occupied by the entire syllable, and then provide the duration model with various information such as the context of this syllable in the lyrics, the first semantic features of this syllable itself, and the singer, and extract the features of each piece of information. Finally, let the duration model predict the phoneme corresponding to each audio frame, making the acquisition of phoneme duration information fast and convenient. In addition, because the type of phoneme is fixed, it can improve the learning speed of the duration model and shorten the training duration of the duration model.
[0238] In the above steps 504-505, the server determines phoneme duration information based on the musical score, the lyrics, and the syllable duration information. By further predicting the phoneme duration information on the basis of the syllable duration information given in the musical score, since the phoneme duration information can represent the number of audio frames occupied by each phoneme, the accuracy of the acoustic model is no longer limited to the rough syllable level, but can reach the precise control at the phoneme level, greatly improving the naturalness of singing synthesis.
[0239] 506. The server obtains the second semantic features of the multiple audio frames based on the lyrics, the musical score, and the phoneme duration information. The second semantic features are used to represent the semantics when the singer sings the corresponding phoneme at the corresponding pitch.
[0240] In the above step 505, the server predicts the phoneme duration information through the duration model and can control the number of audio frames occupied by each phoneme. Under the guidance of this phoneme duration information, the second semantic features of each audio frame at the frame level can be re-extracted in combination with the phoneme sequence, the pitch sequence, and the singer identifier.
[0241] Figure 9 It is a flowchart for obtaining the second semantic features provided by an embodiment of the present application. As Figure 9 shown, when the server obtains the second semantic features for each audio frame, the following steps 5061-5064 can be executed:
[0242] 5061. The server obtains the phoneme features of the multiple phonemes in the lyrics and the pitch features of the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables.
[0243] Among them, the method for obtaining the pitch features is similar to step 5043 above and will not be elaborated here. Additionally, if the pitch features have been obtained in step 5043 above, there is no need to obtain them repeatedly in this step 5061.
[0244] Optionally, when the server obtains the phoneme features, it can perform embedding encoding on each phoneme in the phoneme sequence to obtain the embedding features of each phoneme. Then, based on the sorting information of each phoneme in the phoneme sequence, the position features of each phoneme are obtained. The embedding features and position features of each phoneme are concatenated to obtain the target concatenated features of each phoneme. The target concatenated features of each phoneme are input into the fourth feature extraction model, and the fourth feature extraction model encodes and decodes the target concatenated features of each phoneme to obtain the phoneme features of each phoneme, where the fourth feature extraction model is used to extract the phoneme features of the input phoneme.
[0245] In the above process, the fourth feature extraction model encodes and then decodes the target splicing features of each phoneme, which is beneficial to extracting deeper phoneme features. Compared with simple convolution processing, it can enrich the expression ability of phoneme features.
[0246] Schematically, the fourth feature extraction model can be an FFT (Feed-Forward Transformer). The FFT may include an encoding module and a decoding module. The encoding module is used to encode the features of phonemes, and the decoding module is used to decode the features of phonemes. The encoding module includes one or more feed-forward transformation blocks (FFT blocks), and the decoding module also includes one or more FFT blocks. Usually, the number of FFT blocks in the encoding module and the decoding module is the same. For the encoding module or the decoding module, when the number of FFT blocks is multiple, the multiple FFT blocks are connected in series (i.e., cascaded). That is, the output of the previous FFT block is used as the input of the current FFT block. Similarly, the output of the current FFT block is used as the input of the next FFT block. Among them, the input of the first FFT block in the encoding module is the target splicing feature of each phoneme.
[0247] For example, the server inputs the target splicing feature of each phoneme into the first FFT block of the encoding module. The first FFT block weights the target splicing feature of each phoneme and outputs the first hidden vector of each phoneme. Then, the first hidden vector of each phoneme is input into the second FFT block, and so on, until the last FFT block finishes processing, obtaining the last hidden vector of each phoneme output by the last FFT block, that is, the final hidden state feature of each phoneme. Then, based on a length regulator to adjust the length of the hidden state feature of each phoneme. Next, the adjusted hidden state feature of each phoneme is input into the first FFT block of the decoding module. The first FFT block weights the hidden state feature of each phoneme and outputs the first hidden vector of each phoneme. Then, the first hidden vector of each phoneme is input into the second FFT block, and so on, until the last FFT block finishes processing, obtaining the last hidden vector of each phoneme output by the last FFT block, which is the final phoneme feature of each phoneme.
[0248] Optionally, each FFT block includes a four-layer structure: a multi-head attention layer, a first residual normalization layer, a one-dimensional convolutional layer, and a second residual normalization layer. Among them, multiple attention layers are used to comprehensively extract the correlation relationships between phonemes in the pinyin sequence from multiple expression subspaces. The one-dimensional convolutional layer is similar to the above-mentioned CNN block and will not be elaborated here. Among them, the input of the first residual normalization layer is the first residual feature obtained by concatenating the input of the multi-head attention layer and the output of the multi-head attention layer. The input of the second residual normalization layer is the second residual feature obtained by concatenating the input of the one-dimensional convolutional layer and the output of the one-dimensional convolutional layer. The above two residual normalization layers actually still belong to the normalization layer, but there is a residual connection between the normalization layer and the input and output of the previous layer, so it is called the residual normalization layer. For example, the target concatenated feature of each phoneme is input into the multi-head attention layer of the first FFT block of the FFT encoding module. The multi-head attention layer weights the target concatenated feature of each phoneme to obtain the attention feature of each phoneme. The target concatenated feature and the attention feature of each phoneme are concatenated to obtain the first residual feature of each phoneme. The first residual feature of each phoneme is input into the first residual normalization layer for normalization processing to obtain the first standard feature of each phoneme. Then, the first standard feature of each phoneme is input into the one-dimensional convolutional layer, and a 1×1 convolutional kernel is used to perform convolutional processing on the first standard feature of each phoneme to obtain the convolutional feature map of each phoneme. The first standard feature and the convolutional feature map of each phoneme are concatenated to obtain the second residual feature of each phoneme. The second residual feature of each phoneme is input into the second residual normalization layer for normalization processing to obtain the second standard feature of each phoneme. The second standard feature of each phoneme mentioned above will be used as the input of the second FFT block in the encoding module, and so on. The processing process of the FFT block in the decoding module is the same and will not be elaborated here. Finally, the second standard feature of each phoneme output by the last FFT block in the decoding module is obtained as the phoneme feature of each phoneme.
[0249] In the above process, FFT is used as the fourth feature extraction model. Since the multi-head attention layer is introduced in each FFT block, the phoneme feature can contain the correlation relationships between phonemes in the pinyin sequence. Moreover, the two residual normalization layers in the FFT block avoid losing the original detailed information during the feature extraction process, so that the phoneme feature has stronger expression ability. In some embodiments, in addition to FFT, fourth feature extraction models with structures such as CNN, DNN, and VGG can also be used. The embodiments of the present application do not specifically limit the structure of the fourth feature extraction model.
[0250] 5062. The server obtains the frame-level phoneme features of the multiple audio frames based on the phoneme features of the multiple phonemes and the phoneme duration information, where the frame-level phoneme features are the phoneme features of the phoneme to which the corresponding audio frame belongs.
[0251] In some embodiments, for any one of the multiple phonemes, when the server obtains the corresponding frame-level phoneme features, it may copy the phoneme features of the phoneme a second target number of times to obtain the frame-level phoneme features of the multiple audio frames included in the phoneme, where the second target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme.
[0252] In some embodiments, the server determines the number of audio frames occupied by the current phoneme from the phoneme duration information, subtracts one from the number of audio frames occupied by the current phoneme to obtain the second target number, and then copies the phoneme features of the phoneme the second target number of times. The second target number of copied features and the original phoneme features of the current phoneme are together determined as the frame-level phoneme features of the multiple audio frames. For example, if the first phoneme occupies 3 frames, then the phoneme features of the first phoneme are copied 2 times to obtain 2 features, and the 2 copied features and the original phoneme features of the first phoneme are determined as the frame-level phoneme features of a total of 3 audio frames.
[0253] In the above process, since the duration model has predicted the phoneme duration information in step 505 above, under the indication of the phoneme duration information, the phoneme features of each phoneme can be copied, so that the frame-level phoneme features of each audio frame can be obtained, that is, a sequence of frame-level phoneme features is obtained, which greatly improves the accuracy of subsequent acquisition of acoustic feature parameters, silence parameters, and fundamental frequency parameters.
[0254] 5063. The server obtains the frame-level pitch features of the multiple audio frames based on the pitch features of the multiple pitches and the phoneme duration information, where the frame-level pitch features are the pitch features of the pitch corresponding to the phoneme to which the corresponding audio frame belongs.
[0255] In some embodiments, for any one of the multiple pitches, when the server obtains the corresponding frame-level pitch features, it may copy the pitch features of the pitch a third target number of times to obtain the frame-level pitch features of the multiple audio frames included in the phoneme corresponding to the pitch, where the third target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme corresponding to the pitch.
[0256] In some embodiments, the server first determines the phoneme corresponding to the current pitch, then determines the number of audio frames occupied by the phoneme from the phoneme duration information, subtracts one from the number of audio frames occupied by the phoneme to obtain the third target number of times, and then copies the pitch feature of the current pitch the third target number of times. The third target number of copied features and the original pitch feature of the current pitch are together determined as the frame-level pitch features of the multiple audio frames. For example, the mth (m≥1) pitch corresponds to the mth phoneme, and the mth phoneme occupies 5 frames. Then, the pitch feature of the mth pitch is copied 4 times to obtain 4 features, and the 4 copied features and the original pitch feature of the mth pitch are determined as the frame-level pitch features of a total of 5 audio frames.
[0257] In the above process, since the duration model has predicted the phoneme duration information in step 505 above, under the indication of the phoneme duration information, the pitch features of each pitch can be copied, so that the frame-level pitch features of each audio frame can be obtained, that is, a sequence of frame-level pitch features is obtained, greatly improving the accuracy of subsequent acquisition of acoustic feature parameters, silence parameters, and fundamental frequency parameters.
[0258] 5064. The server splices the frame-level phoneme features of the multiple audio frames with the frame-level pitch features of the multiple audio frames and the singer features of the singer identifier to obtain the second semantic features of the multiple audio frames.
[0259] In some embodiments, for each audio frame, when obtaining the corresponding second semantic features, the server can splice the frame-level phoneme feature of the audio frame, the frame-level pitch feature of the audio frame, and the singer feature to obtain the second semantic feature of the audio frame.
[0260] In some embodiments, for each audio frame, the server can splice the frame-level phoneme feature of the audio frame and the frame-level pitch feature of the audio frame to obtain the second semantic feature of the audio frame, and splice the second semantic features of all audio frames and then splice them with the singer feature to reduce the computational complexity.
[0261] In some embodiments, in addition to splicing, fusion methods such as element-wise addition, element-wise multiplication, and bilinear pooling can also be used. The embodiments of the present application do not specifically limit the fusion method of frame-level phoneme features, frame-level pitch features, and singer features.
[0262] In steps 5061-5064 above, the server can obtain a sequence of frame-level phoneme features and a sequence of frame-level pitch features through phoneme duration information. Compared with traditional singing synthesis schemes, the second semantic features of each audio frame at the frame level are extracted, greatly improving the naturalness of singing synthesis.
[0263] 507. The server encodes the second semantic features of the multiple audio frames to obtain intermediate features of the multiple audio frames.
[0264] In some embodiments, after the server obtains the second semantic features of each audio frame, it can input the second semantic features of each audio frame into an acoustic parameter prediction model. Optionally, the acoustic parameter prediction model includes a first linear layer, an encoding module, a decoding module, a second linear layer, and a post-processing layer. For example, the second semantic features of each audio frame are input into the first linear layer, and the first linear layer performs a linear transformation on the second semantic features of each audio frame to obtain the second semantic features of each audio frame after the linear transformation. Then, the second semantic features of each audio frame after the linear transformation are input into the encoding module, and the encoding module encodes the second semantic features of each audio frame after the linear transformation to obtain intermediate features of each audio frame.
[0265] 508. The server decodes the intermediate features of the multiple audio frames to obtain third semantic features of the multiple audio frames.
[0266] In some embodiments, the server inputs the intermediate features of each audio frame into the decoding module, and the decoding module decodes the intermediate features of each audio frame to obtain the third semantic features of each audio frame.
[0267] Schematically, the acoustic parameter prediction model can have a structure similar to that of the FFT. That is, the encoding module of the acoustic parameter prediction model includes one or more FFT blocks, and the decoding module also includes one or more FFT blocks. Usually, the number of FFT blocks in the encoding module and the decoding module is the same. The number of FFT blocks in the acoustic parameter prediction model can be the same as or different from the number of FFT blocks in the fourth feature extraction model. The embodiments of the present application do not specifically limit this. The connection manner, internal structure, and processing process of the FFT blocks in the acoustic parameter prediction model are similar to those of the FFT blocks in the fourth feature extraction model, and will not be elaborated here.
[0268] In the above steps 507-508, by first encoding and then decoding the second semantic features of each audio frame, it is equivalent to further extracting features on the basis of the second semantic features to obtain the third semantic features of each audio frame. These third semantic features can be input into the prediction process of the acoustic feature parameters to improve the accuracy of the acoustic feature parameters.
[0269] 509. The server processes the third semantic features of the multiple audio frames to obtain acoustic feature parameters of the multiple audio frames.
[0270] Among them, the acoustic feature parameter is used to characterize the acoustic feature of the corresponding audio frame.
[0271] Optionally, the acoustic feature parameter may be at least one of the MGC parameter or the BAP parameter of each audio frame. Of course, the acoustic feature parameter may also be the MFCC parameter of each audio frame, etc. The embodiments of the present application do not specifically limit the type of the acoustic feature parameter.
[0272] In some embodiments, the server may input the third semantic feature of each audio frame into the second linear layer, perform a linear transformation on the third semantic feature of each audio frame through the second linear layer to obtain the third semantic feature of each audio frame after the linear transformation, input the third semantic feature of each audio frame after the linear transformation into the post-processing layer, and process the third semantic feature of each audio frame after the linear transformation through the post-processing layer to obtain the standard semantic feature of each audio frame. By splicing the standard semantic feature of each audio frame and the third semantic feature after the linear transformation, the acoustic feature parameter of each audio frame can be obtained.
[0273] Schematically, in the post-processing layer of the acoustic parameter prediction model, one or more 1D convolutional layers and a normalization layer may be included. For example, the post-processing layer includes 5 1D convolutional layers and 1 normalization layer. The 1D convolutional layer and the normalization layer have both been introduced in step 5042 above and will not be elaborated here.
[0274] In the above steps 506-509, the server obtains the acoustic feature parameters of multiple audio frames in the to-be-generated singing voice based on the musical score, the lyrics, and the phoneme duration information. In some embodiments, the acoustic feature parameters of the multiple audio frames may be directly input into the vocoder, and the vocoder synthesizes the singing voice matching the acoustic feature parameters, which can simplify the singing voice generation process. In some embodiments, by performing the following steps 510-513, at least one of the silence parameter or the fundamental frequency parameter is calculated for each audio frame to accurately control the pitch of each audio frame, thereby improving the expressiveness of the synthesized singing voice.
[0275] Figure 10 is a schematic diagram of the principle of a singing voice generation system provided by the embodiments of the present application. As Figure 10 shown, the singing voice generation system includes a duration model 1010, a pitch prediction model 1020, an acoustic parameter prediction model 1030, and a vocoder 1040. Among them, the input of the duration model 1010 is the musical score information 1001, and the output is the number of frames of the phoneme, that is, the phoneme duration information 1011. The structure of the duration model 1010 may be as Figure 8 shown, which will not be elaborated here. Based on the musical score information 1001, the pinyin sequence of the lyrics, that is, the syllable sequence 1002, and the pitch sequence 1003 of the musical score can be obtained.
[0276] Based on the syllable sequence 1002, a phoneme sequence can be obtained. The phoneme sequence is input into the phoneme encoding layer to output the embedding features of each phoneme in the phoneme sequence. The embedding features of each phoneme are concatenated with the position features of each phoneme to obtain the target concatenated features of each phoneme. The target concatenated features of each phoneme are input into the fourth feature extraction model 1050 to output the phoneme features of each phoneme. Figure 10Taking the fourth feature extraction model 1050 as FFT as an example, the FFT includes N (N≥1) FFT blocks. The phoneme features and phoneme duration information 1011 of each phoneme are input into the length regulator 1051 to obtain the frame-level phoneme features of each audio frame. Similarly, based on the pitch sequence 1003 of the musical score and combined with the phoneme duration information 1011, the frame-level pitch features 1013 of each audio frame can also be obtained. The given singer identifier 1004 is obtained, and the singer identifier 1004 is subjected to embedding processing to obtain the singer feature 1014. The frame-level phoneme features, frame-level pitch features 1013, and singer features 1014 of each audio frame are concatenated to obtain the second semantic feature of each audio frame. The second semantic feature of each audio frame is input into the acoustic parameter prediction model 1030. The acoustic parameter prediction model 1030 includes a first linear layer 1031, an encoding and decoding module 1032, a second linear layer 1033, and a post-processing layer 1034. Among them, the encoding and decoding module 1032 includes an encoding module and a decoding module. Schematically, the encoding module and the decoding module together include N FFT blocks. The second semantic feature of each audio frame is linearly transformed by the first linear layer 1031 to obtain the linearly transformed second semantic feature of each audio frame. The linearly transformed second semantic feature of each audio frame is input into the encoding module, and the encoding module encodes the linearly transformed second semantic feature of each audio frame to obtain the intermediate feature of each audio frame. The intermediate feature of each audio frame is input into the decoding module, and the decoding module decodes the intermediate feature of each audio frame to obtain the third semantic feature of each audio frame. On the one hand, the third semantic feature of each audio frame is input into the second linear layer 1033, and the second linear layer 1033 linearly transforms the third semantic feature of each audio frame to obtain the linearly transformed third semantic feature of each audio frame. The linearly transformed third semantic feature of each audio frame is input into the post-processing layer 1034, and the post-processing layer 1034 processes the linearly transformed third semantic feature of each audio frame to obtain the standard semantic feature of each audio frame. The standard semantic feature of each audio frame and the linearly transformed third semantic feature are concatenated to obtain the acoustic feature parameter 1061 of each audio frame. On the other hand, the third semantic feature of each audio frame and the frame-level pitch feature 1013 of each audio frame are input into the pitch prediction model 1020 together, and the mute parameter vuv and the fundamental frequency parameter lf0 1062 of each audio frame are output. The internal structure of the pitch prediction model 1020 will be introduced in the following steps 510-513 and will not be elaborated here.Next, the acoustic feature parameters 1061 of each audio frame, the voice activity parameter vuv, and the fundamental frequency parameter lf0 1062 are concatenated to obtain the target acoustic features of each audio frame. The target acoustic features of each audio frame are input into the vocoder 1040, and the vocoder 1040 outputs the audio signal 1063 of the finally synthesized singing voice. Among them, the process of generating the audio signal 1063 refers to the following steps 514-515, which will not be elaborated here.
[0277] It should be noted that Figure 10 only one example of the singing voice synthesis system is given, but it is not limited to this. For example, the pitch prediction model 1020 can be embedded inside the acoustic parameter prediction model 1030, so that the acoustic parameter prediction model 1030 can output the parameters of each audio frame at one time: the acoustic feature parameters 1061, the voice activity parameter vuv, and the fundamental frequency parameter lf0 1062.
[0278] It should be noted that Figure 10 only the FFT of the encoding and decoding module 1032 is taken as an example for illustration. The FFT is based on the attention mechanism and can have different emphases when predicting the acoustic feature parameters of each audio frame. Of course, the encoding and decoding module with the CNN or RNN structure can also be used. In this case, the model training speed and the singing voice synthesis speed may be slightly lost, but the model can learn more context information to achieve a better singing voice synthesis effect.
[0279] In the above process, by embedding the data exchange with the pitch prediction model 1020 outside the acoustic parameter prediction model 1030, the information related to the fundamental frequency f0 (including vuv and lf0) can be obtained through the pitch prediction model 1020. Combining the singer information and the output of the encoding and decoding module 1032, multi-dimensional acoustic feature parameters can be generated, which not only enables the full use of the pitch information in the music score during training, but also ensures that there will be no conflict when generating the audio by combining the prediction results of the acoustic parameter prediction model 1030 and the pitch prediction model 1020. Moreover, since the acoustic feature parameters 1061, the voice activity parameter vuv, and the fundamental frequency parameter lf0 1062 are combined and sent to the vocoder 1040 to synthesize the audio signal 1063 of the singing voice, it can be ensured that the pitch of each frame in the predicted audio signal 1063 conforms to the pitch of the corresponding note in the music score, and the audio signal 1063 will not be out of tune.
[0280] 510. The server linearly processes the third semantic features of the multiple audio frames to obtain multiple third target features.
[0281] In some embodiments, the server may input the third semantic features of the multiple audio frames and the frame-level pitch features of the multiple audio frames into a pitch prediction model, which includes a third linear layer, a fourth linear layer, and a one-dimensional convolutional layer. Optionally, the third semantic features of each audio frame are input into the third linear layer, and the third linear layer performs linear processing on the third semantic features of each audio frame to obtain the third target features of each audio frame.
[0282] 511. The server performs linear processing on the frame-level pitch features of the multiple audio frames to obtain multiple fourth target features.
[0283] In some embodiments, the server may input the frame-level pitch features of each audio frame into the fourth linear layer, and the fourth linear layer performs linear processing on the frame-level pitch features of each audio frame to obtain the fourth target features of each audio frame.
[0284] 512. The server respectively concatenates the multiple third target features and the multiple fourth target features to obtain multiple second concatenated features.
[0285] In some embodiments, for each audio frame, the third target feature and the fourth target feature of the audio frame may be concatenated to obtain the second concatenated feature of the audio frame, and the above operation is repeatedly performed to obtain the second concatenated feature of each audio frame.
[0286] 513. The server performs convolutional processing on the multiple second concatenated features to obtain at least one of the silence parameters or fundamental frequency parameters of each of the multiple audio frames. The silence parameter is used to characterize whether the corresponding audio frame is a silent segment, and the fundamental frequency parameter is used to characterize the logarithm of the fundamental frequency of the corresponding audio frame.
[0287] In some embodiments, the second concatenated feature of each audio frame is input into the one-dimensional convolutional layer, and the one-dimensional convolutional layer performs convolutional processing on the second concatenated feature of each audio frame to obtain at least one of the silence parameter vuv or the fundamental frequency parameter lf0 of each audio frame.
[0288] In the above process, through the pitch prediction model, at least one of the silence parameter or the fundamental frequency parameter is additionally obtained for each audio frame, so that at least one of the silence parameter or the fundamental frequency parameter can be input into the vocoder together with the acoustic feature parameters after concatenation to perform precise pitch control on the generated singing voice.
[0289] Figure 11 is a schematic diagram of the principle of a pitch prediction model provided by an embodiment of the present application, as Figure 11As shown in the figure, the pitch prediction model includes a third linear layer 1111, a fourth linear layer 1112, and a 1D convolutional layer 1113. Using the phoneme duration information, the pitch sequence given in the musical score can be converted into a frame-level musical score pitch sequence 1101. After performing an embedding process on the frame-level musical score pitch sequence 1101, the frame-level pitch feature 1102 of each audio frame is obtained. Of course, the phoneme duration information can also be directly used to directly obtain the frame-level pitch feature 1102 of each audio frame from the pitch feature of each pitch in the pitch sequence, thus eliminating the need to perform the step of obtaining the frame-level musical score pitch sequence 1101 and simplifying the singing synthesis process. In addition, obtaining the third semantic feature of each audio frame obtained in step 508 above is also the intermediate output 1103 of the acoustic parameter prediction model. The intermediate output 1103 of the acoustic parameter prediction model is input into the third linear layer 1111 to output the third target feature of each audio frame; the frame-level pitch feature 1102 of each audio frame is input into the fourth linear layer 1112 to output the fourth target feature of each audio frame. The third target feature and the fourth target feature of each audio frame are concatenated to obtain the second concatenated feature of each audio frame, and the second concatenated feature of each audio frame is input into the 1D convolutional layer 1113 to output the voice unvoiced parameter vuv and the fundamental frequency parameter lf0 of each audio frame.
[0290] In the above process, after obtaining the phoneme duration information through the duration model, the pitch feature of each pitch in the pitch sequence of the musical score can be obtained by simply copying to obtain the frame-level pitch feature of each audio frame. This frame-level pitch feature can be used as one of the inputs of the pitch prediction model. Combining with the third semantic feature of each audio frame obtained in step 508 above (i.e., the intermediate output of the acoustic parameter prediction model), specific fundamental frequency parameter lf0 prediction and voice unvoiced parameter vuv prediction are performed for each frame.
[0291] By using lf0 as the fundamental frequency parameter, that is, taking the logarithm of the fundamental frequency f0, it is because the musical score pitches increase in multiples. After taking the logarithm, it can ensure that the gap between different scales is not too large, and it can make the vocoder pay more attention to the low-frequency part that the human ear is concerned about. In addition, since the fundamental frequency parameter lf0 is a continuous line segment, and an additional voice unvoiced parameter vuv is predicted, it can make the vocoder pay more attention to the truly pronounced parts in the audio, and also make the synthesized audio signal more in line with the situation of the mixture of voiceless segments and voiced segments in human singing, so as to improve the expressiveness of the audio signal.
[0292] 514. The server concatenates the acoustic feature parameters of the multiple audio frames and at least one of the voice unvoiced parameters or fundamental frequency parameters of the multiple audio frames respectively to obtain the target acoustic features of the multiple audio frames.
[0293] In some embodiments, the acoustic feature parameters include MGC parameters and BAP parameters. At this time, for each audio frame, the MGC parameters, BAP parameters, silence parameters, and fundamental frequency parameters of the audio frame can be concatenated to obtain the target acoustic features of the audio frame. These target acoustic features characterize the acoustic features of the audio frame to be synthesized from multiple dimensions, which helps to reduce the model pressure of the vocoder. Moreover, the silence parameters and fundamental frequency parameters can control the pitch of each audio frame in the audio signal finally synthesized by the vocoder, ensuring that the fundamental frequency of the audio frame in the synthesized audio signal is highly similar to the silence parameters and fundamental frequency parameters, so as to achieve the purpose of precise pitch control.
[0294] 515. The server inputs the target acoustic features of the multiple audio frames into the vocoder, and synthesizes the audio signal of the to-be-generated singing voice through the vocoder.
[0295] Optionally, the vocoder can be a WORLD vocoder or a vocoder with other structures. The embodiments of the present application do not specifically limit the structure of the vocoder.
[0296] In the above process, for the vocoder, not only the acoustic feature parameters of each audio frame are passed, but also at least one of the silence parameters or fundamental frequency parameters of each audio frame is provided, so that not only can the number of audio frames occupied by each phoneme be accurately controlled by introducing the phoneme duration information, but also the pitch of each phoneme can be accurately controlled by introducing at least one of the silence parameters or fundamental frequency parameters, greatly improving the naturalness and expressiveness of the output singing voice.
[0297] In the above steps 510-515, based on the acoustic feature parameters of the multiple audio frames, the to-be-generated singing voice is generated. Since the pitch prediction model is introduced and combined with the acoustic parameter prediction model to predict the silence parameters and fundamental frequency parameters of each audio frame, the pitch of the singing voice synthesized by the vocoder can be accurately controlled. Of course, the pitch prediction model can also not be introduced during the singing voice synthesis process, which can simplify the singing voice synthesis process.
[0298] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present disclosure, which will not be elaborated one by one here.
[0299] The method provided by the embodiments of the present application further predicts the phoneme duration information on the basis of the syllable duration information given in the musical score. Since the phoneme duration information can represent the number of audio frames occupied by each phoneme, the accuracy of the acoustic model is no longer limited to the rough syllable level, but can reach the precise control at the phoneme level, greatly improving the naturalness of the singing voice synthesis.
[0300] Figure 12 It is a schematic structural diagram of a singing voice generation device provided by the embodiments of the present application, as Figure 12As shown, the device includes:
[0301] A first determination module 1201, configured to determine syllable duration information based on a musical score and the lyrics corresponding to the musical score, where the syllable duration information is used to characterize the number of audio frames occupied by each of the multiple syllables in the lyrics;
[0302] A second determination module 1202, configured to determine phoneme duration information based on the musical score, the lyrics, and the syllable duration information, where the phoneme duration information is used to characterize the number of audio frames occupied by each of the multiple phonemes included in the multiple syllables;
[0303] An acquisition module 1203, configured to acquire acoustic feature parameters of multiple audio frames in the to-be-generated singing voice based on the musical score, the lyrics, and the phoneme duration information, where the acoustic feature parameters are used to characterize the acoustic features of the corresponding audio frames;
[0304] A generation module 1204, configured to generate the to-be-generated singing voice based on the acoustic feature parameters of the multiple audio frames.
[0305] The device provided by the embodiments of the present application further predicts phoneme duration information on the basis of the syllable duration information given in the musical score. Since the phoneme duration information can characterize the number of audio frames occupied by each phoneme, the accuracy of the acoustic model is no longer limited to the rough syllable level, but can reach the precise control at the phoneme level, greatly improving the naturalness of singing voice synthesis.
[0306] In a possible implementation manner, based on Figure 12 the composition of the device, the second determination module 1202 includes:
[0307] A first acquisition sub-module, configured to acquire the multiple syllables and the multiple phonemes in the lyrics;
[0308] The first acquisition sub-module is further configured to acquire the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables;
[0309] A second acquisition sub-module, configured to acquire a first semantic feature of the multiple syllables based on the multiple syllables, the multiple phonemes, the multiple pitches, and the singer identifier, where the first semantic feature is used to characterize the semantics when the singer sings the corresponding syllable at the corresponding pitch;
[0310] A determination sub-module, configured to determine the phoneme duration information based on the syllable duration information and the first semantic feature of the multiple syllables.
[0311] In a possible implementation manner, based on Figure 12 the composition of the device, the determination sub-module includes:
[0312] A first acquisition unit, configured to acquire initial features of the multiple audio frames based on the syllable duration information and first semantic features of the multiple syllables, where the initial features are first semantic features of the syllables to which the corresponding audio frames belong;
[0313] A first determination unit, configured to determine multiple phonemes respectively corresponding to the multiple audio frames based on the initial features of the multiple audio frames and the position features of the multiple audio frames;
[0314] A second determination unit, configured to determine the phoneme duration information based on the correspondence between the multiple audio frames and the multiple phonemes.
[0315] In a possible implementation manner, based on Figure 12 the composition of the device, the first determination unit includes:
[0316] A splicing subunit, configured to splice the initial features of the multiple audio frames and the position features of the multiple audio frames respectively to obtain multiple first splicing features;
[0317] A convolution subunit, configured to perform convolution processing on the multiple first splicing features to obtain multiple first target features;
[0318] A weighting subunit, configured to perform weighting processing on the multiple first target features to obtain multiple second target features;
[0319] A prediction subunit, configured to predict multiple phonemes respectively corresponding to the multiple audio frames based on the multiple second target features.
[0320] In a possible implementation manner, the prediction subunit is configured to:
[0321] Perform a fully connected process on any one of the multiple second target features to obtain multiple prediction probabilities of the audio frame corresponding to the second target feature, where the prediction probability is used to characterize the possibility that the audio frame corresponds to a phoneme;
[0322] Determine the phoneme with the highest prediction probability as the phoneme corresponding to the audio frame.
[0323] In a possible implementation manner, the first acquisition unit is configured to:
[0324] For any one of the multiple syllables, copy the first semantic feature of the syllable for a first target number of times to obtain the initial features of the multiple audio frames included in the syllable, where the first target number is a value obtained by subtracting 1 from the number of audio frames occupied by the syllable.
[0325] In a possible implementation manner, based on Figure 12 the composition of the device, the acquisition module 1203 includes:
[0326] A third acquisition sub-module, configured to acquire a second semantic feature of the plurality of audio frames based on the lyrics, the musical score, and the phoneme duration information, where the second semantic feature is used to characterize the semantics when the singer sings a corresponding phoneme at a corresponding pitch;
[0327] An encoding sub-module, configured to encode the second semantic feature of the plurality of audio frames to obtain an intermediate feature of the plurality of audio frames;
[0328] A decoding sub-module, configured to decode the intermediate feature of the plurality of audio frames to obtain a third semantic feature of the plurality of audio frames;
[0329] A processing sub-module, configured to process the third semantic feature of the plurality of audio frames to obtain an acoustic feature parameter of the plurality of audio frames.
[0330] In a possible implementation manner, based on Figure 12 the composition of the device, the third acquisition sub-module includes:
[0331] A second acquisition unit, configured to acquire a phoneme feature of the plurality of phonemes in the lyrics and a pitch feature of a plurality of pitches corresponding to a plurality of musical notes in the musical score, where the plurality of musical notes correspond to the plurality of syllables;
[0332] A third acquisition unit, configured to acquire a frame-level phoneme feature of the plurality of audio frames based on the phoneme feature of the plurality of phonemes and the phoneme duration information, where the frame-level phoneme feature is the phoneme feature of the phoneme to which the corresponding audio frame belongs;
[0333] A fourth acquisition unit, configured to acquire a frame-level pitch feature of the plurality of audio frames based on the pitch feature of the plurality of pitches and the phoneme duration information, where the frame-level pitch feature is the pitch feature of the pitch corresponding to the phoneme to which the corresponding audio frame belongs;
[0334] A splicing unit, configured to splice the frame-level phoneme features of the plurality of audio frames with the frame-level pitch features of the plurality of audio frames and the singer feature of the singer identifier respectively to obtain a second semantic feature of the plurality of audio frames.
[0335] In a possible implementation manner, the third acquisition unit is configured to:
[0336] For any phoneme in the plurality of phonemes, replicate the phoneme feature of the phoneme a second target number of times to obtain the frame-level phoneme features of the plurality of audio frames included in the phoneme, where the second target number is a value obtained by subtracting one from the number of audio frames occupied by the phoneme.
[0337] In a possible implementation manner, the fourth acquisition unit is configured to:
[0338] For any one of the multiple pitches, the pitch characteristics of the pitch are copied a third target number of times to obtain the frame-level pitch characteristics of multiple audio frames included in the phoneme corresponding to the pitch, where the third target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme corresponding to the pitch.
[0339] In a possible implementation manner, based on Figure 12 the composition of the device, the device further includes:
[0340] A linear processing module, configured to perform linear processing on the third semantic features of the multiple audio frames to obtain multiple third target features;
[0341] The linear processing module is further configured to perform linear processing on the frame-level pitch characteristics of the multiple audio frames to obtain multiple fourth target features;
[0342] A splicing module, configured to splice the multiple third target features and the multiple fourth target features respectively to obtain multiple second spliced features;
[0343] A convolution module, configured to perform convolution processing on the multiple second spliced features to obtain at least one of the silence parameters or fundamental frequency parameters of each of the multiple audio frames, where the silence parameter is used to characterize whether the corresponding audio frame is a silent segment, and the fundamental frequency parameter is used to characterize the logarithm of the fundamental frequency of the corresponding audio frame.
[0344] In a possible implementation manner, the generation module 1204 is configured to:
[0345] Splice the acoustic feature parameters of the multiple audio frames and at least one of the silence parameters or fundamental frequency parameters of each of the multiple audio frames to obtain the target acoustic features of the multiple audio frames;
[0346] Input the target acoustic features of the multiple audio frames into a vocoder, and synthesize the audio signal of the to-be-generated singing voice through the vocoder.
[0347] All the above optional technical solutions can be combined arbitrarily to form alternative embodiments of the present disclosure, which will not be elaborated here one by one.
[0348] It should be noted that: when the singing voice generation device provided in the above embodiment generates a singing voice, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the singing voice generation device provided in the above embodiment and the singing voice generation method embodiment belong to the same concept, and the specific implementation process can be seen in the singing voice generation method embodiment, which will not be elaborated here.
[0349] Figure 13It is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 13 shown, taking the computer device as the terminal 1300 as an example for illustration, at this time, the singing synthesis method can be completed locally by the terminal 1300, that is, the terminal 1300 downloads the duration model, pitch prediction model, acoustic parameter prediction model, and vocoder to the local, so that the singing synthesis operation for any musical score can be completed without interacting with the server. Optionally, the device type of the terminal 1300 includes: smart phones, tablet computers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers or desktop computers. The terminal 1300 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0350] Generally, the terminal 1300 includes a processor 1301 and a memory 1302.
[0351] Optionally, the processor 1301 includes one or more processing cores, such as a 4-core processor, an 8-core processor, etc. Optionally, the processor 1301 is implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). In some embodiments, the processor 1301 includes a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 integrates a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 further includes an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0352] In some embodiments, the memory 1302 includes one or more computer-readable storage media, optionally, the computer-readable storage media is non-transitory. Optionally, the memory 1302 further includes high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 is used to store at least one program code, and the at least one program code is used to be executed by the processor 1301 to implement the singing generation method provided in each embodiment of the present application.
[0353] In some embodiments, the terminal 1300 may further optionally include: a peripheral device interface 1303 and at least one peripheral device. The processor 1301, the memory 1302, and the peripheral device interface 1303 can be connected through a bus or signal lines. Each peripheral device can be connected to the peripheral device interface 1303 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1309.
[0354] The peripheral device interface 1303 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 are implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0355] The radio frequency circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1304 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. Optionally, the radio frequency circuit 1304 communicates with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1304 further includes a circuit related to NFC (Near Field Communication), which is not limited in this application.
[0356] The display screen 1305 is used to display the UI (User Interface). Optionally, the UI includes graphics, text, icons, videos, and any combination thereof. When the display screen 1305 is a touch display screen, the display screen 1305 also has the ability to collect touch signals on or above the surface of the display screen 1305. The touch signal can be input to the processor 1301 as a control signal for processing. Optionally, the display screen 1305 is also used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there is one display screen 1305, which is arranged on the front panel of the terminal 1300; in other embodiments, there are at least two display screens 1305, which are respectively arranged on different surfaces of the terminal 1300 or in a folding design; in still other embodiments, the display screen 1305 is a flexible display screen, which is arranged on the curved surface or folding surface of the terminal 1300. Even more optionally, the display screen 1305 is set to an irregular non-rectangular shape, that is, a special-shaped screen. Optionally, the display screen 1305 is prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0357] The camera assembly 1306 is used to collect images or videos. Optionally, the camera assembly 1306 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1306 further includes a flash. Optionally, the flash is a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to a combination of a warm light flash and a cold light flash for light compensation under different color temperatures.
[0358] In some embodiments, the audio circuit 1307 includes a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 1301 for processing, or input them to the radio frequency circuit 1304 to achieve voice communication. For the purpose of stereo collection or noise reduction, there are multiple microphones, which are respectively disposed at different parts of the terminal 1300. Optionally, the microphone is an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signals from the processor 1301 or the radio frequency circuit 1304 into sound waves. Optionally, the speaker is a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1307 further includes a headphone jack.
[0359] The power supply 1309 is used to supply power to each component in the terminal 1300. Optionally, the power supply 1309 is an alternating current, a direct current, a disposable battery or a rechargeable battery. When the power supply 1309 includes a rechargeable battery, the rechargeable battery supports wired charging or wireless charging. The rechargeable battery is also used to support fast charging technology.
[0360] In some embodiments, the terminal 1300 further includes one or more sensors 1310. The one or more sensors 1310 include but are not limited to: an acceleration sensor 1311, a gyroscope sensor 1312, a pressure sensor 1313, an optical sensor 1315, and a proximity sensor 1316.
[0361] In some embodiments, the acceleration sensor 1311 detects the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal 1300. For example, the acceleration sensor 1311 is used to detect the components of the gravitational acceleration on the three coordinate axes. Optionally, the processor 1301 controls the display screen 1305 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1311. The acceleration sensor 1311 is also used to collect game or user movement data.
[0362] In some embodiments, the gyroscope sensor 1312 detects the body direction and rotation angle of the terminal 1300, and the gyroscope sensor 1312 cooperates with the acceleration sensor 1311 to collect the 3D actions of the user on the terminal 1300. The processor 1301 implements the following functions according to the data collected by the gyroscope sensor 1312: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0363] Optionally, the pressure sensor 1313 is disposed on the side frame of the terminal 1300 and / or the lower layer of the display screen 1305. When the pressure sensor 1313 is disposed on the side frame of the terminal 1300, it can detect the holding signal of the user on the terminal 1300, and the processor 1301 performs left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1313. When the pressure sensor 1313 is disposed on the lower layer of the display screen 1305, the processor 1301 controls the operable controls on the UI interface according to the pressure operation of the user on the display screen 1305. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0364] The optical sensor 1315 is used to collect the ambient light intensity. In one embodiment, the processor 1301 controls the display brightness of the display screen 1305 according to the ambient light intensity collected by the optical sensor 1315. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1305 is increased; when the ambient light intensity is low, the display brightness of the display screen 1305 is decreased. In another embodiment, the processor 1301 also dynamically adjusts the shooting parameters of the camera module 1306 according to the ambient light intensity collected by the optical sensor 1315.
[0365] The proximity sensor 1316, also known as a distance sensor, is typically disposed on the front panel of the terminal 1300. The proximity sensor 1316 is used to collect the distance between the user and the front of the terminal 1300. In one embodiment, when the proximity sensor 1316 detects that the distance between the user and the front of the terminal 1300 is gradually decreasing, the processor 1301 controls the display screen 1305 to switch from the lit state to the off state; when the proximity sensor 1316 detects that the distance between the user and the front of the terminal 1300 is gradually increasing, the processor 1301 controls the display screen 1305 to switch from the off state to the lit state.
[0366] Those skilled in the art can understand that Figure 13 the structure shown in does not constitute a limitation on the terminal 1300, and it can include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0367] Figure 14 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 1400 may vary greatly due to different configurations or performances. The computer device 1400 includes one or more processors (Central Processing Units, CPUs) 1401 and one or more memories 1402. Among them, at least one computer program is stored in the memory 1402, and the at least one computer program is loaded and executed by the one or more processors 1401 to implement the singing voice generation methods provided by the above various embodiments. Optionally, the computer device 1400 also has components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The computer device 1400 also includes other components for implementing the functions of the device, which will not be elaborated here.
[0368] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including at least one computer program. The at least one computer program can be executed by a processor in the terminal to complete the singing voice generation methods in the above various embodiments. For example, the computer-readable storage medium includes a ROM (Read-Only Memory), a RAM (Random-Access Memory), a CD-ROM (Compact Disc Read-Only Memory), magnetic tapes, floppy disks, and optical data storage devices, etc.
[0369] In an exemplary embodiment, a computer program product or a computer program is further provided, including one or more program codes stored in a computer-readable storage medium. One or more processors of the computer device can read the one or more program codes from the computer-readable storage medium, and the one or more processors execute the one or more program codes, so that the computer device can execute to complete the singing voice generation method in the above embodiment.
[0370] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. Optionally, the program is stored in a computer-readable storage medium. Optionally, the above-mentioned storage medium is a read-only memory, a disk, an optical disc, etc.
[0371] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A singing voice generation method, characterized in that, the method includes: Based on the musical score and the lyrics corresponding to the musical score, determine syllable duration information, where the syllable duration information is used to represent the number of audio frames occupied by each of the multiple syllables in the lyrics; Obtain the multiple syllables in the lyrics and the multiple phonemes included in the multiple syllables; Obtain the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables; Based on the multiple syllables, the multiple phonemes, the multiple pitches, and the singer identifier, obtain the first semantic features of the multiple syllables, where the first semantic features are used to represent the semantics when the singer sings the corresponding syllable at the corresponding pitch; Based on the syllable duration information and the first semantic features of the multiple syllables, obtain the initial features of the multiple audio frames in the to-be-generated singing voice, where the initial features are the first semantic features of the syllables to which the corresponding audio frames belong; Based on the initial features of the multiple audio frames and the position features of the multiple audio frames, determine the multiple phonemes corresponding to the multiple audio frames respectively; Based on the correspondence between the multiple audio frames and the multiple phonemes, determine phoneme duration information, where the phoneme duration information is used to represent the number of audio frames occupied by each of the multiple phonemes; Based on the musical score, the lyrics, and the phoneme duration information, obtain the acoustic feature parameters of the multiple audio frames, where the acoustic feature parameters are used to represent the acoustic features of the corresponding audio frames; Generate the to-be-generated singing voice based on the acoustic feature parameters of the multiple audio frames.
2. The method according to claim 1, characterized in that, the determining the multiple phonemes corresponding to the multiple audio frames respectively based on the initial features of the multiple audio frames and the position features of the multiple audio frames includes: Concatenate the initial features of the multiple audio frames with the position features of the multiple audio frames respectively to obtain multiple first concatenated features; Perform convolutional processing on the multiple first concatenated features to obtain multiple first target features; Perform weighted processing on the multiple first target features to obtain multiple second target features; Predict the multiple phonemes corresponding to the multiple audio frames respectively based on the multiple second target features.
3. The method according to claim 2, characterized in that, the predicting the multiple phonemes corresponding to the multiple audio frames respectively based on the multiple second target features includes: Perform fully connected processing on any one of the multiple second target features to obtain multiple prediction probabilities of the audio frame corresponding to the second target feature, where the prediction probabilities are used to represent the possibility that the audio frame corresponds to a phoneme; Determine the phoneme with the highest prediction probability as the phoneme corresponding to the audio frame.
4. The method according to claim 1, characterized in that, the obtaining the initial features of the multiple audio frames based on the syllable duration information and the first semantic features of the multiple syllables includes: For any one of the multiple syllables, the first semantic feature of the syllable is copied a first target number of times to obtain the initial features of the multiple audio frames included in the syllable, where the first target number is the value obtained by subtracting one from the number of audio frames occupied by the syllable.
5. The method according to claim 1, wherein, the obtaining of the acoustic feature parameters of the multiple audio frames in the to-be-generated singing voice based on the musical score, the lyrics, and the phoneme duration information includes: obtaining a second semantic feature of the multiple audio frames based on the lyrics, the musical score, and the phoneme duration information, where the second semantic feature is used to represent the semantics when the singer sings the corresponding phoneme at the corresponding pitch; encoding the second semantic feature of the multiple audio frames to obtain intermediate features of the multiple audio frames; decoding the intermediate features of the multiple audio frames to obtain a third semantic feature of the multiple audio frames; processing the third semantic feature of the multiple audio frames to obtain the acoustic feature parameters of the multiple audio frames.
6. The method according to claim 5, wherein, the obtaining of the second semantic feature of the multiple audio frames based on the lyrics, the musical score, and the phoneme duration information includes: obtaining the phoneme features of the multiple phonemes in the lyrics and the pitch features of the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables; obtaining the frame-level phoneme features of the multiple audio frames based on the phoneme features of the multiple phonemes and the phoneme duration information, where the frame-level phoneme features are the phoneme features of the phoneme to which the corresponding audio frame belongs; obtaining the frame-level pitch features of the multiple audio frames based on the pitch features of the multiple pitches and the phoneme duration information, where the frame-level pitch features are the pitch features of the pitch corresponding to the phoneme to which the corresponding audio frame belongs; concatenating the frame-level phoneme features of the multiple audio frames with the frame-level pitch features of the multiple audio frames and the singer features of the singer identifier respectively to obtain the second semantic feature of the multiple audio frames.
7. The method according to claim 6, wherein, the obtaining of the frame-level phoneme features of the multiple audio frames based on the phoneme features of the multiple phonemes and the phoneme duration information includes: For any one of the multiple phonemes, the phoneme feature of the phoneme is copied a second target number of times to obtain the frame-level phoneme features of the multiple audio frames included in the phoneme, where the second target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme.
8. The method according to claim 6, wherein, the obtaining of the frame-level pitch features of the multiple audio frames based on the pitch features of the multiple pitches and the phoneme duration information includes: For any one of the multiple pitches, the pitch feature of the pitch is copied a third target number of times to obtain the frame-level pitch features of the multiple audio frames included in the phoneme corresponding to the pitch, where the third target number is the value obtained by subtracting one from the number of audio frames occupied by the phoneme corresponding to the pitch.
9. The method according to claim 5, wherein, After decoding the intermediate features of the multiple audio frames to obtain the third semantic features of the multiple audio frames, the method further includes: Performing linear processing on the third semantic features of the multiple audio frames to obtain multiple third target features; Performing linear processing on the frame-level pitch features of the multiple audio frames to obtain multiple fourth target features; Respectively concatenating the multiple third target features and the multiple fourth target features to obtain multiple second concatenated features; Performing convolutional processing on the multiple second concatenated features to obtain at least one of the silence parameters or fundamental frequency parameters of each of the multiple audio frames, where the silence parameter is used to characterize whether the corresponding audio frame is a silent segment, and the fundamental frequency parameter is used to characterize the logarithm of the fundamental frequency of the corresponding audio frame.
10. The method according to claim 9, wherein, generating the to-be-generated singing voice based on the acoustic feature parameters of the multiple audio frames includes: Concatenating the acoustic feature parameters of the multiple audio frames and at least one of the silence parameters or fundamental frequency parameters of each of the multiple audio frames to obtain the target acoustic features of the multiple audio frames; Inputting the target acoustic features of the multiple audio frames into a vocoder, and synthesizing the audio signal of the to-be-generated singing voice through the vocoder.
11. A singing voice generating device, wherein, the device includes: A first determination module, configured to determine syllable duration information based on a musical score and the lyrics corresponding to the musical score, where the syllable duration information is used to characterize the number of audio frames occupied by each of the multiple syllables in the lyrics; A second determination module, including: A first acquisition sub-module, configured to acquire the multiple syllables in the lyrics and the multiple phonemes included in the multiple syllables; The first acquisition sub-module is further configured to acquire the multiple pitches corresponding to the multiple notes in the musical score, where the multiple notes correspond to the multiple syllables; A second acquisition sub-module, configured to acquire the first semantic features of the multiple syllables based on the multiple syllables, the multiple phonemes, the multiple pitches, and a singer identifier, where the first semantic features are used to characterize the semantics when the singer sings the corresponding syllable at the corresponding pitch; A determination sub-module, including: A first acquisition unit, configured to acquire the initial features of multiple audio frames in the to-be-generated singing voice based on the syllable duration information and the first semantic features of the multiple syllables, where the initial features are the first semantic features of the syllables to which the corresponding audio frames belong; A first determination unit, configured to determine the multiple phonemes corresponding to each of the multiple audio frames based on the initial features of the multiple audio frames and the position features of the multiple audio frames; A second determination unit, configured to determine phoneme duration information based on the correspondence between the multiple audio frames and the multiple phonemes, where the phoneme duration information is used to characterize the number of audio frames occupied by each of the multiple phonemes; An acquisition module, configured to acquire the acoustic feature parameters of the multiple audio frames based on the musical score, the lyrics, and the phoneme duration information, where the acoustic feature parameters are used to characterize the acoustic features of the corresponding audio frames. A generation module, configured to generate the to-be-generated singing voice based on the acoustic feature parameters of the plurality of audio frames.
12. The apparatus according to claim 11, wherein, the first determination unit includes: A splicing subunit, configured to splice the initial features of the plurality of audio frames and the position features of the plurality of audio frames respectively to obtain a plurality of first spliced features; A convolution subunit, configured to perform convolution processing on the plurality of first spliced features to obtain a plurality of first target features; A weighting subunit, configured to perform weighting processing on the plurality of first target features to obtain a plurality of second target features; A prediction subunit, configured to predict a plurality of phonemes corresponding to the plurality of audio frames respectively based on the plurality of second target features.
13. The apparatus according to claim 12, wherein, the prediction subunit is configured to: Perform a fully connected process on any one of the plurality of second target features to obtain a plurality of prediction probabilities of the audio frame corresponding to the second target feature, and the prediction probability is used to represent the possibility that the audio frame corresponds to a phoneme; Determine the phoneme with the highest prediction probability as the phoneme corresponding to the audio frame.
14. The apparatus according to claim 11, wherein, the first obtaining unit is configured to: For any one of the plurality of syllables, copy the first semantic feature of the syllable for a first target number of times to obtain the initial features of the plurality of audio frames included in the syllable, where the first target number of times is a value obtained by subtracting one from the number of audio frames occupied by the syllable.
15. The apparatus according to claim 11, wherein, the obtaining module includes: A third obtaining sub-module, configured to obtain the second semantic features of the plurality of audio frames based on the lyrics, the musical score, and the phoneme duration information, and the second semantic features are used to represent the semantics when the singer sings the corresponding phoneme at the corresponding pitch; An encoding sub-module, configured to encode the second semantic features of the plurality of audio frames to obtain intermediate features of the plurality of audio frames; A decoding sub-module, configured to decode the intermediate features of the plurality of audio frames to obtain third semantic features of the plurality of audio frames; A processing sub-module, configured to process the third semantic features of the plurality of audio frames to obtain the acoustic feature parameters of the plurality of audio frames.
16. The apparatus according to claim 15, wherein, the third obtaining sub-module includes: A second obtaining unit, configured to obtain the phoneme features of the plurality of phonemes in the lyrics and the pitch features of the plurality of pitches corresponding to the plurality of notes in the musical score, where the plurality of notes correspond to the plurality of syllables; A third obtaining unit, configured to obtain the frame-level phoneme features of the plurality of audio frames based on the phoneme features of the plurality of phonemes and the phoneme duration information, and the frame-level phoneme features are the phoneme features of the phoneme to which the corresponding audio frame belongs; A fourth obtaining unit, configured to obtain the frame-level pitch features of the plurality of audio frames based on the pitch features of the plurality of pitches and the phoneme duration information, and the frame-level pitch features are the pitch features of the pitch corresponding to the phoneme to which the corresponding audio frame belongs; A splicing unit, configured to splice the frame-level phoneme features of the plurality of audio frames with the frame-level pitch features of the plurality of audio frames and the singer features of the singer identification respectively, so as to obtain the second semantic features of the plurality of audio frames.
17. The apparatus according to claim 16, wherein, the third obtaining unit is configured to: For any one of the plurality of phonemes, duplicate the phoneme feature of the phoneme for a second target number of times to obtain the frame-level phoneme features of the plurality of audio frames included in the phoneme, where the second target number of times is a value obtained by subtracting one from the number of audio frames occupied by the phoneme.
18. The apparatus according to claim 16, wherein, the fourth obtaining unit is configured to: For any one of the plurality of pitches, duplicate the pitch feature of the pitch for a third target number of times to obtain the frame-level pitch features of the plurality of audio frames included in the phoneme corresponding to the pitch, where the third target number of times is a value obtained by subtracting one from the number of audio frames occupied by the phoneme corresponding to the pitch.
19. The apparatus according to claim 15, wherein, the apparatus further includes: a linear processing module, configured to perform linear processing on the third semantic features of the plurality of audio frames to obtain a plurality of third target features; the linear processing module is further configured to perform linear processing on the frame-level pitch features of the plurality of audio frames to obtain a plurality of fourth target features; a splicing module, configured to splice the plurality of third target features and the plurality of fourth target features respectively to obtain a plurality of second splicing features; a convolution module, configured to perform convolution processing on the plurality of second splicing features to obtain at least one of the silence parameters or fundamental frequency parameters of each of the plurality of audio frames, where the silence parameter is used to characterize whether the corresponding audio frame is a silent segment, and the fundamental frequency parameter is used to characterize the logarithm of the fundamental frequency of the corresponding audio frame.
20. The apparatus according to claim 19, wherein, the generation module is configured to: Splice the acoustic feature parameters of the plurality of audio frames and at least one of the silence parameters or fundamental frequency parameters of each of the plurality of audio frames to obtain the target acoustic features of the plurality of audio frames; Input the target acoustic features of the plurality of audio frames into a vocoder, and synthesize the audio signal of the to-be-generated singing voice through the vocoder.
21. A computer device, wherein, the computer device includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the singing voice generation method according to any one of claims 1 to 10.
22. A storage medium, wherein, at least one computer program is stored in the storage medium, and the at least one computer program is loaded and executed by a processor to implement the singing voice generation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Song synthesis method and device, readable medium and electronic equipment
CN111583900A
Audio synthesis method and device, electronic equipment and computer readable storage medium
CN112802446A