Singing synthesis methods, computer equipment and computer-readable storage media

By training the singing features to determine the model, and combining the emotional features of the lyrics with the original vocal information of the vocal elements, the pitch and sound intensity are adjusted, which solves the problem of insufficient adaptability between the synthesized singing and the song content, and realizes dynamic adaptation and rich expression of the singing and the emotional content of the lyrics.

CN116825057BActive Publication Date: 2026-04-03TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the synthesized vocals are not well-suited to the content of the songs, resulting in the synthesized vocals failing to accurately express the emotional characteristics of the lyrics.

Method used

The model is determined by training the vocal features. Based on the similarity between the lyrics and multiple preset emotional types, the singing emotional features of the lyrics are determined. Based on the singing emotional features and the original vocal information of the vocal elements, the pitch and sound intensity of the audio frames are adjusted to generate a synthesized vocal sound.

Benefits of technology

It improves the compatibility between synthesized vocals and song content, enabling synthesized vocals to dynamically adjust pitch and volume according to changes in the emotional characteristics of the lyrics, thereby enhancing the appeal and dynamic richness of the vocals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116825057B_ABST
    Figure CN116825057B_ABST
Patent Text Reader

Abstract

This application relates to a singing voice synthesis method, computer equipment, and storage medium. The method includes: acquiring the original vocal information of each vocal element in the lyrics; determining the singing emotional features of the lyrics based on the similarity between the lyrics and multiple preset emotional types using a trained singing voice feature determination model; determining the pitch adjustment result and sound intensity of the audio frames corresponding to each vocal element based on the singing emotional features and the original vocal information of each vocal element; and outputting singing voice spectrum features based on the pitch adjustment result and sound intensity of each audio frame; and generating a synthesized singing voice corresponding to the lyrics based on the singing voice spectrum features output by the singing voice feature determination model. In this solution, by adjusting the pitch and sound intensity in conjunction with the singing emotional features of the lyrics, the singing voice features can exhibit variations in pitch and sound intensity according to changes in the singing emotional features of the lyrics, effectively improving the adaptability of the synthesized singing voice to the song content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vocal synthesis technology, and in particular to a vocal synthesis method, computer device, and computer-readable storage medium. Background Technology

[0002] With the development of computer technology, vocal synthesis technology has gradually become popular. Vocal synthesis refers to simulating the output of singing voices through computers and other equipment.

[0003] In related technologies, the melody of the final synthesized vocals can be determined based on the song's score. That is, after obtaining the content of the synthesized vocals based on the lyrics, the overall melody and performance style of the synthesized vocals are controlled based on the song's score. However, the compatibility between the synthesized vocals obtained through this method and the song's content still needs improvement. Summary of the Invention

[0004] Therefore, it is necessary to provide a vocal synthesis method, computer device, and computer-readable storage medium that can improve the compatibility of synthesized vocals with song content, in order to address the aforementioned technical problems.

[0005] Firstly, this application provides a method for synthesizing singing voices. The method includes:

[0006] Obtain the original pronunciation information of each phoneme in the lyrics;

[0007] The trained vocal feature determination model determines the singing emotion features of the lyrics based on the similarity between the lyrics and multiple preset emotion types. Based on the singing emotion features and the original vocal information of each vocal element, the model determines the pitch adjustment result and sound intensity of the audio frame corresponding to each vocal element, and outputs the vocal spectrum features based on the pitch adjustment result and sound intensity of each audio frame.

[0008] Based on the vocal characteristics, the spectral characteristics of the vocal output by the model are determined, and the synthesized vocals corresponding to the lyrics are generated.

[0009] In one embodiment, determining the singing emotional characteristics of the lyrics based on the similarity between the lyrics and multiple preset emotional types includes:

[0010] The lyrics are encoded to obtain the lyric text encoding;

[0011] Obtain the corresponding emotion type codes for each of the multiple emotion types, and determine the weight of the emotion type codes for each emotion type based on the similarity between the lyrics text codes and the emotion type codes of each emotion type;

[0012] Based on the emotion type codes of each emotion type and the weights of the emotion type codes of each emotion type, the target code of the lyrics is determined, and the target code is determined as the singing emotion feature of the lyrics.

[0013] In one embodiment, the original vocal information of each vocal element includes the original vocal information of the audio frame corresponding to each vocal element, and the singing feature determination model includes a fundamental frequency prediction module.

[0014] Based on the singing emotional characteristics and the original vocal information of each vocal element, the pitch adjustment result of the audio frame corresponding to each vocal element is determined, including:

[0015] The singing emotion features and the original vocal information of the multiple audio frames are input into the fundamental frequency prediction module. The fundamental frequency prediction module determines the fundamental frequency change features of the multiple audio frames based on the singing emotion features, and adjusts the pitch in the original vocal information of the multiple audio frames according to the fundamental frequency change features, so as to obtain the pitch adjustment result of the audio frame corresponding to each vocal element.

[0016] In one embodiment, the original vocal information of each vocal element includes the original vocal information of the audio frame corresponding to each vocal element, and the singing feature determination model includes a sound energy prediction module.

[0017] Based on the singing emotional characteristics and the original vocal information of each vocal element, the sound intensity of the audio frame corresponding to each vocal element is determined, including:

[0018] The singing emotion features and the original vocal information of the multiple audio frames are input into the sound energy prediction module. Based on the singing emotion features and the original vocal information of the multiple audio frames, the sound energy prediction module determines the target sound energy of the multiple audio frames as the sound intensity of the multiple audio frames.

[0019] In one embodiment, obtaining the original vocal information of each vocal molecule of the lyrics includes:

[0020] Based on the pre-acquired lyrics and the corresponding musical score, determine the pitch and singing technique information of each phoneme of the lyrics in the musical score;

[0021] The pitch and singing technique information of each vocal element are encoded, and the original vocal information of each vocal element is obtained based on the encoding results.

[0022] In one embodiment, encoding the pitch and singing technique information of each vocal element, and obtaining the original vocal information of each vocal element based on the encoding result, includes:

[0023] The pitch and singing technique information of each phoneme are encoded to obtain the phoneme level code of each phoneme;

[0024] Based on the number of audio frames for each voice element, the phoneme-level encoding of the voice element is copied to obtain multiple frame-level encodings for the voice element.

[0025] The multiple frame-level encodings of each of the aforementioned sound elements are used as the original sound information of each sound element.

[0026] In one embodiment, the singing feature determination model includes a decoder;

[0027] The output of the vocal spectral characteristics based on the pitch adjustment results and sound intensity of each audio frame includes:

[0028] The pitch adjustment results and sound intensity of each audio frame are input into the decoder to obtain the Mel spectrum output by the decoder, and the Mel spectrum is used as the spectral feature of the singing voice.

[0029] In one embodiment, the singing feature determination model is trained through the following steps:

[0030] The sample dry audio spectral features and the original vocal information of each vocal element of the sample lyrics are obtained; the sample dry audio is obtained by the user singing according to the sample lyrics and the musical score associated with the sample lyrics.

[0031] The model to be trained determines the sample singing emotion features of the sample lyrics. Based on the sample singing emotion features and the sample original vocal information of each vocal element of the sample lyrics, the sample pitch adjustment result and sample sound intensity of the audio frame corresponding to each vocal element of the sample lyrics are determined. Based on the sample pitch adjustment result and sample sound intensity, the sample singing spectrum features are determined.

[0032] The model loss value is determined based on the difference between the sample singing spectral features and the dry sound spectral features, and the difference between the emotion type corresponding to the sample singing emotion features and the preset emotion label of the sample singing.

[0033] The model parameters of the singing features to be trained are adjusted according to the model loss value until the training termination condition is met, thus obtaining a trained singing feature determination model.

[0034] Secondly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0035] Obtain the original pronunciation information of each phoneme in the lyrics;

[0036] The trained vocal feature determination model determines the singing emotion features of the lyrics based on the similarity between the lyrics and multiple preset emotion types. Based on the singing emotion features and the original vocal information of each vocal element, the model determines the pitch adjustment result and sound intensity of the audio frame corresponding to each vocal element, and outputs the vocal spectrum features based on the pitch adjustment result and sound intensity of each audio frame.

[0037] Based on the vocal characteristics, the spectral characteristics of the vocal output by the model are determined, and the synthesized vocals corresponding to the lyrics are generated.

[0038] Thirdly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0039] Obtain the original pronunciation information of each phoneme in the lyrics;

[0040] The trained vocal feature determination model determines the singing emotion features of the lyrics based on the similarity between the lyrics and multiple preset emotion types. Based on the singing emotion features and the original vocal information of each vocal element, the model determines the pitch adjustment result and sound intensity of the audio frame corresponding to each vocal element, and outputs the vocal spectrum features based on the pitch adjustment result and sound intensity of each audio frame.

[0041] Based on the vocal characteristics, the spectral characteristics of the vocal output by the model are determined, and the synthesized vocals corresponding to the lyrics are generated.

[0042] Fourthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0043] Obtain the original pronunciation information of each phoneme in the lyrics;

[0044] The trained vocal feature determination model determines the singing emotion features of the lyrics based on the similarity between the lyrics and multiple preset emotion types. Based on the singing emotion features and the original vocal information of each vocal element, the model determines the pitch adjustment result and sound intensity of the audio frame corresponding to each vocal element, and outputs the vocal spectrum features based on the pitch adjustment result and sound intensity of each audio frame.

[0045] Based on the vocal characteristics, the spectral characteristics of the vocal output by the model are determined, and the synthesized vocals corresponding to the lyrics are generated.

[0046] The aforementioned singing voice synthesis method, apparatus, computer equipment, and storage medium can acquire the original vocal information of each phoneme in the lyrics. Then, a trained singing voice feature determination model determines the singing emotional features of the lyrics based on the similarity between the lyrics and multiple preset emotional types. Based on the singing emotional features and the original vocal information of each phoneme, the model determines the pitch adjustment result and sound intensity of the corresponding audio frame for each phoneme, and outputs the singing voice spectrum features based on the pitch adjustment result and sound intensity of each audio frame. Furthermore, the model can determine the singing voice features output by the singing voice feature determination model to generate the synthesized singing voice corresponding to the lyrics. In this scheme, by adjusting the pitch and sound intensity in conjunction with the singing emotional features of the lyrics, the singing voice features can exhibit variations in pitch and sound intensity according to changes in the singing emotional features of the lyrics, effectively improving the adaptability of the synthesized singing voice to the song content. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating a singing voice synthesis method in one embodiment;

[0048] Figure 2 This is a flowchart illustrating one step in determining the emotional characteristics of lyrics during singing, as shown in one embodiment.

[0049] Figure 3 This is a schematic diagram of the structure of a text emotion encoder in one embodiment;

[0050] Figure 4 This is a schematic diagram of the structure of a fundamental frequency prediction module in one embodiment;

[0051] Figure 5 This is a flowchart illustrating one step in determining original vocal information in one embodiment;

[0052] Figure 6 This is a flowchart illustrating the steps of training a singing voice feature determination model in one embodiment.

[0053] Figure 7 A schematic diagram of the structure of a singing voice feature determination model in one embodiment;

[0054] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0056] In one embodiment, such as Figure 1 As shown, a method for synthesizing singing voices is provided. This embodiment illustrates the application of this method to a server. In this embodiment, the server can obtain the original vocal information of each vocal molecule in the lyrics, and then determine the singing emotion features of the lyrics using a trained singing voice feature determination model. Based on the singing emotion features, the fundamental frequency and sound intensity of multiple audio frames in the original vocal information are adjusted, and the singing voice features are determined based on the adjusted vocal information. Furthermore, the singing voice features output by the model can be determined based on the singing voice features to generate the synthesized singing voice corresponding to the lyrics.

[0057] It is understood that this method can also be applied to terminals, and to systems including terminals and servers, and implemented through interaction between the terminals and servers. For example, a trained vocal feature determination model can be deployed in a vocal synthesis application, and the vocal synthesis application can be installed on the terminal. Upon receiving a vocal synthesis request, the terminal can determine the lyrics and musical score for which the request is directed, and instruct the vocal synthesis application with the deployed vocal feature determination model to perform vocal synthesis based on the vocal synthesis method provided in this application, generating a synthesized vocal version corresponding to the lyrics.

[0058] The server can be a standalone server or a server cluster consisting of multiple servers; the terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices.

[0059] In this embodiment, the method may include the following steps:

[0060] S101, obtain the original vocal information of each vocal element in the lyrics.

[0061] Lyrics can be composed of multiple phonemes, each of which can consist of a phoneme and the manner in which that phoneme is pronounced during the singing and vocalization process. A phoneme, as referred to in the context of phonemes, is the smallest unit of speech defined based on the natural attributes of speech. Based on the articulation actions within a syllable, one articulation action can correspond to one phoneme. The manner in which a phoneme is pronounced can be determined through vocalization information, which can be unadjusted vocalization information.

[0062] In practical applications, the original vocal information of each vocal element in the lyrics can be obtained. This original vocal information can be obtained solely based on the melody information and / or singing technique information recorded in the score. In other words, the original vocal information can also be understood as: the vocalization method of each vocal element in the lyrics when singing according to the score.

[0063] S102: The trained vocal feature determination model determines the singing emotion features of the lyrics based on the similarity between the lyrics and multiple preset emotion types. Based on the singing emotion features and the original vocal information of each vocal element, it determines the pitch adjustment result and sound intensity of the audio frame corresponding to each vocal element, and outputs the vocal spectrum features based on the pitch adjustment result and sound intensity of each audio frame.

[0064] As an example, the emotional characteristics of singing can characterize the emotional features of lyrics when they are sung. For example, some sad love songs often have rather melancholic lyrics, so such songs can have the emotional characteristics of "sorrow" when sung; while some sweet love songs or festive songs usually have rather joyful and celebratory lyrics, so such songs can have the emotional characteristics of "joy" or "happiness" when sung; for revolutionary songs, the emotional characteristics of singing can be "neutral" or "positive and uplifting".

[0065] The sound intensity of an audio frame can be the loudness corresponding to the sound signal in the audio frame. In some embodiments, the sound intensity can be adjusted by adjusting the energy of the sound.

[0066] Specifically, different lyrics may express different emotions. For example, in lyric adaptations (such as adaptations of Mandarin / Cantonese lyrics, or Chinese and foreign language lyrics), songs like "A Thousand Songs" and "Song of the Setting Sun," or "The First Dream" and "Riding on the Back of a Silver Dragon," may share the same melody, but the emotions expressed in the lyrics differ. It's understandable that in song synthesis, if the vocal synthesis method is determined solely based on the melody while ignoring the lyrics, the resulting synthesized vocals may express emotions different from the actual emotions expressed in the lyrics, resulting in a mismatch between the vocal style and the emotional content of the lyrics.

[0067] To address this, this application pre-trains a vocal feature determination model and uses this model to analyze the textual content of the lyrics, thereby obtaining the emotional features of the lyrics. Specifically, after obtaining the lyrics, the lyrics can be input into the vocal feature determination model, which then compares and matches the lyrics with multiple emotional types to determine the similarity between the lyrics and these types. Based on this similarity, the emotional features of the singing are then determined.

[0068] In some optional embodiments, the singing emotion characteristics can be determined by all the lyrics of a song. That is, for a song, a singing emotion characteristic can be obtained by combining all the lyrics of the song. In this way, the singing feature model determines the singing emotion characteristics relatively quickly. Alternatively, in other embodiments, all the lyrics of a song can be divided, such as by sentence, to obtain multiple lyric segments. The singing feature determination model can determine the singing emotion characteristics of each lyric segment separately, so that the fundamental frequency and sound intensity of multiple audio frames in each lyric segment can be adjusted in a fine-grained manner based on the singing emotion characteristics of each lyric segment, thereby improving the adjustment accuracy.

[0069] In practical applications, singers can adjust the emotions they wish to express by controlling the pitch and intensity of their singing (such as the range and speed of pitch and intensity changes). For example, when singing a song with the emotional characteristics of "joy," the pitch and intensity of the former can be higher than those of the latter to express a cheerful and exciting feeling. Furthermore, in some embodiments, the original vocal information of each vocal element in the lyrics can include pitch, meaning each vocal element has a corresponding pitch when sung. The original vocal information can be divided into the pitches of multiple audio frames. In other embodiments, the original vocal information can also include intensity, meaning the vocal elements can be sung according to the intensity in the original vocal information. This intensity can be an empirical value, or it can be uniformly set to 0, meaning the intensity of each vocal element is temporarily uncertain and further configured by a vocal feature determination model later.

[0070] Based on this, the singing feature determination model in this application, after obtaining the singing emotion features, can determine the pitch adjustment result and sound intensity of the audio frame corresponding to each voice element based on the singing emotion features and the original vocal information of each voice element, and determine the singing spectral features of the synthesized song based on the pitch adjustment result and sound intensity of each audio frame, and then output the model prediction result.

[0071] Therefore, based on the original vocal information obtained from the musical score, the pitch and volume can be adjusted in combination with the singing emotional characteristics of the lyrics, so that the final vocal characteristics can not only match the melody of the song, but also show changes in pitch and volume intensity as the singing emotional characteristics change.

[0072] S103, determine the spectral characteristics of the singing output by the model based on the singing characteristics, and generate the synthesized singing corresponding to the lyrics.

[0073] After obtaining the singing features output by the singing feature determination model, a synthesized singing voice corresponding to the lyrics and its musical score can be generated based on the singing features. This allows for a singing voice that dynamically changes as the input lyrics are modified. For example, an audio waveform can be obtained based on the singing features, and the synthesized singing voice can be determined based on the audio waveform.

[0074] In the aforementioned vocal synthesis method, the original vocal information of each phoneme in the lyrics can be obtained. Then, a trained vocal feature determination model determines the singing emotional features of the lyrics based on the similarity between the lyrics and multiple preset emotional types. Based on the singing emotional features and the original vocal information of each phoneme, the pitch adjustment result and sound intensity of the corresponding audio frame for each phoneme are determined. Based on the pitch adjustment result and sound intensity of each audio frame, the vocal spectrum features are output. Furthermore, the vocal features output by the model can be used to generate the synthesized vocal sound corresponding to the lyrics. In this scheme, by adjusting the pitch and sound intensity in conjunction with the singing emotional features of the lyrics, the vocal features can exhibit variations in pitch and sound intensity according to changes in the singing emotional features of the lyrics, effectively improving the adaptability of the synthesized vocal sound to the song content.

[0075] In each vocal synthesis process, this scheme obtains corresponding vocal features based on the emotional characteristics of the lyrics, and then obtains the synthesized vocal features based on these vocal features. This can increase the dynamic richness of the synthesized vocal as it changes with the lyrics, as well as the appeal of the synthesized vocal.

[0076] In one embodiment, such as Figure 2 As shown, determining the singing emotional characteristics of lyrics based on the similarity between lyrics and multiple preset emotional types can include the following steps:

[0077] S201, Encode the lyrics to obtain the lyric text encoding.

[0078] In practical applications, the lyrics text can be obtained. For example, the lyrics text may include lyrics text in one or more languages, such as Chinese, English, Japanese, etc. Then, the lyrics text can be input into a trained vocal feature determination model, which encodes the lyrics text, and the encoding result is determined as the lyrics text encoding.

[0079] S202, obtain the emotion type codes corresponding to each of the multiple emotion types, and determine the weight of the emotion type codes of each emotion type based on the similarity between the lyrics text code and the emotion type codes of each emotion type.

[0080] In this embodiment, an attention mechanism can be incorporated for encoding adjustments. Specifically, the singing feature determination model may include a text sentiment encoder, whose input is lyrics and output is singing emotion features.

[0081] like Figure 3 As shown, the text emotion encoder can consist of two parts: a text encoder and an attention module. The text encoder encodes the lyrics text to obtain the lyric text encoding, while the attention module determines the singing emotion features of the lyrics based on the attention mechanism and the lyric text encoding.

[0082] Specifically, for different emotion types, a corresponding code can be pre-set for each emotion type, serving as the emotion type code for each emotion type. After obtaining the lyrics text code, the lyrics text code can be compared with the codes for each emotion type to obtain the similarity between the lyrics text code and the corresponding emotion type codes for each emotion type. Then, based on each similarity, the weight of the emotion type code for each emotion type can be obtained.

[0083] S203. Based on the emotion type codes of each emotion type and the weights of the emotion type codes of each emotion type, determine the target code of the lyrics and define the target code as the singing emotion feature of the lyrics.

[0084] After obtaining the weights of the emotion type codes for each emotion type, the final target code can be determined by weighted calculation, combining the emotion type codes of each emotion type with their weights. Specifically, a weighted summation can be performed based on the emotion type codes of each emotion type and their weights, and the resulting weighted summation code can be used as the target code for the lyrics, thus defining the singing emotional characteristics of the lyrics.

[0085] In one example, lyrics text encoding, emotion type encoding, and singing emotion features can be represented using embedding. Embedding is a way to transform discrete variables into continuous vector representations. In neural networks, embedding can not only reduce the spatial dimensionality of discrete variables, but also meaningfully represent those variables.

[0086] To facilitate understanding of the above steps, an example is provided below to illustrate the embodiments of this application. However, it should be understood that the embodiments of this application are not limited thereto.

[0087] like Figure 3As shown, taking the pre-acquired emotion types, which include four emotion types, as an example, for the four emotion types "neutral", "joy", "anger" and "sorrow", the corresponding codes can be determined as emotion type codes. The dimension of each emotion type code is 256. When processing based on the attention mechanism, the emotion type codes of the four input emotion types can be transformed into matrix transformations to obtain the corresponding key (K) matrix and value (V) matrix. Both the key matrix and the value matrix are 4*256 matrices.

[0088] Then, the lyrics can be input into the text encoder to obtain a 1*256 dimensional lyrics text encoding. Based on this lyrics text encoding, the query (Q) matrix is ​​obtained, and the target encoding is determined by combining the following attention formula. This target encoding can also be called global emotion embedding.

[0089]

[0090] Where Attention is the attention computation, Attention(Q,K,V) is the final target encoding, and d k Let T be the dimension number, and T denote the matrix transpose. This can be understood as the weight of the emotion type code corresponding to each emotion type. This weight is a normalized weight. For example, the similarity of the emotion type codes of the four emotion types "neutral", "joy", "anger" and "sorrow" calculated in the above way is 0, 0.5, 0.2 and 0.3 respectively.

[0091] In the examples above, the weights of the emotion type encodings corresponding to each emotion type are predicted based on the text emotion encoder. This method can effectively ensure the stability of the final singing emotion features. In other examples, the weights can also be determined manually, thereby improving the controllability and flexibility of the text emotion encoder.

[0092] In this embodiment, by determining the target code of the lyrics based on the emotion type codes of each emotion type and the weights of the emotion type codes of each emotion type, and determining the target code as the singing emotion feature of the lyrics, the singing emotion feature of the lyrics can be determined quickly and accurately.

[0093] In one embodiment, the singing feature determination model includes a fundamental frequency prediction module, and the original vocal information of each vocal element includes the original vocal information of the audio frames corresponding to each vocal element; in step S102, adjusting the fundamental frequency of multiple audio frames in the original vocal information based on singing emotion features may include the following steps:

[0094] The singing emotion characteristics and the original vocal information of multiple audio frames are input into the fundamental frequency prediction module. The fundamental frequency prediction module determines the fundamental frequency change characteristics of multiple audio frames based on the singing emotion characteristics, and adjusts the pitch in the original vocal information of multiple audio frames according to the fundamental frequency change characteristics, so as to obtain the pitch adjustment result of the audio frame corresponding to each vocal element.

[0095] In practical applications, the original vocal information of each voice element can be expanded at the frame level to obtain the original vocal information of the audio frame corresponding to each voice element, thereby acquiring the original vocal information of multiple audio frames. After obtaining the singing emotion features, the singing emotion features and the original vocal information of multiple audio frames can be used as input to the fundamental frequency prediction module. In an optional embodiment, such as... Figure 4 As shown, the fundamental frequency prediction module can be structured as three convolutional layers (convolutional layer 1, convolutional layer 2, and convolutional layer 3) and one linear layer. Each convolutional layer can include a one-dimensional convolutional function conv1D and a normalization function (such as the layer norm function). The original sound information of each sound element can be expanded frame by frame to obtain the original sound information of multiple audio frames. The original sound information of each audio frame can include the sound element of that audio frame (i.e., the sound element of that audio frame). Figure 4 Phoneme information in the middle), pitch (i.e. Figure 4 (Note information in the text), tuplet markings (a type of singing technique information).

[0096] Then, the fundamental frequency prediction module can predict a set of frame-level fundamental frequency change feature sequences based on the singing emotional characteristics. The fundamental frequency change sequence includes fundamental frequency change features of multiple audio frames. Each fundamental frequency change feature can be sequentially associated with the pitch of multiple audio frames in the original vocal information. Thus, the fundamental frequency change features of multiple audio frames in the original vocal information can be obtained through the fundamental frequency prediction module. In one example, the fundamental frequency change feature can specifically be the fundamental frequency residual ΔF0s.

[0097] Since the pitch of a sound is determined by its fundamental frequency, and there is a corresponding relationship between pitch and fundamental frequency, after obtaining the fundamental frequency variation characteristics of multiple audio frames in the original sound information, the pitch of each audio frame can be adjusted based on these characteristics, and the adjusted pitch (or fundamental frequency) of each audio frame can be obtained based on the adjustment results. In one example, the pitch adjustment method for each audio frame can be as follows:

[0098] F0 = notes + ΔF0s

[0099] Here, notes represent the pitch of each audio frame, which can be expressed by the notes in the audio frame or the frequencies corresponding to the notes, and ΔF0s represents the fundamental frequency variation characteristics.

[0100] In this embodiment, the fundamental frequency prediction module determines the fundamental frequency variation characteristics of multiple audio frames in the original vocal information based on the emotional characteristics of singing. Based on the fundamental frequency variation characteristics, the pitch of multiple audio frames in the original vocal information is adjusted so that the pitch of the singing can change dynamically with the changes in the lyrics. This effectively improves the adaptability of the synthesized song's pitch to the lyrics, while also increasing the naturalness and richness of the pitch changes in the synthesized song, avoiding the pitch changes in the synthesized song being too stiff and rigid.

[0101] In one embodiment, the singing feature determination model may include a sound energy prediction module, and the original vocal information of each vocal element includes the original vocal information of the audio frame corresponding to each vocal element; in step S102, based on the singing emotion features and the original vocal information of each vocal element, determining the sound intensity of the audio frame corresponding to each vocal element may include the following steps:

[0102] The singing emotional characteristics and the original vocal information of multiple audio frames are input into the sound energy prediction module. Based on the singing emotional characteristics and the original vocal information of multiple audio frames, the sound energy prediction module determines the target sound energy of multiple audio frames as the sound intensity of multiple audio frames.

[0103] Specifically, the singing emotional characteristics and the original vocal information of multiple audio frames can be used as input to the sound energy prediction module. The sound energy prediction module can predict a set of frame-level target sound energy sequences based on the input. The target sound energy sequence includes the target sound energy of multiple audio frames, and each target sound energy can be associated with multiple audio frames in the original vocal information.

[0104] After obtaining the target sound energy of each audio frame in multiple audio frames, the target sound energy of the multiple audio frames can be used as the sound intensity of the multiple audio frames in the original sound information. In one embodiment, if each sound energy predicted by the sound energy prediction module is the energy amplitude of the sound, and the value of the energy amplitude is within the range of [0,1], the energy amplitude can be converted, and the sound intensity can be obtained based on the conversion result.

[0105] In this embodiment, the sound energy prediction module determines the target sound energy of multiple audio frames based on the emotional characteristics of singing, and obtains the sound intensity of multiple audio frames in the original vocal information based on the target sound energy of multiple audio frames. This allows the sound intensity (or loudness) of the singing to change with the lyrics, effectively improving the compatibility between the loudness of the synthesized singing and the lyrics, while also making the loudness changes of the synthesized singing more rich and natural.

[0106] In one embodiment, S101 acquires the original vocal information of each vocal element in the lyrics, such as... Figure 5As shown, it may include the following steps:

[0107] S501, based on the pre-acquired lyrics and the corresponding musical score, determine the pitch and singing technique information of each phoneme of the lyrics in the musical score.

[0108] In one example, singing technique information can characterize the way the tongue moves or the way of breathing during vocalization, and may include at least one of the following singing technique information: legato marking, vibrato marking, crescendo marking, diminuendo marking.

[0109] It is understandable that a song is formed by combining lyrics and sheet music, with each element corresponding to the other. In this step, the sheet music can be obtained, which records the melody used to sing the lyrics. Based on this pre-obtained sheet music, the pitch and singing techniques of each vocal element in the lyrics can be determined within the sheet music.

[0110] S502 encodes the pitch and singing technique information of each vocal element, and obtains the original vocal information of each vocal element based on the encoding results.

[0111] After obtaining the pitch and singing technique information of each vocal element, the pitch and singing technique information of each vocal element can be input into the encoder for encoding, and the original vocal information of each vocal element can be obtained based on the encoding result.

[0112] In this embodiment, by encoding based on the pitch and singing technique information of the vocal elements, the original vocal information obtained by encoding can accurately represent the specific vocalization method of the vocal elements when singing according to the score, providing an accurate and reliable basis for subsequent adjustment of the pitch and intensity of the voice in combination with the emotional characteristics of the singing.

[0113] In one embodiment, S502 encodes the pitch and singing technique information of each vocal element, and obtains the original vocal information of each vocal element based on the encoding result, which may include the following steps:

[0114] The pitch and singing technique information of each phoneme are encoded to obtain the phoneme level code of each phoneme; based on the number of audio frames of each phoneme, the phoneme level code of the phoneme is copied to obtain multiple frame level codes of the phoneme; the multiple frame level codes of each phoneme are used as the original vocal information of each phoneme.

[0115] Specifically, after inputting the pitch and singing technique information of each phoneme into the encoder, the encoder can output a phoneme-level code for each phoneme. This phoneme-level code can carry the pitch and singing technique information of the corresponding phoneme. For example, for the phoneme "ian", the pitch of the phoneme can be determined as "C4" based on the notes and tuplets in the score, along with the corresponding tuplet markings (e.g., 0 represents non-tuplet, 1 represents tuplet). Then, the phoneme "ian", its pitch, and the tuplet markings can be input into the encoder to obtain the phoneme-level code for "ian".

[0116] After obtaining the phoneme-level encoding, the number of audio frames for each phoneme can be further determined. The number of audio frames for each phoneme can be determined based on the duration of the phoneme's sound and the preset frame duration of the audio frames. The duration of the phoneme's sound can be determined based on the duration of the notes in the musical score.

[0117] Then, the phoneme-level encoding can be expanded based on the number of audio frames. Specifically, for each phoneme, the phoneme-level encoding of that phoneme can be copied according to the number of audio frames it has. The multiple copied encodings are used as multiple frame-level encodings for the phoneme, and then the multiple frame-level encodings of each phoneme can be used as the original vocal information of each phoneme. For example, the phoneme "ian" has 3 audio frames. If the tethered note is marked as 1, then the phoneme, its pitch, and the tethered note can be expanded as "ian ian ian", "C4 C4 C4", and "1 1 1", respectively.

[0118] In this embodiment, the phoneme-level encoding of each phoneme is copied based on the number of audio frames of each phoneme to obtain multiple frame-level encodings of the phoneme. These multiple frame-level encodings of each phoneme are used as the original vocal information of each phoneme. This allows for precise and detailed determination of the original vocal information of each phoneme in each audio frame, facilitating subsequent high-precision adjustment of pitch and sound intensity based on the emotional characteristics of the singing, thereby enhancing the richness of detail in the synthesized singing voice.

[0119] In one embodiment, the singing feature determination model includes a decoder; step S103, based on the pitch adjustment results and sound intensity of each audio frame, outputs the singing spectral features, which may include the following steps:

[0120] The pitch adjustment results and sound intensity of each audio frame are input into the decoder to obtain the Mel spectrum output by the decoder, and the Mel spectrum is used as the spectral feature of the singing voice.

[0121] In this step, the pitch adjustment results and sound intensity of each audio frame can be input into the decoder for decoding to obtain the spectral characteristics output by the decoder. Specifically, the pitch of the adjusted multiple audio frames output by the fundamental frequency prediction module and the sound energy of the adjusted multiple audio frames output by the sound energy prediction module can be input into the decoder for encoding to obtain the Mel spectrum (also known as Mel-Frequency Spectrum, MFC) output by the decoder.

[0122] In the Mel spectrum, the frequency bands are uniformly distributed on the Mel scale. In other words, the frequency bands on the Mel spectrum are closer to the nonlinear human auditory system than the commonly seen linear cepstrum representation. Moreover, the number of frequency bands in the Mel spectrum is much smaller than that in the linear spectrum, making it more suitable for acoustic modeling.

[0123] In this embodiment, by inputting the adjusted vocal information into the decoder for decoding, a Mel spectrum that matches the singing emotional characteristics of the lyrics can be obtained, which facilitates subsequent simulation of the changes in human voice under different lyrical emotions based on the Mel spectrum.

[0124] In one embodiment, such as Figure 6 As shown, the singing voice feature determination model can be trained through the following steps:

[0125] S601, obtain the dry sound spectrum features corresponding to the dry sound of the sample, as well as the original vocal information of each vocal element of the sample lyrics.

[0126] The sample dry audio is obtained by the user singing according to the sample lyrics and the associated musical score. In other words, the sample dry audio is the audio of the sample song, which is sung by the user according to the sample lyrics and the associated musical score.

[0127] Specifically, sample dry audio can be obtained. For example, with user authorization, audio uploaded by the user to a karaoke app can be used as sample dry audio. Then, feature extraction can be performed on the sample dry audio to obtain the corresponding dry audio spectral features, such as obtaining the Mel spectrum of the sample dry audio. Furthermore, the original vocal information of each phoneme in the sample lyrics can also be obtained as the original vocal information of the sample.

[0128] S602, the model to be trained determines the sample singing emotion features of the sample lyrics. Based on the sample singing emotion features and the original vocal information of each vocal element of the sample lyrics, the sample pitch adjustment result and sample sound intensity of the audio frame corresponding to each vocal element of the sample lyrics are determined. Based on the sample pitch adjustment result and sample sound intensity, the sample singing spectrum features are determined.

[0129] In this step, a vocal feature determination model to be trained can be obtained. In one example, the vocal feature determination model can be a neural network. Then, the vocal feature determination model to be trained can be used to analyze the text content of the sample lyrics to obtain the singing emotion features of the sample lyrics. The method for determining the singing emotion features of the sample lyrics can be the same as the method for determining the singing emotion features of lyrics mentioned earlier. Specifically, the vocal feature determination model to be trained can be equipped with an emotion type encoding providing module. This module can be used to provide emotion type codes for each emotion type. The vocal feature determination model to be trained can determine the singing emotion features of the sample lyrics by comparing the text encoding of the sample lyrics with the emotion type codes provided by the emotion type encoding providing module. Then, based on the singing emotion features and the original vocal information of each vocal element of the sample lyrics, the sample pitch adjustment result and sample sound intensity of the audio frames corresponding to each vocal element of the sample lyrics can be determined, and the sample vocal spectral features of the sample lyrics can be determined based on the sample pitch adjustment result and sample sound intensity of each audio frame.

[0130] like Figure 7 As shown, the vocal feature determination model to be trained may include an encoder, a text sentiment encoder, a style predictor, and a decoder. The style predictor consists of a fundamental frequency predictor (i.e., a fundamental frequency prediction module) and an energy predictor (voice energy prediction module). In one embodiment, the original vocal information of the sample may include specific vocal elements (i.e., phoneme information), the pitch of the vocal elements (i.e., note information), and truncation markings. The original vocal information of the sample can be input into the encoder for encoding and expanded according to the number of audio frames of the vocal elements to obtain multiple frame-level codes for each vocal element.

[0131] Simultaneously, sample lyrics associated with the musical score can be input into the Global Emotion Encoder to determine the sample singing emotion features of the sample lyrics. Then, the sample singing emotion features and multiple frame-level codes for each vocal element can be used as inputs to the fundamental frequency predictor and the energy predictor. Subsequently, the fundamental frequency predictor can determine the fundamental frequency variation features of multiple audio frames for each vocal element based on the sample singing emotion features, and adjust the pitch in the frame-level codes of multiple audio frames according to the fundamental frequency variation features to obtain the sample pitch adjustment results. Furthermore, the energy predictor can determine the target sound energy of multiple audio frames based on the sample singing emotion features, and obtain the sample sound intensity of multiple audio frames based on the target sound energy of multiple audio frames.

[0132] Furthermore, based on the sample pitch adjustment results and sample sound intensity of multiple audio frames, the Mel spectrum can be determined, and this Mel spectrum can be used as the spectral feature of the sample singing voice.

[0133] S603. Based on the difference between the spectral features of the sample singing voice and the spectral features of the dry voice, and the difference between the emotional type corresponding to the emotional features of the sample singing voice and the preset emotional label of the sample singing voice, determine the model loss value.

[0134] On one hand, after obtaining the features of the sample singing voice, a first loss value can be determined based on the difference between the sample singing voice features and the dry voice features. On the other hand, a pre-defined emotion label (emotion ID) can be obtained from the sample singing voice. Then, the obtained sample singing emotion features are input into a trained emotion classifier to obtain the emotion type output by the emotion classifier. Furthermore, a multi-classification loss value can be calculated by comparing the emotion type corresponding to the sample singing emotion features with the pre-defined emotion label of the sample singing voice to obtain a second loss value. Finally, the first and second loss values ​​can be combined to determine the model loss value.

[0135] In one embodiment, the emotional label of the sample singing can be obtained from the emotional labels of the corresponding songs in the music library. For example, the emotional label of a sad love song is "sorrow", the emotional label of a sweet love song is "joy", and the emotional label of some revolutionary songs can correspond to the emotional label "neutral". The emotional type and classification rules can be personalized according to the needs. Then, during training, the song corresponding to the sample singing (e.g., a sample singing of a segment or a sample singing corresponding to a complete song) can be determined, and the emotional label of that song can be determined as the emotional label of the sample singing.

[0136] S604, adjust the singing features to be trained according to the model loss value to determine the model parameters until the training termination condition is met, and obtain the trained singing features to determine the model.

[0137] After obtaining the model loss value, the model parameters of the singing feature determination model to be trained can be adjusted according to the model loss value and the backpropagation algorithm until the training termination condition is met. The current singing feature determination model can then be used as the trained singing feature determination model.

[0138] In this embodiment, the singing spectral feature determination model can be trained in a supervised manner using pre-acquired dry sound spectral features and emotional labels of sample singing voices. This enables the singing feature determination model to encode and decode more accurately, generate meaningful singing emotional features and singing voice features, establish the relationship between the singing emotional features of lyrics and the fundamental frequency and energy of the song, and achieve dynamic control of singing voice more accurately and quickly.

[0139] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0140] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores song data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a song synthesis method.

[0141] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0142] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0143] Obtain the original pronunciation information of each phoneme in the lyrics;

[0144] The singing emotion features of the lyrics are determined by a trained singing feature determination model. Based on the singing emotion features, the pitch and sound intensity of multiple audio frames in the original vocal information are adjusted, and the singing features are determined based on the adjusted vocal information.

[0145] Based on the vocal features, the vocal features output by the model are determined, and the synthesized vocal features corresponding to the lyrics are generated.

[0146] In one embodiment, the processor also performs the steps described in the other embodiments when executing the computer program.

[0147] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0148] Obtain the original pronunciation information of each phoneme in the lyrics;

[0149] The singing emotion features of the lyrics are determined by a trained singing feature determination model. Based on the singing emotion features, the pitch and sound intensity of multiple audio frames in the original vocal information are adjusted, and the singing features are determined based on the adjusted vocal information.

[0150] Based on the vocal features, the vocal features output by the model are determined, and the synthesized vocal features corresponding to the lyrics are generated.

[0151] In one embodiment, the computer program, when executed by a processor, also implements the steps described in the other embodiments above.

[0152] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0153] Obtain the original pronunciation information of each phoneme in the lyrics;

[0154] The singing emotion features of the lyrics are determined by a trained singing feature determination model. Based on the singing emotion features, the pitch and sound intensity of multiple audio frames in the original vocal information are adjusted, and the singing features are determined based on the adjusted vocal information.

[0155] Based on the vocal features, the vocal features output by the model are determined, and the synthesized vocal features corresponding to the lyrics are generated.

[0156] In one embodiment, the computer program, when executed by a processor, also implements the steps described in the other embodiments above.

[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0158] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0159] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0160] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for synthesizing singing voice, characterized in that, The method includes: Obtain the original pronunciation information of each phoneme in the lyrics; The trained vocal feature determination model determines the singing emotion features of the lyrics based on the similarity between the lyrics and multiple preset emotion types. Based on the singing emotion features and the original vocal information of each vocal element, the model determines the pitch adjustment result and sound intensity of the audio frame corresponding to each vocal element, and outputs the vocal spectrum features based on the pitch adjustment result and sound intensity of each audio frame. Based on the vocal characteristics, the spectral features of the vocal output by the model are determined, and the synthesized vocals corresponding to the lyrics are generated.

2. The method according to claim 1, characterized in that, The step of determining the singing emotional characteristics of the lyrics based on the similarity between the lyrics and multiple preset emotional types includes: The lyrics are encoded to obtain the lyric text encoding; Obtain the corresponding emotion type codes for each of the multiple emotion types, and determine the weight of the emotion type codes for each emotion type based on the similarity between the lyrics text codes and the emotion type codes of each emotion type; Based on the emotion type codes of each emotion type and the weights of the emotion type codes of each emotion type, the target code of the lyrics is determined, and the target code is determined as the singing emotion feature of the lyrics.

3. The method according to claim 1, characterized in that, The original vocal information of each vocal element includes the original vocal information of the audio frame corresponding to each vocal element, and the singing feature determination model includes a fundamental frequency prediction module. Based on the singing emotional characteristics and the original vocal information of each vocal element, the pitch adjustment result of the audio frame corresponding to each vocal element is determined, including: The singing emotion features and the original vocal information of the multiple audio frames are input into the fundamental frequency prediction module. The fundamental frequency prediction module determines the fundamental frequency change features of the multiple audio frames based on the singing emotion features, and adjusts the pitch in the original vocal information of the multiple audio frames according to the fundamental frequency change features, so as to obtain the pitch adjustment result of the audio frame corresponding to each vocal element.

4. The method according to claim 1, characterized in that, The original vocal information of each vocal element includes the original vocal information of the audio frame corresponding to each vocal element, and the singing feature determination model includes a sound energy prediction module. Based on the singing emotional characteristics and the original vocal information of each vocal element, the sound intensity of the audio frame corresponding to each vocal element is determined, including: The singing emotion features and the original vocal information of the multiple audio frames are input into the sound energy prediction module. Based on the singing emotion features and the original vocal information of the multiple audio frames, the sound energy prediction module determines the target sound energy of the multiple audio frames as the sound intensity of the multiple audio frames.

5. The method according to claim 1, characterized in that, The process of obtaining the original vocal information of each vocal molecule in the lyrics includes: Based on the pre-acquired lyrics and the corresponding musical score, determine the pitch and singing technique information of each phoneme of the lyrics in the musical score; The pitch and singing technique information of each vocal element are encoded, and the original vocal information of each vocal element is obtained based on the encoding results.

6. The method according to claim 5, characterized in that, The process of encoding the pitch and singing technique information of each vocal element, and obtaining the original vocal information of each vocal element based on the encoding result, includes: The pitch and singing technique information of each phoneme are encoded to obtain the phoneme level code of each phoneme; Based on the number of audio frames for each voice element, the phoneme-level encoding of the voice element is copied to obtain multiple frame-level encodings for the voice element. The multiple frame-level encodings of each of the aforementioned sound elements are used as the original sound information of each sound element.

7. The method according to claim 1, characterized in that, The vocal feature determination model includes a decoder; The output of the vocal spectral characteristics based on the pitch adjustment results and sound intensity of each audio frame includes: The pitch adjustment results and sound intensity of each audio frame are input into the decoder to obtain the Mel spectrum output by the decoder, and the Mel spectrum is used as the spectral feature of the singing voice.

8. The method according to any one of claims 1 to 7, characterized in that, The singing voice feature determination model is trained through the following steps: The sample dry audio spectral features and the original vocal information of each vocal element of the sample lyrics are obtained; the sample dry audio is obtained by the user singing according to the sample lyrics and the musical score associated with the sample lyrics. The model to be trained determines the sample singing emotion features of the sample lyrics. Based on the sample singing emotion features and the sample original vocal information of each vocal element of the sample lyrics, the sample pitch adjustment result and sample sound intensity of the audio frame corresponding to each vocal element of the sample lyrics are determined. Based on the sample pitch adjustment result and sample sound intensity, the sample singing spectrum features are determined. The model loss value is determined based on the difference between the sample singing spectral features and the dry sound spectral features, and the difference between the emotion type corresponding to the sample singing emotion features and the preset emotion label of the sample singing. The model parameters of the singing features to be trained are adjusted according to the model loss value until the training termination condition is met, thus obtaining a trained singing feature determination model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speech synthesis method and device, readable medium and electronic equipment

    CN112489621A

  • Speech synthesis method and device, computer equipment and storage medium

    CN116168678A