A training method, device, equipment and storage medium for a singing voice synthesis model

By collecting and processing the voice information of the target user, retraining the singing vocal synthesis model to enable it to synthesize songs with the tone of the target user, solving the problem that the existing song synthesis model cannot synthesize personalized vocals, and achieving the effect of personalized singing vocal synthesis.

CN115881086BActive Publication Date: 2025-06-24BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211586680.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-10
Publication Date
2025-06-24
Estimated Expiration
2042-12-10

AI Technical Summary

Technical Problem

The existing song synthesis model cannot synthesize songs with the personalized tone of ordinary users, and has high requirements for training samples and requires songs recorded by professional singers.

Method used

By obtaining the pre-trained singing vocal synthesis model, the voice information of the target user is collected, the fundamental frequency and duration of each syllable in the voice information are extracted, and the singing vocal synthesis model is retrained based on this information so that it can synthesize songs with the tone of the target user.

Benefits of technology

Based on the basic singing vocal synthesis model, a singing vocal synthesis model that can synthesize songs with the target user's tone is achieved using the voice information of the target user tone, which meets the needs of personalized tone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881086B_ABST
    Figure CN115881086B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, device, equipment and storage medium for a singing synthesis model. The method includes: obtaining a pre-trained singing synthesis model; collecting voice information input by a target user; obtaining text information corresponding to the voice information, and extracting the fundamental frequency and duration corresponding to each syllable in the voice information; re-training the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, using the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user. Based on the basic singing synthesis model, the embodiment of the present invention can train a singing synthesis model that can synthesize a song with the timbre of the target user by using the voice information of the target user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly to a training method, device, equipment and storage medium for a singing synthesis model. Background Art

[0002] With the continuous progress of technology, audio processing technology has also been continuously developed. End-to-end speech synthesis systems have gradually matured and been widely used in the field of speech synthesis. However, with the continuous change of user needs, in some electronic consumer products and entertainment applications, users not only hope to synthesize speech but also hope to synthesize songs, which requires training a model with the ability to synthesize songs.

[0003] When training a song synthesis model, musical score information and lyrics are input into the song synthesis model as conditions, so that the song synthesis model outputs a predicted song. Then, the real song corresponding to the musical score information and the predicted song output by the song synthesis model are used to adjust the song synthesis model. Finally, a high-quality song synthesis model is obtained through a large amount of training. However, currently, when training a song synthesis model, the requirements for training samples are relatively high. Real songs need to be recorded by professional singers. If songs recorded by ordinary users are used to train the model, due to the lack of professional training of ordinary users, the singing effect is poor and the timbre is unstable, resulting in poor quality of the recorded songs, and further causing extremely poor synthesis effects of the song synthesis model.

[0004] In this way, the existing song synthesis model can only be used to synthesize songs and cannot synthesize singing voices with the personalized timbre of ordinary users. Summary of the Invention

[0005] The main purpose of the present invention is to propose a training method, device, equipment and storage medium for a singing synthesis model, aiming to solve the problem that the existing song synthesis model cannot synthesize singing voices with the personalized timbre of ordinary users.

[0006] To solve the above technical problems, the present invention is implemented through the following technical solutions:

[0007] An embodiment of the present invention provides a training method for a singing synthesis model, including: obtaining a pre-trained singing synthesis model; collecting voice information input by a target user; obtaining text information corresponding to the voice information, and extracting the fundamental frequency and duration corresponding to each syllable in the voice information; re-training the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, the acoustic features included in the voice information are used as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

[0008] Among them, the singing synthesis model includes: an input unit and an acoustic model connected to each other; the retraining of the singing synthesis model includes: enabling the input unit to generate musical score information according to the fundamental frequency and duration corresponding to each syllable in the voice information; enabling the acoustic model to map acoustic features according to the text information and the musical score information; extracting acoustic features from the voice information in advance, and determining the loss value of the singing synthesis model according to the acoustic features extracted from the voice information and the acoustic features mapped by the acoustic model; if the loss value of the singing synthesis model is greater than a preset loss threshold, adjusting the parameters in the singing synthesis model and continuing to train the singing synthesis model; otherwise, stopping training the singing synthesis model when the singing synthesis model meets the preset convergence condition.

[0009] Among them, the enabling the input unit to generate musical score information according to the fundamental frequency and duration corresponding to each syllable in the voice information includes: for each syllable in the voice information, determining the pitch corresponding to the syllable according to the fundamental frequency corresponding to the syllable, and determining the time value corresponding to the syllable according to the duration corresponding to the syllable; generating musical score information according to the pitch and time value corresponding to each syllable in the voice information.

[0010] Among them, the acoustic model includes: an encoder, a duration model, and a decoder connected in sequence; the enabling the acoustic model to map acoustic features according to the text information and the musical score information includes: enabling the encoder to fuse the text information and the musical score information into song score information; enabling the duration model to allocate durations to each phoneme of each syllable in the song score information; enabling the decoder to map acoustic features according to the song score information and the durations corresponding to each phoneme of each syllable in the song score information.

[0011] Among them, the obtaining the text information corresponding to the voice information includes: performing speech recognition processing on the voice information to obtain and acquire the text information corresponding to the voice information; or acquiring the text information used by the target user when inputting the voice information.

[0012] Among them, after the retraining of the singing synthesis model, it further includes: receiving target song score information; the target song score information includes: lyric text and the pitch and time value corresponding to each syllable in the lyric text; inputting the target song score information into the retrained singing synthesis model so that the singing synthesis model synthesizes a song audio with the timbre of the target user according to the target song score information.

[0013] An embodiment of the present invention further provides a training device for a singing synthesis model, including: an acquisition module, configured to acquire a pre-trained singing synthesis model; a collection module, configured to collect voice information input by a target user; an extraction module, configured to obtain text information corresponding to the voice information, and extract the fundamental frequency and duration corresponding to each syllable in the voice information; a training module, configured to re-train the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, taking the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

[0014] Wherein, the singing synthesis model includes: an input unit and an acoustic model connected to each other; the training module is configured to: enable the input unit to generate score information according to the fundamental frequency and duration corresponding to each syllable in the voice information; enable the acoustic model to map acoustic features according to the text information and the score information; extract acoustic features from the voice information in advance, and determine the loss value of the singing synthesis model according to the acoustic features extracted from the voice information and the acoustic features mapped by the acoustic model; if the loss value of the singing synthesis model is greater than a preset loss threshold, adjust the parameters in the singing synthesis model and continue to train the singing synthesis model; otherwise, stop training the singing synthesis model when the singing synthesis model meets the preset convergence condition.

[0015] An embodiment of the present invention further provides a training device for a singing synthesis model. The training device for the singing synthesis model includes a processor and a memory; the processor is configured to execute a training program for the singing synthesis model stored in the memory to implement the training method for the singing synthesis model according to any one of the above.

[0016] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the training method for the singing synthesis model according to any one of the above.

[0017] The beneficial effects of the present invention are as follows:

[0018] In an embodiment of the present invention, a singing synthesis model is pre-trained as a basis, and then voice information input by a target user is collected; text information corresponding to the voice information is obtained, and the fundamental frequency and duration corresponding to each syllable in the voice information are extracted; the singing synthesis model is re-trained according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, the acoustic features included in the voice information are used as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user. Based on the basic singing synthesis model, the embodiment of the present invention can train a singing synthesis model that can synthesize a song with the timbre of the target user by using the voice information of the target user. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0020] Figure 1 is a flowchart of a method for training a singing synthesis model according to an embodiment of the present invention;

[0021] Figure 2 is a flowchart of the steps for re-training a singing synthesis model according to an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of the re-training of a singing synthesis model according to an embodiment of the present invention;

[0023] Figure 4 is a structural diagram of a training device for a singing synthesis model according to an embodiment of the present invention;

[0024] Figure 5 is a structural diagram of a training device for a singing synthesis model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] According to an embodiment of the present invention, a method for training a singing synthesis model is provided. As Figure 1 shown, it is a flowchart of a method for training a singing synthesis model according to an embodiment of the present invention.

[0027] Step S110, obtain a pre-trained singing synthesis model.

[0028] The singing synthesis model may be a song synthesis model trained using songs recorded by professional singers.

[0029] Step S120: Collect the voice information input by the target user.

[0030] The target user is a user with a personalized voice of a singing synthesis model.

[0031] The voice information can be the audio input by the target user according to the lyrics content.

[0032] Step S130: Obtain the text information corresponding to the voice information, and extract the fundamental frequency and duration corresponding to each syllable in the voice information.

[0033] The content of the text information corresponds to the content of the voice information. Further, the text information can be the lyrics. The voice information can be the voice information corresponding to the lyrics.

[0034] The fundamental frequency corresponding to a syllable refers to the lowest oscillation frequency of the syllable.

[0035] The duration corresponding to a syllable refers to the duration of the syllable.

[0036] When obtaining the text information corresponding to the voice information, voice recognition processing can be performed on the voice information to obtain and acquire the text information corresponding to the voice information; or, obtain the text information used by the target user when inputting the voice information. Among them, voice recognition processing is used to recognize voice information as text information.

[0037] When extracting the fundamental frequency and duration corresponding to each syllable in the voice information, a preset duration extraction tool can be used to extract the duration of each phoneme in the voice information. Among them, the duration extraction tool can be tools such as MFA (Montreal Forced Aligner), HTK (HMM Toolkit), Kaldi (a voice recognition tool), etc. A phoneme refers to the smallest speech unit. Further, multiple syllables in the voice information are regularly combined by at least one phoneme. For example: in Chinese, a syllable includes an initial consonant and a final consonant, and the initial consonant is at the beginning. Therefore, according to the regular characteristics of syllables, word boundary recognition can be performed; the word boundary is the boundary of the syllable (the duration of the syllable), then word boundary recognition is: dividing each phoneme in the voice information into multiple syllables, and word boundary recognition can be performed by means of voice recognition or song recognition; according to the duration corresponding to each phoneme, determine the duration corresponding to each syllable; according to each syllable divided in the voice information, divide the spectrum corresponding to each syllable in the overall spectrum of the voice information, and then according to the spectrum corresponding to each syllable, determine the fundamental frequency corresponding to each syllable.

[0038] Step S140: Retrain the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information. Among them, use the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

[0039] The fundamental frequency corresponding to a syllable can reflect the pitch of the syllable in a song.

[0040] The duration corresponding to a syllable can reflect the duration value of the syllable in a song.

[0041] The acoustic features include not only the text content of the voice but also content such as the timbre features of the target user.

[0042] Use the text information and the fundamental frequency and duration corresponding to each syllable in the voice information as the input of the singing synthesis model, so that the singing synthesis model outputs a predicted song audio. Since the embodiments of the present invention hope that the singing synthesis model outputs a song audio with the timbre of the target user, and the acoustic features of the target user are included in the voice information, when retraining the singing synthesis model, use the acoustic features of the target user as the learning target of the singing synthesis model, and determine the loss value of the singing synthesis model based on the acoustic features. In this way, a singing synthesis model that can synthesize the singing with the timbre of the target user can be trained.

[0043] In the embodiments of the present invention, since the amount of voice information of the target user is limited, the voice information of the target user can be intercepted into multiple segments of voice information, thereby increasing the number of training samples. In this way, the singing synthesis model can be retrained using the text information corresponding to each segment of voice information and the fundamental frequency and duration corresponding to each syllable extracted from this segment of voice information.

[0044] In the embodiments of the present invention, first pre-train a singing synthesis model as a basis, and then collect the voice information input by the target user; obtain the text information corresponding to the voice information, and extract the fundamental frequency and duration corresponding to each syllable in the voice information; retrain the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information. Among them, use the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user. Based on the basic singing synthesis model, the embodiments of the present invention can train a singing synthesis model that can synthesize a song with the timbre of the target user by using the voice information of the target user.

[0045] In an embodiment of the present invention, the singing synthesis model includes: an input unit and an acoustic model connected to each other. Of course, the singing synthesis model also includes a vocoder. The vocoder is connected to the acoustic model.

[0046] The retraining process of the singing synthesis model will be described below. As Figure 2 shown, it is a flowchart of the steps for retraining the singing synthesis model according to an embodiment of the present invention. As Figure 3 shown, it is a schematic diagram of the retraining of the singing synthesis model according to an embodiment of the present invention.

[0047] Step S210: Make the input unit generate score information according to the fundamental frequency and duration corresponding to each syllable in the voice information.

[0048] The score information is the musical score, which is used to record the pitch and duration of music.

[0049] For each syllable in the voice information, determine the pitch (Note) corresponding to the syllable according to the fundamental frequency corresponding to the syllable, and, determine the duration (NoteDur) corresponding to the syllable according to the duration corresponding to the syllable; generate score information according to the pitch and duration corresponding to each syllable in the voice information.

[0050] Furthermore, the Librosa tool can be used to convert the fundamental frequency into pitch. The duration unit can be milliseconds (ms), then the duration value can be equal to the duration divided by 20ms, or the duration divided by 25ms, and the value is rounded or rounded to an integer to improve the robustness of the voice synthesis model. Among them, 20ms and 25ms are adjustable parameters, and this parameter is an empirical value or a value obtained through experiments.

[0051] Step S220: Make the acoustic model map acoustic features according to the text information and the score information.

[0052] The acoustic feature is a physical quantity representing the acoustic characteristics of speech, and it is also a general term for the acoustic manifestations of various elements of sound. For example: the energy concentration area, formant frequency, formant intensity and bandwidth representing timbre, and the duration, fundamental frequency, average speech power, etc. representing the prosodic characteristics of speech.

[0053] The acoustic model includes: an encoder, a duration model, and a decoder connected in sequence. During the process of training the voice synthesis model, the duration model is also being trained.

[0054] Cause the encoder to fuse the text information (lyrics) and the musical score information into musical score information. The musical score information is a combination of lyrics and musical scores. Further, the text information includes the text content of the syllables and the time information of the syllable pronunciation. For example: the first syllable is pronounced at the 10th second, indicating that the first 9 seconds are the prelude music. Superimpose the text information and the musical score information to achieve the fusion of the text information and the musical score information.

[0055] Cause the duration model to allocate durations for each phoneme of each syllable in the musical score information. For example: in Chinese, allocate the durations for each initial consonant and each final consonant respectively. During the process of retraining the singing synthesis model, since the durations corresponding to each syllable in the speech information need to be extracted first by extracting the durations corresponding to each phoneme in the speech information, the duration model can allocate durations for each phoneme in the musical score information according to the durations corresponding to each phoneme extracted in advance, and the duration model learns during the process of allocating durations so that after the retraining is completed, it can predict the durations corresponding to each phoneme in other musical score information and allocate durations according to the prediction results.

[0056] Cause the decoder to map acoustic features according to the musical score information and the durations corresponding to each phoneme of each syllable in the musical score information. The acoustic features are Mel spectrograms (Mel frequency spectrograms).

[0057] The acoustic model outputs the acoustic features to the vocoder. Among them, the decoder in the acoustic model can perform frame expansion on the acoustic features and then output them to the vocoder. Frame expansion is used to make the number of frames input to the decoder and the number of frames output consistent.

[0058] The vocoder synthesizes the song audio according to the acoustic features (acoustic features), the fundamental frequency (F0) corresponding to each syllable in the speech information, and the voiceless / voiced (UV decision) corresponding to each phoneme in each syllable. Among them, only when it is determined that the phoneme is voiced, assign the fundamental frequency corresponding to the syllable to which the phoneme belongs.

[0059] Step S230, extract acoustic features in the speech information in advance, and determine the loss value of the singing synthesis model according to the acoustic features extracted in the speech information and the acoustic features mapped by the acoustic model.

[0060] After the target user inputs speech information and before calculating the loss value of the singing synthesis model, acoustic features can be extracted from the speech information by using a preset feature extraction tool. This feature extraction tool can be the Librosa tool.

[0061] If the loss value of the singing synthesis model is greater than a preset loss threshold, adjust the parameters in the singing synthesis model and continue to train the singing synthesis model; otherwise, when the singing synthesis model meets the preset convergence condition, stop training the singing synthesis model. Among them, the loss threshold can be an empirical value or a value obtained through experiments.

[0062] Step S240, determine whether the loss value of the singing synthesis model is greater than a preset loss threshold; if yes, execute step S250; if no, execute step S260.

[0063] Step S250, adjust the parameters in the singing synthesis model and jump to step S120 to continue collecting the voice information input by the target user and train the singing synthesis model.

[0064] After jumping to step S120, according to the new voice information, execute the retraining step again.

[0065] Step S260, determine whether the singing synthesis model meets the preset convergence condition; if yes, execute step S270; if no, execute step S250.

[0066] The convergence condition includes: the parameters in the singing synthesis model tend to be stable, and the number of retraining iterations of the singing synthesis model reaches the iteration threshold. That the parameters in the singing synthesis model tend to be stable means that the difference between the same parameter in the model before and after adjustment is within a preset range.

[0067] Step S270, when the singing synthesis model meets the preset convergence condition, stop training the singing synthesis model.

[0068] In the embodiment of the present invention, the singing synthesis model is retrained using the voice information of the target user, so that the singing synthesis model has the ability to synthesize the singing voice of the target user's timbre. The singing synthesis model already has the ability to synthesize high-quality songs before retraining, which ensures the stability of the model. After retraining, the singing synthesis model realizes timbre migration, enabling the singing synthesis model to use the timbre of the target user to synthesize high-quality singing voices, thereby realizing the personalized needs of singing synthesis.

[0069] In the embodiment of the present invention, after retraining the singing synthesis model, target music score information can also be received; the target music score information includes: the lyric text and the pitch and duration corresponding to each syllable in the lyric text; input the target music score information into the retrained singing synthesis model so that the singing synthesis model synthesizes a song audio with the timbre of the target user according to the target music score information.

[0070] An embodiment of the present invention also provides a training device for a singing synthesis model. As Figure 4 shown, it is a structural diagram of a training device for a singing synthesis model according to an embodiment of the present invention.

[0071] The training device for the singing synthesis model includes:

[0072] An acquisition module 410, configured to acquire a pre-trained singing synthesis model.

[0073] A collection module 420, configured to collect voice information input by a target user.

[0074] An extraction module 430, configured to obtain text information corresponding to the voice information, and extract the fundamental frequency and duration corresponding to each syllable in the voice information.

[0075] A training module 440, configured to re-train the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, the acoustic features included in the voice information are used as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

[0076] Wherein, in the singing synthesis model, there are an input unit and an acoustic model connected to each other; the training module 440 is configured to: enable the input unit to generate score information according to the fundamental frequency and duration corresponding to each syllable in the voice information; enable the acoustic model to map acoustic features according to the text information and the score information; extract acoustic features from the voice information in advance, and determine the loss value of the singing synthesis model according to the acoustic features extracted from the voice information and the acoustic features mapped by the acoustic model; if the loss value of the singing synthesis model is greater than a preset loss threshold, adjust the parameters in the singing synthesis model and continue to train the singing synthesis model; otherwise, stop training the singing synthesis model when the singing synthesis model meets the preset convergence condition.

[0077] Wherein, the training module 440 is configured to: for each syllable in the voice information, determine the pitch corresponding to the syllable according to the fundamental frequency corresponding to the syllable, and determine the duration value corresponding to the syllable according to the duration corresponding to the syllable; generate score information according to the pitch and duration value corresponding to each syllable in the voice information.

[0078] Among them, the acoustic model includes: an encoder, a duration model, and a decoder connected in sequence; the training module 440 is configured to: enable the encoder to fuse the text information and the music score information into song score information; enable the duration model to allocate a corresponding duration to each phoneme of each syllable in the song score information; enable the decoder to map acoustic features according to the song score information and the corresponding duration of each phoneme of each syllable in the song score information.

[0079] Among them, the obtaining module 410 is configured to: perform speech recognition processing on the speech information to obtain and acquire the text information corresponding to the speech information; or, acquire the text information used by the target user when inputting the speech information.

[0080] Among them, the device further includes an application module (not shown in the figure), and the application module is configured to, after retraining the singing synthesis model, receive target song score information; the target song score information includes: lyric text and the pitch and time value corresponding to each syllable in the lyric text; input the target song score information into the retrained singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user according to the target song score information.

[0081] The functions of the device according to the embodiments of the present invention have been described in the above method embodiments. Therefore, for the parts not described in detail in this embodiment, reference may be made to the relevant descriptions in the foregoing embodiments, and details are not repeated here.

[0082] This embodiment provides a training device for a singing synthesis model. As Figure 5 shown, it is a structural diagram of a training device for a singing synthesis model according to an embodiment of the present invention.

[0083] In this embodiment, the training device for the singing synthesis model includes but is not limited to: a processor 510 and a memory 520.

[0084] The processor 510 is configured to execute a training program for the singing synthesis model stored in the memory 520 to implement the above-mentioned training method for the singing synthesis model.

[0085] Specifically, the processor 510 is configured to execute a training program of a singing synthesis model stored in the memory 520 to implement the following steps: obtaining a pre-trained singing synthesis model; collecting voice information input by a target user; obtaining text information corresponding to the voice information, and extracting the fundamental frequency and duration corresponding to each syllable in the voice information; re-training the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, using the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

[0086] Wherein, the singing synthesis model includes: an input unit and an acoustic model connected to each other; the re-training of the singing synthesis model includes: enabling the input unit to generate musical score information according to the fundamental frequency and duration corresponding to each syllable in the voice information; enabling the acoustic model to map acoustic features according to the text information and the musical score information; pre-extracting acoustic features from the voice information, and determining a loss value of the singing synthesis model according to the acoustic features extracted from the voice information and the acoustic features mapped by the acoustic model; if the loss value of the singing synthesis model is greater than a preset loss threshold, adjusting the parameters in the singing synthesis model and continuing to train the singing synthesis model; otherwise, stopping training the singing synthesis model when the singing synthesis model meets a preset convergence condition.

[0087] Wherein, the enabling the input unit to generate musical score information according to the fundamental frequency and duration corresponding to each syllable in the voice information includes: for each syllable in the voice information, determining the pitch corresponding to the syllable according to the fundamental frequency corresponding to the syllable, and determining the time value corresponding to the syllable according to the duration corresponding to the syllable; generating musical score information according to the pitch and time value corresponding to each syllable in the voice information.

[0088] Wherein, the acoustic model includes: an encoder, a duration model, and a decoder connected in sequence; the enabling the acoustic model to map acoustic features according to the text information and the musical score information includes: enabling the encoder to fuse the text information and the musical score information into song score information; enabling the duration model to allocate durations to each phoneme of each syllable in the song score information; enabling the decoder to map acoustic features according to the song score information and the durations corresponding to each phoneme of each syllable in the song score information.

[0089] Wherein, the obtaining the text information corresponding to the voice information includes: performing speech recognition processing on the voice information to obtain and acquire the text information corresponding to the voice information; or, acquiring the text information used by the target user when inputting the voice information.

[0090] After retraining the singing synthesis model, the method further includes: receiving target musical score information; the target musical score information includes: lyric text and the pitch and duration corresponding to each syllable in the lyric text; inputting the target musical score information into the retrained singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user according to the target musical score information.

[0091] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The computer-readable storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive; the memory may also include a combination of the above types of memory.

[0092] When one or more programs in the computer-readable storage medium can be executed by one or more processors to implement the above-mentioned training method of the singing synthesis model.

[0093] Specifically, the processor is configured to execute the training program of the singing synthesis model stored in the memory to implement the following steps: obtaining a pre-trained singing synthesis model; collecting voice information input by the target user; obtaining the text information corresponding to the voice information, and extracting the fundamental frequency and duration corresponding to each syllable in the voice information; retraining the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; wherein, using the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

[0094] Among them, the singing synthesis model includes: an input unit and an acoustic model connected to each other; retraining the singing synthesis model includes: enabling the input unit to generate musical score information according to the fundamental frequency and duration corresponding to each syllable in the voice information; enabling the acoustic model to map acoustic features according to the text information and the musical score information; extracting acoustic features from the voice information in advance, and determining the loss value of the singing synthesis model according to the acoustic features extracted from the voice information and the acoustic features mapped by the acoustic model; if the loss value of the singing synthesis model is greater than a preset loss threshold, adjusting the parameters in the singing synthesis model and continuing to train the singing synthesis model; otherwise, stopping training the singing synthesis model when the singing synthesis model meets the preset convergence condition.

[0095] Among them, making the input unit generate score information according to the fundamental frequency and duration corresponding to each syllable in the voice information includes: for each syllable in the voice information, determining the pitch corresponding to the syllable according to the fundamental frequency corresponding to the syllable, and determining the time value corresponding to the syllable according to the duration corresponding to the syllable; generating score information according to the pitch and time value corresponding to each syllable in the voice information.

[0096] Among them, the acoustic model includes: an encoder, a duration model, and a decoder connected in sequence; making the acoustic model map acoustic features according to the text information and the score information includes: making the encoder fuse the text information and the score information into song score information; making the duration model allocate durations to each phoneme of each syllable in the song score information; making the decoder map acoustic features according to the song score information and the durations corresponding to each phoneme of each syllable in the song score information.

[0097] Among them, obtaining the text information corresponding to the voice information includes: performing speech recognition processing on the voice information to obtain and acquire the text information corresponding to the voice information; or acquiring the text information used by the target user when inputting the voice information.

[0098] Among them, after retraining the singing synthesis model, it further includes: receiving target song score information; the target song score information includes: lyric text and the pitch and time value corresponding to each syllable in the lyric text; inputting the target song score information into the retrained singing synthesis model so that the singing synthesis model synthesizes a song audio with the timbre of the target user according to the target song score information.

[0099] The above are only embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A training method for a singing voice synthesis model, characterized in that Comprising: Obtaining a pre-trained singing synthesis model; The singing synthesis model includes an input unit and an acoustic model connected to each other; the acoustic model includes an encoder, a duration model, and a decoder connected in sequence; Collecting voice information input by a target user; Obtaining text information corresponding to the voice information, and extracting the fundamental frequency and duration corresponding to each syllable in the voice information; Retraining the singing synthesis model according to the text information, the voice information, and the fundamental frequency and duration corresponding to each syllable in the voice information; The retraining of the singing synthesis model includes: enabling the input unit to generate score information according to the fundamental frequency and duration corresponding to each syllable in the voice information, enabling the acoustic model to map acoustic features according to the text information and the score information, extracting acoustic features from the voice information in advance, determining the loss value of the singing synthesis model according to the acoustic features extracted from the voice information and the acoustic features mapped by the acoustic model, if the loss value of the singing synthesis model is greater than a preset loss threshold, adjusting the parameters in the singing synthesis model and continuing to train the singing synthesis model, otherwise, stopping training the singing synthesis model when the singing synthesis model meets the preset convergence condition; The enabling the acoustic model to map acoustic features according to the text information and the score information includes: enabling the encoder to fuse the text information and the score information into score information, enabling the duration model to allocate durations to each phoneme of each syllable in the score information, and enabling the decoder to map acoustic features according to the score information and the durations corresponding to each phoneme of each syllable in the score information; Wherein, using the acoustic features included in the voice information as the learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

2. The method according to claim 1, wherein The enabling the input unit to generate score information according to the fundamental frequency and duration corresponding to each syllable in the voice information includes: For each syllable in the voice information, determining the pitch corresponding to the syllable according to the fundamental frequency corresponding to the syllable, and determining the time value corresponding to the syllable according to the duration corresponding to the syllable; Generating score information according to the pitch and time value corresponding to each syllable in the voice information.

3. The method according to claim 1, wherein The obtaining the text information corresponding to the voice information includes: Performing speech recognition processing on the voice information to obtain and acquire the text information corresponding to the voice information; or, acquiring the text information used by the target user when inputting the voice information.

4. The method according to claim 1, wherein After the retraining of the singing synthesis model, it further includes: Receiving target score information; the target score information includes: lyric text and the pitch and time value corresponding to each syllable in the lyric text; Inputting the target score information into the retrained singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user according to the target score information.

5. A training device for a singing voice synthesis model, characterized in that, Comprising: An acquisition module, configured to acquire a pre-trained singing synthesis model, where the singing synthesis model includes: an input unit and an acoustic model connected to each other; the acoustic model includes: an encoder, a duration model, and a decoder connected in sequence; a collection module, configured to collect speech information input by a target user; An extraction module, configured to obtain text information corresponding to the speech information, and extract the fundamental frequency and duration corresponding to each syllable in the speech information; A training module, configured to re-train the singing synthesis model according to the text information, the speech information, and the fundamental frequency and duration corresponding to each syllable in the speech information; the re-training of the singing synthesis model includes: enabling the input unit to generate score information according to the fundamental frequency and duration corresponding to each syllable in the speech information, enabling the acoustic model to map acoustic features according to the text information and the score information, extracting acoustic features from the speech information in advance, determining a loss value of the singing synthesis model according to the acoustic features extracted from the speech information and the acoustic features mapped by the acoustic model, if the loss value of the singing synthesis model is greater than a preset loss threshold, adjusting parameters in the singing synthesis model, and continuing to train the singing synthesis model, otherwise, stopping training the singing synthesis model when the singing synthesis model meets a preset convergence condition; the enabling the acoustic model to map acoustic features according to the text information and the score information includes: enabling the encoder to fuse the text information and the score information into score information, enabling the duration model to allocate durations to each phoneme of each syllable in the score information, enabling the decoder to map acoustic features according to the score information and the durations corresponding to each phoneme of each syllable in the score information; wherein, taking the acoustic features included in the speech information as a learning target of the singing synthesis model, so that the singing synthesis model synthesizes a song audio with the timbre of the target user.

6. The apparatus according to claim 5, wherein in the singing synthesis model, it includes: an input unit and an acoustic model connected to each other; the training module is configured to: enable the input unit to generate score information according to the fundamental frequency and duration corresponding to each syllable in the speech information; enable the acoustic model to map acoustic features according to the text information and the score information; extract acoustic features from the speech information in advance, and determine a loss value of the singing synthesis model according to the acoustic features extracted from the speech information and the acoustic features mapped by the acoustic model; if the loss value of the singing synthesis model is greater than a preset loss threshold, adjust parameters in the singing synthesis model, and continue to train the singing synthesis model; otherwise, stop training the singing synthesis model when the singing synthesis model meets a preset convergence condition.

7. A training device for a singing voice synthesis model, characterized in that, The training device of the singing synthesis model includes a processor and a memory; the processor is configured to execute the training program of the singing synthesis model stored in the memory to implement the training method of the singing synthesis model according to any one of claims 1 to 4.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the training method of the singing synthesis model according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Singing synthesis method, device and equipment

    CN112750422A

  • Voice synthesis method, voice synthesis device, and storage medium

    US20190251950A1