A training method, device, equipment and storage medium for a song synthesis model
By first using the voice sample library to train the basic model and then using the song sample library to retrain, the problem of high training cost of song synthesis model is solved, and efficient and low-cost song synthesis model training is achieved.
Patent Information
- Application Number
- CN202211583385.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-10
AI Technical Summary
The existing song synthesis model training method is costly, mainly due to the small number of high-quality song samples and high price, which leads to the high cost of building the sample library.
By setting up the speech sample library and the song sample library, first use the speech sample library to train the basic model until the model converges, and then use the song sample library to retrain the basic model until the model converges again, and obtain a song synthesis model with the ability to synthesize songs.
The cost and difficulty of training the song synthesis model is reduced. By using a large number of speech samples to train the stability of the basic model, and then retraining it with a small number of song samples, efficient training of the song synthesis model is achieved.
Smart Images

Figure CN115881066B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and in particular, to a method, apparatus, device, and storage medium for training a song synthesis model. Background Art
[0002] With the continuous progress of technology, in addition to song recording (such as recording songs sung by singers), song audio has a new source form, namely: song synthesis.
[0003] Currently, song synthesis can be completed using a song synthesis model. A stable song synthesis model needs to be obtained through training with a large number of song samples. The song samples are song audio with high audio quality and singing quality. However, due to problems such as the small number, high price, and difficulty in obtaining song audio with high quality, the construction cost of the song sample library is relatively high, making it not easy to train a stable song synthesis model.
[0004] Therefore, providing a method for training a song synthesis model with a lower cost has become an urgent problem in this field. Summary of the Invention
[0005] The main object of the present invention is to propose a method, apparatus, device, and storage medium for training a song synthesis model, aiming to solve the problem of the relatively high cost of the existing training method for song synthesis models.
[0006] To achieve the above technical problem, the present invention is implemented through the following technical solutions:
[0007] An embodiment of the present invention provides a method for training a song synthesis model, including: respectively setting a speech sample library and a song sample library; wherein, the number of speech samples in the speech sample library is greater than the number of song samples in the song sample library; using the speech sample library to train a basic model until the basic model converges; using the song sample library to retrain the basic model until the basic model converges again, to obtain a song synthesis model.
[0008] Wherein, the speech sample library includes multiple speech samples; wherein, each speech sample includes: corresponding sample text and sample speech; the song sample library includes multiple song samples; wherein, each song sample includes: corresponding lyric text, song audio, and song score.
[0009] Among them, training the basic model by using the speech sample library includes: sequentially obtaining a speech sample from the speech sample library; determining the pitch and duration corresponding to each character in the sample text based on the sample text and the sample speech in the speech sample; inputting the sample text and the pitch and duration corresponding to each character in the sample text into the basic model, and obtaining the predicted speech output by the basic model; determining the loss value of the basic model according to the sample speech in the speech sample and the predicted speech output by the basic model; adjusting the basic model according to the loss value of the basic model, and continuing to obtain speech samples from the speech sample library to train the basic model until the basic model converges.
[0010] Among them, determining the pitch and duration corresponding to each character in the sample text based on the sample text and the sample speech in the speech sample includes: performing a phoneme alignment operation on the corresponding sample text and sample speech; extracting the fundamental frequency and duration corresponding to each character in the sample speech; for each character, determining the pitch corresponding to the character according to the fundamental frequency corresponding to the character, and determining the duration corresponding to the character according to the duration corresponding to the character.
[0011] Among them, retraining the basic model by using the song sample library includes: sequentially obtaining a song sample from the song sample library; where the song sample includes: corresponding lyric text, song audio, and song score; the song score includes pitch and beat information; determining the pitch and duration corresponding to each character in the lyric text according to the pitch and beat information in the song score; inputting the lyric text and the pitch and duration corresponding to each character in the lyric text into the basic model, and obtaining the predicted audio output by the basic model; determining the loss value of the basic model according to the song audio in the song sample and the predicted audio output by the basic model; adjusting the basic model according to the loss value of the basic model, and continuing to obtain song samples from the song sample library to train the basic model until the basic model converges, to obtain a song synthesis model.
[0012] Among them, after obtaining the song synthesis model, it further includes: receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text; inputting the target lyric text and the pitch and duration corresponding to each character in the target lyric text into the song synthesis model, and obtaining the synthesized song audio output by the song synthesis model; displaying a synthesized song icon on a preset display interface; where the synthesized song icon is used to trigger the playback of the synthesized song audio.
[0013] After receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text, it further includes: in the display interface, displaying the target lyric text and the pitch and duration corresponding to each character in the target lyric text; detecting a song modification instruction; and modifying the target lyric text, and / or modifying the pitch corresponding to a specified character in the target lyric text, and / or modifying the duration corresponding to a specified character in the target lyric text according to the detected song modification instruction.
[0014] An embodiment of the present invention further provides a training device for a song synthesis model, including: a setting module for respectively setting a speech sample library and a song sample library; wherein, the number of speech samples in the speech sample library is greater than the number of songs in the song sample library; a training module for training a basic model using the speech sample library until the basic model converges; a retraining module for retraining the basic model using the song sample library until the basic model converges again to obtain a song synthesis model.
[0015] An embodiment of the present invention further provides a training device for a song synthesis model. The training device for the song synthesis model includes a processor and a memory; the processor is configured to execute a training program for the song synthesis model stored in the memory to implement the training method for the song synthesis model described above.
[0016] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the training method for the song synthesis model described in any one of the above.
[0017] The beneficial effects of the present invention are as follows:
[0018] In the embodiment of the present invention, the basic model is first trained using the speech sample library, which can make the basic model have stability in speech synthesis. However, at this time, the basic model only outputs speech with high synthesis accuracy rather than songs. By means of the basic model with stable speech synthesis in the embodiment of the present invention, the basic model is retrained using the song sample library, so that the output speech of the basic model has a musical rhythm, thereby enabling the basic model to have the ability to synthesize songs. Since the embodiment of the present invention first uses a large number of speech samples to train the basic model to ensure the stability of the synthesis effect of the basic model, and then retrains the basic model with a very small song sample library, the cost of constructing the sample library and the training difficulty of the song synthesis model are greatly reduced. Description of the Drawings
[0019] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0020] Figure 1 is a flowchart of a method for training a song synthesis model according to an embodiment of the present invention;
[0021] Figure 2 is a flowchart of the preliminary training steps of a basic model according to an embodiment of the present invention;
[0022] Figure 3 is a flowchart of the retraining steps of a basic model according to an embodiment of the present invention;
[0023] Figure 4 is a flowchart of the steps of song synthesis processing according to an embodiment of the present invention;
[0024] Figure 5 is a structural diagram of a training device for a song synthesis model according to an embodiment of the present invention;
[0025] Figure 6 is a structural diagram of a training device for a song synthesis model according to an embodiment of the present invention. Detailed Embodiments
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] According to an embodiment of the present invention, a method for training a song synthesis model is provided. As Figure 1 shown, it is a flowchart of a method for training a song synthesis model according to an embodiment of the present invention.
[0028] Step S110: Set up a voice sample library and a song sample library respectively; wherein, the number of voice samples in the voice sample library is greater than the number of song samples in the song sample library.
[0029] The voice sample library includes a plurality of voice samples; wherein, each voice sample includes: corresponding sample text and sample voice. Further, the voice sample library can include a large number of voice samples. The sample voice in the voice sample is a voice without a melody, which can be obtained by recording or by voice synthesis technology. Recording voice content is relatively easy, and voice synthesis technology is relatively mature and low-cost.
[0030] The song sample library includes multiple song samples; among them, each song sample includes: corresponding lyric text, song audio, and song score. Further, the song audio is audio with a melody. The song sample library can be set with a relatively small number of song samples, which can effectively reduce the cost of building the sample library.
[0031] Step S120, use the speech sample library to train the basic model until the basic model converges.
[0032] After training, the basic model is a speech synthesis model and can be used for speech synthesis.
[0033] Step S130, use the song sample library to retrain the basic model until the basic model converges again to obtain a song synthesis model.
[0034] After retraining, the basic model is transformed into a song synthesis model and can be used for song synthesis.
[0035] In the embodiment of the present invention, first, a speech sample library is used to train the basic model, which can make the basic model have stability in speech synthesis. However, at this time, the basic model only outputs speech with high synthesis accuracy rather than songs. In the embodiment of the present invention, by virtue of the basic model with stable speech synthesis, the song sample library is used to retrain the basic model, so that the output speech of the basic model has a musical rhythm, thereby enabling the basic model to have the ability to synthesize songs. Since the embodiment of the present invention first uses a large number of speech samples to train the basic model to ensure the stability of the synthesis effect of the basic model, and then retrains the basic model with a very small song sample library (such as two songs), the cost of building the sample library and the training difficulty of the song synthesis model are greatly reduced.
[0036] First, the process of training the basic model using the speech sample library will be described below.
[0037] As Figure 2 shown, it is a flowchart of the preliminary training steps of the basic model according to an embodiment of the present invention.
[0038] Step S210, sequentially obtain a speech sample from the speech sample library.
[0039] Step S220, based on the sample text and sample speech in the speech sample, determine the pitch and duration corresponding to each character in the sample text.
[0040] Pitch and duration are music theory concepts. The characters in the sample text correspond to syllables in music theory.
[0041] Step S1, perform phoneme alignment operations on the corresponding sample text and sample speech.
[0042] A preset phoneme alignment tool can be used to perform phoneme alignment operations on corresponding sample texts and sample voices.
[0043] Step S2: In the sample voice, extract the fundamental frequency and duration corresponding to each character.
[0044] In a voice, both the fundamental frequency and the duration are basic features of the voice.
[0045] The fundamental frequency of a character can reflect the lowest oscillation frequency of the character.
[0046] The duration of a character can reflect the pronunciation duration of the character.
[0047] In this way, it can be understood that the fundamental frequency of a character can correspond to the pitch of a phoneme, and the duration of a character can correspond to the time value of a phoneme.
[0048] Step S3: For each of the characters, determine the pitch corresponding to the character according to the fundamental frequency corresponding to the character, and determine the time value corresponding to the character according to the duration corresponding to the character.
[0049] The time value (unit: millisecond) corresponding to a character (syllable) can be determined by the following formula:
[0050] Time value = ROUND(1000d i / 20);
[0051] where d i represents the duration of the i-th syllable (unit: second); ROUND is a rounding function; 20 is an adjustable parameter that can be set according to experience or obtained through experiments.
[0052] The pitch corresponding to a character (syllable) can be determined by the following formula:
[0053]
[0054] where p represents the distance (pitch) between the pitch marked in the numbered musical notation information and the A note on middle C, and its unit is semitone; f0 represents the fundamental frequency of the character; 440 represents the frequency (unit: HZ) emitted by the A note on middle C.
[0055] Step S230: Input the sample text and the pitch and time value corresponding to each character in the sample text into the basic model, and obtain the predicted voice output by the basic model.
[0056] Step S240: Determine the loss value of the basic model according to the sample voice in the voice sample and the predicted voice output by the basic model.
[0057] Use a preset first loss function to determine the loss value of the basic model.
[0058] Step S250: Adjust the base model according to the loss value of the base model, continue to obtain speech samples in the speech sample library, and train the base model until the base model converges.
[0059] During the process of training the base model, determine whether the base model has met a preset first convergence condition; in the case where it is determined that the base model meets this first convergence condition, determine that the base model has converged; in the case where it is determined that the base model does not meet this first convergence condition, continue to obtain speech samples in the speech sample library and train the base model.
[0060] The convergence condition may include at least one of the following conditions:
[0061] 1. The loss value of the base model is less than a first loss threshold.
[0062] 2. The parameters in the base model tend to be stable. Tending to be stable means that the difference between the same parameter before and after adjustment is within a preset range.
[0063] 3. The number of training times of the base model reaches a preset first iteration number threshold.
[0064] Next, the retraining process of the base model using the song sample library will be described.
[0065] As Figure 3 shown, it is a flowchart of the retraining steps of the base model according to an embodiment of the present invention.
[0066] Step S310: Sequentially obtain a song sample in the song sample library. Among them, the song sample includes: corresponding lyric text, song audio, and song score; the song score includes pitch and beat information.
[0067] The beat information in the song score includes: the number of beats per minute in the song score and the beats of each syllable (corresponding to the characters in the lyric text).
[0068] Step S320: Determine the pitch and duration corresponding to each character in the lyric text according to the pitch and beat information in the song score.
[0069] Because there is a corresponding relationship between the song score and the lyric text, the pitch corresponding to each character in the lyric text can be directly extracted from the song score, and the duration corresponding to each character in the lyric text can be calculated according to the beat information in the song score.
[0070] For example: The beat of the syllable can be converted into duration through the following formula:
[0071]
[0072] Time value = ROUND(1000d i / 20);
[0073] where d i represents the duration of the i-th syllable (in seconds); tmpo is the number of beats per minute in the simplified musical notation of the song; dnote i represents the beat of the i-th syllable; ROUND is the rounding function; 20 is an adjustable parameter.
[0074] Step S330: Input the lyric text and the pitch and time value corresponding to each character in the lyric text into the basic model, and obtain the predicted audio output by the basic model.
[0075] Step S340: Determine the loss value of the basic model based on the song audio in the song sample and the predicted audio output by the basic model.
[0076] Determine the loss value of the basic model using a preset second loss function.
[0077] Step S350: Adjust the basic model according to the loss value of the basic model, continue to obtain song samples in the song sample library, and train the basic model until the basic model converges, thus obtaining a song synthesis model.
[0078] During the process of retraining the basic model, determine whether the basic model has met a preset second convergence condition; if it is determined that the basic model meets this second convergence condition, it is determined that the basic model has converged; if it is determined that the basic model does not meet this second convergence condition, continue to obtain speech samples in the speech sample library and train this basic model.
[0079] This second convergence condition may include at least one of the following conditions:
[0080] 1. The loss value of the basic model is less than a second loss threshold.
[0081] 2. The parameters in the basic model tend to be stable. Tending to be stable means that the difference between the same parameter before and after adjustment is within a preset range.
[0082] 3. The number of training times of the basic model reaches a preset second iteration number threshold.
[0083] In an embodiment of the present invention, since the number of song samples is small, in order to expand the number of samples, each song sample can be incrementally processed. Further, for each song sample, the lyric text, the simplified musical score of the song, and the song audio in the song sample are respectively intercepted into a plurality of segments with the same quantity and corresponding content. For example: first, the lyric text is intercepted into a preset number of segments, the corresponding segments of the simplified musical score of the song are respectively intercepted in the simplified musical score of the song, and the corresponding segments of the song audio for each lyric text segment are intercepted in the song audio. The mutually corresponding lyric text segments, simplified musical score segments, and song audio segments are used as a song sample and set in the song sample library for retraining the basic model.
[0084] After obtaining the song synthesis model, the song synthesis model can be used to perform song synthesis processing.
[0085] As Figure 4 shown, it is a step flow chart of song synthesis processing according to an embodiment of the present invention.
[0086] Step S410, receive the target lyric text and the pitch and duration corresponding to each character in the target lyric text.
[0087] After receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text, the target lyric text and the pitch and duration corresponding to each character in the target lyric text can also be displayed on the display interface; detect a song modification instruction; according to the detected song modification instruction, modify the target lyric text, and / or modify the pitch corresponding to a specified character in the target lyric text, and / or modify the duration corresponding to a specified character in the target lyric text. Further, the song modification instruction is used to indicate the modification of the text content, pitch, and / or duration in the target lyric text.
[0088] The modified lyric text, the pitch and duration corresponding to each character in the lyric text are used as the target lyric text and the pitch and duration corresponding to each character in the target lyric text.
[0089] For example: the user selects the text to be modified on the display interface, enters the modified text to be displayed in the popped-up text box, clicks OK, the song modification instruction is triggered, and the text to be modified is replaced with the modified text according to the song modification instruction.
[0090] Step S420, input the target lyric text and the pitch and duration corresponding to each character in the target lyric text into the song synthesis model, and obtain the synthesized song audio output by the song synthesis model.
[0091] Step S430: Display a synthesized song icon in a preset display interface; wherein, the synthesized song icon is used to trigger the playback of the synthesized song audio.
[0092] Through this embodiment, the user can use the song synthesis model to synthesize the required songs and can modify the songs before synthesis, meeting the user's personalized needs and increasing the fun of song synthesis.
[0093] The embodiment of the present invention also provides a training device for a song synthesis model. As Figure 5 shown, it is a structural diagram of a training device for a song synthesis model according to an embodiment of the present invention.
[0094] The training device for the song synthesis model includes:
[0095] A setting module 510, configured to respectively set a speech sample library and a song sample library; wherein, the number of speech samples in the speech sample library is greater than the number of songs in the song sample library.
[0096] A training module 520, configured to use the speech sample library to train a basic model until the basic model converges.
[0097] A retraining module 530, configured to use the song sample library to retrain the basic model until the basic model converges again to obtain a song synthesis model.
[0098] The functions of the device described in the embodiment of the present invention have been described in the above method embodiment. Therefore, for the parts not elaborated in the description of this embodiment, reference can be made to the relevant descriptions in the foregoing embodiments and will not be repeated here.
[0099] This embodiment provides a training device for a song synthesis model. As Figure 6 shown, it is a structural diagram of a training device for a song synthesis model according to an embodiment of the present invention.
[0100] In this embodiment, the training device for the song synthesis model includes but is not limited to: a processor 610 and a memory 620.
[0101] The processor 610 is configured to execute the training program of the song synthesis model stored in the memory 620 to implement the above-mentioned song synthesis model training method.
[0102] Specifically, the processor 610 is used to execute the training program of the song synthesis model stored in the memory 620 to implement the following steps: respectively set a speech sample library and a song sample library; wherein, the number of speech samples in the speech sample library is greater than the number of song samples in the song sample library; use the speech sample library to train a basic model until the basic model converges; use the song sample library to retrain the basic model until the basic model converges again to obtain a song synthesis model.
[0103] Among them, the speech sample library includes a plurality of speech samples; wherein, each speech sample includes: a corresponding sample text and sample speech; the song sample library includes a plurality of song samples; wherein, each song sample includes: a corresponding lyric text, song audio and song score.
[0104] Among them, the step of using the speech sample library to train the basic model includes: sequentially obtaining a speech sample in the speech sample library; based on the sample text and sample speech in the speech sample, determining the pitch and duration corresponding to each character in the sample text; inputting the sample text and the pitch and duration corresponding to each character in the sample text into the basic model, and obtaining the predicted speech output by the basic model; determining the loss value of the basic model according to the sample speech in the speech sample and the predicted speech output by the basic model; adjusting the basic model according to the loss value of the basic model, and continuing to obtain the speech samples in the speech sample library to train the basic model until the basic model converges.
[0105] Among them, the step of determining the pitch and duration corresponding to each character in the sample text based on the sample text and sample speech in the speech sample includes: performing a phoneme alignment operation on the corresponding sample text and sample speech; extracting the fundamental frequency and duration corresponding to each character in the sample speech; for each character, determining the pitch corresponding to the character according to the fundamental frequency corresponding to the character, and determining the duration corresponding to the character according to the duration corresponding to the character.
[0106] Among them, the retraining of the basic model by using the song sample library includes: sequentially obtaining a song sample from the song sample library; wherein, the song sample includes: corresponding lyric text, song audio, and song score; the song score includes pitch and beat information; according to the pitch and beat information in the song score, determining the pitch and duration corresponding to each word in the lyric text; inputting the lyric text and the pitch and duration corresponding to each word in the lyric text into the basic model, and obtaining the predicted audio output by the basic model; determining the loss value of the basic model according to the song audio in the song sample and the predicted audio output by the basic model; adjusting the basic model according to the loss value of the basic model, and continuing to obtain the song samples in the song sample library to train the basic model until the basic model converges, so as to obtain a song synthesis model.
[0107] Among them, after obtaining the song synthesis model, it further includes: receiving the target lyric text and the pitch and duration corresponding to each word in the target lyric text; inputting the target lyric text and the pitch and duration corresponding to each word in the target lyric text into the song synthesis model, and obtaining the synthesized song audio output by the song synthesis model; displaying a synthesized song icon on a preset display interface; wherein, the synthesized song icon is used to trigger the playback of the synthesized song audio.
[0108] Among them, after receiving the target lyric text and the pitch and duration corresponding to each word in the target lyric text, it further includes: displaying the target lyric text and the pitch and duration corresponding to each word in the target lyric text on the display interface; detecting a song modification instruction; modifying the target lyric text, and / or modifying the pitch corresponding to a specified word in the target lyric text, and / or modifying the duration corresponding to a specified word in the target lyric text according to the detected song modification instruction.
[0109] The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs herein. Among them, the computer-readable storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include a combination of the above types of memory.
[0110] When one or more programs in the computer-readable storage medium can be executed by one or more processors to implement the above-mentioned method for training a song synthesis model.
[0111] Specifically, the processor is used to execute the training program of the song synthesis model stored in the memory to implement the following steps: respectively set a speech sample library and a song sample library; wherein, the number of speech samples in the speech sample library is greater than the number of song samples in the song sample library; use the speech sample library to train the basic model until the basic model converges; use the song sample library to retrain the basic model until the basic model converges again to obtain the song synthesis model.
[0112] Among them, the speech sample library includes a plurality of speech samples; wherein, each speech sample includes: a corresponding sample text and sample speech; the song sample library includes a plurality of song samples; wherein, each song sample includes: a corresponding lyric text, song audio, and song score.
[0113] Among them, the using the speech sample library to train the basic model includes: sequentially obtaining a speech sample in the speech sample library; based on the sample text and sample speech in the speech sample, determining the pitch and duration corresponding to each character in the sample text; inputting the sample text and the pitch and duration corresponding to each character in the sample text into the basic model, and obtaining the predicted speech output by the basic model; determining the loss value of the basic model according to the sample speech in the speech sample and the predicted speech output by the basic model; adjusting the basic model according to the loss value of the basic model, and continuing to obtain the speech samples in the speech sample library to train the basic model until the basic model converges.
[0114] Among them, the determining the pitch and duration corresponding to each character in the sample text based on the sample text and sample speech in the speech sample includes: performing a phoneme alignment operation on the corresponding sample text and sample speech; extracting the fundamental frequency and duration corresponding to each character in the sample speech; for each character, determining the pitch corresponding to the character according to the fundamental frequency corresponding to the character, and determining the duration corresponding to the character according to the duration corresponding to the character.
[0115] Among them, the retraining of the basic model by using the song sample library includes: sequentially obtaining a song sample from the song sample library; wherein, the song sample includes: corresponding lyric text, song audio, and song score; the song score includes pitch and beat information; determining the pitch and duration corresponding to each character in the lyric text according to the pitch and beat information in the song score; inputting the lyric text and the pitch and duration corresponding to each character in the lyric text into the basic model, and obtaining the predicted audio output by the basic model; determining the loss value of the basic model according to the song audio in the song sample and the predicted audio output by the basic model; adjusting the basic model according to the loss value of the basic model, and continuing to obtain song samples from the song sample library to train the basic model until the basic model converges, so as to obtain a song synthesis model.
[0116] Among them, after obtaining the song synthesis model, it further includes: receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text; inputting the target lyric text and the pitch and duration corresponding to each character in the target lyric text into the song synthesis model, and obtaining the synthesized song audio output by the song synthesis model; displaying a synthesized song icon on a preset display interface; wherein, the synthesized song icon is used to trigger the playback of the synthesized song audio.
[0117] Among them, after receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text, it further includes: displaying the target lyric text and the pitch and duration corresponding to each character in the target lyric text on the display interface; detecting a song modification instruction; modifying the target lyric text, and / or modifying the pitch corresponding to a specified character in the target lyric text, and / or modifying the duration corresponding to a specified character in the target lyric text according to the detected song modification instruction.
[0118] The above are only embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A training method for a song synthesis model, characterized in that, Including: A voice sample library and a song sample library are respectively set; wherein, the number of voice samples in the voice sample library is greater than the number of songs in the song sample library; Using the voice sample library to train a basic model until the basic model converges; Using the song sample library to retrain the basic model until the basic model converges again to obtain a song synthesis model; The voice sample library includes a plurality of voice samples; wherein, each voice sample includes: a corresponding sample text and sample voice; The song sample library includes a plurality of song samples; wherein, each song sample includes: a corresponding lyric text, song audio, and song score; The using the voice sample library to train the basic model includes: Sequentially obtaining a voice sample from the voice sample library; Based on the sample text and sample voice in the voice sample, determining the pitch and duration corresponding to each character in the sample text; Inputting the sample text and the pitch and duration corresponding to each character in the sample text into the basic model, and obtaining the predicted voice output by the basic model; Determining the loss value of the basic model according to the sample voice in the voice sample and the predicted voice output by the basic model; Adjusting the basic model according to the loss value of the basic model, and continuing to obtain voice samples from the voice sample library to train the basic model until the basic model converges; The using the song sample library to retrain the basic model includes: Sequentially obtaining a song sample from the song sample library; wherein, the song sample includes: a corresponding lyric text, song audio, and song score; the song score includes pitch and beat information; According to the pitch and beat information in the song score, determining the pitch and duration corresponding to each character in the lyric text; Inputting the lyric text and the pitch and duration corresponding to each character in the lyric text into the basic model, and obtaining the predicted audio output by the basic model; Determining the loss value of the basic model according to the song audio in the song sample and the predicted audio output by the basic model; Adjusting the basic model according to the loss value of the basic model, and continuing to obtain song samples from the song sample library to train the basic model until the basic model converges to obtain a song synthesis model.
2. The method according to claim 1, characterized in that, The determining the pitch and duration corresponding to each character in the sample text based on the sample text and sample voice in the voice sample includes: Performing a phoneme alignment operation on the corresponding sample text and sample voice; Extracting the fundamental frequency and duration corresponding to each character in the sample voice; For each character, determining the pitch corresponding to the character according to the fundamental frequency corresponding to the character, and determining the duration corresponding to the character according to the duration corresponding to the character.
3. The method according to any one of claims 1 to 2, characterized in that, After obtaining the song synthesis model, it further includes: Receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text; Input the target lyric text, as well as the pitch and duration corresponding to each character in the target lyric text, into the song synthesis model, and obtain the synthesized song audio output by the song synthesis model; On a preset display interface, display a synthesized song icon; wherein, the synthesized song icon is used to trigger the playback of the synthesized song audio.
4. The method according to claim 3, wherein After receiving the target lyric text and the pitch and duration corresponding to each character in the target lyric text, it further includes: On the display interface, display the target lyric text and the pitch and duration corresponding to each character in the target lyric text; Detect a song modification instruction; According to the detected song modification instruction, modify the target lyric text, and / or modify the pitch corresponding to a specified character in the target lyric text, and / or modify the duration corresponding to a specified character in the target lyric text.
5. A training device for a song synthesis model, characterized in that, It includes: A setting module for respectively setting a voice sample library and a song sample library; wherein, the number of voice samples in the voice sample library is greater than the number of songs in the song sample library; A training module for training a basic model using the voice sample library until the basic model converges; A retraining module for retraining the basic model using the song sample library until the basic model converges again to obtain a song synthesis model; The song sample library includes multiple song samples; wherein, each song sample includes: a corresponding lyric text, a song audio, and a song score; The training of the basic model using the voice sample library includes: Sequentially obtain a voice sample from the voice sample library; Based on the sample text and sample voice in the voice sample, determine the pitch and duration corresponding to each character in the sample text; Input the sample text and the pitch and duration corresponding to each character in the sample text into the basic model, and obtain the predicted voice output by the basic model; According to the sample voice in the voice sample and the predicted voice output by the basic model, determine the loss value of the basic model; According to the loss value of the basic model, adjust the basic model, and continue to obtain voice samples from the voice sample library to train the basic model until the basic model converges; The retraining of the basic model using the song sample library includes: Sequentially obtain a song sample from the song sample library; wherein, the song sample includes: a corresponding lyric text, a song audio, and a song score; the song score includes pitch and beat information; According to the pitch and beat information in the song score, determine the pitch and duration corresponding to each character in the lyric text; Input the lyric text and the pitch and duration corresponding to each character in the lyric text into the basic model, and obtain the predicted audio output by the basic model; According to the song audio in the song sample and the predicted audio output by the basic model, determine the loss value of the basic model; Adjust the basic model according to the loss value of the basic model, continue to obtain song samples from the song sample library, and train the basic model until the basic model converges to obtain a song synthesis model.
6. A training device for a song synthesis model, characterized in that, The training device of the song synthesis model includes a processor and a memory; the processor is configured to execute the training program of the song synthesis model stored in the memory to implement the training method of the song synthesis model according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the training method of the song synthesis model according to any one of claims 1 to 4.
Citation Information
Patent Citations
Training method and device of song synthesis model and song synthesis method and device
CN115273806A
Electronic musical instrument, electronic musical instrument control method, and storage medium
US20190392799A1