Song generation method and device, electronic equipment and medium
By selecting the target template and using the lyrics imitation model to imitate the lyrics and synthesize them in combination with the melody of the song, the problems of high quality and cost of song generation in the existing technology are solved, and high-quality and low-cost song generation are achieved.
Patent Information
- Application Number
- CN202510125282.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
The quality of existing multimodal models in song generation is difficult to guarantee, and the generation cost is high.
Select the target template through the song theme description, use the lyrics imitation model to imitate the target lyrics text based on the lyrics requirement description, and synthesize the imitation lyrics with the target song melody to generate a new song.
Improves the quality of song generation, reduces the cost of song generation, and realizes automatic generation from song theme to complete songs.
Smart Images

Figure CN120045741A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the fields of natural language processing, content generation, deep learning, and large models. Specifically, it relates to a method for generating songs. Background Art
[0002] The rise of multi-modal large models has brought new breakthroughs to the field of music generation. Songwriting is no longer a traditional, single, time-consuming, and highly talent-dependent process.
[0003] However, currently, multi-modal large models generally generate songs in an end-to-end manner, making it difficult to guarantee the quality of song generation and keeping the cost of song generation high. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, and medium for song generation.
[0005] According to one aspect of the present disclosure, there is provided a method for song generation, the method comprising:
[0006] Selecting a target template from at least two candidate templates according to a song theme description; wherein, the candidate templates are used to describe the correspondence between candidate lyric texts and candidate song melodies in song materials;
[0007] Using a lyric imitation model to imitate the target lyric text in the target template based on a lyric requirement description and output an imitated lyric; wherein, the number of lines and the number of characters per line of the imitated lyric are the same as those of the target lyric text; the lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song corpus;
[0008] Based on the correspondence between the target lyric text and the target song melody in the target template, synthesizing the imitated lyric and the target song melody in the target template to generate a new song.
[0009] According to another aspect of the present disclosure, there is provided a song generation apparatus, the apparatus comprising:
[0010] A template selection module, configured to select a target template from at least two candidate templates according to a song theme description; wherein, the candidate templates are used to describe the correspondence between candidate lyric texts and candidate song melodies in song materials;
[0011] A first imitation module, configured to use a lyric imitation model to imitate the target lyric text in the target template based on a lyric requirement description and output an imitated lyric; wherein, the number of lines and the number of characters per line of the imitated lyric are the same as those of the target lyric text; the lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song corpus;
[0012] A song generation module, configured to synthesize the imitated lyrics and the target song melody in the target template based on the correspondence between the target lyric text and the target song melody in the target template, so as to generate a new song.
[0013] According to another aspect of the present disclosure, there is provided an electronic device, which includes:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the song generation method according to any embodiment of the present disclosure.
[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the song generation method according to any embodiment of the present disclosure.
[0018] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the song generation method according to any embodiment of the present disclosure when executed by a processor.
[0019] The present disclosure can ensure the song generation quality and reduce the song generation cost.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings
[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0022] Figure 1 is a flowchart of a song generation method according to an embodiment of the present disclosure;
[0023] Figure 2 is a flowchart of another song generation method according to an embodiment of the present disclosure;
[0024] Figure 3 is a flowchart of yet another song generation method according to an embodiment of the present disclosure;
[0025] Figure 4It is a schematic structural diagram of a song generation device provided according to an embodiment of the present disclosure;
[0026] Figure 5 A block diagram of an electronic device for implementing the song generation method according to an embodiment of the present disclosure. Detailed implementation manners
[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0028] Figure 1 It is a flowchart of a song generation method provided according to an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of automatically generating songs based on song themes. This method can be executed by a song generation device, and the device can be implemented in software and / or hardware. As Figure 1 shown, the song generation method of this embodiment may include:
[0029] S101, select a target template from at least two candidate templates according to the song theme description; wherein, the candidate templates are used to describe the correspondence between candidate lyric texts and candidate song melodies in song materials.
[0030] S102, based on the lyric requirement description, imitate the target lyric text in the target template through a lyric imitation model and output the imitated lyrics; wherein, the number of lines and the number of characters per line of the imitated lyrics are the same as those of the target lyric text.
[0031] S103, based on the correspondence between the target lyric text and the target song melody in the target template, synthesize the imitated lyrics and the target song melody in the target template to generate a new song.
[0032] Among them, the song theme description is natural language used to describe the song theme. The song theme description can reflect which theme of song the user wants to generate. The song theme description is used as a reference for template selection, and based on the song theme description, a target template can be selected from the candidate templates. There are at least two candidate templates, and each candidate template has a definite song theme. The theme corresponding to the target template matches the song theme description.
[0033] The candidate template is used to describe the correspondence between the candidate lyric text and the candidate song melody in the song material. The lyric requirement description is the natural language used to describe the lyric imitation requirement. The lyric requirement description can reflect the dimensions from which the user wants to imitate the lyrics. The target template includes the target lyric text and the target song melody. The target lyric text serves as the basis for lyric imitation, and the lyric requirement description serves as the guidance for lyric imitation. By inputting the target lyric text and the lyric description requirement into the lyric imitation model, the lyric imitation model can, based on the target lyric text, imitate the lyrics according to the dimensions described in the lyric requirement description and output the imitated lyrics. The number of lines and the number of characters per line of the imitated lyrics are the same as those of the target lyric text.
[0034] Among them, the lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song lyric corpus. By learning the format and structural features of the song lyric corpus, the lyric imitation model has the ability to imitate the format and structure of the target lyric text to perform lyric imitation.
[0035] Among them, the large language model (LLM, Large Language Model) refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks, such as text classification, question answering, dialogue, etc. Through training, the large language model captures knowledge from a large amount of labeled and unlabeled data and stores the knowledge in a large number of parameters, and the model parameters can reach the level of tens of billions or hundreds of billions.
[0036] The target template records the correspondence between the target lyric text and the target song melody. This correspondence is reflected in that each character or phrase in the target lyric text usually matches a certain note or combination of notes in the target song melody. Since the number of lines and the number of characters per line of the imitated lyrics are the same as those of the target lyric text. Optionally, replace the target lyric text with the imitated lyrics and follow the correspondence between the target lyric text and the target song melody. Synthesize the imitated lyrics and the target song melody in the target template to generate a new song.
[0037] Among them, compared with the song material corresponding to the target template, the melody of the new song is the same but the lyric content is different.
[0038] The technical solution of the present disclosure combines template matching with a large language model for song generation. The template matching technology is used to select a target template that matches the song theme description from candidate templates. A lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song lyric corpus. Based on the lyric requirement description, the target lyric text in the target template is imitated to output the imitated lyrics. The imitated lyrics and the target song melody in the target template are synthesized to achieve song generation from the song theme to a complete song. By generating imitated lyrics that are consistent with the number of lines and the number of words per line of the target lyric text, the imitated lyrics can match the target lyric melody, effectively improving the quality of song generation. The lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song lyric corpus, without the need to redesign the model architecture, nor does it involve retraining and optimizing the model, greatly reducing the cost of song generation.
[0039] In an optional embodiment, the lyric imitation model includes a single-line imitation model and a multi-line imitation model; the single-line imitation model and the multi-line imitation model are respectively obtained by performing supervised fine-tuning on a pre-trained large language model using single-line song lyric corpus and multi-line song lyric corpus.
[0040] The imitated lyrics output by the single-line imitation model are consistent with the number of words per line of the target lyric text; the imitated lyrics output by the multi-line imitation model are consistent with the number of lines of the target lyric text, and the context semantics between the lyrics of each line in the imitated lyrics are consistent.
[0041] Among them, the single-line imitation model and the multi-line imitation model are two independent models, which can reduce the information interference between fine-tuning and can fine-tune a better imitation effect with less song lyric corpus. The song lyric corpus used by the single-line lyric imitation model and the multi-line lyric imitation model is different during the process of performing supervised fine-tuning on the pre-trained large language model. The single-line lyric imitation model uses single-line song lyric corpus, and the multi-line lyric imitation model uses multi-line song lyric corpus.
[0042] The multi-line lyric imitation model can ensure the context consistency between the lyrics of each line in the imitated lyrics, and can ensure the semantic integrity of the imitated lyrics to the greatest extent. The imitated lyrics output by the multi-line imitation model are consistent with the number of lines of the target lyric text. The single-line lyric imitation model pays more attention to the consistency between the number of words per line of the imitated lyrics and the target lyric text.
[0043] During the process of lyric imitation, the multi-line lyric imitation model and the single-line lyric imitation model are used in combination, so that the imitated song lyrics are semantically complete, and the number of lines and the number of words per line of the imitated lyrics are consistent with the target lyric text.
[0044] The language used for lyric imitation with the lyric imitation model is not limited here and is specifically determined according to actual business requirements. Here, taking the imitation of Chinese lyrics with the lyric imitation model as an example, the determination process of the lyric imitation model will be explained. Supervisory fine-tuning of a pre-trained large language model can include two links: basic learning and reinforcement learning.
[0045] First, prepare Chinese lyric corpus, then perform text data preprocessing on the Chinese lyric corpus, and then construct a single-line text dataset and a multi-line text dataset. Among them, the single-line text dataset includes single-line lyric corpus, and the multi-line text dataset includes multi-line lyric corpus. The pre-trained large language model is respectively subjected to basic learning using the single-line lyric corpus and the multi-line lyric corpus. After basic learning, the large language model can imitate lyrics in combination with new themes and ensure that the number of words in the lyrics is basically the same as that of the control sample.
[0046] Optionally, after basic learning, a positive and negative sample corpus is constructed using the large language model, and the fine-tuned model is further subjected to reinforcement learning using the positive and negative sample corpus. For example, the DPO algorithm (Direct Preference Optimization) is used to perform reinforcement learning on the large model. The sample labels of the positive and negative samples in the positive and negative sample corpus are determined according to the number of words. Exemplarily, when the control sample is "Hoeing the crops under the midday sun", "Sweat drips onto the soil beneath the crops" is a positive sample, and "Sweat drips onto the crops beneath" is a negative sample. After reinforcement learning, the number of lines and the number of words in the lyrics output by the large language model are strictly the same as those of the control sample. Thus, a single-line imitation model and a multi-line imitation model can be obtained, and the lyric imitation model can be determined.
[0047] In the above technical solution, the pre-trained large language model is respectively subjected to supervisory fine-tuning using the single-line lyric corpus and the multi-line lyric corpus to obtain independent single-line and multi-line imitation models, which can reduce information interference between fine-tuning, can obtain a better imitation effect with less lyric corpus, and is beneficial to reducing the cost of song generation. The multi-line lyric imitation model can ensure the context consistency between the lyrics of each line in the imitated lyrics, and can maximize the semantic integrity of the imitated lyrics. The number of lines of the imitated lyrics output by the multi-line imitation model is the same as that of the target lyric text. The single-line lyric imitation model pays more attention to the consistency of the number of words in each line between the imitated lyrics and the target lyric text. Cooperating the multi-line lyric imitation model and the single-line lyric imitation model for lyric generation is beneficial to improving the accuracy of lyric imitation and the quality of song generation.
[0048] Figure 2 It is a flowchart of another song generation method provided according to an embodiment of the present disclosure; this embodiment is an alternative solution proposed on the basis of the above embodiment.
[0049] See Figure 2 , the song generation method provided in this embodiment includes:
[0050] S201, select a target template from at least two candidate templates according to the song theme description; wherein, the candidate template is used to describe the correspondence between candidate lyric texts and candidate song melodies in song materials.
[0051] S202, extract at least one of the theme element, style element, and emotion element from the lyric requirement description as the imitation guiding parameter.
[0052] Among them, the lyric requirement description can reflect from which dimensions the user wants to imitate the lyrics. The theme element is used to limit the core content or story framework expressed by the imitated lyrics, the style element is used to limit the overall style presented by the imitated lyrics, and the emotion element is used to limit the emotional color contained in the imitated lyrics.
[0053] The imitation guiding parameter includes at least one of the theme element, style element, and emotion element. The imitation guiding parameter is provided to the lyric imitation model to guide the lyric imitation model to generate lyrics that meet specific themes, styles, and emotions.
[0054] S203, call the multi-line imitation model in the lyric imitation model, and use the imitation guiding parameter to imitate the target lyric text to output a draft lyric.
[0055] The draft lyric is generated by the multi-line imitation model in the lyric imitation model based on the imitation guiding parameter. The multi-line imitation model can ensure the semantic integrity of the draft lyric to the greatest extent, and the context semantics between the lyrics in each line of the draft lyric is consistent. The number of lines of the draft lyric is the same as that of the target lyric text. Optionally, the multi-line lyric imitation model is called using an api interface.
[0056] S204, compare the number of characters in each line of the draft lyric with the target lyric text one by one, and determine the line to be adjusted as the line with inconsistent number of characters between the draft lyric and the target lyric text.
[0057] The number of lines of the target lyric text is the same as that of the draft lyric. Compare the number of characters in each line of the target lyric text and the draft lyric one by one, and determine the line to be adjusted in the draft lyric.
[0058] In the line to be adjusted, the number of characters in the draft lyric is inconsistent with that in the target lyric text. It is necessary to adjust the content of the line to be adjusted in the draft lyric.
[0059] S205, call the single-line imitation model in the lyric imitation model, and use the imitation guiding parameter to re-imitate the lyric content corresponding to the line to be adjusted in the target lyric text to output a new lyric.
[0060] The number of words in each line of the lyrics output by the single-line imitation model is the same as that of the target lyrics text. When the line to be adjusted is determined, the single-line imitation model is called based on the API interface, and based on the lyrics content corresponding to the line to be adjusted in the target lyrics text, imitation is performed again according to the imitation guidance parameters.
[0061] The number of words in the new lyrics is the same as that of the lyrics content corresponding to the line to be adjusted in the target lyrics text. Both the new lyrics and the initial draft of the lyrics are obtained by imitation based on the imitation guidance parameters. Therefore, the new lyrics are consistent with the initial draft of the lyrics in terms of theme, style, and mood.
[0062] S206, determine the imitated lyrics based on the new lyrics and the initial draft of the lyrics; wherein, the number of lines and the number of words in each line of the imitated lyrics are the same as those of the target lyrics text.
[0063] Optionally, replace the lyrics content corresponding to the line to be adjusted in the initial draft of the lyrics with the new lyrics to obtain the imitated lyrics. In this way, the number of lines and the number of words in each line of the imitated lyrics are the same as those of the target lyrics text.
[0064] S207, based on the correspondence between the target lyrics text and the target song melody in the target template, synthesize the imitated lyrics and the target song melody in the target template to generate a new song.
[0065] In the technical solution of the present disclosure, first, the multi-line imitation model in the lyrics imitation model is called, and the target lyrics text is imitated using the imitation guidance parameters to output the initial draft of the lyrics. The semantic integrity of the initial draft of the lyrics is guaranteed to the greatest extent, and it is ensured that the number of lyrics lines in the initial draft of the lyrics is the same as that of the target lyrics text. If there are lines in the initial draft of the lyrics with a different number of words from those in the target lyrics text, the single-line imitation model in the lyrics imitation model is called, and the lyrics content corresponding to the line to be adjusted in the target lyrics text is re-imitated using the imitation guidance parameters to output the new lyrics. It is ensured that the number of words in the new lyrics is the same as that of the lyrics content corresponding to the line to be adjusted in the target lyrics text. Based on the new lyrics and the initial draft of the lyrics, the imitated lyrics are determined. The technical solution of the present disclosure provides a practical lyrics imitation solution. Through the mutual cooperation of the multi-line lyrics imitation model and the single-line imitation model, lyrics imitation that conforms to a specific theme, style, and mood is realized. It is ensured that the number of lines and the number of words in each line of the imitated lyrics are the same as those of the target lyrics text. This makes it possible to synthesize the imitated lyrics and the target song melody in the target template subsequently, enabling the imitated lyrics to match the target lyrics melody, which is beneficial to improving the quality of song generation.
[0066] In an optional embodiment, the method further includes: if the number of words in each line of the initial draft of the lyrics is the same as that of the target lyrics text, determine the imitated lyrics based on the initial draft of the lyrics.
[0067] The number of lines in the initial draft of the lyrics is the same as that of the target lyrics text. If the number of words in each line of the initial draft of the lyrics is the same as that of the target lyrics text, then the number of lines and the number of words in each line of the initial draft of the lyrics are the same as those of the target lyrics text. Meeting the necessary conditions for synthesis with the target song melody in the target template can ensure that the initial draft of the lyrics can match the target lyrics melody. Optionally, the initial draft of the lyrics is used as the imitated lyrics.
[0068] The above technical solution provides a practical lyrics imitation solution, which is applicable to the situation where the number of words in each line of the initial draft of the lyrics is the same as that of the target lyrics text, and expands the applicable scenarios of the lyrics imitation solution.
[0069] In an optional embodiment, the method further includes: obtaining user feedback on the imitated lyrics; modifying the imitated lyrics based on the modification suggestions in the user feedback to obtain a modification result; and updating the imitated lyrics based on the modification result.
[0070] Optionally, after the lyrics imitation model outputs the imitated lyrics, the imitated lyrics are displayed. The user feedback is used to determine the user's satisfaction with the imitated lyrics. Optionally, the user feedback includes modification suggestions, where the modification suggestions are used to modify the imitated lyrics. Optionally, the modification suggestions are in natural language.
[0071] In the case where the user feedback includes modification suggestions, the imitated lyrics are modified using the modification suggestions to obtain a modification result. The imitated lyrics are updated using the modification result. The above technical solution supports users to modify according to their needs after the lyrics imitation model generates the imitated lyrics, which can improve the user's satisfaction with the imitated lyrics and is beneficial to improving the user experience.
[0072] In an optional embodiment, the target template is constructed based on template metadata, the template metadata is in JSON format, and the template metadata includes a timestamp array and a composite array;
[0073] Among them, the array elements in the timestamp array are the timestamps corresponding to the target lyrics text; the array elements in the composite array are obtained by combining the target lyrics text and the target note sequence based on the note durations corresponding to the target note sequence; among them, the target note sequence and the note durations corresponding to the target note sequence are obtained by performing note feature extraction on the target song melody.
[0074] Among them, the template metadata is in JSON format. JSON is a lightweight data interchange format, and the template metadata is used to construct a song template. The number of elements in the array elements of the timestamp array is determined according to the number of lines of the target lyric text. The number of elements in the array elements of the combination array is also determined according to the number of lines of the target lyric text. The timestamps corresponding to the target lyric text are used to align the target lyric text and the target note sequence. The target note sequence is the basic component of the target song melody. The target note sequence includes at least two note units. The note duration corresponding to the target note sequence refers to the relative duration between each note unit. Based on the note duration corresponding to the target note sequence, the words in the lyric text fragments can be combined with the note units in the note sequence fragments.
[0075] Optionally, MIDI (Musical Instrument Digital Interface) vector information is extracted from the human voice audio in the target song melody to obtain the target note sequence and the note duration corresponding to the target note sequence. Optionally, a digital audio processing tool such as DAW (Digital Audio Workstation) is used to align the note units and the note duration in the target note sequence. Optionally, after the template metadata is determined, the template metadata is stored in a database for subsequent use.
[0076] The above technical solution provides a feasible song template construction solution. By constructing the target template through the template metadata, the separation between the lyrics and the melody is achieved, making it possible to synthesize the imitated lyrics and the target song melody in the target template to generate a new song later, which is beneficial to improving the song generation quality.
[0077] In an optional embodiment, the method further includes: determining the rest positions of the target song melody based on the note durations of the note units in the target note sequence; determining the sentence-breaking positions of the target lyric text; respectively based on the rest positions and the sentence-breaking positions, splitting the target note sequence and the target lyric text to obtain lyric text fragments and note sequence fragments; based on the note duration corresponding to the note sequence fragments and the timestamps corresponding to the lyric text fragments, combining the lyric text fragments and the note sequence fragments, and writing the obtained combination result as an array element into the combination array.
[0078] Among them, the target note sequence includes at least two note units, and the note duration is used to describe the relative duration between each note unit. The rest position is determined by a rest, and is used to identify the pause of the target song melody. At the rest position of the target song melody, there is no corresponding note unit. Based on the rest position, the target note sequence can be segmented into at least two note sequence segments. The sentence-breaking positions of the target lyric text are usually determined according to the pause, rhythm, and phrase structure of the target song melody. The sentence-breaking positions of the target lyric text correspond to the timestamps corresponding to the target lyric text.
[0079] In the target lyric text, the sentence-breaking positions are often closely related to the pause of the target song melody. Based on the sentence-breaking positions, the target note sequence can be segmented into at least two lyric text segments. The number of note sequence segments is the same as the number of lyric text segments. The length of the note sequence segment is related to the number of note units in the note sequence segment, and the length of the lyric text segment is related to the number of lyric words in the lyric text segment. The length of the note sequence segment is consistent with the length of the lyric text segment.
[0080] Each lyric text segment has a corresponding timestamp, and the timestamp is used to determine the correspondence between the lyric text segment and the note sequence segment. The note duration corresponding to the note sequence segment is used to determine the correspondence between the words in the lyric text segment and the note units in the note sequence segment.
[0081] Optionally, first determine the note sequence segment that matches the lyric text segment based on the timestamp corresponding to the lyric text segment. When the correspondence between the lyric text segment and the note sequence segment is determined, based on the note duration corresponding to the note sequence segment, correspond the words in the lyric text segment with the note units in the note sequence segment to achieve the combination of the lyric text segment and the note sequence segment. Write the combination result as an array element into the combination array.
[0082] Exemplarily, the template metadata corresponding to the target template may include a timestamp array and a combination array. Among them, the timestamp array can be identified by the "lead_time" field, for example, "lead_time":[7.3708,11.2000,15.0229,18.8375], and the combination array can be identified by the "inputs_list" field. The array elements in the combination array correspond one by one to the array elements in the timestamp array. When the timestamp array lead_time includes the above 4 array elements, the combination array inputs_list will also include 4 array elements, corresponding to the 4 array elements in the timestamp array in sequence.
[0083] Taking the first array element in the combined array inputs_list as an example, the content of the array element in the combined array inputs_list is explained, and the contents of the other three array elements in the combined array are not repeated here. The first array element in the combined array inputs_list corresponds to 7.3708 (the first array element) in the timestamp array.
[0084] The first array element in the composite array inputs_list is {
[0085] "input_type":"word","text":"AP The old friend bids farewell to the Yellow Crane Tower","notes":"rest|C4|D4|E4|E4|G4|A3|D4","notes_duration":"0.1000|0.1615|0.3285|0.3358|0.3242|0.2083|0.4708|0.7667"}. Among them, "input_type":"word", is used to identify the sentence segmentation position of the target lyrics text, "text":"AP Guren Xici Huanghelou" refers to the lyrics text segment, "notes":"rest|C4|D4|E4|E4|G4|A3|D4" refers to the audio sequence segment, and "notes_duration":"0.1000|0.1615|0.3285|0.3358|0.3242|0.2083|0.4708|0.7667" refers to the note duration corresponding to the note sequence segment. Among them, AP in the "text" field, rest in the "note" field, and 0.1000 in the "notes_duration" field all correspond to the rest position of the target song melody.
[0086] It is worth noting that the timestamp array and combination array in the above examples are only used as examples and do not limit the technical solutions provided by the present disclosure. The element content and number of array elements in the timestamp array and the content and number of array elements in the combination array are not limited here and are determined based on actual conditions.
[0087] Optionally, a lyrics text array is added to the template metadata corresponding to the target template to store the pure version of the target lyrics text. It is worth noting that the pure version of the target lyrics does not include characters used to indicate the position of sentence breaks and rest positions. Continuing with the above example, the lyrics text array can be identified by the lyrics_sample field, specifically expressed as, "lyrics_sample":["The old friend bids farewell to the Yellow Crane Tower","Fireworks in March in Yangzhou","The lonely sail is far away and the blue sky ends","Only the Yangtze River is seen flowing across the sky"]. Adding the lyrics text array to the template metadata corresponding to the target template provides convenience for subsequent lyrics imitation based on the target lyrics text and proofreading the number of words in each line of lyrics.
[0088] The above technical solution provides a feasible combination array determination solution, which can be used to determine the template metadata corresponding to the target template, and provides data support for the subsequent lyrics imitation based on the target lyrics text in the target template, and the synthesis of the imitated lyrics and the target song melody in the target template to generate a new song.
[0089] In a specific embodiment, the template metadata corresponding to the target template is determined by the following steps:
[0090] 1. Obtain song material, use the vocal separation model to separate the track of the song material, split the song material into vocal audio and background music; at the same time, perform text extraction on the song material to obtain lyrics text. 2. Perform note feature extraction on the vocal audio, for example, MIDI vector information extraction can be performed. The extraction scheme can be to use the DIO algorithm to identify the fundamental frequency F0 of the vocal audio, and mark the starting position and end position of each F0 at the same time. After statistics, use the MIDI processing module to reassemble and generate a MIDI file. The MIDI file includes a note sequence and the note duration corresponding to the note sequence. Among them, the DIO algorithm is a speech signal processing algorithm based on Mel Frequency Cepstrum Transform (MFCC). 3. Use digital audio processing tools such as DAW (Digital Audio Workstation) to calibrate the MIDI file. Specifically, in the MIDI file, align the note units and note durations in the note sequence. 4. Use the calibrated MIDI file and combine it with the lyrics text in step 1 to obtain template metadata. 5. Associate the background music with the template metadata, and then store the template metadata associated with the background music in the database.
[0091] Figure 3 It is a flowchart of another song generation method provided according to an embodiment of the present disclosure; this embodiment is an optional solution proposed on the basis of the above embodiment.
[0092] See also Figure 3, the song generation method provided in this embodiment includes:
[0093] S301, select a target template from at least two candidate templates according to the song theme description; wherein, the candidate template is used to describe the correspondence between the candidate lyric text and the candidate song melody in the song material.
[0094] S302, based on the lyric requirement description, imitate the target lyric text in the target template through a lyric imitation model and output the imitated lyrics; wherein, the number of lines and the number of characters per line of the imitated lyrics are the same as those of the target lyric text.
[0095] S303, based on the template metadata corresponding to the target template, determine the lyric text fragments corresponding to the target lyric text, the note sequence fragments corresponding to the target song melody, as well as the note durations corresponding to the note sequence fragments and the timestamps corresponding to the lyric text fragments.
[0096] The template metadata corresponding to the target template includes a timestamp array and a combination array. The timestamp array includes the timestamps corresponding to the lyric text fragments.
[0097] The array elements in the combination array are obtained by combining the target lyric text and the target note sequence based on the note durations corresponding to the target note sequence.
[0098] S304, split the imitated lyrics according to the sentence-breaking positions in the target lyric text to obtain imitated lyric fragments.
[0099] The number of lines and the number of characters per line of the imitated lyrics are the same as those of the target lyric text, and the sentence-breaking positions of the target lyric text can be reused for the imitated lyrics. Based on the sentence-breaking positions of the target lyric text, the imitated lyrics are split to obtain imitated lyric fragments. Among them, the number of imitated lyric fragments is the same as the number of lyric text fragments obtained by splitting the target lyric text, and the number of characters included in each fragment is also the same.
[0100] S305, based on the note durations corresponding to the note sequence fragments and the timestamps corresponding to the lyric text fragments, replace the lyric text fragments in the template metadata with the imitated lyric fragments to obtain new metadata.
[0101] Based on the timestamps corresponding to the lyric text fragments, determine the note sequence fragments that match the imitated lyric fragments. In the case where the pairs of imitated lyric fragments and note sequence fragments are determined, based on the note durations corresponding to the note sequence fragments, correspond the words in the imitated lyric fragments to the note units in the note sequence fragments to obtain new metadata.
[0102] S306, generate a new song based on the new metadata.
[0103] Optionally, a singing synthesis model is used to synthesize the new metadata into a new human voice audio, and then an audio effect processor is used to mix and paste the new human voice audio and the background music associated with the target template to obtain a new song.
[0104] Continuing with the above example, when the imitated lyrics are "New friends come to the Yueyang Tower by the lake, where the autumn moon reflects in your eyes. The vast waves and cloud shadows disperse, and only the tranquility of Dongting Lake can be heard.", the imitated lyrics can be segmented into 4 imitated lyric segments, namely "New friends come to the Yueyang Tower by the lake", "where the autumn moon reflects in your eyes", "The vast waves and cloud shadows disperse", and "and only the tranquility of Dongting Lake can be heard".
[0105] Based on the note durations corresponding to the note sequence segments and the timestamps corresponding to the lyric text segments, the lyric text segments in the template metadata are replaced with the imitated lyric segments to obtain new metadata. Continuing with the above example, the timestamp array in the new metadata is still identified by the "lead_time" field and is still "lead_time":[7.3708,11.2000,15.0229,18.8375]. The combination array continues to be identified by the "inputs_list" field, and the array elements in the combination array correspond one by one to the array elements in the timestamp array. When the timestamp array lead_time includes the above 4 array elements, the combination array inputs_list will also include 4 array elements, corresponding to the 4 array elements in the timestamp array in sequence.
[0106] Compared with the template metadata of the target template, the array elements in the combination array inputs_list of the new metadata have changed. Specifically, the lyric text segments in the array elements are replaced with the imitated lyric segments.
[0107] Still taking the first array element in the combination array inputs_list as an example, the content of the array elements in the combination array inputs_list is described. The content of the other 3 array elements in the combination array will not be elaborated here.
[0108] For the new metadata, the first array element in the combination array inputs_list becomes {
[0109] "input_type":"word","text":"AP New friend comes to the East Yueyang Tower","notes":"rest|C4|D4|E4|E4|G4|A3|D4","notes_duration":"0.1000|0.1615|0.3285|0.3358|0.3242|0.2083|0.4708|0.7667"}. Compared with the template metadata corresponding to the target template, the content identified by the "text" field in the array element has changed from "AP Old friend takes leave of the Yellow Crane Tower to the west" to "AP New friend comes to the East Yueyang Tower". Other fields such as "input_type", "notes", and "notes_duration" have not changed.
[0110] This disclosure utilizes the characteristic that the imitated lyrics are consistent with the target lyrics in the number of text lines and the number of characters per line. According to the sentence-breaking positions in the target lyrics text, the imitated lyrics are segmented to obtain imitated lyrics segments. Based on the note durations corresponding to the note sequence segments and the timestamps corresponding to the lyrics text segments in the template metadata, the lyrics text segments in the template metadata are replaced with the imitated lyrics segments to obtain new metadata. Based on the new metadata, a new song is generated. This enables the imitated lyrics to match the melody of the target lyrics, effectively improving the quality of song generation.
[0111] Figure 4 It is a schematic structural diagram of a song generation device provided by an embodiment of this disclosure. The embodiments of this disclosure are applicable to the situation of automatically generating songs based on a song theme. This device can be implemented using software and / or hardware, and this device can implement the song generation method described in any embodiment of this disclosure.
[0112] As Figure 4 shown, the song generation device 400 includes:
[0113] A template selection module 401, configured to select a target template from at least two candidate templates according to a song theme description; wherein, the candidate templates are used to describe the correspondence between candidate lyrics text and candidate song melodies in song materials;
[0114] A first imitation module 402, configured to imitate the target lyrics text in the target template based on a lyrics requirement description through a lyrics imitation model and output imitated lyrics; wherein, the imitated lyrics are consistent with the target lyrics text in the number of text lines and the number of characters per line; the lyrics imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song corpus;
[0115] The song generation module 403 is configured to synthesize the imitated lyrics and the target song melody in the target template based on the correspondence between the target lyric text and the target song melody in the target template, so as to generate a new song.
[0116] The technical solution of the present disclosure combines template matching with a large language model for song generation. The template matching technology is used to select a target template that matches the song theme description from candidate templates. The lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song lyric corpus. Based on the lyric requirement description, the target lyric text in the target template is imitated to output imitated lyrics. The imitated lyrics and the target song melody in the target template are synthesized to achieve song generation from the song theme to a complete song. By generating imitated lyrics that are consistent with the number of lines and the number of words per line of the target lyric text, the imitated lyrics can be matched with the target lyric melody, effectively improving the quality of song generation. The lyric imitation model is obtained by performing supervised fine-tuning on a pre-trained large language model using song lyric corpus, without the need to redesign the model architecture, nor does it involve retraining and optimizing the model, greatly reducing the cost of song generation.
[0117] Optionally, the lyric imitation model includes a single-line imitation model and a multi-line imitation model; the single-line imitation model and the multi-line imitation model are respectively obtained by performing supervised fine-tuning on a pre-trained large language model using single-line song lyric corpus and multi-line song lyric corpus; the imitated lyrics output by the single-line imitation model are consistent with the number of words per line of the target lyric text; the imitated lyrics output by the multi-line imitation model are consistent with the number of lines of the target lyric text, and the context semantics between the lyrics of each line in the imitated lyrics are consistent.
[0118] Optionally, the first imitation module 402 includes: a parameter extraction sub-module, configured to extract at least one of a theme element, a style element, and an emotion element from the lyric requirement description as an imitation guidance parameter; a draft output sub-module, configured to call the multi-line imitation model in the lyric imitation model, and use the imitation guidance parameter to imitate the target lyric text to output a lyric draft; a word count comparison sub-module, configured to perform a line-by-line word count comparison on the lyric draft based on the target lyric text, and determine the lines with inconsistent word counts between the lyric draft and the target lyric text as lines to be adjusted; a lyric rewriting sub-module, configured to call the single-line imitation model in the lyric imitation model, and use the imitation guidance parameter to re-imitate the lyric content corresponding to the line to be adjusted in the target lyric text to output new lyrics; a lyric determination sub-module, configured to determine the imitated lyrics based on the new lyrics and the lyric draft.
[0119] Optionally, the device 400 further includes: a second imitation module, specifically configured to determine an imitated lyric based on the initial draft of the lyric if the number of words in each line of the initial draft of the lyric is the same as that of the target lyric text.
[0120] Optionally, the device 400 further includes: a feedback acquisition module, configured to acquire user feedback on the imitated lyric; a lyric modification module, configured to modify the imitated lyric based on the modification suggestions in the user feedback to obtain a modification result; and a lyric update module, configured to update the imitated lyric based on the modification result.
[0121] Optionally, the target template is constructed based on template metadata, the template metadata is in JSON format, and the template metadata includes a timestamp array and a combination array; wherein, the array elements in the timestamp array are the timestamps corresponding to the target lyric text; the array elements in the combination array are obtained by combining the target lyric text and the target note sequence based on the note durations corresponding to the target note sequence; wherein, the target note sequence and the note durations corresponding to the target note sequence are obtained by performing note feature extraction on the target song melody.
[0122] Optionally, the device 400 further includes: a rest position determination module, configured to determine the rest positions of the target song melody based on the note durations of the note units in the target note sequence; a sentence segmentation position determination module, configured to determine the sentence segmentation positions of the target lyric text; a data segmentation module, configured to segment the target note sequence and the target lyric text respectively based on the rest positions and the sentence segmentation positions to obtain a lyric text segment and a note sequence segment; and a data combination module, configured to combine the lyric text segment and the note sequence segment based on the note durations corresponding to the note sequence segment and the timestamps corresponding to the lyric text segment, and write the obtained combination result as an array element into the combination array.
[0123] Optionally, the song generation module 403 includes: a data determination sub-module, configured to determine the lyric text segment corresponding to the target lyric text, the note sequence segment corresponding to the target song melody, the note durations corresponding to the note sequence segment, and the timestamps corresponding to the lyric text segment based on the template metadata corresponding to the target template; a lyric segmentation sub-module, configured to segment the imitated lyric according to the sentence segmentation positions in the target lyric text to obtain imitated lyric segments; a lyric replacement sub-module, configured to replace the lyric text segment in the template metadata with the imitated lyric segments based on the note durations corresponding to the note sequence segment and the timestamps corresponding to the lyric text segment to obtain new metadata; and a song generation sub-module, configured to generate a new song based on the new metadata.
[0124] The song generation device provided by the embodiments of the present disclosure can execute the song generation method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the song generation method.
[0125] In the technical solution of the present disclosure, the song materials and user data involved, such as the collection, storage, use, processing, transmission, provision, and disclosure of song theme descriptions, lyric requirement descriptions, and user feedback, all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0126] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0127] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0128] As Figure 5 shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0129] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0130] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the song generation method. For example, in some embodiments, the song generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the song generation method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the song generation method by any other suitable means (e.g., by means of firmware).
[0131] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0132] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable song generation devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0133] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0134] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0135] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0136] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0137] Artificial intelligence is a discipline that studies the use of computers to simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has technologies at both the hardware and software levels. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0138] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0139] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0140] The above specific implementation manners do not constitute a limitation to the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A song generation method, the method comprising: According to the song theme description, a target template is selected from at least two candidate templates; wherein the candidate template is used to describe the corresponding relationship between the candidate lyrics text in the song material and the candidate song melody; The target lyrics text in the target template is imitated and the imitated lyrics are outputted based on the lyrics requirement description by the lyrics imitation model; wherein the imitated lyrics have the same number of lines and the number of words per line as the target lyrics text; the lyrics imitation model is obtained by fine-tuning the pre-trained large language model using lyrics corpus; Based on the correspondence between the target lyrics text and the target song melody in the target template, the imitation lyrics and the target song melody in the target template are synthesized to generate a new song.
2. The method according to claim 1, wherein: The lyrics imitation model includes a single-line imitation model and a multi-line imitation model; the single-line imitation model and the multi-line imitation model are respectively obtained by fine-tuning a pre-trained large language model using a single-line lyrics corpus and a multi-line lyrics corpus; The imitation lyrics output by the single-line imitation model have the same number of words per line as the target lyrics text; The imitated lyrics output by the multi-line imitation model have the same number of lyrics lines as the target lyrics text, and the contextual semantics between each line of lyrics in the imitated lyrics are consistent.
3. The method according to claim 1, wherein: The method of imitating the target lyrics text in the target template and outputting the imitative lyrics based on the lyrics requirement description by the lyrics imitation model includes: Extracting at least one of a theme element, a style element and an emotion element from the lyrics requirement description as a simulation writing guiding parameter; Calling a multi-line imitation model in the lyrics imitation model, using the imitation guidance parameters to imitate the target lyrics text, and outputting a first draft of the lyrics; Based on the target lyrics text, the word count of the first draft lyrics is compared line by line, and the lines with inconsistent word counts between the first draft lyrics and the target lyrics text are determined as lines to be adjusted; Calling a single-line imitation model in the lyrics imitation model, using the imitation guidance parameters to re-write the lyrics content corresponding to the line to be adjusted in the target lyrics text, and outputting new lyrics; Based on the new lyrics and the first draft of the lyrics, the parody lyrics are determined.
4. The method according to claim 3, further comprising: If the number of words in each line of the lyrics in the first draft of the lyrics is consistent with that in the target lyrics text, the imitated lyrics are determined based on the first draft of the lyrics.
5. The method according to claim 1, further comprising: Obtaining user feedback on the imitation lyrics; Based on the modification suggestions in the user feedback, modify the imitation lyrics to obtain a modification result; Based on the modification result, the imitation lyrics are updated.
6. The method according to claim 1, wherein: The target template is constructed based on template metadata, the template metadata is in JSON format, and the template metadata includes a timestamp array and a combination array; Among them, the array elements in the timestamp array are the timestamps corresponding to the target lyrics text; the array elements in the combination array are obtained by combining the target lyrics text and the target note sequence based on the note duration corresponding to the target note sequence; wherein the target note sequence and the note duration corresponding to the target note sequence are obtained by extracting the note features of the target song melody.
7. The method according to claim 6, further comprising: Determining the rest position of the target song melody based on the note duration of each note unit in the target note sequence; Determine the sentence segmentation position of the target lyrics text; Based on the rest positions and the segmentation positions, the target note sequence and the target lyrics text are segmented to obtain lyrics text segments and note sequence segments; Based on the note duration corresponding to the note sequence segment and the timestamp corresponding to the lyric text segment, the lyric text segment and the note sequence segment are combined, and the obtained combination result is written into the combination array as an array element.
8. The method according to claim 6, wherein: Based on the correspondence between the target lyrics text and the target song melody in the target template, the imitation lyrics and the target song melody in the target template are synthesized to generate a new song, including: Based on the template metadata corresponding to the target template, determine the lyrics text segments corresponding to the target lyrics text and the note sequence segments corresponding to the target song melody, as well as the note durations corresponding to the note sequence segments and the timestamps corresponding to the lyrics text segments; According to the sentence segmentation positions in the target lyrics text, the imitation lyrics are segmented to obtain imitation lyrics segments; Based on the note duration corresponding to the note sequence segment and the timestamp corresponding to the lyrics text segment, the lyrics text segment in the template metadata is replaced with the imitation lyrics segment to obtain new metadata; Based on the new metadata, a new song is generated.
9. A song generation device, comprising: A template selection module, used to select a target template from at least two candidate templates according to the song theme description; wherein the candidate template is used to describe the corresponding relationship between the candidate lyrics text in the song material and the candidate song melody; The first imitation module is used to imitate the target lyrics text in the target template based on the lyrics requirement description through the lyrics imitation model and output the imitated lyrics; wherein the imitated lyrics have the same number of lines and the number of words per line as the target lyrics text; the lyrics imitation model uses the lyrics corpus to supervise and fine-tune the pre-trained large language model; The song generation module is used to synthesize the imitation lyrics and the target song melody in the target template based on the correspondence between the target lyrics text in the target template and the target song melody to generate a new song.
10. The device according to claim 9, wherein: The lyrics imitation model includes a single-line imitation model and a multi-line imitation model; the single-line imitation model and the multi-line imitation model are respectively obtained by fine-tuning a pre-trained large language model using a single-line lyrics corpus and a multi-line lyrics corpus; The imitation lyrics output by the single-line imitation model have the same number of words per line as the target lyrics text; The imitated lyrics output by the multi-line imitation model have the same number of lyrics lines as the target lyrics text, and the contextual semantics between each line of lyrics in the imitated lyrics are consistent.
11. The device according to claim 9, wherein: The first imitation writing module comprises: A parameter extraction submodule, used to extract at least one of a theme element, a style element and an emotion element from the lyrics requirement description as a simulation writing guiding parameter; A first draft output submodule is used to call the multi-line imitation model in the lyrics imitation model, use the imitation guide parameters to imitate the target lyrics text, and output the first draft of the lyrics; A word count comparison submodule is used to compare the word count of the first draft of lyrics line by line based on the target lyrics text, and determine the lines with inconsistent word counts between the first draft of lyrics and the target lyrics text as lines to be adjusted; A lyrics rewriting submodule is used to call the single-line imitation model in the lyrics imitation model, use the imitation guidance parameters to rewrite the lyrics content corresponding to the line to be adjusted in the target lyrics text, and output new lyrics; The lyrics determination submodule is used to determine the imitation lyrics based on the new lyrics and the first draft of the lyrics.
12. The device according to claim 11, further comprising: The second imitation writing module is specifically used to determine the imitation lyrics based on the first draft of lyrics if the number of words in each line of lyrics of the first draft of lyrics is consistent with that of the target lyrics text.
13. The apparatus according to claim 9, further comprising: A feedback acquisition module, used to acquire user feedback on the imitation lyrics; A lyrics modification module, used to modify the imitation lyrics to obtain a modification result based on the modification suggestions in the user feedback; The lyrics updating module is used to update the imitation lyrics based on the modification result.
14. The device according to claim 9, wherein: The target template is constructed based on template metadata, the template metadata is in JSON format, and the template metadata includes a timestamp array and a combination array; Among them, the array elements in the timestamp array are the timestamps corresponding to the target lyrics text; the array elements in the combination array are obtained by combining the target lyrics text and the target note sequence based on the note duration corresponding to the target note sequence; wherein the target note sequence and the note duration corresponding to the target note sequence are obtained by extracting the note features of the target song melody.
15. The device according to claim 14, further comprising: A rest position determination module, used to determine the rest position of the target song melody based on the note duration of each note unit in the target note sequence; A sentence segmentation position determination module, used to determine the sentence segmentation position of the target lyrics text; A data segmentation module, used for segmenting the target note sequence and the target lyrics text based on the rest position and the sentence segmentation position, respectively, to obtain lyrics text segments and note sequence segments; The data combination module is used to combine the lyrics text segment and the note sequence segment based on the note duration corresponding to the note sequence segment and the timestamp corresponding to the lyrics text segment, and write the obtained combination result as an array element into the combination array.
16. The device according to claim 14, wherein: Song generation module, including: A data determination submodule, for determining, based on the template metadata corresponding to the target template, lyrics text segments corresponding to the target lyrics text and note sequence segments corresponding to the target song melody, as well as note durations corresponding to the note sequence segments and timestamps corresponding to the lyrics text segments; A lyrics segmentation module is used to segment the imitation lyrics according to the sentence segmentation positions in the target lyrics text to obtain imitation lyrics segments; A lyrics replacement submodule, for replacing the lyrics text segment in the template metadata with the imitation lyrics segment based on the note duration corresponding to the note sequence segment and the timestamp corresponding to the lyrics text segment, to obtain new metadata; The song generation submodule is used to generate a new song based on the new metadata.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the song generation method according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the song generation method according to any one of claims 1-8.
19. A computer program product, comprising a computer program, which, when executed by a processor, implements the song generation method according to any one of claims 1 to 8.
Citation Information
Cited By
Song imitation writing method and device, electronic equipment and storage medium
CN116453489A
A method, apparatus, electronic device, and storage medium for song imitation
CN116453489B