Lyric or song creation method and system based on large language model
By building a track library and quality discriminator, the generation process of large language models is solved, and the semantic deviation and rhythmic adaptation problems in AI lyrics creation are achieved, high-quality lyrics or song generation is achieved, adapting to a specific music style and automatically evaluating quality.
Patent Information
- Application Number
- CN202510512184.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
The existing AI lyrics creation has semantic deviations from the theme, templated emotional expression, and difficult to adapt to the rhythmic requirements of a specific musical style. Quality assessment relies on manual review, resulting in the lack of literary and musicality of generated lyrics, which makes it difficult to reach the level of human artists.
Build a library of tracks and classify the style of the song, optimize the generation process of the large language model through the sparse attention mechanism and quality discriminator, filter candidate tracks using dynamic sparse encoding and similarity calculation, and build multiple rounds of prompt word templates for iterative generation until the quality is achieved.
The quality and adaptability of the lyrics output from the large language model are improved, ensuring that the generated content matches the music style, and automatic quality evaluation and iterative optimization are realized, and the generated lyrics or songs reach the level of human artists.
Smart Images

Figure CN120386892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital music development and production, and specifically discloses a method and system for lyric or song creation based on a large language model. Background Art
[0002] The core value of lyric creation lies in constructing a balance between phonetic beauty and emotional resonance through refined language: it is necessary to satisfy the literary narrative of "beginning, development, turning point, and conclusion", and also need to adapt to the "level and oblique tones rhyming" rules of the musical score rhythm. Traditional manual lyric creation has always had the characteristics of "high threshold and great difficulty". However, with the development of AI (Artificial Intelligence) technology and LLM (large language model) technology in recent years, such characteristics are gradually changing, and more and more people are using LLM to create lyrics. However, when using AI technology to create lyrics currently, there are still the following problems: First, restricted by the homogenization of training data, traditional large language models are prone to problems such as semantic deviation from the theme and templatized emotional expression; second, the generated content lacks the rhythm design of "beginning, rhyme, turning point, and conclusion", and it is difficult to adapt to the rhythm requirements of specific music styles; third, the quality assessment relies on manual review. This results in the defect that AI-generated lyrics often have "flowery language but empty content", and it is difficult to reach the level of artistic conception construction of human artists. At the same time, since lyrics need to be paired with musical scores to become a complete song, thus creating lyrics with musical scores has also become one of the future development trends. The SongComposer technology in the field of AI lyric creation integrates lyrics, melodies, and paired databases, converts them into Token sequences and uses LLM for pre-training, and then performs SFT fine-tuning (Supervised Fine-Tuning) through 10,000 manually constructed question-and-answer pairs, and finally realizes the generation of lyrics and songs. However, this technology relies on high-cost LLM pre-training and fine-tuning, and needs to be re-trained as the base model is upgraded; at the same time, music-specific training may damage the original capabilities of the base model, resulting in semantic logic problems in the generated lyrics.
[0003] In view of this, the present invention provides a method and system for lyric or song creation based on a large language model, constructs a complete LLM lyric quality iteration pipeline, and improves the quality of lyrics output by the large language model. Summary of the Invention
[0004] The object of the present invention is to provide a method for creating lyrics or songs based on a large language model, including: constructing a repertoire library and classifying the repertoires in the repertoire library based on musical styles; the repertoire library includes a lyrics library and a song library; obtaining a plurality of candidate repertoires related to the input instruction from the repertoire library; constructing a first prompt word template and inputting the first prompt word template into the large language model to generate a first repertoire; the first prompt word template at least includes a plurality of candidate repertoires; discriminating the first repertoire through a quality discriminator to obtain a discrimination result; the discrimination result includes passing the discrimination and poor quality; when the quality is poor, constructing a second prompt word template and inputting the second prompt word template into the large language model to generate a second repertoire; the second prompt word template at least includes a plurality of candidate repertoires, the first repertoire and the reason for poor quality; the reason for poor quality is generated by the quality discriminator; repeating the discrimination and generation operations of the generated repertoire until the generated repertoire passes the discrimination to obtain the final created repertoire.
[0005] Further, the lyric sequence in the lyrics library is in text format, and the content is lyrics; the song sequence in the song library is in text format, and the content includes the lyrics themselves, pitch, lyric duration, and rest duration.
[0006] Further, the input instruction at least includes an intended category and an intended theme, and obtaining a plurality of candidate repertoires related to the input instruction from the repertoire library; including: determining a matching library according to the intended category; respectively performing dynamic sparse coding on the input sequence and the repertoire sequences in the matching library to obtain a plurality of corresponding sparse attention weights; calculating the modified input sequence coding and a plurality of modified repertoire sequence codings respectively by encoding the input sequence and a plurality of repertoire sequences with the corresponding sparse attention weights; calculating the similarity between the modified input sequence coding and a plurality of modified repertoire sequence codings respectively to obtain the plurality of candidate repertoires.
[0007] Further, performing dynamic sparse coding on the input sequence and the repertoire sequence based on a pre-trained model of the sparse attention mechanism, including: respectively encoding the input sequence and a plurality of repertoire sequences to obtain an input sequence coding and a plurality of repertoire sequence codings; h = Encoder(x); h i = Encoder(d i ); where h represents the input sequence coding; Encoder(*) represents the encoding function; x represents the input sequence; h i represents the repertoire sequence coding; i represents the repertoire variable, and the value range is [1, N], and N represents the total number of repertoires in the matching library; d iRepresents a track sequence; the input sequence encoding and multiple track sequence encodings are dynamically weighted respectively through an attention mechanism to generate a sparse matrix; A = Sparsify(Attention(Q, K, V)); where A represents the sparse matrix; Sparsify() represents the sparsification process, Attention() represents the calculation of the attention weight matrix; Q represents the query vector, which is the focus to be concerned currently; K represents the key vector; V represents the value vector.
[0008] Further, the focus to be concerned currently is a keyword, and the types of keywords can include nouns, phrases, and / or sentences.
[0009] Further, the input sequence encoding and multiple track sequence encodings are respectively multiplied by the corresponding sparse attention weights to obtain a corrected input sequence encoding and multiple corrected track sequence encodings; h′ = h ⊙ A; h′ i = h i ⊙ A i ; where h′ represents the corrected input sequence encoding; h represents the input sequence encoding; ⊙ represents element-wise multiplication; A represents the sparse matrix of the input sequence encoding; h′ i represents the corrected track sequence encoding; h i represents the track sequence encoding; i represents the track variable, and the value range is [1, N], where N represents the total number of tracks in the matching library; A i represents the sparse matrix of the track sequence encoding.
[0010] Further, calculate the similarity values between the corrected input sequence encoding and multiple corrected track sequence encodings respectively, sort the multiple tracks in the matching library according to the similarity values from high to low, and take the tracks ranked before and at the preset number as the candidate tracks; where sim represents the calculation of similarity; h′ represents the corrected input sequence encoding; h′ i represents the corrected track sequence encoding; h represents the input sequence encoding; h i represents the track sequence encoding; ||*|| represents the calculation of the absolute value.
[0011] Further, it also includes: concatenating the intended category, intended theme, and generation requirements to obtain an input instruction; concatenating the input instruction with multiple candidate tracks to obtain the first prompt word template; concatenating the intended category, intended theme, generation requirements, the first track, and the reason for inferiority to obtain a new input instruction; concatenating the new input instruction with multiple candidate tracks to obtain the second prompt word template.
[0012] Further, the quality discriminator is trained through deep learning technology, including: using the tracks in the matching library as multiple positive samples; inputting the input instruction into the large language model to obtain multiple negative samples; the input instruction refers to the unoptimized prompt words for guiding track generation; classifying and labeling the multiple negative samples according to the inferiority reasons of the tracks in the negative samples; training the classifier based on the positive samples and negative samples to obtain the quality discriminator.
[0013] The object of the present invention also lies in providing a lyrics or song creation system based on a large language model, including a track library construction module, a track matching module, a template construction module, a discrimination module, and a loop module; the track library construction module is used to construct a track library and classify the tracks in the track library based on the music genre; the track library includes a lyrics library and a song library; the track matching module is used to obtain multiple candidate tracks related to the input instruction from the track library; the template construction module is used to construct a first prompt word template and input the first prompt word template into the large language model to generate a first track; the first prompt word template at least includes multiple candidate tracks; when it is discriminated as inferior, construct a second prompt word template and input the second prompt word template into the large language model to generate a second track; the second prompt word template at least includes multiple candidate tracks, the first track, and the inferiority reason; the inferiority reason is generated by the quality discriminator; a quality discriminator is set in the discrimination module, and the first track is discriminated through the quality discriminator to obtain a discrimination result; the discrimination result includes passing the discrimination and being discriminated as inferior; the loop module is used to repeatedly perform the discrimination and generation operations on the generated track until the generated track passes the discrimination to obtain the final created track.
[0014] The present invention improves the quality of the tracks output by the large language model by screening candidate tracks and submitting the information of the lyrics or songs with the same intended theme for the large language model to refer to.
[0015] The present invention has completed a three-stage work of optimizing the generation effect of the LLM: in the first stage, collect the track library; in the second stage, match and obtain tracks with high similarity (such as TOP100); in the third stage, fill the track content into the prompt word template so that the LLM obtains more relevant prompt words, improving the generation quality of the LLM. At the same time, when calculating the similarity, we use the sparse coding matrix and the weighted similarity calculation method to obtain a more accurate similarity matching result.
[0016] The present invention has also completed a quality iteration Agent pipeline (intelligent agent process). By training the quality discriminator, the quality discriminator determines the quality of the LLM generation result and identifies the deficiencies of the works with poor quality. For the works with poor quality, fill their inferiority reasons into the generation instruction and let the LLM generate again. Through these three steps, the Agent work pipeline for making the LLM iteratively generate lyrics or songs until the quality meets the standard is completed. Brief Description of the Drawings
[0017] Figure 1 This is an exemplary schematic diagram of a method for creating lyrics or songs based on a large language model provided by some embodiments of the present invention. Detailed Description of the Embodiments
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations.
[0019] The present invention constructs a closed-loop Agent creation system of "data-driven - retrieval enhancement - discrimination iteration", and obtains a well-quality lyrics or music sequence through a cyclic workflow and returns it to the user until then.
[0020] Figure 1 This is an exemplary schematic diagram of a method for creating lyrics or songs based on a large language model provided by some embodiments of the present invention. As Figure 1 shown, the present invention proposes a method for creating lyrics or songs based on a large language model, including the following:
[0021] Construct a track library and classify the tracks in the track library based on the music genre; the track library includes a lyrics library and a song library.
[0022] A track library refers to a structured collection of track data systematically organized for training a model, usually including a lyric library and a song library. The lyric library refers to a systematically organized collection of pure lyric data. The song library refers to a systematically organized collection of data where lyrics and melodies match. The lyric sequences in the lyric library are in text format, with the content being lyrics; the song sequences in the song library are in text format, each musical phrase includes multiple units, and the content includes the lyrics themselves, pitch, lyric duration, and rest duration. For example, for a piece of music, its content is "sun comes shining through", its sequence in the lyric library is "sun;comes;shining;through;", and its sequence in the song library is "sun,<D#5>,<119>,<79>;comes,<D#5>,<122>,<79>;shining,<D#5>,<123>,<79>;through,<D#5>,<120>,<245>;", where "sun,<D#5>,<119>,<79>" indicates that the lyric of this unit is "sun", the pitch is D#5, the duration is 119 ms, and the rest duration is 79 ms. Based on the music genre, tracks can be classified into types such as pop, Hippop, folk, and rock.
[0023] Obtain multiple candidate tracks related to the input instruction from the track library. In some embodiments, according to the user's input instruction, semantic similarity matching can be used to obtain the top pre-set number of tracks in terms of similarity. For example, obtain the top 100 tracks in terms of similarity. The input instruction can refer to a text instruction related to track generation entered by the user. Candidate tracks can refer to multiple lyrics or songs screened from the track library according to the user's needs. The input instruction includes at least the intended category and the intended theme. The intended category refers to the type of track the user needs to output. For example, whether to generate pure lyrics or a complete song. The intended theme is related to the content of the track the user needs to output. For example, "The Wind in Summer" or "Gravity", etc. The user can determine whether they currently need to generate the content of pure text lyrics or a song sequence with paired lyrics and melodies through words such as "lyrics" or "song". Obtain multiple candidate tracks related to the input instruction from the track library, including:
[0024] Determine the matching library according to the intended category. The matching library refers to a track library that meets the user's intended category requirements. When the user's intended category is lyrics, the matching library is the lyric library; when the user's intended category is a song, the matching library is the song library.
[0025] Perform dynamic sparse coding on the input sequence and the track sequences in the matching library respectively to obtain multiple corresponding sparse attention weights. The input sequence refers to the structured time-series data of the interaction instructions or creative materials provided by the user to the model. The track sequence refers to the structured time-series data obtained by uniformly encoding music elements (melody, chord, rhythm, etc.) and / or lyric texts along the time axis. The sparse attention weights are used to represent the importance of different parts in the sequences (lyric sequence, song sequence, and user input sequence). In some embodiments, the input sequence and the track sequences can be dynamically sparsely encoded based on the pre-trained model DeBERTa-v3 of the Splade (Sparse Lexical and Expansion Model, an information retrieval model) sparse attention mechanism (the third-generation DeBERTa model based on gradient decoupled embedding sharing). For example, assume the user input theme is "The gentle wind in summer blows by". The model processing flow is as follows: DeBERTa-v3 encodes the input and the descriptions of songs such as "The Wind in Summer" and "Ningxia" in the song library. The Splade mechanism highlights the weights of keywords such as "summer" and "wind". Since the weight of "wind" in the input encoding is relatively high, "The Wind in Summer" in the song library obtains a higher similarity due to containing the same keyword; "Ningxia" has a lower similarity as it lacks "wind" but contains "summer". Therefore, after sorting by cosine similarity, "The Wind in Summer" is returned as the first result, and "Quiet Summer" is the second result. The dynamic sparse coding includes: encoding the input sequence and multiple track sequences respectively to obtain the input sequence encoding and multiple track sequence encodings;
[0026] h = Encoder(x);
[0027] h i = Encoder(d i );
[0028] Among them, h represents the input sequence encoding; Encoder(*) represents the encoding function; x represents the input sequence; h i represents the track sequence encoding; i represents the track variable, and its value range is [1, N], where N represents the total number of tracks in the matching library; d i represents the track sequence.
[0029] Dynamically assign weights to the input sequence encoding and multiple track sequence encodings respectively through the attention mechanism to generate a sparse matrix;
[0030] A = Sparsify(Attention(Q, K, V));
[0031] Among them, A represents a sparse matrix; Sparsify() represents performing sparsification processing, Attention() represents calculating an attention weight matrix; Q represents a query vector, which is the focus to be concerned about currently; K represents a key vector; V represents a value vector.
[0032] The input sequence encoding and multiple track sequence encodings are respectively calculated with the corresponding sparse attention weights to obtain a corrected input sequence encoding and multiple corrected track sequence encodings. The input sequence encoding and the track sequence encoding respectively refer to the content obtained after encoding the input sequence and the track sequence. The corrected input sequence encoding and the corrected track sequence encoding respectively refer to the content obtained after weight correction of the input sequence encoding and the track sequence encoding. By performing weight correction on the sequence encoding, the importance of different parts of the sequence can be allocated, enhancing the recognition degree of the subsequent model for important content (such as nouns), and thus improving the naturalness and relevance of the subsequent generated tracks. Multiply the input sequence encoding and multiple track sequence encodings with the corresponding sparse attention weights respectively to obtain a corrected input sequence encoding and multiple corrected track sequence encodings;
[0033] h′ = h ⊙ A;
[0034] h′ i = h i ⊙ A i ;
[0035] Among them, h′ represents the corrected input sequence encoding; h represents the input sequence encoding; ⊙ represents element-wise multiplication; A represents the sparse matrix of the input sequence encoding; h′ i represents the corrected track sequence encoding; h i represents the track sequence encoding; i represents a track variable, and its value range is [1, N], where N represents the total number of tracks in the matching library; A i represents the sparse matrix of the track sequence encoding.
[0036] Calculate the similarity between the corrected input sequence encoding and multiple corrected track sequence encodings to obtain multiple of the candidate tracks. Calculate the similarity values between the corrected input sequence encoding and multiple corrected track sequence encodings, sort the multiple tracks in the matching library according to the similarity values from high to low, and take the tracks ranked before and at the preset number as the candidate tracks;
[0037]
[0038] Among them, sim represents calculating similarity; h′ represents the corrected input sequence encoding; h′ i represents the corrected track sequence encoding; h represents the input sequence encoding; h i represents the track sequence encoding; ||*|| represents calculating the absolute value.
[0039] Construct the first prompt template, and input the first prompt template into the large language model to generate the first track; the first prompt template includes at least multiple candidate tracks; splice the intended category, intended theme, and generation requirements to obtain an input instruction; splice the input instruction with multiple candidate tracks to obtain the first prompt template. The generation requirements can be used to require the content and format of the lyrics, etc. For example, it can require "use concise spoken language expression without mixing classical Chinese" and / or "do not contain sexual and pornographic related words", etc.
[0040] Identify the first track through a quality discriminator to obtain an identification result; the identification result includes passing the identification and being of poor quality. When the quality is poor, output the reason for the poor quality, and the reasons for the poor quality include logical errors before and after, being too verbose, and not rhyming, etc. The quality discriminator is used to identify whether the current lyrics or song is created by humans or generated by AI. The quality discriminator is trained through deep learning technology and includes:
[0041] Use the tracks in the matching library as multiple positive samples. For example, use the collected pure lyrics library or the melody-lyrics matching library as positive training materials.
[0042] Input the input instruction into the large language model to obtain multiple negative samples. The input instruction refers to the unoptimized prompt for guiding track generation. For example, use the lyrics or songs generated by the LLM without any prompt optimization as negative training materials.
[0043] Classify and label the multiple negative samples according to the reasons for the poor quality of the tracks in the negative samples. For example, the poor-quality tracks can be divided into three categories (such as, logical errors before and after are grouped into one category, being too verbose is grouped into the second category, and not rhyming is grouped into the third category). Respectively label the positive samples, negative samples with logical errors before and after, negative samples that are too verbose, and negative samples that do not rhyme as 0, 1, 2, and 3.
[0044] Train the classifier based on the positive samples and negative samples to obtain the quality discriminator. For example, the positive samples and negative samples can be divided into a training set, a test set, and a validation set, and the classifier is trained through the training set, the test set, and the validation set to obtain the quality discriminator in this application (the accuracy rate of the finally obtained discriminator is above 90%). The model used by the classifier is a pre-trained Bert model. Use the Bert model to classify the lyrics sequence, y = bert(x), where x is the input track sequence and Y is the output category, making full use of the semantic ability of the Bert model itself and the context modeling ability of the Transformer model.
[0045] When identifying inferior quality, construct a second prompt template, input the second prompt template into the large language model to generate a second piece of music; the second prompt template includes at least multiple candidate pieces of music, the first piece of music, and the reason for inferior quality; the reason for inferior quality is generated by the quality discriminator. Concatenate the intended category, intended theme, generation requirements, and the reason for inferior quality to obtain a new input instruction; concatenate the new input instruction with multiple candidate pieces of music and the first piece of music to obtain the second prompt template.
[0046] Repeat the identification and generation operations of the generated piece of music until the generated piece of music passes the identification to obtain the final creative piece of music. The generated piece of music refers to the piece of music generated by the large language model, including the first piece of music and the second piece of music. Finally, return the generated creative piece of music to the user.
[0047] The present invention also provides a lyrics or song creation system based on a large language model, including a music library construction module, a music matching module, a template construction module, an identification module, and a loop module.
[0048] The music library construction module is used to construct a music library and classify the music in the music library based on the music genre; the music library includes a lyrics library and a song library. For more content about the music library construction module, see Figure 1 and its related descriptions.
[0049] The music matching module is used to obtain multiple candidate pieces of music related to the input instruction from the music library. For more content about the music matching module, see Figure 1 and its related descriptions.
[0050] The template construction module is used to construct a first prompt template, input the first prompt template into the large language model to generate a first piece of music; the first prompt template includes at least multiple candidate pieces of music; when identifying inferior quality, construct a second prompt template, input the second prompt template into the large language model to generate a second piece of music; the second prompt template includes at least multiple candidate pieces of music, the first piece of music, and the reason for inferior quality; the reason for inferior quality is generated by the quality discriminator. For more content about the template construction module, see Figure 1 and its related descriptions.
[0051] The identification module is provided with a quality discriminator to identify the first piece of music through the quality discriminator to obtain an identification result; the identification result includes passing the identification and identifying inferior quality. For more content about the identification module, see Figure 1 and its related descriptions.
[0052] The loop module is used to repeat the identification and generation operations of the generated piece of music until the generated piece of music passes the identification to obtain the final creative piece of music. For more content about the loop module, see Figure 1 and its related descriptions.
[0053] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for creating lyrics or songs based on large language models, characterized in that, Including: Construct a track library and classify the tracks in the track library based on the music genre; the track library includes a lyric library and a song library; Obtain multiple candidate tracks related to the input instruction from the track library; Construct a first prompt word template and input the first prompt word template into a large language model to generate a first track; The first prompt word template includes at least multiple candidate tracks; Discriminate the first track through a quality discriminator to obtain a discrimination result; the discrimination result includes passing the discrimination and being of poor quality; When the quality is poor, construct a second prompt word template and input the second prompt word template into the large language model to generate a second track; The second prompt word template includes at least multiple candidate tracks, the first track, and the reason for poor quality; the reason for poor quality is generated by the quality discriminator; Repeat the discrimination and generation operations of the generated tracks until the generated track passes the discrimination to obtain the final creative track.
2. The method for creating lyrics or songs based on a large language model according to claim 1, wherein The lyric sequence in the lyric library is in text format, and the content is lyrics; the song sequence in the song library is in text format, and the content includes the lyrics themselves, pitch, lyric duration, and rest duration.
3. The method for creating lyrics or songs based on a large language model according to claim 1, wherein The input instruction includes at least an intended category and an intended theme, and obtains multiple candidate tracks related to the input instruction from the track library; including: Determine a matching library according to the intended category; Perform dynamic sparse coding on the input sequence and the track sequences in the matching library respectively to obtain multiple corresponding sparse attention weights; Calculate the input sequence encoding and multiple track sequence encodings respectively with the corresponding sparse attention weights to obtain a corrected input sequence encoding and multiple corrected track sequence encodings; Calculate the similarity between the corrected input sequence encoding and multiple corrected track sequence encodings to obtain multiple of the candidate tracks.
4. The method for creating lyrics or songs based on a large language model according to claim 3, wherein, Performing dynamic sparse coding on the input sequence and the track sequence based on a pre-trained model of the sparse attention mechanism, including: Encode the input sequence and multiple track sequences respectively to obtain an input sequence encoding and multiple track sequence encodings; h = Encoder(x); h i = Encoder(d i ); Among them, h represents the encoding of the input sequence; Encoder(*) represents the encoding function; x represents the input sequence; h i represents the encoding of the track sequence; i represents the track variable, and its value range is [1, N], where N represents the total number of tracks in the matching library; d i represents the track sequence; Dynamically allocate weights to the input sequence encoding and multiple track sequence encodings respectively through the attention mechanism to generate a sparse matrix; A = Sparsify(Attention(Q, K, V)); Where A represents the sparse matrix; Sparsify() represents performing sparsification processing, Attention() represents calculating the attention weight matrix; Q represents the query vector, which is the focus to be concerned about currently; K represents the key vector; V represents the value vector.
5. The method for creating lyrics or songs based on a large language model according to claim 4, wherein The focus to be concerned about currently is the keyword.
6. The method for creating lyrics or songs based on a large language model according to claim 3, wherein Multiply the input sequence encoding and multiple track sequence encodings respectively with the corresponding sparse attention weights to obtain a corrected input sequence encoding and multiple corrected track sequence encodings; h' = h ⊙ A; h′ i = h i ⊙A i ; Among them, h′ represents the corrected input sequence encoding; h represents the input sequence encoding; ⊙ represents element-wise multiplication; A represents the sparse matrix of the input sequence encoding; h′ i represents the corrected track sequence encoding; h i represents the track sequence encoding; i represents the track variable, with a value range of [1, N], and N represents the total number of tracks in the matching library; A i represents the sparse matrix of the track sequence encoding.
7. The method for creating lyrics or songs based on a large language model according to claim 3, wherein, Calculate the similarity values between the corrected input sequence encoding and multiple corrected track sequence encodings, sort the multiple tracks in the matching library according to the similarity values from high to low, and use the tracks ranked before and at the preset number as the candidate tracks; Among them, sim represents calculating similarity; h' represents correcting the encoding of the input sequence; h' i represents correcting the encoding of the track sequence; h represents the encoding of the input sequence; h i represents the encoding of the track sequence; ||*|| represents calculating the absolute value.
8. The method for creating lyrics or songs based on a large language model according to claim 1, wherein Also including: Concatenate the intended category, intended theme, and generation requirements to obtain the input instruction; Concatenate the input instruction with multiple candidate tracks to obtain the first prompt word template; Concatenate the intended category, intended theme, generation requirements, and reasons for inferiority to obtain a new input instruction; Concatenate the new input instruction with multiple candidate tracks to obtain the second prompt word template.
9. The method for creating lyrics or songs based on a large language model according to claim 1, wherein The quality discriminator is trained through deep learning technology and includes: Use the tracks in the matching library as multiple positive samples; Input the input instruction into the large language model to obtain multiple negative samples; the input instruction refers to the unoptimized prompt word for guiding track generation; Classify and label the multiple negative samples according to the reasons for inferiority of the tracks in the negative samples; Train the classifier based on the positive samples and negative samples to obtain the quality discriminator.
10. A lyrics or song creation system based on a large language model, characterized in that, It includes a track library construction module, a track matching module, a template construction module, a discrimination module, and a loop module; The track library construction module is used to construct a track library and classify the tracks in the track library based on the music style; the track library includes a lyric library and a song library; The track matching module is used to obtain multiple candidate tracks related to the input instruction from the track library; The template construction module is used to construct the first prompt word template and input the first prompt word template into the large language model to generate the first track; The first prompt word template includes at least multiple candidate tracks; when discriminating inferiority, construct the second prompt word template and input the second prompt word template into the large language model to generate the second track; the second prompt word template includes at least multiple candidate tracks, the first track, and the reason for inferiority; the reason for inferiority is generated by the quality discriminator; A quality discriminator is set in the discrimination module, and the first track is discriminated through the quality discriminator to obtain a discrimination result; the discrimination result includes passing the discrimination and discriminating inferiority; The loop module is used to repeatedly perform the discrimination and generation operations of the generated track until the generated track passes the discrimination to obtain the final creative track.