Melody generation method and device, computer equipment and storage medium
The structural similarity between the target lyrics and melody motivation is evaluated through various matching methods, and the optimal melody motivation is selected to generate melody, which solves the problem of mismatch between melody and lyrics and improves the quality of melody generation.
Patent Information
- Application Number
- CN202510243501.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the melody generated based on artificial intelligence often does not match the lyrics, and there are problems of poor melody quality and poor listening experience.
By obtaining the target lyrics, using multiple matching methods to match the target lyrics with the melody motivation in the motivation library, calculate the lyric structure similarity score, select the melody motivation with the highest total score as the target motivation, and finally generate the melody based on the target motivation.
Improves the fit between melody and lyrics, improves the quality of melody generation, and provides a reliable technical route for automatic composition based on lyrics.
Smart Images

Figure CN120048232A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of music generation, and particularly to a melody generation method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, a variety of applications have been derived, including automatically arranging music according to the lyrics input by users. However, in traditional technologies, the melodies generated based on artificial intelligence often do not match the lyrics, resulting in problems such as poor melody quality and unpleasant listening experience. Summary of the Invention
[0003] The purpose of this application is to at least solve one of the above technical defects, especially the defect that the melody does not match the lyrics in the prior art.
[0004] In a first aspect, this application provides a melody generation method, including:
[0005] Obtain target lyrics;
[0006] For any melody motive to be matched in the motive library, respectively use a variety of matching methods to match the target statement in the target lyrics with the melody motive, and obtain multiple lyric structure similarity scores corresponding to the melody motive;
[0007] According to the multiple lyric structure similarity scores corresponding to the melody motive, obtain the corresponding total structure similarity score;
[0008] Determine the melody motive with the highest total structure similarity score as the target motive;
[0009] Generate the melody corresponding to the target lyrics according to the target motive.
[0010] In one embodiment, generating the melody corresponding to the target lyrics according to the target motive includes:
[0011] Determine a melody generation mode combination; the melody generation mode combination includes a preset number of melody generation modes, and the melody generation mode is at least one selected from a variety of optional methods;
[0012] Sequentially and cyclically select the melody generation modes in the melody generation mode combination, and generate melodies for each statement in the target lyrics in sequence according to the selected melody generation mode and the target motive.
[0013] In one embodiment, the melody motive includes relative pitch information, and the relative pitch information includes the relative pitch of the notes included in each word of the motive lyrics; generating melodies for each statement in the target lyrics in sequence according to the selected melody generation mode and the target motive includes:
[0014] Determine the key of the target lyrics;
[0015] Obtain the absolute pitch information according to the key and the relative pitch information;
[0016] Generate a melody for each statement in the target lyrics in sequence according to the selected melody generation mode and the absolute pitch information.
[0017] In one embodiment, the optional techniques include repetition, sequence progression, anadiplosis, and symmetry.
[0018] In one embodiment, the melody motive includes motive lyrics, the matching method includes word count similarity, and the lyric structure similarity score includes the word count similarity score; use multiple matching methods respectively to match the target statement in the target lyrics with the melody motive, and obtain multiple lyric structure similarity scores corresponding to the melody motive, including:
[0019] Obtain the word count similarity score according to the difference between the word count of the motive lyrics and the word count of the target statement.
[0020] In one embodiment, the melody motive includes motive lyrics, the matching method includes segmentation structure similarity, and the lyric structure similarity score includes the segmentation structure similarity score; use multiple matching methods respectively to match the target statement in the target lyrics with the melody motive, and obtain multiple lyric structure similarity scores corresponding to the melody motive, including:
[0021] Segment the motive lyrics and the target statement respectively to obtain the corresponding segmentation structure sequences;
[0022] Use the dynamic time warping method to process the segmentation structure sequences corresponding to the motive lyrics and the target statement, and obtain the segmentation structure similarity score.
[0023] In one embodiment, the melody motive includes a rhythm structure sequence, the matching method includes rhythm structure similarity, and the lyric structure similarity score includes the rhythm structure similarity score; use multiple matching methods respectively to match the target statement in the target lyrics with the melody motive, and obtain multiple lyric structure similarity scores corresponding to the melody motive, including:
[0024] Determine the rhythm structure sequence of the target statement;
[0025] Use the dynamic time warping method to process the rhythm structure sequences corresponding to the melody motive and the target statement, and obtain the rhythm structure similarity score.
[0026] In one embodiment, the melody motive further includes a filtering attribute, and the melody generation method further includes:
[0027] Determine the filtering conditions corresponding to the target lyrics;
[0028] Determine the melodic motives in the motive library whose filtering attributes meet the filtering conditions as the melodic motives to be matched.
[0029] In a second aspect, the present application provides a melody generation device, including:
[0030] A data acquisition module, configured to acquire target lyrics;
[0031] A motive matching module, configured to, for any melodic motive to be matched in the motive library, respectively use a variety of matching methods to match the target statement in the target lyrics with the melodic motive, and obtain multiple lyric structure similarity scores corresponding to the melodic motive;
[0032] A data analysis module, configured to obtain a corresponding total structure similarity score according to the multiple lyric structure similarity scores corresponding to the melodic motive;
[0033] A target motive selection module, configured to determine the melodic motive with the highest total structure similarity score as the target motive;
[0034] A melody generation module, configured to generate a melody corresponding to the target lyrics according to the target motive.
[0035] In a third aspect, the present application provides a computer device, including one or more processors and a memory. When the computer-readable instructions stored in the memory are executed by the one or more processors, the steps of the melody generation method in any of the above embodiments are executed.
[0036] In a fourth aspect, the present application provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the melody generation method in any of the above embodiments.
[0037] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0038] Based on the melody generation method in this embodiment, first obtain the target lyrics for which the melody needs to be generated. Then, for each melody motive to be matched in the motive library, use multiple matching methods to calculate the structural similarity scores of the target sentence in the target lyrics and the melody motive on each matching attribute respectively. According to these individual scores, comprehensively calculate the total structural similarity score with the target lyrics. Next, select the melody motive with the highest total score as the target motive. Finally, based on this target motive, develop the complete melody of the target lyrics. This solution comprehensively evaluates the structural similarity between the target lyrics and the melody motive by adopting multiple structural matching methods, ensuring that the optimal matching target motive is selected as the material for melody generation. The melody generated based on the target motive has a higher degree of fit with the target lyrics, improves the quality of melody generation, provides a reliable technical route for automatic songwriting based on lyrics, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of the melody generation method provided by an embodiment of the present application;
[0041] Figure 2 It is the internal structure diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0043] The present application provides a melody generation method. Please refer to Figure 1 , including steps S102 to S110.
[0044] S102, obtain the target lyrics.
[0045] It can be understood that the target lyrics are the lyrics given by the user during automatic music arrangement. This embodiment is used to generate a corresponding melody for the target lyrics. Obtaining the target lyrics is a basic step in generating the melody of a song, and the lyrics play an important role in music arrangement. The target lyrics here can be lyrics written by the user himself, or can be automatically generated by a software application according to the user's needs, or can be obtained by the user revising the lyrics generated by artificial intelligence.
[0046] S104. For any melody motive to be matched in the motive library, use multiple matching methods to match the target sentence in the target lyrics with the melody motive respectively, and obtain multiple lyric structure similarity scores corresponding to the melody motive.
[0047] It can be understood that the motive library is a data set storing multiple melody motives. The melody motive is the core material for the development of the melody and is the most representative small unit in the whole melody. For example, the first four notes in Beethoven's Symphony No. 5 are the most important melody motives in the whole melody. The melody motive includes pitch information, which reflects which notes the melody motive consists of. In this embodiment, the melody motive that best matches the target lyrics will be selected from the motive library as the basis for melody generation. Therefore, the melody motives in the motive library will also include multiple matching attributes, and each matching attribute is related to the lyric structure of the melody motive. The matching attributes can include the motive lyrics corresponding to the melody motive, the number of words in the motive lyrics, the rhythm structure sequence of the motive lyrics, etc. These can be added to the motive library together when collecting melody motives as needed.
[0048] Since the target lyrics are the complete lyrics of the whole song, but the melody motive is only a fragment of it, directly matching the melody motive with the target lyrics will cause the matching result to be affected by the unimportant parts in the target lyrics. Therefore, in this embodiment, a representative sentence will be selected from the whole target lyrics as the matching object, that is, the target sentence. The selection of the target sentence can be marked by the user himself, or can be determined from the target lyrics according to the default position. Generally speaking, the first sentence of the target lyrics can be selected as the target sentence.
[0049] Due to the complexity of lyric matching, in order to find the most suitable motive in the motive library for the target statement, the similarity between the target statement and each melody motive in terms of lyric structure will be calculated from multiple different perspectives. Specifically, the choice of matching method mainly depends on the types of matching attributes obtainable from the user's input and the matching attributes possessed by the melody motives in the motive library. Each matching method can be used to calculate the similarity between the corresponding matching attributes. For example, if the user's input only contains the target lyrics, after determining the target statement, matching attributes directly reflected by the lyrics themselves, such as the number of words and the word segmentation structure, can be extracted from it. When the motive lyrics are also included in the melody motives in the motive library, similarity matching can be performed at least from the perspectives of the number of words and the word segmentation structure. If other matching attributes not directly reflected by the lyrics are included in both the user's input and the melody motives, the corresponding matching method can be used for similarity matching. Finally, the matching result of each matching method is a lyric structure similarity score. Each lyric structure similarity score reflects the similarity between the target statement and the melody motive in the corresponding matching attribute.
[0050] In addition, when performing the matching, all the melody motives in the motive library can be directly used as the melody motives to be matched with the target statement. However, a preliminary filtering step can also be added before the matching. That is, when collecting the melody motives, filtering-related attributes, such as the playing style and the playing mood, can be added to the melody motive data. When the user inputs the target lyrics, the filtering conditions can be input together. Then, before starting the matching, the melody motives in the motive library whose filtering attributes meet the filtering conditions can be found as the melody motives to be matched.
[0051] S106, obtain the corresponding total structure similarity score according to the multiple lyric structure similarity scores corresponding to the melody motive.
[0052] It can be understood that the total structure similarity score is the overall matching score obtained by synthesizing multiple lyric structure similarity scores. The level of this score will determine the final motive selection. Through step S104, the lyric structure similarity between each melody motive to be matched and the target statement has been evaluated from multiple perspectives, and multiple lyric structure similarity scores have been obtained. These individual scores can only reflect the matching degree between the target statement and the melody motive at a specific level and cannot comprehensively measure the overall structural fit. Therefore, it is necessary to reasonably synthesize multiple scores to obtain a total structure similarity score to evaluate the overall structural matching level of the melody motive and the target lyrics. Specifically, weights can be preset in advance for each lyric structure similarity score according to the degree of emphasis. The higher the degree of emphasis, the higher the corresponding weight. By performing a weighted sum of multiple lyric structure similarity scores, the total structure similarity score can be obtained.
[0053] S108. Determine the target motive as the melody motive with the highest overall structural similarity score.
[0054] It can be understood that after the previous steps, the structural matching degree between all the melody motives to be matched in the motive library and the target lyrics has been evaluated, and each melody motive to be matched has obtained an overall structural similarity score. Now, based on these scores, the best melody motive needs to be selected as the final target motive. The higher the overall structural similarity score, the more closely the melody motive matches the target statement in terms of the lyrics structure, and the higher the degree of fit between the generated melody and the target lyrics. Therefore, the melody motive with the highest score should be selected as the target motive. This can not only make full use of the high-quality materials in the motive library but also ensure that the generated melody conforms to the structural characteristics of the lyrics.
[0055] S110. Generate the melody corresponding to the target lyrics based on the target motive.
[0056] It can be understood that through the previous steps, we have selected the target motive from the motive library that best matches the structure of the target lyrics. The next task is to generate a complete melody phrase based on this target motive so that it can fit the entire target lyrics. The core idea of generating the melody is to start with the pitch information of the target motive and develop the target motive through the melody generation mode, ultimately generating a melody line that fully covers the entire lyrics.
[0057] Based on the melody generation method in this embodiment, first, obtain the target lyrics for which the melody needs to be generated. Then, for each melody motive to be matched in the motive library, use multiple matching methods to calculate the structural similarity scores of the target statement in the target lyrics and the melody motive in various matching attributes respectively. Based on these individual scores, comprehensively calculate the overall structural similarity score with the target lyrics. Next, select the melody motive with the highest total score as the target motive. Finally, based on this target motive, develop the complete melody of the target lyrics. This solution comprehensively evaluates the structural similarity between the target lyrics and the melody motive by adopting multiple structural matching methods, ensuring that the optimal matching target motive is selected as the material for melody generation. The melody generated based on the target motive has a higher degree of fit with the target lyrics, improving the quality of melody generation, providing a reliable technical route for automatic songwriting based on lyrics, and having broad application prospects.
[0058] In one embodiment, generating the melody corresponding to the target lyrics based on the target motive includes:
[0059] (1) Determine the combination of melody generation modes. The combination of melody generation modes includes a preset number of melody generation modes, and the melody generation mode is at least one selected from multiple optional methods.
[0060] It can be understood that the melody generation mode is a method of generating a melody segment for a sentence of the target lyrics based on the pitch information of the target motive. The melody generation mode is at least one selected from multiple optional methods. Common optional methods include repetition, sequence, anadiplosis, symmetry, etc. The melody generation mode combination includes a preset number of melody generation modes. In the melody generation mode combination, the type and number of melody generation modes are different. Each selected melody generation mode can be used only once in the combination, or used twice or more, as long as there are a total of preset number of melody generation modes in the final combination. For example, if the preset number is 4, the combination can contain only one melody generation mode, then the 4 melody generation modes are the same, such as only including the repetition method, then the melody generation mode combination can be [repetition, repetition, repetition, repetition]. The combination can contain two melody generation modes, one with a quantity of 1 and the other with a quantity of 3. For example, including the anadiplosis and symmetry methods, then the melody generation mode combination can be [anadiplosis, anadiplosis, symmetry, anadiplosis]. After determining the target motive, at least one method can be selected from multiple optional methods according to needs and actual situations, and the selected methods can be arranged as needed to form a melody generation mode combination that contains a total of preset number of melody generation modes.
[0061] The melody generation mode combination can be designed by a composition expert based on experience, and corresponding method combinations can be designed for different playing styles, playing emotions, etc. Then, according to the user's input, the corresponding method combination can be selected. It can also be based on a large number of labeled corpora, and methods such as supervised or reinforcement learning are used to train the model to automatically learn the optimal method combination strategy. Finally, the method combination can be automatically generated based on the semantics, lyrics structure, etc. of the target sentence.
[0062] (2) Select the melody generation modes in the melody generation mode combination in sequence and loop, and generate melodies for each sentence in the target lyrics in sequence according to the selected melody generation mode and the target motive.
[0063] It can be understood that after obtaining the combination of techniques, it is necessary to actually apply these techniques next to generate corresponding melody segments for each statement in the target lyrics. In this embodiment, each technique in the combination of techniques is sequentially and cyclically selected, and appropriate melodies are generated for each statement of the target lyrics according to the characteristics of the current technique and the target motive. For example, if the melody generation mode combination is [repetition, symmetry, anadiplosis, sequence], and the target lyrics have 8 sentences, then the melody generation modes corresponding to these 8 sentences are: repetition - symmetry - anadiplosis - sequence - repetition - symmetry - anadiplosis - sequence. By developing the target motive according to the melody generation mode corresponding to each statement in the target lyrics, the melody segment corresponding to this statement can be obtained, and these melody segments are connected to form the complete melody of the target lyrics. In this way, through the organic transformation of techniques and the basic support of the target motive, a complete melody that can well fit the target lyrics and is rich in detailed changes can be generated.
[0064] In one embodiment, the melody motive includes relative pitch information, and the relative pitch information includes the relative pitch of the notes contained in each word of the motive lyrics. This embodiment uses the relative harmony theory as the theoretical framework. In this theoretical framework, chords are usually represented by their distances and relationships with the tonic chord. This method can make the structure of chord progressions easier to understand and can transplant chord progressions in different keys. That is, this theoretical framework pays more attention to the relative relationships of chord progressions rather than the specific keys in which they are located. That is, for the same harmonic progression, although they may eventually migrate to different keys, their relative interval relationships are equivalent. Therefore, the pitch information carried by the melody motive in this embodiment is specifically relative pitch information, which reflects the relative pitch of the notes contained in each word of the motive lyrics. Relative pitch is the pitch information obtained by subtracting the original absolute pitch from the pitch of the tonic of the key. For example, the tonic of a song is C4, the absolute pitch of C4 in MIDI format is 60, and the absolute pitch of a chord corresponding to a word is represented as [62, 63, 66]. Based on this, the relative pitch of the chord corresponding to this word is [2, 3, 6].
[0065] In one embodiment, according to the selected melody generation mode and the target motive, generating melodies for each statement in the target lyrics in sequence includes:
[0066] (1) Determine the key of the target lyrics.
[0067] It can be understood that tonality refers to the pitch position of the tonic in a musical work and the pitch system constructed around this tonic. It determines the basic pitch framework of the music, giving the music a specific sense of tonality and color. In principle, different tonalities can bring completely different auditory sensations and emotional implications. For example, major keys usually give people a feeling of brightness, cheerfulness, and openness, while minor keys often convey melancholy, sadness, and introverted emotions. For instance, C major has the characteristics of brightness and stability, and its interval relationship is relatively harmonious and simple, suitable for expressing cheerful emotions. If the theme of the lyrics tends to be deep, plaintive, or mysterious, it may tend to be in a minor key, such as A minor, which has a unique melancholy color. The tonality of the target lyrics can be set by the user himself, which is generally applicable to users with music theory foundation. The tonality of the target lyrics can also be determined according to the emotional expression of the target lyrics and the first mapping relationship. The emotional expression can be set by the user himself, or the target lyrics can be input into a large language model, and the large language model uses its own semantic understanding ability to select from multiple preset emotions. And the first mapping relationship reflects the corresponding relationship between emotional expression and tonality.
[0068] (2)Based on the tonality and relative pitch information, obtain the absolute pitch information.
[0069] It can be understood that absolute pitch information refers to the precise and fixed pitch value that a note has in music, usually measured by a specific pitch standard. For example, in the MIDI format, the pitch is represented by a number. Relative pitch information, on the other hand, describes the relative height relationship between notes based on a reference pitch (such as the pitch of the root note corresponding to the tonality). For example, in C major with C as the root note, the absolute pitch of C is 60 in the MIDI format. If the relative pitch of a note in a certain melodic motive with respect to the root note C of C major is 3, then the absolute pitch of this note is 63.
[0070] (3)According to the selected melody generation mode and absolute pitch information, generate melodies for each statement in the target lyrics in sequence.
[0071] It can be understood that the melody generation mode is a specific operation method for constructing a complete melody based on the target motivation. Common melody generation modes such as repetition, sequence, anadiplosis, symmetry, etc. all have their unique musical principles and creative effects. The repetition technique is to repeatedly present the target motivation in the melody, strengthening the musical theme through repetition and enhancing the unity and coherence of the melody. Sequence is to regularly move the target motivation at different pitches. This technique can create the development and variation of the melody while maintaining the basic form of the melody motivation, making the melody have a sense of dynamics and propulsion. The anadiplosis technique is to make the ending note of the previous musical phrase the same or similar to the beginning note of the next musical phrase, forming a musical connection and coherence. The symmetry technique is to construct a left-right symmetric melody structure around a certain central note or central melody segment, giving people a sense of balance and harmony. In this embodiment, the melody for each statement of the target lyrics will be gradually generated according to the melody generation modes sequentially selected from the melody generation mode combination, and finally these melody segments will be connected to form a complete melody.
[0072] In one embodiment, the melody motivation includes motivation lyrics, the matching method includes word count similarity, and the lyrics structure similarity score includes the word count similarity score. Using various matching methods respectively, the target statement in the target lyrics is matched with the melody motivation to obtain multiple lyrics structure similarity scores corresponding to the melody motivation, including: obtaining the word count similarity score according to the difference between the word count of the motivation lyrics and the word count of the target statement.
[0073] It can be understood that the melody motivation includes melody information and lyrics information, and the motivation lyrics are the lyrics information corresponding to the melody motivation. Word count similarity, as a matching method, aims to measure the degree of fit between the target statement and the motivation lyrics from the dimension of word count. The lyrics structure similarity score is a quantitative index comprehensively reflecting the matching situation of the target statement and the melody motivation at each level of the lyrics structure, and the word count similarity score among them focuses on the specific numerical manifestation of the matching degree of the two in terms of word count. The reason for taking word count similarity as a matching method and obtaining the corresponding score through the word count difference is that the word count structure of the lyrics will affect the construction of the melody and the overall rhythm arrangement to a certain extent. In music creation, statements with different word counts often need to be adapted to different rhythm patterns and melody trends. For example, statements with fewer words may be more suitable for concise, lively, and rhythmically compact melody segments, while statements with more words may require relatively more stretched melodies with more note changes to match them. By comparing the word count difference between the motivation lyrics and the target statement, the initial judgment of the structural fit between the two can be made, providing a reference basis for screening the melody motivation that best matches the target lyrics in one dimension. When the word count difference is smaller, it means that from this perspective, the structures of the two are more similar, and the corresponding melody adaptability may be higher; on the contrary, if the word count difference is too large, it may be difficult to achieve a good fit in subsequent melody generation.
[0074] In one embodiment, the melodic motive includes motive lyrics, the matching method includes the similarity of the word segmentation structure, and the lyric structure similarity score includes the similarity score of the word segmentation structure. Using various matching methods respectively, the target sentence in the target lyrics is matched with the melodic motive to obtain multiple lyric structure similarity scores corresponding to the melodic motive, including: segmenting the motive lyrics and the target sentence respectively to obtain the corresponding word segmentation structure sequences. Using the dynamic time warping method, process the word segmentation structure sequences corresponding to the motive lyrics and the target sentence to obtain the similarity score of the word segmentation structure.
[0075] It can be understood that the word segmentation structure sequence is the manifestation of the arrangement order of words formed after segmenting the motive lyrics and the target sentence according to certain language rules. The similarity score of the word segmentation structure in the lyric structure similarity score is a quantitative evaluation index specifically given for the matching situation of the word segmentation structures of the two, which intuitively reflects their similarity degree from the dimension of the word segmentation structure.
[0076] The word segmentation structure is of great significance in lyrics, which is related to the rhythm and rhyme of sentences as well as the semantic expression logic. Different word segmentation methods will shape different sentence structure characteristics, thereby affecting the melody trend adapted to them. The dynamic time warping method (DTW) is introduced here because it has unique advantages in dealing with the similarity comparison of sequence data. The word segmentation structure sequences corresponding to the motive lyrics and the target sentence essentially belong to sequence data, and their lengths, word arrangement orders, etc. may be different. The dynamic time warping method can find the optimal matching path between the two under the condition of allowing a certain time axis stretching (that is, the flexible matching of word order and quantity), so as to scientifically measure their similarity degree and obtain the similarity score of the word segmentation structure.
[0077] For example, the motive lyrics are "Spring flowers bloom all over the hillside". According to the Chinese language habits and word segmentation rules, it can be segmented into "Spring day", "flowers bloom", and "all over the hillside", and a word segmentation structure sequence [2, 2, 3] is generated according to the number of characters in each segmentation. In this way, the word segmentation structure sequence corresponding to the motive lyrics is obtained. Then, the same word segmentation work is carried out on the target sentence. Suppose the target sentence is "The autumn wind blows across the fields", and after segmentation, it is obtained as "autumn wind", "blows across", and "fields", and a word segmentation structure sequence [2, 2, 3] is generated according to the number of characters in each segmentation. Then, the dynamic time warping method is used to process these two groups of word segmentation structure sequences. The core of the dynamic time warping method is to construct a cost matrix. The rows of the matrix correspond to the word segmentation structure sequence of the motive lyrics, and the columns correspond to the word segmentation structure sequence of the target sentence. By calculating the cost at each position (such as the cost of word mismatch, the cost of order difference, etc.), an optimal path from the upper left corner to the lower right corner is found to minimize the total cost. This path reflects the best matching situation between the two. Along this optimal path, the word segmentation structure similarity score is determined according to the pre-set calculation rules. For example, it can be set that the lower the total cost, the higher the score. If the total cost is 0, that is, a perfect match, the word segmentation structure similarity score is full marks (such as 100 points), and as the total cost increases, the score is deducted proportionally accordingly.
[0078] In one embodiment, the melody motive includes a rhythm structure sequence, the matching method includes rhythm structure similarity, and the lyric structure similarity score includes a rhythm structure similarity score. Using various matching methods respectively, the target sentence in the target lyrics is matched with the melody motive to obtain multiple lyric structure similarity scores corresponding to the melody motive, including: determining the rhythm structure sequence of the target sentence. Using the dynamic time warping method, the rhythm structure sequences corresponding to the melody motive and the target sentence are processed to obtain the rhythm structure similarity score.
[0079] It can be understood that the rhythmic structure similarity, as a matching method, focuses on measuring the degree of fit between the melodic motive and the target sentence from the perspective of the rhythmic structure. The rhythmic structure similarity score is a quantitative evaluation index for the matching situation between the two at the rhythmic structure level, aiming to accurately present the degree of their mutual matching from the rhythmic perspective. Rhythm plays a crucial role in music creation. It is one of the core elements in shaping the music style, creating the emotional atmosphere, and constructing the overall framework of the melody. The dynamic time warping method is introduced in the matching process because the rhythmic structure sequences corresponding to the melodic motive and the target sentence belong to time series data, and there may be differences in their lengths, the arrangement order of rhythmic elements, and the distribution of specific rhythm patterns. The dynamic time warping method can skillfully handle the misalignment problem of this time series data in the time dimension. By allowing certain rhythm stretching, deformation, and other operations, it finds the optimal matching path between the two in the rhythmic structure, and then scientifically and reasonably measures their similarity degree, thereby obtaining the rhythmic structure similarity score. The rhythmic structure sequence can specifically be formed by the number of beats sustained by each segmented word, such as [0.5, 0.5, 1, 0.25]. The motives in the motive library carry rhythmic structure sequences, and the rhythmic structure sequence of the target sentence needs to be regenerated.
[0080] In one embodiment, the melodic motive further includes a filtering attribute, and the melody generation method further includes: determining a filtering condition corresponding to the target lyrics. Determining the melodic motives in the motive library whose filtering attributes meet the filtering condition as the melodic motives to be matched.
[0081] It can be understood that the filtering attribute is a characteristic description used to classify and screen melodic motives. It can include various attribute information such as the playing style (such as classical style, pop style, rock style, etc.), playing emotion (different emotion categories such as cheerful, sad, exciting, etc.). Through these attributes, different melodic motives can be further refined and distinguished. The filtering condition corresponding to the target lyrics is a screening criterion set according to factors such as the music style and emotional expression that the target lyrics themselves are expected to match, and is used to select the parts that meet specific requirements from numerous melodic motives. The melodic motives to be matched are the set of melodic motives that remain after screening the melodic motives in the motive library according to the filtering condition and are expected to achieve a good match with the target lyrics and can participate in the subsequent matching process. In addition, for the melodic motives to be matched, further screening can be performed based on the avoidance notes, and the melodic motives in which the avoidance notes appear are excluded.
[0082] In one embodiment, the determination of the avoidance note specifically includes:
[0083] (1) Obtaining the emotional parameters of the target lyrics in at least one emotional dimension.
[0084] It can be understood that the target lyrics are the lyrics given by the user during the automatic arrangement. The emotional dimension is a dimension used to quantify the emotional characteristics of a given input. It provides a way to map subjective emotional experience to objective values, and the emotional parameter is the quantitative result evaluated using the quantitative method corresponding to the emotional dimension. In order to expand the richness of emotional expression, two or more emotional dimensions can be used. Commonly used emotional dimensions can be two emotional dimensions related to the VA (Valence-Arousal) emotional model, namely, the pleasure (Valence) dimension and the arousal (Arousal) dimension. The emotional parameter corresponding to the pleasure dimension indicates the degree of positivity of the emotion. The higher the pleasure, the more positive the corresponding emotion, and the lower the pleasure, the more negative the corresponding emotion. The emotional parameter corresponding to the arousal dimension indicates the intensity or excitement of the emotion. The higher the arousal, the more excited and excited the emotion, and the lower the arousal, the calmer and calmer the emotion. The following text will take the emotional dimensions, namely the pleasure dimension and the arousal dimension, as an example for explanation.
[0085] Obtaining the emotional parameters of the target lyrics in at least one emotional dimension can be done manually, but it is time-consuming and labor-intensive. Therefore, it is generally automatic evaluation using rule dictionaries, deep learning models, etc. Taking the rule dictionary solution as an example, specifically, a certain number of emotional seed words are first obtained, and the emotional parameters of the emotional seed words in the selected emotional dimension are evaluated. Then Word2Vec is used to calculate the similarity between each non-seed word and the emotional seed word. The highest similarity is used as the judgment basis, and the words with similarity higher than the similarity threshold are also added to the emotional dictionary, and the emotional parameters of the most similar emotional seed words are assigned to the non-emotional seed words newly added to the emotional dictionary. The emotional parameters of the non-emotional seed words can also be fine-tuned manually. After the emotional dictionary is constructed, the target lyrics are segmented. Word2Vec is used to calculate the similarity between each segmented word and the emotional words in the emotional dictionary. If there are emotional words with similarity greater than the matching threshold, the emotional parameters of the corresponding emotional words are assigned to the corresponding segmented words. All segmented words in the target lyrics that are assigned emotional parameters are regarded as emotional segmented words. Each sentiment segmentation word is assigned a weight (the sum of all weights is 1) according to the preset correspondence between its part of speech and its role in expressing emotion. For example, nouns are assigned a higher weight. The weight of each sentiment segmentation word and the assigned emotion parameter are weighted and summed to obtain the emotion parameter of the target lyrics in at least one emotion dimension.
[0086] (2) Input the emotional parameters into the harmony progression prediction model to obtain the target harmony progression.
[0087] It can be understood that the harmonic progression reflects the connection relationship between multiple chords, which is the manifestation of the change of harmony during performance, and can show the functional relationship and sound color between chords. Although an individual's emotional experience is affected by factors such as culture, background, and personal experience, some harmonic progressions will trigger common emotional reactions under normal circumstances. Therefore, in music works, the emotional ups and downs are often shown by designing the harmonic progression. For example, the [I - IV - V - I] harmonic progression is generally considered to have a stable and satisfying emotion, and is commonly found in happy and celebratory music. The [i - iv - V - i] harmonic progression is generally considered to have a sad and melancholy emotion, especially when the chords contain minors, such as Am - Dm - E - Am.
[0088] After evaluating the emotional parameters of the target lyrics, the relevant information about the emotion that the user wants to express has been obtained, which can guide the emotional color of music creation. This step starts from the perspective of designing the harmonic progression to ensure that the generated music matches the user's emotional expression needs. Specifically, in order to establish the correlation between the emotional parameters and the harmonic progression, considering that there are relatively strict music theory rules in the production of existing music works, the emotions expressed by using the harmonic progression design can be accepted by the public. Therefore, a machine learning model can be used to learn the mapping relationship between the harmonic progression and the emotion parameters in existing music works. Through continuous training and parameter adjustment, the model can capture the mapping relationship between the emotional parameters and the harmonic progression that meets the emotional expression requirements. The model obtained in this way is the harmonic progression prediction model. The input of the harmonic progression prediction model is the emotion parameter, and the output is the target harmonic progression. The target harmonic progression is the harmonic progression predicted by the harmonic progression prediction model after learning the mapping relationship between the harmonic progression design of a large number of existing music works and the emotion parameters. The target harmonic progression matches the emotion parameter, and the emotion parameter is generated according to the target lyrics. Therefore, the obtained target harmonic progression will match the emotional expression needs of the target lyrics.
[0089] (3) Input the target harmonic progression into the pitch prediction model to obtain the pitch probability sequence of each chord in the target harmonic progression.
[0090] It can be understood that during the process of automatically generating music, after generating the harmonic progression, it is necessary to select a suitable texture for the harmonic progression. However, the texture combinations generated by artificial intelligence may have problems such as out-of-tune and disharmony. Due to the lack of relevant criteria, it is difficult to determine whether the notes involved in the texture combination are qualified. Therefore, this embodiment needs to further confirm the pitch range of the target harmonic progression. Specifically, considering that existing music works will focus on note selection during production to avoid out-of-tune and disharmony problems, a machine learning model can be used to learn the preference relationship between each chord included in the harmonic progression and the corresponding pitch in existing music works. Through continuous training and parameter tuning, the model can capture the preference relationship between the harmonic progression and the harmonious and in-tune note selection. The model obtained thereby is the pitch prediction model. The input of the pitch prediction model is the target harmonic progression, and the output is the pitch probability sequence of each chord in the target harmonic progression. The pitch probability sequence corresponds one-to-one with each chord in the target harmonic progression, and it reflects the selection probability of each alternative note being selected as the note composition of the corresponding chord. For example, assuming that the samples learned by the pitch prediction model and the harmonic progression prediction model are tonally unified, such as unified to C major, the alternative notes will only be the 12 notes corresponding to C major. At this time, the pitch probability sequence is the probability of these 12 notes being selected as the chord composition corresponding to each chord. In one embodiment, the input target harmonic progression is (F, G, Dm, Am), and the pitch probability sequence of the F chord is [0.072, 0.095, 0.081, 0.025, 0.123, 0.065, 0.042, 0.117, 0.036, 0.048, 0.119, 0.177].
[0091] (4) Determine the avoidance notes corresponding to each chord in the target harmonic progression according to the pitch probability sequence.
[0092] It can be understood that specifically, the selection of the pitch range can be refined into recommended notes and avoidance notes. Recommended notes are the notes recommended to be used, but no additional processing is required when there are recommended notes in the melody. Avoidance notes are the notes that should not be used or are rarely used. Improper use of avoidance notes will be considered disharmonious and out-of-tune. The pitch probability information can provide good suggestions for the final selection of melody notes. Among them, the notes with higher selection probabilities are the preferred choices in harmonious works, and the notes with lower selection probabilities are the choices to be avoided as much as possible in harmonious works. The avoidance notes obtained based on the pitch probability information delimit the range for pitch selection, and can provide a criterion for whether the melody of the generated music is harmonious, so that timely processing can be carried out before outputting the generated music to the user to generate a more reasonable and natural melody line.
[0093] In one embodiment, the training process of the harmonic progression prediction model includes:
[0094] (1)Obtain a first training set with first annotation information. The first training set includes multiple first training harmonic progressions, and the first annotation information includes the emotional parameters of the first training harmonic progressions in at least one emotional dimension.
[0095] It can be understood that the first training harmonic progressions are harmonic progressions extracted from existing music works for training a harmonic progression prediction model. A machine learning model needs a large amount of annotated data for training to learn the mapping relationship between the input and the output. The knowledge that the harmonic progression prediction model in this embodiment needs to learn is the mapping from emotional parameters to harmonic progressions. Therefore, it is necessary to obtain harmonic progression training data with emotional parameter annotations. For example, the selected emotional dimensions include pleasantness and arousal, and the value ranges of both emotional dimensions are 1-9. There is a first training harmonic progression (Am7, Fmaj7, Dm7, E7), and its corresponding first annotation information is (pleasantness: 1, arousal: 1). There is also a first training harmonic progression (C, G, Am, Em7), and its corresponding first annotation information is (pleasantness: 7, arousal: 7).
[0096] In addition, the number of first training harmonic progressions in each emotional classification quadrant of the first training set is greater than a preset lower limit, and the difference in the number of first training harmonic progressions in different emotional classification quadrants is less than a deviation threshold.
[0097] (2)Unify the key of each first training harmonic progression in the first training set to the target key.
[0098] It can be understood that the key is the general term of the tonic of the key and the mode category. For example, C major is based on C as the tonic and major as the mode category. In this embodiment, in order to simplify the learning task of the model, the relative harmony theory is referred to as the theoretical framework. In this theoretical framework, chords are usually represented by their distance and relationship with the tonic chord. This method can make the structure of the chord progression easier to understand and can transplant the chord progression in different keys. That is, this theoretical framework pays more attention to the relative relationship of the chord progression rather than the specific key in which they are located. Therefore, for the same harmonic progression, they may eventually migrate to different keys, but their relative interval relationships are equivalent. Unifying the keys of all first training harmonic progressions to the target key, then only the relative relationship of the chord progression will be reflected in all first training harmonic progressions, eliminating the additional training burden brought by different keys. In some embodiments, the target key is C major.
[0099] (3)Train an initial harmonic progression prediction model using the first training set with unified key to obtain a harmonic progression prediction model.
[0100] It can be understood that the initial harmonic progression prediction model refers to the original harmonic progression prediction model using default parameters, and the first training set with unified key is the training material for the initial harmonic progression prediction model. The training process of this model is a typical supervised learning process. Taking the first annotation information as input, the predicted harmonic progression is obtained. The initial harmonic progression prediction model is adjusted with the goal of reducing the difference between the predicted harmonic progression and the first training harmonic progression corresponding to the first annotation information, continuously improving the prediction performance of the model until the error requirement is finally met, and the final harmonic progression prediction model is obtained. The structure selected by the initial harmonic progression prediction model can be a transformer structure with 4 layers of encoders and 4 layers of decoders. This structure can achieve good results when performing the sequence-to-sequence (seq2seq) prediction task from emotion parameters to harmonic progressions.
[0101] In one embodiment, obtaining the first training set with the first annotation information includes:
[0102] (1) Obtaining the first training harmonic progression and its corresponding initial emotion parameter set. The initial emotion parameter set includes the initial emotion parameters annotated for the first training harmonic progression by multiple annotation subjects in at least one emotion dimension.
[0103] It can be understood that the collection of the first training set requires two parts, namely the collection of training samples and the annotation of training samples. Among them, the first training harmonic progression can be extracted from the MIDI file of the piano accompaniment work. After obtaining the first training harmonic progression, it is necessary to evaluate the first training harmonic progression from at least one selected emotion dimension to obtain the corresponding emotion parameters. Since the evaluation of emotions is highly subjective, in order to ensure the generality of the annotation information, in this embodiment, multiple annotation subjects are used to annotate each first training harmonic progression together. The annotation results obtained by these annotation subjects are the initial emotion parameters, and the set formed by the initial emotion parameters is the initial emotion parameter set. The initial emotion parameter set needs to be further comprehensively analyzed to obtain the final first annotation information. In addition, the annotation subject here can be an annotator or an annotation model. That is, the initial emotion parameter set can be manually annotated, semi-automatically annotated, or fully automatically annotated.
[0104] (2) For any emotion dimension, filtering out the invalid annotations in the initial emotion parameters, and obtaining the emotion parameter corresponding to this emotion dimension according to the first average value of the initial emotion parameters after filtering.
[0105] It can be understood that since the initial emotion parameter set may include initial emotion parameters corresponding to more than one emotion dimension, each emotion dimension needs to be screened separately. Screening is to exclude the initial emotion parameters that deviate from other annotation subjects beyond the set conditions and retain the initial emotion parameters that can reach a consensus, so that the finally obtained first annotation information is relatively objective and more in line with the preferences of the public. The first average obtained after averaging the screened initial emotion parameters is the emotion parameter corresponding to this emotion dimension. Among them, the determination of invalid annotations can be to first calculate the average of all initial emotion parameters under each emotion dimension, that is, to obtain the second average, and also calculate the standard deviation of all initial emotion parameters. Then, each initial emotion parameter to be screened is compared with the second average respectively. If the absolute value of the difference from the second average is greater than the standard deviation, it can be determined that the degree of deviation of this initial emotion parameter from the annotation results of other annotation subjects is relatively large and does not meet the requirements. This initial emotion parameter can be determined as an invalid annotation.
[0106] In one of the embodiments, the training process of the pitch prediction model includes:
[0107] (1) Obtain a second training set with second annotation information. The second training set includes a plurality of second training harmonic progressions with the tonality unified to the target tonality, and the second annotation information includes the notes included in each chord in the second training harmonic progression.
[0108] It can be understood that the second training harmonic progression is a harmonic progression extracted from existing music works for training the pitch prediction model. The machine learning model needs a large amount of annotated data for training to learn the mapping relationship between the input and the output. The knowledge that the pitch prediction model in this embodiment needs to learn is the mapping between the harmonic progression and the preference selection of chord notes. Therefore, it is necessary to obtain the harmonic progression training data containing note selection annotations. For example, if the second training harmonic progression is (Am, C, Am, G), the corresponding second annotation information needs to reflect the notes included in each chord. The second annotation information can be represented in a vectorized form. In the vector, each bit corresponds to one of the 12 alternative notes respectively. If the corresponding bit is 1, it means that the chord contains the note at the corresponding position. If the corresponding bit is 0, it means that the chord does not contain the note at the corresponding position. For example, Am: [0, 0, 1, 0, 1, 0, 1, 0, 0, 0, 1, 0]. Since the harmonic progression prediction model focuses more on the relative relationship of the chord progression rather than the specific tonality in the relative harmony theory framework, in order to be consistent with the previous model and reduce the learning difficulty, the second training harmonic progression in this embodiment also needs to unify the tonality to the target tonality. The second training harmonic progression here can be extracted from a piece of music different from the first training harmonic progression, or it can be the same.
[0109] (2)Train the initial pitch prediction model using the second training set to obtain the pitch prediction model.
[0110] It can be understood that the initial pitch prediction model refers to the original pitch prediction model with default parameters, and the second training set is the training material for the initial pitch prediction model. Different from the sequence-to-sequence prediction model of the harmonic progression prediction model, the initial pitch prediction model is a multi-classification prediction model. Its input is the harmonic progression, and the output is the predicted pitch probability sequence corresponding to the harmonic progression. The training of this type of model is generally based on the second annotation information and the predicted pitch probability sequence to calculate the cross-entropy loss function. By adjusting the parameters of the initial pitch prediction model until the value of the cross-entropy loss function converges or meets the requirements, the final pitch prediction model is obtained. The structure selected for the initial pitch prediction model can be a two-layer bilstm model.
[0111] In one embodiment, the recommended notes corresponding to each chord in the target harmonic progression can also be determined according to the pitch probability sequence. Specifically, it includes: determining the in-chord notes and out-of-chord notes according to the corresponding chord types. The in-chord notes among the top three notes with the highest probability in the pitch probability sequence and the out-of-chord notes among the top five notes with the highest probability are used as the recommended notes.
[0112] It can be understood that the in-chord notes refer to the constituent notes of the chord, which are determined by the chord type of the chord. The out-of-chord notes refer to other notes in the target key except the in-chord notes. For example, in the Cmaj chord in C major, the three notes C, E, and G are the in-chord notes. The notes among the twelve notes corresponding to C major except C, E, and G are the out-of-chord notes.
[0113] The pitch probability sequence reflects the frequency of alternative notes appearing as the chord notes in the training set. A high probability means a high chance of being accepted and used in a large amount of music. As the constituent notes of the chord, the in-chord notes should be used most frequently in the arrangement. Therefore, the in-chord notes among the top three notes with the highest probability are used as the recommended notes. The out-of-chord notes can be further divided into notes that can be used to increase the complexity and expressiveness of the chord, and notes that damage the harmony. The former type of out-of-chord notes can make the chord more rich and are recommended to be used in the chord. The latter type of out-of-chord notes should be avoided in the chord. The former type of out-of-chord notes will also be used relatively frequently in this chord, and this frequent use will also be reflected in the pitch probability sequence, making it have a relatively large selection probability. Therefore, the out-of-chord notes among the top five notes with the highest probability are also used as the recommended notes.
[0114] In one embodiment, determining the avoidance notes corresponding to each chord in the target harmonic progression according to the pitch probability sequence includes:
[0115] (1)Determine the out-of-key notes according to the target key.
[0116] It can be understood that each target tonality will correspond to twelve notes, and the notes outside these twelve notes are out-of-tune notes. Out-of-tune notes themselves are dissonant parts in the structure of the music, and their use needs to have a clear purpose and meaning.
[0117] (2) Determine the chord tones based on the corresponding chord type.
[0118] It can be understood that the chord tones refer to the constituent tones of the chord, which are determined by the chord type of the chord. The chord tones outside the chord refer to the other notes in the target tonality except the chord tones. For example, in the Cmaj chord in C major, the three notes C, E, and G are the chord tones. The notes other than C, E, and G in the twelve notes corresponding to C major are the chord tones outside the chord.
[0119] (3) The out-of-tune tones and the out-of-chord tones in the first two notes with the smallest probability among the notes other than the out-of-tune tones in the probability sequence of pitch are taken as avoidance tones.
[0120] It can be understood that the pitch probability sequence reflects the frequency of candidate notes appearing as notes of the chord in the training set. A high probability means a high probability of being accepted and used in a large amount of music. The chord tones can be further divided into notes that can be used to increase the complexity and expressiveness of the chord, and notes that destroy the harmony. The former chord tones can make the chord richer and are recommended for use in the chord. The latter chord tones should be avoided in the chord. The latter chord tones will rarely appear in the chord, and this cautious use will also be reflected in the pitch probability sequence, making it have a smaller selection probability. Therefore, the chord tones in the smallest first two notes are used as avoidance tones. In addition, if there is a conflict between the avoidance tone and the selection and recommendation of the tone, the note is preferentially determined as the avoidance tone.
[0121] The present application provides a melody generation device, including a data acquisition module, a motive matching module, a data analysis module, a target motive selection module and a melody generation module.
[0122] The data acquisition module is used to obtain the target lyrics. The motive matching module is used to match the target sentence in the target lyrics with any melody motive to be matched in the motive library using multiple matching methods, and obtain multiple lyrics structure similarity scores corresponding to the melody motive. The data analysis module is used to obtain the corresponding total structure similarity score based on the multiple lyrics structure similarity scores corresponding to the melody motive. The target motive selection module is used to determine the melody motive with the highest total structure similarity score as the target motive. The melody generation module is used to generate the melody corresponding to the target lyrics based on the target motive.
[0123] For the specific limitations of the melody generation device, reference can be made to the limitations on the melody generation method in the foregoing text, which will not be elaborated here. Each module in the above melody generation device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0124] The present application provides a computer device, including one or more processors, and a memory. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more processors, the steps of the melody generation method in any of the above embodiments are executed.
[0125] Schematically, as Figure 2 shown, Figure 2 is a schematic internal structure diagram of a computer device provided by an embodiment of the present application. Referring to Figure 2 , the computer device 200 includes a processing component 202, which further includes one or more processors, and memory resources represented by the memory 201 for storing instructions executable by the processing component 202, such as application programs. The application programs stored in the memory 201 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 202 is configured to execute instructions to perform the steps of the melody generation method in any of the above embodiments.
[0126] The computer device 200 may further include a power supply component 203 configured to perform power management of the computer device 200, a wired or wireless model interface 204 configured to connect the computer device 200 to the model, and an input / output (I / O) interface 205.
[0127] The present application provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the melody generation method in any of the above embodiments.
[0128] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0129] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0130] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A melody generation method, characterized in that: include: Get target lyrics; For any melody motive to be matched in the motive library, a plurality of matching methods are used to match the target sentence in the target lyrics with the melody motive, so as to obtain a plurality of lyrics structure similarity scores corresponding to the melody motive; Obtaining a corresponding total structural similarity score according to the plurality of lyrics structural similarity scores corresponding to the melody motive; determining the melody motive with the highest total structural similarity score as the target motive; According to the target motive, a melody corresponding to the target lyrics is generated.
2. The melody generation method according to claim 1, characterized in that: Generating a melody corresponding to the target lyrics according to the target motive includes: Determine a melody generation mode combination; the melody generation mode combination includes a preset number of melody generation modes, and the melody generation mode is at least one selected from a plurality of optional modes; The melody generation modes in the melody generation mode combination are selected in turn in a cyclic manner, and a melody is generated for each sentence in the target lyrics in turn according to the selected melody generation mode and the target motive.
3. The melody generation method according to claim 2, characterized in that: The melody motive includes relative pitch information, and the relative pitch information includes the relative pitch of the notes contained in each word in the motive lyrics; the melody is generated for each sentence in the target lyrics in sequence according to the selected melody generation mode and the target motive, including: Determining the tonality of the target lyrics; Obtaining absolute pitch information according to the tonality and the relative pitch information; According to the selected melody generation mode and the absolute pitch information, a melody is generated for each sentence in the target lyrics in turn.
4. The melody generation method according to claim 2, characterized in that: The optional techniques include repetition, modulation, topology, and symmetry.
5. The melody generation method according to claim 1, characterized in that: The melody motive includes motive lyrics, the matching method includes word count similarity, and the lyrics structure similarity score includes a word count similarity score; the target sentence in the target lyrics is matched with the melody motive by using multiple matching methods to obtain multiple lyrics structure similarity scores corresponding to the melody motive, including: The word count similarity score is obtained according to the difference between the word count of the motivation lyrics and the word count of the target sentence.
6. The melody generation method according to claim 1, characterized in that: The melody motive includes motive lyrics, the matching method includes word segmentation structure similarity, and the lyrics structure similarity score includes a word segmentation structure similarity score; the target sentence in the target lyrics is matched with the melody motive by using multiple matching methods to obtain multiple lyrics structure similarity scores corresponding to the melody motive, including: Segmenting the motivation lyrics and the target sentence respectively to obtain corresponding segmentation structure sequences; The dynamic time warping method is used to process the word segmentation structure sequence corresponding to the motivation lyrics and the target sentence to obtain the word segmentation structure similarity score.
7. The melody generation method according to claim 1, characterized in that: The melody motive includes a rhythm structure sequence, the matching method includes a rhythm structure similarity, and the lyrics structure similarity score includes a rhythm structure similarity score; the target sentence in the target lyrics is matched with the melody motive by using multiple matching methods to obtain multiple lyrics structure similarity scores corresponding to the melody motive, including: determining the rhythmic structure sequence of the target sentence; The rhythm structure sequence corresponding to the melody motive and the target sentence is processed by using a dynamic time warping method to obtain the rhythm structure similarity score.
8. The melody generation method according to claim 1, characterized in that: The melody motive further includes a filtering attribute, and the melody generation method further includes: Determine the filtering condition corresponding to the target lyrics; The melody motive in the motive library whose filtering attribute meets the filtering condition is determined as the melody motive to be matched.
9. A melody generating device, characterized in that: include: A data acquisition module, used to acquire target lyrics; A motive matching module is used to match a target sentence in the target lyrics with any melody motive to be matched in the motive library by using multiple matching methods to obtain multiple lyrics structure similarity scores corresponding to the melody motive; A data analysis module, configured to obtain a corresponding total structural similarity score according to the plurality of lyrics structural similarity scores corresponding to the melody motive; A target motive selection module, used for determining the melody motive with the highest total structural similarity score as the target motive; The melody generation module is used to generate a melody corresponding to the target lyrics according to the target motive.
10. A computer device, characterized in that: The invention comprises one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the melody generation method according to any one of claims 1 to 8 are executed.
11. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the melody generation method according to any one of claims 1 to 8.