Methods, apparatus, equipment and readable storage media for generating music

By acquiring chord sequences, reference textures of the target real music, and texture data generated by the deep model REMI, analyzing and updating the texture structure, and generating music, the problem of high manual costs in existing technologies is solved, and efficient automation and innovation in music composition are achieved.

CN119028301BActive Publication Date: 2025-10-31GUANGZHOU QUYAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411269883.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-10-31
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

The existing technology for generating new textures is costly and heavily reliant on music arrangement experts, resulting in a high barrier to entry for music creation.

Method used

By acquiring chord sequences, reference textures of the target real music, and multiple texture data generated by the deep model REMI, the texture structure is analyzed and updated to generate the target texture, which is then combined with the chord sequence to generate the music.

Benefits of technology

It reduces labor costs, improves the innovation, harmony, and audience appeal of music, and reduces the reliance on experts in music creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119028301B_ABST
    Figure CN119028301B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and readable storage medium for generating music. The method acquires a chord sequence, a reference texture of a target real-world piece of music, and multiple texture data generated based on the deep model REMI. For each texture structure, the method analyzes the texture structure, updates the texture data with the highest matching degree to generate a target texture, combines the target textures to generate a musical texture for interpreting the chord sequence, and generates a piece of music based on the musical texture and chord sequence. Since the texture structures referenced by the combined target textures belong to the same reference texture, the harmony of the generated new music is improved. Therefore, this application can adjust the texture data with reference to a real-world texture, combining the musical texture with the chord sequence to generate a new piece of music, reducing reliance on experts in the music composition process, and improving the innovation of the musical texture while ensuring its harmony, layering, and audience recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method, apparatus, device, and readable storage medium for generating music. Background Technology

[0002] Listening to music is a common form of leisure and entertainment, while traditional music creation mainly relies on music producers.

[0003] The creation of musical compositions requires music producers to possess extensive music theory knowledge, resulting in a high barrier to entry. To address this issue, existing technologies offer a method for generating new textures based on texture templates created by music arrangers, and then using these new textures to create new musical compositions. However, generating texture templates requires music arrangers to understand the textures of different types of music and their key aspects, which is time-consuming, labor-intensive, and heavily reliant on music arrangers. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, device and readable storage medium for generating music, in order to solve the disadvantage of high manual cost in generating new textures in the prior art.

[0005] To achieve the above objectives, the following solution is proposed:

[0006] A method for generating music, comprising:

[0007] Acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music and are composed of multiple texture structures.

[0008] For each texture structure of the reference texture, the texture structure is analyzed, and the texture data with the highest matching degree with the texture structure in each texture data is updated to generate the target texture corresponding to the texture structure;

[0009] The target textures are combined to generate a musical texture for interpreting the chord sequence;

[0010] A musical piece is generated based on the musical texture and the chord sequence.

[0011] Optionally, multiple texture data generated based on the deep model REMI, including:

[0012] Multiple text prompt parameters, MIDI data labeled with text prompt parameters, or random seeds are input into the REMI to obtain multiple output data;

[0013] Each output data is segmented based on a preset set of time value thresholds to obtain multiple texture data.

[0014] Optionally, obtain the reference texture of the target real music, including:

[0015] Determine the type of music to be generated;

[0016] Select real music pieces from multiple real music pieces that match the target music piece type as the target real music piece;

[0017] Use the texture of the target real music as the reference texture.

[0018] Optionally, the target music type may consist of any one or more of the following: target music style, target music mood, target music section type, and target bpm.

[0019] Optionally, the texture structure is analyzed, and the texture data with the highest matching degree to the texture structure among all texture data is updated to generate the target texture corresponding to the texture structure, including:

[0020] Determine the first musical attribute of the texture structure, and the second musical attribute of each texture data;

[0021] Select the target music attribute that has the highest similarity to the first music attribute from each of the second music attributes, and the texture data corresponding to the target music attribute is the texture data that has the highest matching degree with the texture structure;

[0022] Based on the derivation method of the texture structure, the texture data with the highest matching degree with the texture structure is updated to form the target texture corresponding to the texture structure.

[0023] Optionally, determining a first musical attribute of the texture structure includes:

[0024] The first time value, first density, first stability, first rhythm, first segment trend, first pitch width, and first pitch range of the texture structure are determined.

[0025] The first musical attribute is composed of the first time value, the first density, the first stability, the first rhythm, the first paragraph trend, the first pitch width, and the first pitch range.

[0026] Optionally, the deduction method based on the texture structure updates the texture data with the highest matching degree to form the target texture corresponding to the texture structure, including:

[0027] Based on the note interpretation method of the texture structure, the note interpretation method of the texture data with the highest matching degree with the texture structure is adjusted, and the target texture corresponding to the texture structure is obtained after adjustment.

[0028] A musical texture generation device, comprising:

[0029] The acquisition module is used to acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music, and the reference textures are composed of multiple texture structures.

[0030] The analysis module is used to analyze the texture structure for each texture structure of the reference texture, update the texture data with the highest matching degree to the texture structure in each texture data, and generate the target texture corresponding to the texture structure.

[0031] The combination module is used to combine the various target textures to generate a musical texture for interpreting the chord sequence;

[0032] The generation module is used to generate a musical piece based on the musical texture and the chord sequence.

[0033] A musical texture generation device, including a memory and a processor;

[0034] The memory is used to store programs;

[0035] The processor is used to execute the program to implement the various steps of the above-described music generation method.

[0036] A readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the various steps of the above-described music generation method.

[0037] As can be seen from the above technical solution, the music generation method provided in this application can obtain chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference texture is used to characterize the chord interpretation of the target real music, and the reference texture is composed of multiple texture structures. For each texture structure of the reference texture, the texture structure is analyzed, and the texture data with the highest matching degree among the various texture data is updated to generate the target texture corresponding to the texture structure. Based on this, the target texture of this application is the data obtained after updating the texture data generated by REMI. The texture data generated by REMI has strong randomness, thereby improving the innovation of the target texture and avoiding the homogenization and mechanical listening experience of music textures. At the same time, when generating the target texture, the texture of the real song is used as a reference to learn the note arrangement method of the real texture and apply the note arrangement method of the real texture to the texture data generated by REMI, ensuring the harmony, hierarchy, and audience recognition of the target texture. Meanwhile, the target songs are readily available, requiring no additional input from arrangement experts, and the texture data generated by REMI also necessitates no further input from arrangement experts, reducing labor costs. Subsequently, the various target textures can be combined to generate a musical texture for interpreting the chord sequence; based on the musical texture and the chord sequence, a musical piece is generated. On this basis, this application can combine musical textures and chord sequences to form new musical pieces, completing the construction of new musical pieces. Furthermore, since the textures referenced by the various target textures combined in this application belong to the same reference texture, the adaptability of each target texture is strong, thereby improving the progression, coordination, and audience resonance of the musical textures obtained by combining the various target textures, further enhancing the hierarchy and harmony of the chord sequence interpretation. It is evident that this application can adjust the texture data generated by REMI with reference to real textures to generate musical textures, and combine the musical textures with chord sequences to form new musical pieces. Thus, in the process of constructing new musical pieces, labor costs are reduced, and reliance on experts in the music creation process is decreased. At the same time, this application can also enhance the innovation of newly generated music while ensuring its coherence, layering, and audience acceptance. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0039] Figure 1 This is a flowchart of a music generation method disclosed in an embodiment of this application;

[0040] Figure 2 This is a structural block diagram of a music generation device disclosed in an embodiment of this application;

[0041] Figure 3 This is a hardware structure block diagram of a music generation device disclosed in an embodiment of this application. Detailed Implementation

[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] This application provides a method for generating music. This method can be applied to various music creation systems or music exchange systems, and can also be applied to various computer terminals or smart terminals. The executing entity can be the processor or server of the computer terminal or smart terminal. The flowchart of the method is shown below. Figure 1 As shown, it specifically includes:

[0044] Step S1: Obtain the chord sequence, the reference texture of the target real music, and multiple texture data generated based on the deep model REMI.

[0045] Specifically, the chord sequence can consist of a combination of multiple target chords, and each target chord can be obtained in a variety of ways. For example, it can be obtained in response to the user's chord specification operation, using the user-specified chord as the target chord; or it can be obtained using the chord recommended by the chord generation system as the target chord.

[0046] You can select the target real music from various authorized songs and obtain the texture of the target real music as the reference texture.

[0047] The target real music can be selected in several ways. For example, it can select a specific song from various authorized songs as the target real music in response to the user's specification; or it can select the target real music from various authorized songs according to the actual target music type requirements.

[0048] The reference texture is used to represent the instrument playing style and chord interpretation style corresponding to the target real music.

[0049] The reference texture is composed of multiple texture structures.

[0050] Different ways of playing musical instruments have different effects on the emotions of the audience.

[0051] Each texture data can consist of multiple monotone data, used to construct chords.

[0052] Each note's data can include pitch, loudness, note start time, and note duration.

[0053] Each texture data may contain the necessary structure to form a chord, such as the root, third, and fifth.

[0054] REMI is an open-source model, and the texture distribution of its output data lacks consistency. Therefore, the output data needs to be processed to obtain the texture data.

[0055] Step S2: For each texture structure of the reference texture, analyze the texture structure, update the texture data with the highest matching degree to the texture structure in each texture data, and generate the target texture corresponding to the texture structure.

[0056] Specifically, machine algorithms can be used to learn each texture structure and update the matching texture data to obtain the target texture.

[0057] Step S3: Combine the target textures to generate a musical texture for interpreting the chord sequence.

[0058] Specifically, the target textures can be combined according to the combination method of each texture structure in the reference texture to generate the musical texture.

[0059] The music texture can be a MIDI file, a WAV file, or an MP3 file.

[0060] Step S4: Generate a musical piece based on the musical texture and the chord sequence.

[0061] Specifically, the chord sequence can be played according to the musical texture to form a piece of music.

[0062] As can be seen from the above technical solutions, the music generation method provided in this application can obtain chord sequences, reference textures of target real music, and multiple texture data generated based on the deep model REMI. The reference texture is used to characterize the chord interpretation of the target real music, and the reference texture is composed of multiple texture structures. For each texture structure of the reference texture, the texture structure is analyzed, and the texture data with the highest matching degree among the various texture data is updated to generate the target texture corresponding to the texture structure. Based on this, the target texture of this application is the data obtained after updating the texture data generated by REMI. The texture data generated by REMI has strong randomness, thereby improving the innovation of the target texture and avoiding the homogenization and mechanical listening experience of music textures. At the same time, when generating the target texture, the texture of the real song is used as a reference to learn the note arrangement method of the real texture and apply the note arrangement method of the real texture to the texture data generated by REMI, ensuring the harmony, hierarchy, and audience recognition of the target texture. Meanwhile, the target songs are readily available, requiring no additional input from arrangement experts, and the texture data generated by REMI also necessitates no further input from arrangement experts, reducing labor costs. Subsequently, the various target textures can be combined to generate a musical texture for interpreting the chord sequence; based on the musical texture and the chord sequence, a musical piece is generated. On this basis, this application can combine musical textures and chord sequences to form new musical pieces, completing the construction of new musical pieces. Furthermore, since the textures referenced by the various target textures combined in this application belong to the same reference texture, the adaptability of each target texture is strong, thereby improving the progression, coordination, and audience resonance of the musical textures obtained by combining the various target textures, further enhancing the hierarchy and harmony of the chord sequence interpretation. It is evident that this application can adjust the texture data generated by REMI with reference to real textures to generate musical textures, and combine the musical textures with chord sequences to form new musical pieces. Thus, in the process of constructing new musical pieces, labor costs are reduced, and reliance on experts in the music creation process is decreased. At the same time, this application can also enhance the innovation of newly generated music while ensuring its coherence, layering, and audience acceptance.

[0063] In some embodiments of this application, the process of obtaining multiple texture data generated based on the depth model REMI in step S1 is described in detail, and the steps are as follows:

[0064] S10. Input the text prompt parameters, MIDI data with text prompt parameters marked, or random seed into the REMI multiple times to obtain multiple output data.

[0065] Specifically, data can be repeatedly input into REMI to obtain output data;

[0066] Each input can be a text prompt parameter, MIDI data with text prompt parameters, or a random seed.

[0067] Text hint parameters can be used to characterize texture type requirements.

[0068] The duration of the MIDI data can be set according to actual needs; for example, it can be set to 3 minutes.

[0069] The random seed can be a randomly generated parameter.

[0070] S11. Based on a preset set of time value thresholds, each output data is segmented to obtain multiple texture data.

[0071] Specifically, the time-value threshold set can contain multiple time-value thresholds.

[0072] For each time value threshold, each output data can be segmented according to that time value threshold to obtain multiple texture data, and the time values ​​of texture data corresponding to the same time value threshold are the same.

[0073] For example, when the time value threshold is 1 beat, each output data can be segmented to obtain texture data where the time value is 1 beat.

[0074] It can clean the data of each texture.

[0075] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining multiple texture data. Through the above method, various types of data can be input into REMI and the output data of REMI can be processed to obtain multiple texture data, thereby improving the randomness and diversity of texture data.

[0076] In some embodiments of this application, the process of obtaining the reference texture of the target real music in step S1 is described in detail, and the steps are as follows:

[0077] S12. Determine the type of music to be generated.

[0078] Specifically, it can respond to users' creative needs and determine the type of target music to be generated.

[0079] The target music genre is determined by multi-dimensional information.

[0080] S13. Select real music pieces that match the target music type from multiple real music pieces as the target real music pieces.

[0081] Specifically, one can select a real song that matches the target song type from multiple authorized real songs as the target real song.

[0082] S14. Use the texture of the target real music as the reference texture.

[0083] Specifically, the texture of the target real music can be obtained and used as a reference texture.

[0084] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining reference textures. The above method can extract reference textures that meet the creative requirements and better complete texture generation.

[0085] In some embodiments of this application, the music type is determined by multi-dimensional information, and different music types may contain the same dimensional information.

[0086] Music genres can be composed of one or more dimensions of information, including musical style, musical mood, musical section type, and bpm.

[0087] The target music type to be generated can include any one or more dimensions of information such as target music style, target music mood, target music section type, and target bpm.

[0088] In some embodiments of this application, the process of step S2—analyzing the texture structure for each texture structure of the reference texture, updating the texture data with the highest matching degree to the texture structure among the various texture data, and generating the target texture corresponding to the texture structure—is described in detail as follows:

[0089] S20. Determine the first musical attribute of the texture structure, and the second musical attribute of each texture data.

[0090] Specifically, the musical properties of the texture can be used as the first musical property;

[0091] The musical attribute of each texture data can be used as a second musical attribute, and each texture data has a unique corresponding second musical attribute.

[0092] The first musical attribute can be used to characterize the first parameter value of the corresponding texture structure in each dimension.

[0093] Each second musical attribute can be used to characterize the second parameter value of the corresponding texture data in each dimension.

[0094] S21. Select the target music attribute that has the highest similarity to the first music attribute from each of the second music attributes. The texture data corresponding to the target music attribute is the texture data that has the highest matching degree with the texture structure.

[0095] Specifically, the similarity between the first music attribute and each of the second music attributes can be calculated; the texture data with the highest similarity among the various texture data can be selected as the texture data with the highest matching with the texture structure.

[0096] S22. Based on the derivation method of the texture structure, update the texture data with the highest matching degree with the texture structure to form the target texture corresponding to the texture structure.

[0097] Specifically, machine algorithms can be used to learn how to interpret texture structures, adjust the interpretation method of the texture data with the highest matching degree, and obtain the target texture corresponding to the texture structure after adjustment.

[0098] As can be seen from the above technical solution, this embodiment provides an optional method for generating target textures based on texture structure. Through the above method, the similarity between the musical attributes of texture structure and the musical attributes of texture data can be combined to identify the texture data with the highest matching degree with texture structure, thereby improving the reliability of the searched texture data.

[0099] In some embodiments of this application, the process of determining the first musical attribute of the texture structure in step S20 is described in detail, and the steps are as follows:

[0100] S300, determine the first time value, first density, first stability, first rhythm, first segment trend, first pitch width, and first pitch range of the texture structure.

[0101] Specifically, the first time value, first density, first stability, first rhythm, first paragraph trend, first pitch width, and first pitch range can constitute the first musical attribute;

[0102] The time value of the texture structure can be used as the first time value.

[0103] The first time value can be used to represent the time value of the corresponding texture structure;

[0104] The value of the time value varies depending on the actual music being played. Generally, the time value can be 0.0, 1.0, 2.0, 3.0, or 4.0.

[0105] The density of notes in the texture structure can be calculated, and the density of the texture structure can be used as the first density.

[0106] Density can be calculated using the following expression:

[0107]

[0108] Where x1 represents density; ClickNum represents the number of times the note appears; and dur represents duration.

[0109] The first density can be used to represent the density of the playing instrument in the corresponding texture structure.

[0110] The stability of the texture structure can be calculated and used as the first stability.

[0111] The stability can be calculated using the following expression:

[0112]

[0113] Where x2 represents stability; ColumnNum represents the number of notes belonging to the column structure; there are more than three notes played simultaneously in the column structure.

[0114] The first stability factor is used to represent the proportion of columnar structures in the corresponding texture structure.

[0115] The rhythmicity of the texture structure can be determined, and the rhythmicity of the texture structure can be taken as the primary rhythmicity.

[0116] Rhythm represents the special rhythmic patterns formed by the timing / density of the notes struck in the texture structure. The arrangement and combination are composed of uppercase and lowercase letters and Arabic numerals. There are a total of 81 rhythmic combinations for each beat (including empty cases), and there are 81 to the power of 4 (approximately 43 million) rhythmic combinations for a measure.

[0117] The first rhythmic property can be used to represent the rhythmic form of the corresponding texture structure.

[0118] The trend of the texture structure can be taken as the first trend.

[0119] Trends can be categorized into three types: flat, upward, and downward.

[0120] The first trend can be used to represent the tonal changes of the corresponding texture structure.

[0121] The pitch width of the texture can be used as the first pitch width.

[0122] The range of pitch width can be [1, B8-C2].

[0123] The first pitch width can be used to represent the number of semitones between the lowest and highest pitches of a corresponding texture.

[0124] The range of the texture structure can be taken as the first range of the pitch.

[0125] The range of values ​​for the vocal range can be [1, 64].

[0126] The vocal range can be calculated using the following expression:

[0127]

[0128] Where x3 represents the pitch range; MinDomain is the lowest pitch; and MaxDomain is the highest pitch.

[0129] The first pitch range can be used to represent the pitch range interval of the corresponding texture structure.

[0130] As can be seen from the above technical solution, this embodiment provides an optional method for calculating the first musical attribute. Through the above method, the similarity of the texture structure and texture data can be evaluated by comprehensively considering the time value, density, stability, rhythm, trend, pitch width and pitch range. Through multi-dimensional evaluation, the matching degree of the texture structure and texture data can be improved.

[0131] In addition, the second time value, second density, second stability, second rhythm, second segment trend, second pitch width, and second pitch range of each texture data can be calculated.

[0132] The second time value, second density, second stability, second rhythm, second paragraph trend, second pitch width, and second pitch range can constitute the second musical attribute;

[0133] The second density can be used to represent the density of the instruments played in the corresponding texture data; the second stability can be used to represent the proportion of columnar structures in the corresponding texture data; the second rhythm can be used to represent the rhythmic form of the corresponding texture data; the second trend can be used to represent the melody changes of the corresponding texture data; the second pitch width can be used to represent the number of semitones between the lowest and highest pitches of the corresponding texture data; the second pitch range can be used to represent the pitch range of the corresponding texture data.

[0134] In some embodiments of this application, the process of updating the texture data with the highest matching degree to the texture structure based on the deduction method of the texture structure to form the target texture corresponding to the texture structure is described in detail, and the steps are as follows:

[0135] S220. Based on the note interpretation method of the texture structure, adjust the note interpretation method of the texture data with the highest matching degree with the texture structure, and obtain the target texture corresponding to the texture structure after adjustment.

[0136] Specifically, machine algorithms can be used to analyze and learn the interpretation of texture structure, generate random seeds, and input the random seeds into the machine algorithm. Based on the interpretation of texture structure and random seeds, the machine algorithm can adjust the interpretation of notes in the matched texture data, and the target texture can be obtained after adjustment.

[0137] As can be seen from the above technical solution, this embodiment provides an interpretation method based on the texture structure, and an optional method for updating the texture data with the highest matching degree to the texture structure. Through the above method, the note combination method of the texture structure can be combined to specifically update and adjust the note combination method of the texture data, thereby further improving the harmony and sense of hierarchy of the target texture generated by this application.

[0138] Next, we will combine Figure 2 The music generation apparatus provided in this application will be described in detail. The music generation apparatus described below can be compared with the music generation method described above.

[0139] See Figure 2 It can be observed that the music generation device may include:

[0140] The acquisition module 10 is used to acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music, and the reference textures are composed of multiple texture structures.

[0141] Analysis module 20 is used to analyze the texture structure for each texture structure of the reference texture, update the texture data with the highest matching degree to the texture structure in each texture data, and generate the target texture corresponding to the texture structure.

[0142] The combination module 30 is used to combine the various target textures to generate a musical texture for interpreting the chord sequence;

[0143] The generation module 40 is used to generate a musical piece based on the musical texture and the chord sequence.

[0144] Furthermore, the acquisition module may include:

[0145] The output data acquisition unit is used to input text prompt parameters, MIDI data labeled with text prompt parameters, or random seeds into the REMI multiple times to obtain multiple output data.

[0146] The texture data acquisition unit is used to cut each output data based on a preset set of time value thresholds to obtain multiple texture data.

[0147] Furthermore, the acquisition module may also include:

[0148] The target music type acquisition unit is used to determine the target music type to be generated;

[0149] The target real music selection unit is used to select real music that matches the target music type from multiple real music pieces as the target real music piece;

[0150] The reference texture acquisition unit is used to use the texture of the target real music as a reference texture.

[0151] The target music type consists of any one or more of the following: target music style, target music mood, target music section type, and target bpm.

[0152] Furthermore, the analysis module may include:

[0153] A music attribute determination unit is used to determine a first music attribute of the texture structure and a second music attribute of each texture data.

[0154] A music attribute matching unit is used to select the target music attribute that has the highest similarity to the first music attribute from each second music attribute, wherein the texture data corresponding to the target music attribute is the texture data that has the highest matching degree with the texture structure;

[0155] The target texture generation unit is used to update the texture data that has the highest matching degree with the texture structure based on the derivation method of the texture structure, so as to form the target texture corresponding to the texture structure.

[0156] Furthermore, the music attribute determination unit may include:

[0157] A numerical determination component is used to determine the first time value, first density, first stability, first rhythm, first segment trend, first pitch width, and first pitch range of the texture structure;

[0158] The first musical attribute is composed of the first time value, the first density, the first stability, the first rhythm, the first paragraph trend, the first pitch width, and the first pitch range.

[0159] Furthermore, the target texture generation unit may include:

[0160] The texture data adjustment component is used to adjust the note interpretation method of the texture data that best matches the texture structure based on the note interpretation method of the texture structure, and obtain the target texture corresponding to the texture structure after adjustment.

[0161] The music generation device provided in this application embodiment can be applied to music generation equipment, such as PC terminals, cloud platforms, servers, and server clusters. Optionally, Figure 3 The hardware structure block diagram of the music generation device is shown, with reference to... Figure 3 The hardware structure of the music generation device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0162] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0163] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0164] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0165] The memory stores a program, which the processor can call. The program is used for:

[0166] Acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music and are composed of multiple texture structures.

[0167] For each texture structure of the reference texture, the texture structure is analyzed, and the texture data with the highest matching degree with the texture structure in each texture data is updated to generate the target texture corresponding to the texture structure;

[0168] The target textures are combined to generate a musical texture for interpreting the chord sequence;

[0169] A musical piece is generated based on the musical texture and the chord sequence.

[0170] Optionally, the refined and extended functions of the program can be referred to the above description.

[0171] This application embodiment also provides a readable storage medium that can store a program suitable for execution by a processor, the program being used for:

[0172] Acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music and are composed of multiple texture structures.

[0173] For each texture structure of the reference texture, the texture structure is analyzed, and the texture data with the highest matching degree with the texture structure in each texture data is updated to generate the target texture corresponding to the texture structure;

[0174] The target textures are combined to generate a musical texture for interpreting the chord sequence;

[0175] A musical piece is generated based on the musical texture and the chord sequence.

[0176] Optionally, the refined and extended functions of the program can be referred to the above description.

[0177] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0178] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0179] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. The various embodiments of this application can be combined with each other. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating music, characterized in that, include: Acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music and are composed of multiple texture structures. For each texture structure of the reference texture, the texture structure is analyzed, and the texture data with the highest matching degree with the texture structure in each texture data is updated to generate the target texture corresponding to the texture structure; The target textures are combined to generate a musical texture for interpreting the chord sequence; A musical piece is generated based on the musical texture and the chord sequence.

2. The music generation method according to claim 1, characterized in that, Obtain multiple texture data generated based on the deep model REMI, including: Multiple text prompt parameters, MIDI data labeled with text prompt parameters, or random seeds are input into the REMI to obtain multiple output data; Each output data is segmented based on a preset set of time value thresholds to obtain multiple texture data.

3. The music generation method according to claim 1, characterized in that, Obtain the reference texture of the target real music piece, including: Determine the type of music to be generated; Select real music pieces from multiple real music pieces that match the target music piece type as the target real music piece; Use the texture of the target real music as the reference texture.

4. The music generation method according to claim 3, characterized in that, The target music type consists of any one or more of the following: target music style, target music mood, target music section type, and target bpm.

5. The music generation method according to claim 1, characterized in that, The step of analyzing the texture structure, updating the texture data with the highest matching degree among all texture data, and generating the target texture corresponding to the texture structure includes: Determine the first musical attribute of the texture structure, and the second musical attribute of each texture data; Select the target music attribute that has the highest similarity to the first music attribute from each of the second music attributes, and the texture data corresponding to the target music attribute is the texture data that has the highest matching degree with the texture structure; Based on the derivation method of the texture structure, the texture data with the highest matching degree with the texture structure is updated to form the target texture corresponding to the texture structure.

6. The music generation method according to claim 5, characterized in that, Determining the first musical property of the texture structure includes: The first time value, first density, first stability, first rhythm, first segment trend, first pitch width, and first pitch range of the texture structure are determined. The first musical attribute is composed of the first time value, the first density, the first stability, the first rhythm, the first paragraph trend, the first pitch width, and the first pitch range.

7. The music generation method according to claim 5, characterized in that, The deduction method based on the texture structure updates the texture data with the highest matching degree to the texture structure, forming the target texture corresponding to the texture structure, including: Based on the note interpretation method of the texture structure, the note interpretation method of the texture data with the highest matching degree with the texture structure is adjusted, and the target texture corresponding to the texture structure is obtained after adjustment.

8. A musical texture generation device, characterized in that, include: The acquisition module is used to acquire chord sequences, reference textures of the target real music, and multiple texture data generated based on the deep model REMI. The reference textures are used to characterize the chord interpretation of the target real music, and the reference textures are composed of multiple texture structures. The analysis module is used to analyze the texture structure for each texture structure of the reference texture, update the texture data with the highest matching degree to the texture structure in each texture data, and generate the target texture corresponding to the texture structure. The combination module is used to combine the various target textures to generate a musical texture for interpreting the chord sequence; The generation module is used to generate a musical piece based on the musical texture and the chord sequence.

9. A musical texture generation device, characterized in that, Including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the music generation method as described in any one of claims 1-7.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the music generation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Textile combination method and device, computer equipment and storage medium

    CN118397987A

  • System and methods for automatically generating a muscial composition having audibly correct form

    GB202104696D0