Music generation method based on natural language prompt
Through the music generation method based on natural language prompts, and using technical means such as the Transformer encoder and multi-head attention mechanism, the problem of insufficient quality and diversity of music generation in the existing technology is solved, and efficient and high-precision music generation is achieved, which meets music theory and user needs.
Patent Information
- Application Number
- CN202411925587.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art lacks quality and diversity in the field of music generation, making it difficult to generate personalized music according to the specific needs of users, and lacks an in-depth understanding of the music structure and rhythm.
The music generation method based on natural language prompts is adopted, and natural language is embedded into high-dimensional vectors through the Transformer encoder. Combined with a tree network encoder and a multi-head attention mechanism, music elements are encoded and spliced, and notes are generated using the BP network until the music generation is over.
The speed and quality of music generation are improved, the model's understanding of the relationship between music elements is enhanced, and the generated music is more in line with music theory and user intentions, and the degree of diversity and personalization is improved.
Smart Images

Figure BDA0005209177570000031 
Figure BDA0005209177570000032 
Figure BDA0005209177570000033
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio technology, and in particular to a music generation method based on natural language prompts. Background Art
[0002] Music generation technology is an important application in the field of computer technology. With the rapid development of artificial intelligence technology, natural language processing technology is widely used in music generation, using machine learning models to automatically generate music data based on text data.
[0003] Existing technologies in the field of music generation mainly rely on rule-driven or simple statistical models, and have the following technical problems: insufficient quality and diversity of generated music; difficulty in generating personalized music based on the specific needs of users; and lack of in-depth understanding of music structure and rhythm.
[0004] Under the premise of the rapid development of artificial intelligence technology, the present invention proposes a music generation method based on natural language prompts, aiming to solve the above problems and achieve efficient and high-precision music generation. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a music generation method based on natural language prompts, which solves the problem that the quality and diversity of generated music are insufficient to meet user needs.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a music generation method based on natural language prompts, comprising the following steps:
[0007] S1. Obtain the user input natural language prompt words from the front-end / mobile text box and send them to the back-end for processing;
[0008] S2. Embed it into a high-dimensional vector based on the Transformer encoder;
[0009] S3. Initialize the music <songstart>Marking, as the starting point for music generation;
[0010] S4. First encode and concatenate the notes in the music, then use a multi-head attention mechanism or replace them with a tree network, then concatenate the note encoding vector with the natural language vectorization vector, and then generate notes at the end of the music based on the concatenated vector through the BP network, and finally repeat these steps until <songend>Mark the end of music generation;
[0011] S5. Return the music to the front end in the form of ABC score or midi, and display it using the abc.js component.
[0012] Preferably, the S4 step is specifically as follows:
[0013] S41. Encode each note in the music, including pitch and position encoding, and perform splicing;
[0014] S42. Use a multi-head attention mechanism for the notes in music, or replace it with a tree network;
[0015] S43. Concatenate the vector obtained by encoding the musical note with the vector obtained by vectorizing the natural language;
[0016] S44. Generate a note and insert it into the end of the music through a BP network according to the vector obtained by splicing;
[0017] S45. Repeat steps S41 to S44 until a <songend>Mark, end music generation.
[0018] Preferably, the pitch encoding in step S41 is one-hot encoding using relative tuning, wherein the pitch dimension is at most 24, and the position encoding is spliced rhythm rotation position encoding and multi-layer absolute position encoding based on beats.
[0019] Preferably, the calculation method of the rhythm rotation position encoding is as follows:
[0020] a. Rotation formula
[0021] Given a position p and word vector x, the formula for rhythmic rotation position encoding (BRoPE) is as follows:
[0022]
[0023] Among them, α p is the rotation angle of position p, x a and x b are the different dimensions of word vectors;
[0024] b. Time signature and duration calculation rotation angle α p
[0025] Rotation angle α p The calculation formula is:
[0026]
[0027] Where T is the total number of notes in a time signature cycle; p is the position index, the position of the note head, that is, the start time of the note; θ total It is the total rotation angle in one cycle, which is 360 degrees.
[0028] Preferably, the calculation method of the beat-based multi-layer absolute position encoding is as follows:
[0029] Positional Encoding P x (t) describes the absolute position at time t relative to the structural unit x, and its calculation formula is as follows:
[0030]
[0031] Among them, t represents a moment, for a note, it can be its start time and duration; x represents the duration of a musical structure unit such as a beat or a measure.
[0032] Preferably, a music generation system based on natural language prompts comprises an application server, an AI server, a data server and a front end, wherein the application server and the data server exchange data via a communication connection, the application server and the AI server perform model training and reasoning via a communication connection, and the front end and the application server interact via the Internet;
[0033] The AI server is used to provide computing power for the music generation model, optimize model performance, and realize music generation;
[0034] The data server is used to store user data, music model parameters and generation results to ensure safe and efficient data access;
[0035] The application server is used to receive natural language prompts input by the user, display an interactive interface, and coordinate data processing between the AI server and the data server.
[0036] The present invention provides a music generation method based on natural language prompts. It has the following beneficial effects:
[0037] 1. The present invention improves the speed and quality of music generation through a tree network encoder and a multi-head attention mechanism. The application of the tree network encoder makes the music generation process more structured and hierarchical. The application of the multi-head attention mechanism improves the model's understanding of the relationship between music elements.
[0038] 2. The present invention adopts rhythmic rotation position encoding (BRoPE) and multi-layer absolute position encoding. The combination of BroPE and multi-layer absolute position encoding enables the model to better process music rhythm and structure, making the generated music more in line with music theory and user intention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] Example:
[0041] The embodiment of the present invention provides a method for generating music based on natural language prompts, comprising the following steps:
[0042] S1. Obtain the user input natural language prompt words from the front-end / mobile text box and send them to the back-end for processing;
[0043] S2. Embed it into a high-dimensional vector based on the Transformer encoder;
[0044] S3. Initialize the music <songstart>Marking, as the starting point for music generation;
[0045] S4. First encode and concatenate the notes in the music, then use a multi-head attention mechanism or replace them with a tree network, then concatenate the note encoding vector with the natural language vectorization vector, and then generate notes at the end of the music based on the concatenated vector through the BP network, and finally repeat these steps until <songend>Mark the end of music generation;
[0046] S5. Return the music to the front end in the form of ABC score or midi, and display it using the abc.js component.
[0047] The S4 steps are as follows:
[0048] S41. Encode each note in the music, including pitch and position encoding, and perform splicing;
[0049] S42. Use a multi-head attention mechanism for the notes in music, or replace it with a tree network;
[0050] S43. Concatenate the vector obtained by encoding the musical note with the vector obtained by vectorizing the natural language;
[0051] S44. Generate a note and insert it into the end of the music through a BP network according to the vector obtained by splicing;
[0052] S45. Repeat steps S41 to S44 until a <songend>Mark, end music generation.
[0053] The pitch encoding in step S41 is a one-hot encoding using relative tuning, where the pitch dimension is at most 24, and the position encoding is a spliced rhythm rotation position encoding and a multi-layer absolute position encoding based on the beat.
[0054] Rotational position encoding shows significant advantages in processing periodic sequence data. In the field of music creation, the beat of music has obvious periodicity. To this end, a BRoPE encoding method based on music time information is proposed, which rotates the word vector at each beat to adapt to the rhythm of the music. By applying RoPE within a measure, the time signature properties of the music can be significantly emphasized, ensuring that the generated music meets the required time signature. The calculation method of rhythm rotational position encoding is as follows:
[0055] a. Rotation formula
[0056] Given a position p and word vector x, the formula for rhythmic rotation position encoding (BRoPE) is as follows:
[0057]
[0058] Among them, α p is the rotation angle of position p, x a and x b are different dimensions of word vectors; in addition, considering that the main melody of a pop song usually varies within two octaves, there are actually only 24 types of pitch. Therefore, in the case of one-hot encoding, the pitch dimension of the melody vector only needs 24 dimensions at most to be fully represented, which makes them very suitable for rotation encoding using RoPE, which can be quickly calculated or cached using memoization functions. If the pitch is further embedded, the vector dimension of the pitch can be reduced to a smaller scale while retaining the relative relationship between the pitches and the musical characteristics.
[0059] b. Time signature and duration calculation rotation angle α p
[0060] Rotation angle α p The calculation formula is:
[0061]
[0062] Where T is the total number of notes in a time signature cycle, for example, in 4 / 4 time, T = 4; p is the position index, the position of the note head, that is, the start time of the note; θ total It is the total rotation angle in one cycle, which is 360 degrees.
[0063] Taking 4 / 4 time as an example, set the position vector corresponding to each quarter note to rotate 90 degrees to ensure that each measure can complete a full 360-degree rotation. For eighth notes, the rotation angle is 45 degrees, for sixteenth notes it is 22.5 degrees, and so on. According to different note values and time signatures, the above formula is used to adjust the rotation angle, and the corresponding dimensions of the word vector are rotated using the rotation matrix and the rotation angle, which can effectively integrate the music rhythm information into the word vector, thereby improving the model's ability to process time series data in music creation.
[0064] c. Calculate the rotated vector of each unique pitch
[0065] Based on the circle of fifths, the farthest pitch pairs are selected to define the rotation plane. Specifically, if x is set a represents C (with the encoding subscript 0), then x b will correspond to F# (coded as 6). Thus, within the octave range, six rotation planes can be divided, each corresponding to a note. For any pitch, the pitch subscript that forms the rotation plane with it can be calculated by the formula m=(n+6)mod12.
[0066] Select C, G, D, A, E, F in C major as x a The pitch of their corresponding x b The pitches are Gb, Ab, Eb, Bb, B. For example, for C, F#:
[0067] <x C , x F# > = BRoPE(x C , x F# , p)
[0068] Here<xC,xF#> The rotation encoding vector representing the composition of the C pitch and the F# pitch is rotationally encoded at a given position p through the BRoPE function to embed the position information into the pitch information.
[0069] The calculation method of the multi-layer absolute position encoding based on the beat is as follows:
[0070] Positional Encoding P x (t) describes the absolute position at time t relative to the structural unit x, and its calculation formula is as follows:
[0071]
[0072] Among them, t represents a moment, for a note, it can be its start time and duration; x represents the duration of a musical structure unit such as a beat or a measure.
[0073] As shown above, in order to more accurately characterize the timing characteristics of the sound in music, a hierarchical absolute position encoding method is used. For the starting point (note head) of the sound, its absolute position relative to the following musical structural units is calculated: tick (the smallest time unit, usually corresponding to a clock pulse in the MIDI file), beat, measure, phrase, and segment. This encoding method allows the model to capture extremely delicate time details in music. As for the duration of the sound, only its relative position within the measure is calculated. The values of these positions are then concatenated after the vector as position encoding.
[0074] For the tick period, the model captures the precise moment of the note attack in the music stream through sine and cosine encoding, thereby perceiving the most subtle time fluctuations.
[0075] Encoding at the beat period level helps the model to firmly grasp the rhythmic sense of the rhythm.
[0076] The bar cycle helps the model understand and express the syntactic structure of music and provides a framework for the construction of melody.
[0077] The encoding of phrases and periodicities takes a macro perspective, allowing the model to lay out and shape the dynamic changes and turning points of the melody within the broader musical narrative.
[0078] Embodiment 2:
[0079] The embodiment of the present invention provides a music generation system based on natural language prompts, which is characterized by comprising an application server, an AI server, a data server and a front end, wherein the application server and the data server exchange data through a communication connection, the application server and the AI server perform model training and reasoning through a communication connection, and the front end and the application server interact through the Internet;
[0080] The AI server is used to provide computing power for the music generation model, optimize model performance, and realize music generation;
[0081] The data server is used to store user data, music model parameters and generation results, ensuring secure and efficient data access;
[0082] The application server is used to receive natural language prompts input by users, display the interactive interface, and coordinate data processing between the AI server and the data server.
[0083] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.< / songend> < / songend> < / songstart> < / songend> < / songend> < / songstart>
Claims
1. A music generation method based on natural language prompts, characterized in that: The following steps are involved: S1. Obtain the user input natural language prompt words from the front-end / mobile text box and send them to the back-end for processing; S2. Embed it into a high-dimensional vector based on the Transformer encoder; S3. Initialize the music <songstart> Marking, as the starting point for music generation;< / songstart> S4. First encode and concatenate the notes in the music, then use a multi-head attention mechanism or replace them with a tree network, then concatenate the note encoding vector with the natural language vectorization vector, and then generate notes at the end of the music based on the concatenated vector through the BP network, and finally repeat these steps until <songend> Mark the end of music generation;< / songend> S5. Return the music to the front end in the form of ABC score or midi, and display it using the abc.js component.
2. A method for generating music based on natural language prompts according to claim 1, characterized in that: The S4 step is specifically as follows: S41. Encode each note in the music, including pitch and position encoding, and perform splicing; S42. Use a multi-head attention mechanism for the notes in music, or replace it with a tree network; S43. Concatenate the vector obtained by encoding the musical note with the vector obtained by vectorizing the natural language; S44. Generate a note and insert it into the end of the music through a BP network according to the vector obtained by splicing; S45. Repeat steps S41 to S44 until a <songend> Mark, end music generation.< / songend> 3. The method for generating music based on natural language prompts according to claim 2, characterized in that: The pitch encoding in the step S41 is a one-hot encoding using relative tuning, wherein the pitch dimension is at most 24, and the position encoding is a spliced rhythm rotation position encoding and a multi-layer absolute position encoding based on the beat.
4. The method for generating music based on natural language prompts according to claim 3, characterized in that: The calculation method of the rhythm rotation position encoding is as follows: a. Rotation formula Given a position p and word vector x, the formula for rhythmic rotation position encoding (BRoPE) is as follows: Among them, α p is the rotation angle of position p, x a and x b are the different dimensions of word vectors; b. Calculate the rotation angle α for time signature and duration p Rotation angle α p The calculation formula is: Where T is the total number of notes in a time signature cycle; p is the position index, the position of the note head, that is, the start time of the note; θ total It is the total rotation angle in one cycle, which is 360 degrees.
5. The method for generating music based on natural language prompts according to claim 1, characterized in that: The calculation method of the multi-layer absolute position encoding based on the beat is as follows: Positional Encoding P x (t) describes the absolute position at time t relative to the structural unit x, and its calculation formula is as follows: Among them, t represents a moment, for a note, it can be its start time and duration; x represents the duration of a musical structure unit such as a beat or a measure.
6. A music generation system using the music generation method based on natural language prompts according to any one of claims 1 to 5, characterized in that: It includes an application server, an AI server, a data server and a front end. The application server and the data server exchange data through a communication connection, the application server and the AI server perform model training and reasoning through a communication connection, and the front end and the application server interact through the Internet; The AI server is used to provide computing power for the music generation model, optimize model performance, and realize music generation; The data server is used to store user data, music model parameters and generation results to ensure safe and efficient data access; The application server is used to receive natural language prompts input by the user, display an interactive interface, and coordinate data processing between the AI server and the data server.