Music generation method based on ViT-LSTM model

By constructing a music generation method based on ViT-LSTM model, the problem of music generation relying on training data quality and diversity in the prior art is solved, and higher creativity and expressiveness are achieved, and music sequences that meet different styles and emotional needs can be generated.

CN119943009APending Publication Date: 2025-05-06INST OF INT RELATIONS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411895539.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the performance of musical works generated by neural networks relies heavily on the quality and diversity of training data, resulting in the generated music that may lack innovation, single style or deviate from expected music genres.

Method used

Using the music generation method based on the ViT-LSTM model, the ViT-LSTM model is constructed by converting the music data in MIDI format into one-hot encoding, and using the model to generate music sequences, and finally converting the music sequences into MIDI files.

Benefits of technology

This method can capture the long-term dependence of time series data and focus on key features in music, improve the creativity and expressiveness of music generation, enhance the generalization ability and applicability of the model, and generate music sequences that meet different styles and emotional needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943009A_ABST
    Figure CN119943009A_ABST
Patent Text Reader

Abstract

The invention provides a music generation method based on a ViT-LSTM model, and solves the problem that the performance of generated music works greatly depends on the quality and diversity of training data. Comprising the following steps: step 1, converting music data in an MIDI format into one-hot codes to form an independent attribute sequence, and dividing the coded sequence data into training samples and labels for predicting future notes; 2, constructing a ViT-LSTM model, inputting the sequence data generated in the step 1 into the model for training, and finally obtaining a deep learning model capable of generating a music sequence; and step 3, generating a music sequence by using the ViT-LSTM model constructed in the step 2, and converting the music sequence into an MIDI file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and in particular relates to a music generation method based on a ViT-LSTM model. Background Art

[0002] The rapid development of artificial neural networks is gradually blurring the line between art and science. Many research results show how fields that were previously considered to be entirely human domains can be transformed into algorithmic methods. Music is one of these fields. With the advent of artificial neural networks in the late 1980s, new methods for music generation began to emerge. These early works laid the foundation for the deep learning-based music generation techniques we see today. Deep learning has gradually come into the spotlight as a method for music generation, surpassing traditional prediction and classification tasks.

[0003] Recurrent Neural Networks (RNNs) and Long Short-Term Memory Networks (LSTMs), especially LSTMs, have been proven to be effective tools for music generation. In music generation, LSTM networks can generate new music with a certain degree of creativity by learning the patterns and regularities of music sequence data. Its long-term dependency and memory capacity make it suitable for processing complex structures and musical development in music, and can generate more coherent and emotional music works. Therefore, LSTM has been widely used in the field of music generation and has become a core component of many current music generation models; the music generated by LSTM-based models lacks innovation and is extremely dependent on the quality and diversity of training data. By adopting the ViT-LSTM model, it can capture time series data, pay attention to the key features in music, and effectively process long-sequence music data. In music generation, the use of the MAESTRO dataset can further improve the training effect of the model. This synergistic effect makes the model more effective and accurate in generating complex and dynamically changing music works.

[0004] In the prior art, patent CN202410713493.7 discloses a symbolic music generation method, device, equipment and medium based on random forest. It is a method that uses a random forest algorithm to train multiple decision trees to process pitch, length, loudness and style consistency features through a large amount of learning of sample data, and can accurately capture the detailed features of music in different areas; patent CN202110931804.3 discloses a recursive jump connection deep learning music automatic generation method based on layer standardization, which consists of collecting musical instrument digital interface data, preprocessing the training set, building a music automatic generation network, training the music automatic generation network, and automatically generating music files; patent CN201811554945.2 discloses a music automatic generation method based on deep learning The music arrangement method loads a deep learning operating environment, creates a machine learning library, imports a local music database or a network music database to train an arrangement prediction model, identifies the required music style, and then extracts corresponding music elements from the arrangement database according to the identified music style, and synthesizes an audio file based on the arrangement prediction model; Patent CN202110210275.8 discloses a method for automatic generation and evaluation of music based on deep learning, which collects MIDI data and preprocesses the MIDI data, trains a GPT-2 model using the preprocessed MIDI data, and inputs several notes into the trained GPT-2 model to obtain a music melody, and completes the music evaluation by performing mathematical and statistical evaluation, music theory evaluation, and musical test evaluation on the music melody.

[0005] However, the music generated by such neural networks in the prior art usually depends on the training data, and the performance depends greatly on the quality and diversity of the training data. In this case, if the data set is biased or limited, the generated music may reflect these limitations, may lack real innovation or personality, and lead to a single style or deviate from the expected music type. Summary of the invention

[0006] Based on this, the present invention proposes a music generation method based on the ViT-LSTM model, which can not only capture the long-term dependency of time series data, but also focus on the key features in music, solving the problem that the performance of generated music works greatly depends on the quality and diversity of training data.

[0007] The present invention is achieved through the following technical solutions.

[0008] A music generation method based on the ViT-LSTM model includes the following steps:

[0009] Step 1: Convert the music data in MIDI format into one-hot encoding to form an independent attribute sequence, and divide the encoded sequence data into training samples and labels for predicting future notes;

[0010] Step 2: Build a ViT-LSTM model and input the sequence data generated in step 1 into the model for training, and finally obtain a deep learning model that can generate music sequences;

[0011] Step 3: Use the ViT-LSTM model built in step 2 to generate music sequences and convert the music sequences into MIDI files.

[0012] Beneficial effects of the present invention:

[0013] 1. This invention combines the long short-term memory network (LSTM) with the visual attention transformer (ViT) by building a deep learning model, and cooperates with the MAESTRO dataset to generate rhythmic music; by combining these two powerful neural network architectures, it can simultaneously capture the time series characteristics of music data and the relationship between data blocks, thereby achieving higher expressiveness and creativity when generating music;

[0014] 2. Since both the self-attention mechanism of ViT and the gating unit of LSTM can effectively process long sequence data, the present invention can better adapt to music sequences of different lengths, thereby improving the generalization ability and applicability of the model. At the same time, ViT can provide an overall grasp of different music elements, while LSTM can maintain the coherence of the music sequence, which is helpful to generate novel and expressive music works, thereby enriching the possibilities of music creation;

[0015] 3. The ViT-LSTM model shows strong adaptability when dealing with diverse music styles, and is able to generate music sequences that meet different styles and emotional needs. Its robustness enables the model to cope with complex music data, provide more stable generation results, and can be effectively adjusted and optimized in a changing creative environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flow chart of the music generation method based on the ViT-LSTM model of the present invention. DETAILED DESCRIPTION

[0017] The exemplary embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the accompanying drawings are only exemplary and are intended to illustrate the principles and spirit of the present invention, rather than to limit the scope of the present invention.

[0018] The implementation principle of the present invention is as follows: First, the data of the MIDI file is extracted into note, volume and time sequences, and these sequences are preprocessed by One-hot encoding. Next, the encoded data is divided into training samples xi and tags where n prev is the sequence length, which is used to predict future notes. Then, by combining the Vision Transformer (ViT) and the Long Short-Term Memory Network (LSTM), a ViT-LSTM network is constructed for training by integrating the advantages of the two models. ViT effectively captures global dependencies through its self-attention mechanism, while LSTM can model the timing information in music. The combination of the two can significantly improve the quality and diversity of generated music. Finally, the trained ViT-LSTM model is used to generate music sequences and convert them into MIDI files for further editing and playback.

[0019] like Figure 1 As shown, a music generation method based on the ViT-LSTM model of the present invention specifically includes the following steps:

[0020] Step 1: Convert the music data in MIDI format into one-hot encoding to form an independent attribute sequence, and divide the encoded sequence data into training samples x i and tags where n prev is the sequence length, used to predict future notes;

[0021] Specifically: the music data includes notes, velocity and time; the time information in the music data set is discretized, the time value is divided into multiple intervals, and the time is converted into discrete type variables according to the intervals; secondly, for the pitch, volume and discretized time value of each note, one-hot encoding is used to convert them into three one-hot vectors respectively, and these three vectors are then merged into one vector.

[0022] Through this step, the model can try to predict future sequences of musical features by observing previous sequences of notes, dynamics, and time, thereby generating music or performing other music-related tasks.

[0023] Step 2: Build a ViT-LSTM model and input the sequence data generated in step 1 into the model for training, and finally obtain a deep learning model that can generate music sequences; specifically:

[0024] In the prior art, the long short-term memory network (LSTM) is a variant of the recurrent neural network (RNN), which aims to solve the problem of gradient vanishing or gradient exploding of traditional RNN on long sequence data. LSTM effectively captures long-term dependencies through cell states and gating mechanisms (forget gate, input gate and output gate), and can retain and forget information for a long time. This structure enables LSTM to perform well in sequence data processing tasks, such as natural language processing, time series prediction and music generation. ViT is a transformer model based on the attention mechanism, which was originally used to process visual data. By applying it to music generation, the model is able to extract and focus on important notes, rhythms and melodic elements from music clips, and achieve a global perception of the music structure.

[0025] The ViT-LSTM model construction process is as follows: 30 groups of music data features are input, each group contains notes, note intensity (velocity) and time information (time), these features are arranged into a data sequence, forming a total of 30 such sequences.

[0026] Subsequently, these music data feature sequences are fed into the Long Short-Term Memory (LSTM) network structure, which processes the input data by learning the patterns and dependencies in the music sequence and generates the corresponding output results. After the data sequence is processed by the LSTM structure, the output result is reshaped through a fully connected layer to form a tensor with a shape of (128, 128, 3).

[0027] Next, this tensor of shape (128,128,3) is used as input and input into the ViT structure; set the ViT parameters: the image size is 128x128, the image is divided into small blocks of size 4x4, the output layer size num_classes is 128, the data dimension (dim) is 32, the number of layers (depth) is 6, the number of heads (heads) is 16, the multi-layer perceptron dimension (mlp_dim) is 64, the dropout rate (dropout) is 0.1, and the embedding layer dropout rate (emb_dropout) is 0.1. The ViT structure uses the self-attention mechanism to process the input tensor to capture the global and local information in it; finally, the ViT structure produces a processed output result, which may be applied to music generation or other related tasks.

[0028] This whole process describes that after the music data feature sequence is processed by the LSTM structure, it is reshaped into a tensor of shape (128, 128, 3) through the fully connected layer, and then the tensor is input into the ViT structure for further processing. This combination of LSTM and ViT helps the model better understand the patterns and relationships of music data and provide richer and more valuable information for music generation tasks.

[0029] In the specific implementation, the negative log-likelihood loss function (NLLLLoss) is used as the loss function for model training; this is because NLLLoss is suitable for multi-classification problems, especially when the output is a prediction of the probability of each category. The specific method is as follows:

[0030] The input features of the model are composed of three parts, namely (x1, x2, x3), corresponding to (note, velocity, time), and the label (y1, y2, y3) represents the (note, velocity, time) sequence data at the n+ith moment. The loss function is as follows:

[0031]

[0032] L=L n (x1,y1)+L n (x2,y2)+L n (x3,y3)

[0033] Among them, L n is the NLLLoss loss function, and OneHot is the data encoding method.

[0034] This approach enables the model to use a combination of notes, note intensity, and time information to predict the sequence data for the next moment, thereby more accurately training and optimizing music generation or sequence prediction tasks; by selecting a negative log-likelihood loss function and converting the data into a suitable input and label format, the patterns and features of music sequences can be learned more accurately, providing a more effective training and prediction basis for music generation tasks.

[0035] Step 3: Generate a music sequence using the ViT-LSTM model constructed in step 2, and convert the music sequence into a MIDI file;

[0036] The trained ViT-LSTM model is used to enable in-depth learning and effective generation of notes, volume, and time information in music sequences. This model combines the excellent ability of the visual transformer (ViT) in capturing global features with the advantages of the long short-term memory network (LSTM) in modeling time series data. It can not only extract complex music patterns, but also generate music sequences that are coherent and creative in style. The generated music sequences are further converted into standard MIDI files for easy storage, editing, and playback. This process realizes a closed loop from music data learning to actual generation, providing strong technical support for intelligent music creation, while expanding the application scenarios of artificial intelligence in the field of music.

[0037] Example 1: "Data Sequence" is taken as input, and the data sequence enters the "ViT-LSTM" processing module. From the ViT-LSTM processing module, the data is split into three different outputs: note, velocity, and time. These three outputs are then passed to two different modules, one is "MidiFileFunction" and the other is "MidiTrack Function". These modules process the input and ultimately generate "msg", indicating further data transformation. Finally, these processed messages are integrated into "MIDIFile", i.e., MIDI file.

[0038] In this example, the ViT-LSTM model generates a sequence with certain musical features by learning the music patterns and structures in the dataset. This generated sequence will be converted into a standard MIDI music file format by the MidiFile function and the MidiTrack function, which can be widely used in music creation and appreciation. The generated music sequence has specific musical features based on the model training.

[0039] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0040] It is obvious to those skilled in the art that the embodiments of the present invention are not limited to the details of the above exemplary embodiments, and that the embodiments of the present invention can be implemented in other specific forms without departing from the spirit or basic features of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the embodiments of the present invention is limited by the attached claims rather than the above description, so it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the embodiments of the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units, modules or devices stated in the system, device or terminal claims can also be implemented by the same unit, module or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

[0041] Finally, it should be noted that the above implementation modes are only used to illustrate the technical solutions of the embodiments of the present invention and are not intended to limit them. Although the embodiments of the present invention have been described in detail with reference to the above preferred implementation modes, those skilled in the art should understand that the technical solutions of the embodiments of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A music generation method based on the ViT-LSTM model, characterized in that: The following steps are involved: Step 1: Convert the music data in MIDI format into one-hot encoding to form an independent attribute sequence, and divide the encoded sequence data into training samples and labels for predicting future notes; Step 2: Build a ViT-LSTM model and input the sequence data generated in step 1 into the model for training, and finally obtain a deep learning model that can generate music sequences; Step 3: Use the ViT-LSTM model built in step 2 to generate music sequences and convert the music sequences into MIDI files.

2. A music generation method based on the ViT-LSTM model as claimed in claim 1, characterized in that: The music data includes note, velocity and time; the time information in the music data set is discretized, the time value is divided into multiple intervals, and the time is converted into a discrete type variable according to the interval; secondly, for each note's pitch, volume and discretized time value, they are converted into three one-hot vectors respectively using one-hot encoding, and the three vectors are then merged into one vector.

3. A music generation method based on the ViT-LSTM model as claimed in claim 2, characterized in that: The ViT-LSTM model construction process is as follows: 30 sets of music data features are input, each set contains a note, a note's velocity, and time information. These features are arranged into a data sequence, forming a total of 30 such sequences; Subsequently, these music data feature sequences are fed into the LSTM structure, which processes the input data by learning the patterns and dependencies in the music sequence and generates corresponding output results. After the data sequence is processed by the LSTM structure, the output result is reshaped through a fully connected layer to form a tensor with a shape of (128,128,3); Next, this tensor of shape (128,128,3) is used as input and input into the ViT structure; set the ViT parameters: the image size is 128x128, the image is divided into small blocks of size 4x4, the output layer size num_classes is 128, the data dimension (dim) is 32, the number of layers (depth) is 6, the number of heads (heads) is 16, the multi-layer perceptron dimension (mlp_dim) is 64, the dropout rate (dropout) is 0.1, and the embedding layer dropout rate (emb_dropout) is 0.

1. The ViT structure uses the self-attention mechanism to process the input tensor to capture the global and local information in it; finally, the ViT structure produces a processed output result, which may be applied to music generation or other related tasks.

4. A music generation method based on the ViT-LSTM model as claimed in claim 3, characterized in that: The negative log-likelihood loss function NLLLoss is used as the loss function for model training. The specific method is as follows: the input features of the model are composed of three parts, namely (x1, x2, x3), corresponding to (note, velocity, time), and the label (y1, y2, y3) represents the (note, velocity, time) sequence data at the n+ith moment. The loss function is as follows: <h2 style=";text-align:left;direction:ltr">L=L<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr"> (x1,y1)+L<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr"> (x2,y2)+L<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr"> (x3,y3) Among them, L n is the NLLLoss loss function, and OneHot is the data encoding method.

5. A music generation method based on the ViT-LSTM model as described in claim 3 or 4, characterized in that: "DataSequence" is taken as input, and the data sequence enters the "ViT-LSTM" processing module. From the ViT-LSTM processing module, the data is split into three different outputs: note, velocity, and time. These three outputs are then passed to two different modules, one is "MidiFile Function" and the other is "MidiTrack Function". These modules process the input and ultimately generate "msg", indicating further data transformation. Finally, these processed messages are integrated into "MIDIFile", i.e. MIDI files.

Citation Information

Patent Citations

  • Music composition method and system based on deep learning

    CN109785818A

  • A Deep Learning-Based Method for Automatic Music Generation and Evaluation

    CN112951183B

  • Recursive jump connection deep learning music automatic generation method based on layer standardization

    CN113707112A

  • Method, device, equipment and medium for generating symbolic music based on random forest

    CN118280325B