Two-stage music automatic generation method and system based on multi-level music feature information
By using a two-stage generation network model and a bundle search algorithm for multi-level music feature information in the music automatic generation method, the problem of insufficient utilization of multi-level music feature information in the prior art is solved, and the structure and melody of the generated music are significantly improved.
Patent Information
- Application Number
- CN202510326026.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-27
AI Technical Summary
The existing automatic music generation method fails to make full use of the multi-level music feature information in the music dataset, resulting in the monotony of the generated music, poor melody, and insufficient model output optimization.
A two-stage music automatic generation method based on multi-level music feature information is adopted. By obtaining multi-level music feature information in the music dataset, a two-stage generation network model of the Transformer model is constructed, and a beam search algorithm based on multi-level music feature information is designed for output optimization.
The generated music has better structure, better melody and higher quality, solving the problems of monotony and poor melody, and improving the overall quality of music generation through optimized output.
Smart Images

Figure CN120220627A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a two-stage music automatic generation method and system based on multi-level music feature information, and belongs to the fields of deep learning, signal processing, and music automatic generation. Background Art
[0002] With the continuous improvement of living standards, people's spiritual and cultural needs have also been continuously growing, and music culture and art are a typical demand. All along, music culture and art with strong artistic sense, such as piano music, bamboo flute music, and erhu music, have been deeply loved by people, and excellent music art works often receive extensive attention from people. However, in the past, to obtain a work with a specific music art style, it usually required a well-trained professional composer to spend a lot of time and energy to create. In this regard, realizing the automatic generation of music is very important for the creation and dissemination of music works and for enriching the ways for people to obtain art.
[0003] In recent years, Huang et al. proposed research on using the Transformer model for automatic music generation. It uses the self-attention mechanism to model the distribution characteristics and internal correlations of music note sequences, and then reconstructs the sequence at the output end, constraining this reconstruction process with the maximum likelihood estimation loss. The generated music has higher quality and better coordination than previous methods. However, relying solely on a single model to generate a complete music sequence at once poses a greater learning difficulty for the model and cannot well model the features at multiple levels of music. In 2022, Yi Zou et al. proposed Melons (from the paper "MELONS: GENERATING MELODY WITH LONG-TERM STRUCTURE USING TRANSFORMERS AND STRUCTURE GRAPH"), and for the first time proposed the concept of two-stage generation. This method first extracts the changing structural relationships between musical phrases in the music sequence as the target for the first-stage generation, and then generates a complete music sequence based on the structural relationships. The generated music has a better structure and melody. In 2023, Kejun Zhang et al. proposed Wuyun (from the paper "WuYun: Exploring hierarchical skeleton-guided melody generation using knowledge-enhanced deep learning"), and noticed that some notes in music play an important guiding role in the development of the melody, and the remaining notes are organized based on them. The authors proposed a method to obtain the key notes in music and a two-stage generation method. First, generate the key note sequence, and then complete it based on the sequence generated in the previous stage. This method uses the key notes in music to guide the model training and generates music with higher quality and better melody.
[0004] However, existing automatic music generation methods often only utilize the music features at a certain level in the training music dataset, without fully exploring the multi-level music feature information in the music dataset, resulting in the generated music being monotonous and repetitive, with a chaotic structure, and even generating noise. At the same time, current automatic music generation methods do not make full use of music feature information and do not fully optimize the model generation results. Summary of the Invention
[0005] In view of this, the present invention provides a two-stage music automatic generation method, system, computer device and storage medium based on multi-level music feature information. By obtaining the multi-level music feature information in the music dataset, it can overcome the problems such as monotonous generated music and poor melody caused by insufficient utilization of multi-level music information in music automatic generation, and introduce a beam search algorithm based on multi-level music feature information to optimize the output of the network model, so as to generate music with better structure and coordination.
[0006] The first object of the present invention is to provide a two-stage music automatic generation method based on multi-level music feature information.
[0007] The second object of the present invention is to provide a two-stage music automatic generation system based on multi-level music feature information.
[0008] The third object of the present invention is to provide a computer device.
[0009] The fourth object of the present invention is to provide a computer-readable storage medium.
[0010] The first object of the present invention can be achieved by adopting the following technical solutions:
[0011] A two-stage music automatic generation method based on multi-level music feature information, the method includes:
[0012] Obtain a music dataset, the music dataset includes a Western piano music midi dataset and a Chinese-style music midi dataset;
[0013] Preprocess the music dataset, and obtain the multi-level music feature information in the preprocessed music dataset, the multi-level music feature information includes feature note combinations, note change features within a music measure, and note change relationship features between music measures;
[0014] Construct a two-stage music generation network model based on Transformer, the two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence;
[0015] Design a beam search algorithm based on multi-level music feature information, the beam search algorithm uses the multi-level music feature information to calculate the feature scores of candidate output sequences and determine the final output, the feature scores include output probability scores, feature note combination matching scores, note change feature matching scores within a music measure, and note change relationship feature matching scores between music measures;
[0016] Construct a training set using the music contour sequence and the complete music sequence that contain multi-level music feature information, train the two-stage music generation network model according to the training set, and use the trained two-stage music generation network model as the music automatic generation model;
[0017] Input the random noise sequence into the music automatic generation model and optimize it using the beam search algorithm based on multi-level music feature information to achieve music automatic generation.
[0018] Further, the characteristic note combinations are used to mine multiple combinations with the highest co-occurrence frequency and a length of three or four from the pitch sequences of the music dataset through the byte pair encoding algorithm, obtaining a characteristic pitch combination dictionary. According to the characteristic pitch combination dictionary, various repeated variations of the combinations are matched from the entire music dataset and marked.
[0019] Further, the note change characteristics within the music measure are used to match the characteristics of the pitch first rising and then falling, the pitch first falling and then rising, and three consecutive identical notes from the pitch sequences of the music dataset and mark them.
[0020] Further, the note change relationship characteristics between music measures are used to match and mine the relationship characteristics of repetition between music measures, sequence progression between measures, evolution between measures, and anaphora between measures from the music dataset and mark them.
[0021] Further, the optimization objective of the first-stage generation model is as follows:
[0022] Π t p(o t |O <t )
[0023] where o t represents the music symbol output at time t, and O represents the music contour sequence to be generated;
[0024] The optimization objective of the second-stage generation model is as follows:
[0025] Π t p(m t |M <t ;O <=t )
[0026] where m t represents the music symbol output at time t, and M and O represent the complete music sequence and the music contour sequence respectively;
[0027] The calculation of the characteristic score is as follows:
[0028]
[0029] where p(mt ) represents the probability that the two - stage music generation network model generates music symbols at time t. subword_num represents the number of characteristic note combinations within the search interval (i, j). match_inner and match_between represent the results of matching the characteristics within and between musical measures. Score subword represents the matching score of the characteristic note combination. Score BarInner represents the matching score of the note change characteristics within the musical measure. Score BarBetween represents the matching score of the note change relationship characteristics between musical measures. k1, k2, and k3 are hyperparameters for adjusting the algorithm.
[0030] Furthermore, training the two - stage music generation network model according to the training set and using the trained two - stage music generation network model as the music automatic generation model specifically includes:
[0031] Inputting the music contour sequence into the first - stage generation model for learning, enabling the first - stage generation model to predict the output of the music contour sequence at the next moment, calculating the loss based on the output music contour sequence, performing gradient backpropagation, and updating the parameters of the first - stage generation model;
[0032] Inputting the music contour sequence as context information into the second - stage generation model, and at the same time inputting the complete music sequence as the target into the second - stage generation model, enabling the second - stage generation model to predict the output of the complete music sequence at the next moment, calculating the loss based on the output complete music sequence, performing gradient backpropagation, and updating the parameters of the second - stage generation model;
[0033] Using the trained two - stage music generation network model as the music automatic generation model.
[0034] Furthermore, inputting the random noise sequence into the music automatic generation model and using the beam search algorithm based on multi - level music feature information for optimization to achieve music automatic generation, specifically including:
[0035] Inputting the random factor into the first - stage generation model to obtain the generated music contour sequence;
[0036] Inputting the generated music contour sequence into the second - stage generation model, and adopting the beam search algorithm based on multi - level music feature information in the output space of the second - stage model to obtain the optimized complete music sequence.
[0037] The second objective of the present invention can be achieved by adopting the following technical solutions:
[0038] A two - stage music automatic generation system based on multi - level music feature information, the system includes:
[0039] An acquisition unit for acquiring a music dataset, where the music dataset includes a Western piano music MIDI dataset and a Chinese-style music MIDI dataset;
[0040] A preprocessing unit for preprocessing the music dataset and obtaining multi-level music feature information in the preprocessed music dataset, where the multi-level music feature information includes feature note combinations, note change features within a music measure, and note change relationship features between music measures;
[0041] A construction unit for constructing a two-stage music generation network model based on Transformer, where the two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence;
[0042] A design unit for designing a beam search algorithm based on multi-level music feature information, where the beam search algorithm uses the multi-level music feature information to calculate the feature scores of candidate output sequences and determine the final output, and the feature scores include output probability scores, feature note combination matching scores, note change feature matching scores within a music measure, and note change relationship feature matching scores between music measures;
[0043] A training unit for constructing a training set using the music contour sequence and the complete music sequence containing multi-level music feature information, training the two-stage music generation network model according to the training set, and using the trained two-stage music generation network model as a music automatic generation model;
[0044] An automatic generation unit for inputting a random noise sequence into the music automatic generation model and optimizing it using the beam search algorithm based on multi-level music feature information to achieve music automatic generation.
[0045] The third object of the present invention can be achieved by adopting the following technical solutions:
[0046] A computer device includes a processor and a memory for storing processor-executable programs. When the processor executes the programs stored in the memory, the above two-stage music automatic generation method is implemented.
[0047] The fourth object of the present invention can be achieved by adopting the following technical solutions:
[0048] A computer-readable storage medium stores a program, and when the program is executed by a processor, the above two-stage music automatic generation method is implemented.
[0049] The present invention has the following beneficial effects compared with the prior art:
[0050] 1. The present invention makes full use of music feature information at multiple levels to assist the model in learning music features. It divides the music automatic generation task into two stages by using a two-stage generation method, reducing the learning complexity of each model. And it introduces a beam search algorithm based on multi-level music feature information to construct a search space in the original model output, optimize the output sequence according to the multi-level music feature information, and generate music with better structure, better melody and higher quality.
[0051] 2. Aiming at the problem that the existing methods do not make full use of music information at multiple levels, resulting in poor structure and low quality of the generated music, the present invention can improve the quality of the generated music by mining and utilizing multi-level music feature information.
[0052] 3. Aiming at the problem that the single-stage generation method cannot learn complete music feature information, resulting in low quality of the generated music, the present invention adopts a two-stage music automatic generation method, which can improve the model learning ability and the quality of the generated music.
[0053] 4. Aiming at the problem that the existing methods lack optimization in the model output stage, resulting in poor structure and lack of multi-level music features of the generated music, the present invention introduces a beam search algorithm based on multi-level music feature information, which can improve the quality of the generated music.
[0054] 5. The present invention can realize the automatic generation of high-quality music works, greatly simplify and innovate the way of obtaining and creating music, and is conducive to assisting music creation and promoting music dissemination. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the structures shown in these drawings.
[0056] Figure 1 It is a flowchart of the two-stage music automatic generation method based on multi-level music feature information in Embodiment 1 of the present invention.
[0057] Figure 2 It is an example diagram of obtaining multi-level music feature information in Embodiment 1 of the present invention.
[0058] Figure 3 It is a structural diagram of the two-stage music generation model based on multi-level music feature information in Embodiment 1 of the present invention.
[0059] Figure 4This is the structural block diagram of the two-stage music automatic generation system based on multi-level music feature information in Embodiment 2 of the present invention.
[0060] Figure 5 This is the structural block diagram of the computer device in Embodiment 3 of the present invention. Detailed implementation manners
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] Embodiment 1:
[0063] As Figure 1 shown, this embodiment provides a two-stage music automatic generation method based on multi-level music feature information, and the method includes the following steps:
[0064] S101. Obtain a music data set.
[0065] The music data set in this embodiment includes a Western piano music midi data set and a Chinese-style music midi data set, and the specific obtaining process is as follows:
[0066] Obtain the Western piano music midi data set: Obtain the Western piano music midi data set from existing music sharing websites, and filter out data with too short duration or duplicates, thereby obtaining the Western piano music midi data set.
[0067] Obtain the Chinese-style music midi data set: Collect music midi data with Chinese-style characteristics from channels such as the Chinese National Music Network and the Chinese Music Education Network, and manually enter the midi music data according to the numbered musical notation and the staff notation, thereby creating the Chinese-style music midi data set.
[0068] S101. Preprocess the music data set and obtain multi-level music feature information in the preprocessed music data set.
[0069] Preprocessing the music data set includes: In order to use the Western piano music midi data set and the Chinese-style music midi data set for model training, first convert them into a representation sequence in a unified format, and subsequent model training can directly read it.
[0070] Obtain multi-level music feature information in the preprocessed music dataset: The multi-level music feature information mined from the music dataset includes characteristic note combinations, note change features within a musical measure, and note change relationship features between musical measures, which are the three levels of multi-level music feature information, and the detailed information is shown in Table 1 below.
[0071] Table 1 Information at each level of multi-level music feature information
[0072]
[0073]
[0074] Among them, the characteristic note combinations are used to mine the most co-occurring combinations of length three or four from the pitch sequence of the music dataset through the Byte-Pair-Encoding (BPE) algorithm to obtain a characteristic pitch combination dictionary, and various repeated changes of the combinations are matched from the entire music dataset according to the characteristic pitch combination dictionary and marked; specifically, combinations in the dataset that are the same as the dictionary pitch combination, have the same pitch change, and have the same subsequence are matched.
[0075] Among them, the note change features within a musical measure are used to match and mark the features of pitches rising first and then falling, pitches falling first and then rising, and three consecutive identical notes from the pitch sequence of the music dataset.
[0076] Among them, the note change relationship features between musical measures are used to match and mine the relationship features of repetition, sequence progression, evolution, and anaphora between musical measures from the music dataset and marked; specifically, repetition means that two measures are exactly the same, modulation means that the note sequences of two measures are the same but the starting pitches are different, progression means that the similarity between two measures is greater than 50%, and anaphora means that the last note of the previous measure is the same as the first note of the next measure. It is set that each relationship is mined up to the previous 24 measures of each measure and marked in the sequence.
[0077] S103. Construct a two-stage music generation network model based on Transformer.
[0078] The two-stage music generation network model in this embodiment uses a Transformer decoder as the structure and is constrained by maximum likelihood estimation. It includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence.
[0079] The optimization objective of the first-stage generation model is as follows:
[0080] Π t p(ot |O <t )
[0081] Among them, o t is the musical symbol output at time t, and O represents the musical contour sequence to be generated. That is, the output at time t is predicted based on the information of the musical contour sequence O before time t.
[0082] The optimization objective of the second-stage generation model is as follows:
[0083] Π t p(m t |M <t ; O <=t )
[0084] Among them, m t is the musical symbol output at time t, and M and O represent the complete music sequence and the musical contour sequence respectively. That is, the output at time t is predicted based on the information of the musical contour sequence O at and before time t and the complete music sequence M before time t.
[0085] S104. Design a beam search algorithm based on multi-level music feature information.
[0086] The beam search algorithm of this embodiment uses multi-level music feature information to calculate the feature scores of candidate output sequences and determine the final output.
[0087] The probability distribution of the musical symbol generated by the two-stage music generation network model at time t is:
[0088]
[0089] Among them, eti represents the output value of the i-th musical symbol at time t, Vs represents the set of musical symbols, a certain number of output symbols are temporarily stored during the output process, and the candidate sequence with the highest feature score is selected from the temporarily stored search space as the output sequence after the subsequent output delimiter.
[0090] The feature scores include the output probability score, the feature note combination matching score, the note change feature matching score within the music measure, and the note change relationship feature matching score between music measures, as shown in the following formula:
[0091]
[0092] Among them, p(m t ) represents the probability of the two-stage music generation network model generating a musical symbol at time t, subword_num represents the number of feature note combinations in the search interval (i, j), match_inner and match_between represent the results of matching the features within and between measures, Scoresubword Represents the matching score of characteristic note combinations, Score BarInner Represents the matching score of the note change characteristics within a musical measure, Score BarBetween Represents the matching score of the note change relationship characteristics between musical measures, and k1, k2, and k3 are hyperparameters for adjusting the algorithm.
[0093] Match all consecutive notes in the search space according to the characteristic note combination dictionary. When a match is found, the corresponding candidate output sequence is incremented. If the head of the music contour sequence has in-measure and between-measure structure identifiers, match the candidate sequences in the search space, and compare the note and rhythm information with the already output sequence according to the relationship. Sequences that meet the characteristics are incremented; finally, select the sequence with the highest score from the search space as the current output, and then continue to construct the search space for the next segment until the end symbol is output.
[0094] S105. Use the music contour sequence and the complete music sequence containing multi-level music feature information to construct a training set, train the two-stage music generation network model according to the training set, and use the trained two-stage music generation network model as the music automatic generation model.
[0095] Furthermore, this step S105 is the model training stage, and the specific steps are as follows:
[0096] (1) Model initialization: All parameters to be trained are initialized with a Gaussian distribution with a mean of 0 and a variance of 0.01.
[0097] (2) Set model parameters: The learning factor is 0.0001 in the first 50 epochs, 0.0002 in the subsequent epochs, the training epoch is 800, and the mini-batch data is 2.
[0098] (3) Load the music contour sequence and the complete music sequence containing multi-level music feature information into two models (the first-stage generation model and the second-stage generation model) respectively for learning, with the maximum likelihood estimation as the optimization objective.
[0099] (4) Train the two-stage music generation network model: Input the music contour sequence into the first-stage generation model for learning, enabling the first-stage generation model to predict the output of the music contour sequence at the next moment. Calculate the loss based on the output music contour sequence, perform gradient backpropagation, and update the parameters of the first-stage generation model; Input the music contour sequence as context information into the second-stage generation model, and at the same time input the complete music sequence as the target into the second-stage generation model, enabling the second-stage generation model to predict the output of the complete music sequence at the next moment. Calculate the loss based on the output complete music sequence, perform gradient backpropagation, and update the parameters of the second-stage generation model; Use the trained two-stage music generation network model as the music automatic generation model.
[0100] S106. Input the random noise sequence into the music automatic generation model and optimize it using the beam search algorithm based on multi-level music feature information to achieve music automatic generation.
[0101] In this embodiment, input the random factor into the first-stage generation model to obtain the generated music contour sequence; Input the generated music contour sequence into the second-stage generation model, and adopt the beam search algorithm based on multi-level music feature information in the output space of the second-stage model to obtain the optimized complete music sequence.
[0102] This embodiment conducts experiments on both the Western piano music midi dataset and the Chinese-style music midi dataset, uses the trained network to generate music works, and evaluates the quality of the generated music through indicators such as PR, PCH, GS, ISR, APS, and IOI. PR calculates the size of the pitch distribution interval of the generated samples, which can measure the distance between the generated music and the training dataset in terms of pitch distribution range. PCH calculates the distribution distance between the generated data and the training dataset in terms of the number of notes in each scale, which can measure the similarity between the generated data and the training dataset in terms of scale distribution. GS calculates the difference in the average number of empty beats between the generated data and the training dataset, which can measure the similarity between the generated data and the training dataset in terms of rhythm. ISR calculates the proportion of C major notes among the non-empty notes of the generated music, which can measure the degree of fit between the generated music and the training dataset in terms of scale. APS is the average semitone interval between two consecutive notes, which can measure the melodic fluidity of the music. IOI calculates the time interval between two consecutive notes, which can measure the similarity between the generated music and the training dataset in terms of rhythm and tempo; When making comparisons, calculate the difference between the evaluation value of the generated music and the evaluation value of the training dataset, and the lower the obtained value, the better, indicating that the distribution of the generated data and the training dataset is closer; As can be seen from Tables 2 and 3, this application has better performance in comparison with existing advanced algorithms; It proves from the perspective of quantitatively evaluating the generated music that this application has better performance in the music automatic generation task.
[0103] Table 2 Comparison of Two Methods on the Western Piano Music MIDI Dataset
[0104] PR PCH GS ISR APS IOI Wuyun 1.23 0.29 0.0003 0.0284 0.2486 0.176 The method of this application 0.09 0.02 0.0002 0.0013 0.0109 0.064
[0105] Table 3 Comparison of Two Methods on the Chinese Style Music MIDI Dataset
[0106] PR PCH GS ISR APS IOI Wuyun 12.26 0.17 0.0037 0.0277 1.9195 0.069 The method of this application 7.69 0.02 0.0031 0.0116 1.4473 0.035
[0107] It should be noted that although the method operations of the above embodiments are described in a specific order, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the described steps can be changed in the order of execution. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0108] Embodiment 2:
[0109] As Figure 4 shown, this embodiment provides a two-stage music automatic generation system based on multi-level music feature information. The system includes an acquisition unit 401, a preprocessing unit 402, a construction unit 403, a design unit 404, a training unit 405, and an automatic generation unit 406. The specific descriptions of each unit are as follows:
[0110] The acquisition unit 401 is used to acquire a music dataset, and the music dataset includes a Western piano music MIDI dataset and a Chinese style music MIDI dataset;
[0111] The preprocessing unit 402 is used to preprocess the music dataset and obtain the multi-level music feature information in the preprocessed music dataset. The multi-level music feature information includes feature note combinations, note change features within a music measure, and note change relationship features between music measures;
[0112] The construction unit 403 is used to construct a two-stage music generation network model based on Transformer. The two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence;
[0113] The design unit 404 is used to design a beam search algorithm based on multi-level music feature information. The beam search algorithm uses the multi-level music feature information to calculate the feature scores of candidate output sequences and determine the final output. The feature scores include output probability scores, feature note combination matching scores, note change feature matching scores within a music measure, and note change relationship feature matching scores between music measures;
[0114] A training unit 405, configured to construct a training set by using a music contour sequence and a complete music sequence that contain multi-level music feature information, train a two-stage music generation network model according to the training set, and use the trained two-stage music generation network model as a music automatic generation model;
[0115] An automatic generation unit 406, configured to input a random noise sequence into the music automatic generation model and optimize it by using a beam search algorithm based on multi-level music feature information to achieve music automatic generation.
[0116] It should be noted that the system provided in this embodiment is only illustrated by the above division of each functional unit. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure is divided into different functional units to complete all or part of the functions described above.
[0117] Embodiment 3:
[0118] This embodiment provides a computer device, as Figure 5 shown, which includes a processor 502, a memory, an input device 503, a display device 504, and a network interface 505 connected through a device bus 501. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 506 and an internal memory 507. The non-volatile storage medium 506 stores an operating system, a computer program, and a database. The internal memory 507 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 502 executes the computer program stored in the memory, the two-stage music automatic generation method of the above Embodiment 1 is implemented as follows:
[0119] Obtain a music dataset, where the music dataset includes a Western piano music MIDI dataset and a Chinese-style music MIDI dataset; preprocess the music dataset and obtain multi-level music feature information in the preprocessed music dataset, where the multi-level music feature information includes feature note combinations, note change features within a music measure, and note change relationship features between music measures; construct a two-stage music generation network model based on Transformer, where the two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence; design a beam search algorithm based on multi-level music feature information, where the beam search algorithm uses the multi-level music feature information to calculate the feature scores of candidate output sequences and determine the final output, and the feature scores include output probability scores, feature note combination matching scores, note change feature matching scores within a music measure, and note change relationship feature matching scores between music measures; construct a training set using the music contour sequence and the complete music sequence containing multi-level music feature information, train the two-stage music generation network model according to the training set, and use the trained two-stage music generation network model as a music automatic generation model; input a random noise sequence into the music automatic generation model and optimize it using the beam search algorithm based on multi-level music feature information to achieve music automatic generation.
[0120] Embodiment 4:
[0121] This embodiment provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, it implements the two-stage music automatic generation method of the above Embodiment 1 as follows:
[0122] Obtain a music dataset, where the music dataset includes a Western piano music MIDI dataset and a Chinese-style music MIDI dataset; preprocess the music dataset, and obtain multi-level music feature information in the preprocessed music dataset, where the multi-level music feature information includes feature note combinations, note change features within a music measure, and note change relationship features between music measures; construct a two-stage music generation network model based on Transformer, where the two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence; design a beam search algorithm based on multi-level music feature information, where the beam search algorithm uses the multi-level music feature information to calculate the feature scores of candidate output sequences and determine the final output, and the feature scores include output probability scores, feature note combination matching scores, note change feature matching scores within a music measure, and note change relationship feature matching scores between music measures; construct a training set using the music contour sequence and the complete music sequence containing multi-level music feature information, train the two-stage music generation network model according to the training set, and use the trained two-stage music generation network model as a music automatic generation model; input a random noise sequence into the music automatic generation model and optimize it using the beam search algorithm based on multi-level music feature information to achieve automatic music generation.
[0123] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0124] In this embodiment, a computer-readable storage medium can be any tangible medium that includes or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this embodiment, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program included on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0125] The above computer-readable storage medium can be written in one or more programming languages or combinations thereof for executing the computer program of this embodiment. The above programming languages include object-oriented programming languages - such as Java, Python, C++, and also include conventional procedural programming languages - such as the C language or similar programming languages. The program can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0126] In summary, the present invention makes full use of music feature information at multiple levels to assist the model in learning music features, divides the music automatic generation task into two stages by using a two-stage generation method, reduces the learning complexity of each model, and introduces a beam search algorithm based on multi-level music feature information to construct a search space in the original model output, optimizes the output sequence according to multi-level music feature information, and generates music with better structure, better melody, and higher quality.
[0127] As described above, only the preferred embodiments of the present invention for patents are provided, but the protection scope of the present invention for patents is not limited thereto. Any person skilled in the art within the scope disclosed by the present invention for patents, according to the technical solution of the present invention for patents and its inventive concept, makes equivalent substitutions or changes, and all belong to the protection scope of the present invention for patents.
Claims
1. A two-stage music automatic generation method based on multi-level music feature information, characterized in that: The method comprises: Acquire a music data set, wherein the music data set includes a western piano music midi data set and a Chinese style music midi data set; Preprocessing the music data set and obtaining multi-level music feature information in the preprocessed music data set, wherein the multi-level music feature information includes characteristic note combinations, note change features within a music measure, and note change relationship features between music measures; Constructing a two-stage music generation network model based on Transformer, wherein the two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence; Design a beam search algorithm based on multi-level music feature information, the beam search algorithm uses the multi-level music feature information to calculate the feature scores of the candidate output sequence and determine the final output, the feature scores include output probability score, feature note combination matching score, note change feature matching score within a music measure, and note change relationship feature matching score between music measures; A training set is constructed using a music profile sequence containing multi-level music feature information and a complete music sequence, a two-stage music generation network model is trained based on the training set, and the trained two-stage music generation network model is used as a music automatic generation model; The random noise sequence is input into the music automatic generation model, and the beam search algorithm based on multi-level music feature information is used to optimize it to achieve automatic music generation.
2. The two-stage music automatic generation method according to claim 1, characterized in that: The characteristic note combination mines multiple combinations of length three or four with the highest co-occurrence frequency from the pitch sequence of the music data set through a byte pair encoding algorithm, obtains a characteristic pitch combination dictionary, and matches multiple repeated variations of the combination from the entire music data set based on the characteristic pitch combination dictionary and marks them.
3. The two-stage music automatic generation method according to claim 1 is characterized in that: The note change feature within the music measure matches the features of the pitch rising first and then falling, the pitch falling first and then rising, and three consecutive identical notes from the pitch sequence of the music data set, and marks them.
4. The two-stage music automatic generation method according to claim 1 is characterized in that: The feature of the relationship between the musical notes changes between the musical bars is mined by matching the feature of the relationship between the musical bars repetition, the bar progression, the bar evolution and the bar stitch from the music data set, and then marked.
5. The two-stage music automatic generation method according to claim 1 is characterized in that: The optimization objectives of the first stage generation model are as follows: ∏ t after t |O <t ) Among them, t represents the music symbol output at time t, and O represents the music contour sequence to be generated; The optimization objectives of the second stage generation model are as follows: ∏ t p(m t |M <t ;O <=t ) Among them, m t represents the music symbol output at time t, M and O represent the complete music sequence and music contour sequence respectively; The feature score is calculated as follows: Among them, p(m t ) represents the probability of the two-stage music generation network model generating music symbols at time t, subword_num represents the number of characteristic note combinations in the search interval (i, j), match_inner and match_between represent the results of matching the intra-bar and inter-bar features, Score subword Indicates the matching score of the feature note combination, Score BarInner Indicates the score of note change feature matching within a music measure, Score BarBetween It represents the feature matching score of the note change relationship between music bars, and k1, k2 and k3 are the hyper parameters of the adjustment algorithm.
6. The two-stage music automatic generation method according to any one of claims 1 to 5, characterized in that: The two-stage music generation network model is trained according to the training set, and the trained two-stage music generation network model is used as the automatic music generation model, which specifically includes: Input the music contour sequence into the first-stage generative model for learning, so that the first-stage generative model predicts the output of the music contour sequence at the next moment, calculates the loss according to the output music contour sequence, performs gradient back propagation, and updates the parameters of the first-stage generative model; The music contour sequence is input as context information into the second-stage generative model, and the complete music sequence is input as the target into the second-stage generative model, so that the second-stage generative model predicts the complete music sequence output at the next moment, calculates the loss based on the output complete music sequence, performs gradient backpropagation, and updates the parameters of the second-stage generative model; The trained two-stage music generation network model is used as the automatic music generation model.
7. The two-stage music automatic generation method according to any one of claims 1 to 5, characterized in that: The random noise sequence is input into the music automatic generation model, and the beam search algorithm based on multi-level music feature information is used for optimization to realize the music automatic generation, which specifically includes: The random factors are input into the first-stage generative model to obtain the generated music contour sequence; The generated music contour sequence is input into the second-stage generative model, and a beam search algorithm based on multi-level music feature information is used in the output space of the second-stage model to obtain the optimized complete music sequence.
8. A two-stage music automatic generation system based on multi-level music feature information, characterized in that: The system comprises: An acquisition unit, used for acquiring a music data set, wherein the music data set includes a western piano music MIDI data set and a Chinese style music MIDI data set; A preprocessing unit, used to preprocess the music data set and obtain multi-level music feature information in the preprocessed music data set, wherein the multi-level music feature information includes characteristic note combinations, note change features within a music measure, and note change relationship features between music measures; A construction unit, used to construct a two-stage music generation network model based on Transformer, wherein the two-stage music generation network model includes a first-stage generation model for generating a music contour sequence and a second-stage generation model for generating a complete music sequence from the music contour sequence; A design unit is used to design a beam search algorithm based on multi-level music feature information, wherein the beam search algorithm uses the multi-level music feature information to calculate feature scores of candidate output sequences and determine final outputs, wherein the feature scores include output probability scores, feature note combination matching scores, feature matching scores of note changes within a music measure, and feature matching scores of note change relationships between music measures; A training unit, for constructing a training set using a music profile sequence containing multi-level music feature information and a complete music sequence, training a two-stage music generation network model according to the training set, and using the trained two-stage music generation network model as a music automatic generation model; The automatic generation unit is used to input the random noise sequence into the music automatic generation model and optimize it using a beam search algorithm based on multi-level music feature information to achieve automatic music generation.
9. A computer device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, the two-stage automatic music generation method described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the two-stage automatic music generation method described in any one of claims 1-7 is implemented.