A method for generating symbolic music with emotional conditions

By combining diffusion models with stochastic control guidance, multidimensional emotion space, and non-differentiable rule functions, the problems of subjective bias and low classification accuracy of emotion labels in symbolic music generation in existing technologies are solved, achieving high-precision, diversified emotion music generation and objective evaluation.

CN119832882BActive Publication Date: 2025-11-25TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411901306.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-11-25
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing technologies for generating symbolic music suffer from problems such as subjective bias in emotional labeling, low accuracy in emotional classification, difficulty in learning harmonic variation characteristics, limitation of emotional modeling to four categories, and lack of objective evaluation methods.

Method used

By combining a diffusion model with stochastic control guidance (SCG) and multidimensional emotion space and non-differentiable rule functions, symbolic music that meets various emotion conditions is generated by training a VAE network and a DiT architecture, and the music emotion evaluation index MAV is introduced for objective evaluation.

Benefits of technology

It improves the accuracy and valence expression of music generation, supports modeling of more than four emotion categories, reduces subjective bias, provides flexibility and objective evaluation of emotion generation, and is suitable for multi-domain applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832882B_ABST
    Figure CN119832882B_ABST
Patent Text Reader

Abstract

The application discloses a symbolic music generation method with emotional conditions, comprising the following steps: S1, preparing a data set, the data set comprising a midi data set and a piano roll data set, the piano roll data set being obtained by preprocessing the midi data set, the midi data set and the piano roll data set both containing symbolic music data covering classical music and popular music; S2, training a VAE network for encoding the piano roll data into a hidden space; S3, pre-training a diffusion model; the diffusion model adopts a DiT architecture, and the diffusion model is trained on the piano roll data with a split time sequence length of 10.24s; S4, using a stochastic control guide SCG to perform forward evaluation on a non-differentiable rule function, so as to work with the diffusion model in a plug-and-play manner, and to realize non-differentiable rule guidance on the diffusion model; S5, setting a music emotion evaluation index MAV; and S6, setting a non-differentiable rule function with emotional conditions to guide the diffusion model to generate music segments meeting the emotional conditions in a reverse diffusion process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep neural networks, and in particular, to a method for generating symbolic music with emotional conditions based on deep neural networks. BACKGROUND

[0002] Music is not only an effective tool for regulating emotions and expressing feelings, but also plays an important role in promoting cognitive function and mental health. With the advancement of technology and the popularity of digital music, music has become extremely convenient to access, but at the same time, there has been a rapid increase in demand for music. This demand not only manifests in quantity, but also in the increasing demand for refinement and personalization of music. In this context, the research of music generation technology is particularly important. Emotional music generation not only helps people regulate their emotions in specific situations, but also provides new possibilities for creation. Through emotional music generation technology,

[0003] It is possible to more accurately regulate the emotional expression of music to meet individual emotional needs or specific situational requirements. The development of this technology not only enriches the means of music creation, but also plays an important role in the fields of psychological therapy, education, and entertainment. Therefore, the research and implementation of emotional music generation technology not only represent the frontier of music technology development, but also have a positive driving effect on various aspects of society.

[0004] There are two methods for generating symbolic music with emotional conditions based on deep learning: the first is to convert emotional labels into embeddings and use them as model inputs Madhok et al. [1] generate music emotions using LSTM and combine some emotions represented by facial expressions. Hung et al. [2] proposed the EOMPIA music emotion dataset, manually segmented music and labeled emotions using the Russell 4Q emotion model to ensure that each segment can express consistent emotions. Music was generated using a Transformer model and an LSTM model. Pangestu [3] et al. used the Transformer-XL model to generate music with negative, positive, and neutral emotions. Sulun et al. [4] transformer model using continuous arousal and valence labels, and used a continuous value concatenation model to combine emotional conditions with each music token during generation, ensuring that emotional conditions are evenly applied throughout the sequence. Grekow et al. [5] used a conditional autoencoder to train the model to generate music. Xu et al. [6]Extracting the polyphonic and tonal features of music as emotional features, decoupling the hidden distribution of the two features in the emotional space through a deep latent variable model, the second type is to train an emotional classifier and apply it to the model output to guide the decoding process or the latent space of the variational autoencoder and the generative adversarial network. Ferreira et al. [7] The hidden state of LSTM as a feature representation, training a logistic regression model for emotion classification. Bao and Sun [8] Using BERT to train an automatic emotion recognition model on the Edmonds Dance dataset, then training an end-to-end encoder-decoder model to generate music. Tan and Herremans [9] Modeling corresponding quantifiable low-level attributes to learn high-level feature representations, inferring high-level features from low-level representations through Gaussian Mixture Variational Autoencoder (GM-VAEs), opening up the possibility of learning abstract high-level music quality representations under data-scarce conditions. Tseng et al.

[10] By combining generative adversarial networks (GAN) and variational autoencoders (VAE) to expand music works, the model trains two classifiers (one for music emotion and one for music tonality) for GAN training. In addition to training the classifier to classify music emotions, Qin Yong et al.

[11] A multi-modal music emotion label generation method is proposed, which combines audio mel-spectrogram and EEG spectrogram to predict music emotion according to lyrics text data.

[0005] However, the existing methods still have some shortcomings. 1) The emotional labels given by data annotators may be influenced by objective and subjective factors, resulting in subjective bias in emotional labels. End-to-end methods are difficult to learn the relationship between emotions and music sequences. While the accuracy of the emotion classifier is low, other classification methods require additional labels such as lyrics, which are not suitable for symbolic music generation. 2) The speed and rhythm related to arousal in music and emotion are relatively easy for the model to learn, but the form and frequency of harmonic changes related to valence in emotion are almost not learned by the model, which leads to the generated music based on emotional conditions being very ambiguous in valence. 3) The modeling of these emotional conditions is still a difficult problem, and basically only 4 types of emotional condition generation can be performed. 4) The evaluation of emotion is relatively subjective and the accuracy of emotion classification task for symbolic music is relatively low, making it difficult to perform reliable objective evaluation.

[0006] References:

[0007] [1] Madhok R, Goel S, Garg S. Sentimozart: Music generation based on emotions [C]. In Proceedings of International Conference on Agents and Artificial Intelligence, 2018: 501-506.

[0008] [2] Hung H, Ching J, Doh S, Kim N, Nam J, and Yang Y. EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation [C]. In Proceedings of International Society for Music Information Retrieval Conference, 2021: 318-325.

[0009] [3] Pangestu M, Suyanto S. Generating music with emotion using transformer [C]. In Proceedings of International Conference on Computer Science and Engineering, 2021: 1-6.

[0010] [4] Sulun S, Davies M E P, and Viana P. Symbolic music generation conditioned on continuous-valued emotions [J]. IEEE Access, 2022, 10: 44617-44626.

[0011] [5] Grekow J, Dimitrova-Grekow T. Monophonic music generation with a given emotion using conditional variational autoencoder [J]. IEEE Access, 2021, 9: 129088-129101.

[0012] [6] Xu, Liu, Fu, et al. A specific emotional music generation method and system[P]. Jiangsu: CN202310234247.9, 2023-08-18.

[0013] [7] Ferreira L, Whitehead J. Learning to generate music with sentiment[C]. In Proceedings of International Society for Music Information Retrieval Conference, 2019: 384-390.

[0014] [8] Bao C, Sun Q. Generating music with emotions[J]. IEEE Transactions on Multimedia, 2022, 24(8): 3602-3614

[0015] [9] Tan H H, Herremans D. Music fadernets: Controllable music generation based on high-level features via lowlevel feature modelling[C]. In Proceedings of International Society for Music Information Retrieval Conference, 2020: 109-116.

[0016]

[10] Tseng B W, Shen Y L, Chi T S. Extending Music Based On Emotion And Tonality Via Generative Adversarial Network. International Conference on Acoustics, Speech and Signal Processing, 2021.

[0017]

[11] Qin Y, Wang X C, Li C Q, et al. A music emotion label generation method based on multi-modal fusion[P]. Tianjin: CN202211480294.3, 2023-03-07.

[0018]

[12] Kang C, Lu P, Yu B, Tan X, Ye W, Zhang S, Bian J. EmoGen: Eliminating Subjective Bias in Emotional Music Generation. arXiv preprint arXiv:2307.01229, 2023. SUMMARY

[0019] The purpose of the present application is to overcome the deficiencies in the prior art, provide a symbolic music generation method with emotional conditions, which generates music segments with higher accuracy according to multiple music emotional conditions (more than four categories), reduces subjective bias, and improves the accuracy of valence. And the method contains a new music emotion objective evaluation method (index).

[0020] The purpose of the present application is realized by the following technical solutions:

[0021] A symbolic music generation method with emotional conditions, comprising:

[0022] S1. Prepare a data set, which includes a midi data set and a piano roll data set, the piano roll data set being obtained by preprocessing the midi data set, and the midi data set and the piano roll data set both containing symbolic music data covering classical music and popular music; the midi data set is preprocessed to obtain the piano roll data set, specifically, the midi data is first cut into 10.24s segments, and then the pitch, note start and pedal in the midi data are spliced into the piano roll data;

[0023] S2. Train a VAE network for encoding piano roll data (3x128x1024) to latent space (4x16x128); 3 in the piano roll data (3x128x1024) refers to pitch roll, note start roll and pedal roll; 128 refers to the pitch range; 1024 is the time series length, i.e. 10.24s;

[0024] S3. Pretrain the diffusion model; the diffusion model adopts DiT architecture, and the diffusion model is trained on the piano roll data with a cut time series length of 1024 (10.24s);

[0025] S4. Use random control to guide SCG to perform forward evaluation on non-differentiable rule functions, and work with the diffusion model in a plug-and-play manner to realize non-differentiable rule guidance for the diffusion model;

[0026] S5. Set up the music emotion evaluation index MAV, construct a multi-dimensional emotion space with several music attributes that can affect music emotion, each multi-dimensional emotion space contains several multi-dimensional music emotion spaces, each music emotion is regarded as a music emotion space, and the emotion score of the music segment is obtained by calculating which music emotion space the music segment is in; and map the emotion score to the two-dimensional VA model to obtain the objective evaluation index;

[0027] S6. Set up a non-differentiable rule function with emotional conditions to guide the diffusion model to generate music segments that meet the emotional conditions in the reverse diffusion process.

[0028] Further, the data set includes the MAESTRO data set, the Pop1k7 data set, and the Pop909 data set; wherein the midi data set is preprocessed to obtain a piano roll data set, specifically, the midi data is first cut into 10.24 second segments, and then the pitch, note start, and pedal in the midi data are spliced into a piano roll data.

[0029] Further, in step S3, to generate music segments of arbitrary length, the score function aggregation mechanism of DiffCollage is used to aggregate the score function of each piano roll data.

[0030] Further, in step S4, a random control guide SCG is introduced in the t step of reverse diffusion, and X t-1 is calculated, and then n samples t-1 are randomly sampled from the posterior mean The denoising result of each noisy sample is estimated The sample that minimizes the loss of X0 is calculated, and the corresponding sample is taken as the result of denoising in this step X t-1 .

[0031] Further, in step S5, a point representing general emotional attributes is determined for each multi-dimensional plane through a clustering algorithm, and the similarity between the attributes of the music segment and these general emotional attributes is calculated to obtain the music emotion evaluation index MAV. The calculation formula of the music emotion evaluation index MAV is as follows:

[0032]

[0033] Where x j is the music attribute vector of the music segment with a certain type of emotional label, S i is the clustering of music emotion, C i is the vector of music emotion space i, n i is the number of music segments, and c ik ​​is the kth element of the vector of musical affective space i, m is the dimension of the vector of musical affective space, y k is the kth element of the vector of musical segment attribute, U i is the score on different musical affective space, L is the number of musical affective space, μ i is the mean, σ i is the standard deviation, X i is the feature vector of musical affective space, β i is the weight of musical affective space.

[0034] Further, in step S5, first use jSymbolic to select the music attribute highly related to musical affect (Kang C, Lu P, Yu B, Tan X, Ye W, Zhang S, Bian J. EmoGen: Eliminating Subjective Bias in Emotional Music Generation. arXiv preprint arXiv: 2307.01229, 2023.) to form a multi-dimensional affective space; calculate the music attribute vector of the music segment in the midi data set, then cluster to get different musical affective spaces; get the center value of the musical affective space by calculating the value of the cluster center and represent the general emotion, the generation process of symbolic music is separated from the emotional label to avoid the subjective bias of the emotional label, better control the generation of emotional music, the generation process of symbolic music is also the process of reverse diffusion; to calculate the music affective evaluation index MAV of a music segment, first calculate its music attribute vector X, then the score of the music attribute vector X in different musical affective spaces, and the score is mapped into the two-dimensional VA model to get the objective evaluation index.

[0035] Further, in step S6, the non-differentiable rule function with emotional condition includes three parts:

[0036] First, after the reverse diffusion reaches the step number T, add emotional guidance to the reverse diffusion process of the diffusion model, the prediction result generated by each step of the diffusion model is decoded into midi data through the trained VAE network, where the prediction result generated by each step of the diffusion model is piano roll data; the step number T ranges from 100 to 150 to balance the training time and emotional guidance;

[0037] Then use jSymbolic to obtain each music attribute of the midi data, and calculate the music attribute vector X;

[0038] Finally, calculate the emotional score loss I of the music attribute vector X, which is defined as

[0039]

[0040] c ik is the kth element of the vector of music emotion space i, x k is the kth element of the vector of music attribute X, L is the number of music emotion spaces, m is the dimension of the vector of music emotion space, and γ i is the guide weight of music emotion space i.

[0041] The present application also provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method for generating symbolic music with emotional conditions when executing the program.

[0042] The present application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the method for generating symbolic music with emotional conditions.

[0043] Compared with the prior art, the technical scheme of the present application has the beneficial effects that:

[0044] 1. Improved generation accuracy: the generation method based on the diffusion model can better capture the detailed characteristics of music under emotional conditions, and the generated music is more accurate in expressing the valence and arousal dimensions, improving the learning ability of the diffusion model for valence characteristics and enabling the generation of symbolic music with high quality and authenticity, significantly reducing emotional ambiguity.

[0045] 2. Diversification of emotional modeling: supports modeling of more than four emotional categories, and can flexibly generate conforming symbolic music according to multiple music emotional conditions and other music conditions (without the need to train the network model again); greatly improves the flexibility and adaptability of emotional generation, breaking through the limitations of traditional models on emotional classification.

[0046] 3. Reduction of subjective bias: by introducing the emotional black box rule function and the objective evaluation index MAV based on the multi-dimensional emotion space, the blank of lacking objective analysis of music emotion is filled, providing convenience for music emotion related tasks; the index is more consistent with human emotional cognition, and eliminates the subjective bias of music emotion labels, improving the emotional consistency and credibility of the generated music.

[0047] 4. High efficiency of generation control: using the random control guidance (SCG) technology and the plug-and-play method, the diffusion model can only guide the emotional conditions in the backward diffusion process (quickly), greatly reducing the time and computer resources consumed by retraining, and improving the generation efficiency and effect.

[0048] 5. The scientific nature of emotional evaluation: The invention introduces the objective evaluation index MAV based on the multi-dimensional emotional space, which evaluates the generated music in terms of emotion, ensuring the objectivity and quantification of the evaluation results, and providing a reliable basis for subsequent optimization.

[0049] 6. The universality of the scope of application: The data set covers classical music and popular music, and can generate music segments of various styles, providing solid support for application needs in different fields (such as psychological treatment, education, entertainment, etc.).

[0050] 7. The flexibility of the generated length: The arbitrary length generation of music segments is realized by using the DiffCollage score function aggregation mechanism technology, further meeting the personalized and complex music needs.

[0051] In summary, the invention is superior to the prior art in terms of accuracy, diversity and objective evaluation of emotional expression, providing a new way of thinking and practical approach for the field of emotional music generation.

[0052] An objective evaluation index for deep neural network music emotion is proposed, which also eliminates subjective bias BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 is the flowchart of the symbolic music generation method with emotional conditions of the invention.

[0054] Figure 2 is the symbolic music generation process based on the diffusion model; the diffusion model includes a forward diffusion process and a reverse diffusion process, Figure 2 The forward diffusion is a left dashed arrow, and the reverse is a right solid arrow.

[0055] Figure 3 shows the specific reverse diffusion process using random control guidance and introducing a non-differentiable rule function with emotional conditions.

[0056] Figure 4 is the result graph generated by the method of the invention. DETAILED DESCRIPTION

[0057] The invention will be further described in detail below in conjunction with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and do not limit the invention.

[0058] The symbolic music generation method with emotional conditions provided in this embodiment is shown in Figure 1 , including:

[0059] S1. Prepare dataset: The dataset includes midi dataset and piano roll dataset, the piano roll dataset is obtained by preprocessing the midi dataset, and the midi dataset and the piano roll dataset both contain symbolic music data covering classical music and popular music to generate music of various styles. The MAESTRO dataset is the midi data recorded by the participants of the international piano competition, containing about 1200 classical music pieces, about 200 hours in total. The Pop1k7 dataset contains 1700 popular music midi versions, a total of 108 hours. The Pop909 dataset contains 909 midi of popular music, a total of 60 hours. The midi dataset is preprocessed to obtain the piano roll dataset, specifically, the midi data is cut into 10.24s segments, and then the pitch, note start and pedal in the midi data are spliced into the piano roll data;

[0060] S2. Train a VAE network to encode the piano roll data (3x128x1024) to the latent space (4x16x128) for better training of the diffusion model;

[0061] S3. Pretrain the diffusion model; the diffusion model adopts the DiT architecture. Compared with the standard U-Net, the transformer as the backbone network is better at processing hidden token sequences. In this embodiment, the diffusion model is trained on the piano roll data with a cut-off time sequence length of 1024 (10.24s). In order to generate music segments of arbitrary length, the score function of each piano roll data is aggregated using the aggregation mechanism of the score function of DiffCollage.

[0062] S4. Use the stochastic control guide SCG to forward evaluate the non-differentiable rule function to work with the diffusion model in a plug-and-play manner to realize the rule guidance of the non-differentiable diffusion model;

[0063] Since the music conditions of conditional music generation are usually expressed in the form of symbolic features of notes, such as note density or chord progression, many of which are non-differentiable, which poses a challenge when guiding diffusion using them. Especially the abstract emotional conditions make it very difficult to calculate the gradient in the reverse diffusion. This embodiment uses the stochastic control guide SCG, which is a guidance method without the need to calculate the gradient, only the forward evaluation of the rule function is needed, which can work with the diffusion model in a plug-and-play manner, so as to realize the rule guidance of the non-differentiable diffusion model. The rule function is the concept of SCG, and the rule function plays the same role as the loss function for loss guidance. Generally, the loss function must be derivable, and the rule function in this embodiment is generally not derivable, so it is called a non-differentiable rule function.

[0064] As Figure 2 ,Figure 3 The SCG random control guide is introduced in the t-step of reverse diffusion, and X is calculated t-1 The posterior mean value is obtained, and X is randomly sampled t-1 The posterior mean value is obtained, and n samples are obtained To The de-noising result of each noisy sample is estimated To The sample that minimizes the loss of X0 is calculated, and the corresponding sample is taken as the de-noising result of this step X t-1 .

[0065] S5. Set the music emotion evaluation index MAV:

[0066] For music, how to evaluate the emotion of music is an important problem. At present, the evaluation of music emotion is mostly based on subjective evaluation of people. The embodiment designs a music emotion evaluation index MAV (Music Attribute Value), constructs a multi-dimensional emotion space with several music attributes that can affect music emotion, each multi-dimensional emotion space contains several multi-dimensional music emotion spaces, and each music emotion is regarded as a music emotion space. The emotion score of the music segment is obtained by calculating which music emotion space the music segment is in; however, due to subjective bias of people, these emotion multi-dimensional planes will intersect, and at this time, the same music segment will feel different emotions in different people, different moods, etc. In order to eliminate the influence of this bias, a point representing general emotion attributes is determined for each multi-dimensional music emotion space through a clustering algorithm, and the similarity (loss) between the attributes of the music segment and the general emotion attributes is calculated to obtain the score of the music emotion evaluation index MAV. The MAV formula is as follows:

[0067]

[0068]

[0069] Where x j is the music attribute vector of the music segment of a certain type of emotion label, S i is the clustering of music emotion, C i is the vector of music emotion space i, n i is the number of music segments, c ik is the kth element of the vector of music emotion space i, m is the dimension of the vector of music emotion space, y k is the kth element of the music segment attribute vector, U i is the score on different music emotion spaces, L is the number of music emotion spaces, μ i is the mean, σ i is the standard deviation, X iIt is the feature vector of the musical emotional space, β i It is the weight of the emotional space in music.

[0070] Specifically: First, use jSymbolic to select elements highly correlated with the emotional content of the music.

[12] The musical attributes constitute a multidimensional emotional space. Musical attribute vectors of music segments in the MIDI dataset are calculated, and then clustered to obtain different emotional spaces. The central value of the emotional space is obtained by calculating the cluster center value, which can represent general emotions. The generation process of symbolic music is separated from the emotional label, which can avoid the subjective bias of emotional labels and thus better control the generation of emotional music. The generation process of symbolic music is also a process of backdiffusion. Calculating the MAV of a music segment requires first calculating its musical attribute vector, then calculating the score of the musical attribute vector in different emotional spaces, and mapping the score to a two-dimensional VA model to obtain an objective evaluation index.

[0071] The objective evaluation metric MAV can be combined with subjective evaluation in form, providing an intuitive understanding of the connection between music emotion scores and specific emotions, which is useful for work related to music emotion. It addresses the problems of the limited availability of music datasets with emotion labels, the difficulty in labeling emotion labels, and the presence of subjective bias.

[0072] S6. Set a non-differentiable rule function with emotional conditions to guide the diffusion model to generate music fragments that meet the emotional conditions during the back diffusion process. In this embodiment, the non-differentiable rule function with emotional conditions is a black box rule function, which is different from a general non-differentiable rule function. It is a rule function specifically for emotional conditions.

[0073] First, after a certain number of steps T in the backdiffusion process, emotional guidance is incorporated into the backdiffusion process of the diffusion model. The prediction results (piano slide data) generated by the diffusion model at each step are decoded into MIDI data by a pre-trained VAE network. If the number of steps T is too high, the music segments generated by the pre-trained diffusion model still contain a lot of noise, which cannot effectively guide emotions. If T is set too low, the emotional guidance effect is not obvious. After multiple experiments, this embodiment sets T to 100-150, which can balance training time and emotional guidance.

[0074] Then, jSymbolic is used to obtain the various music attributes of the MIDI data, and the music attribute vector X is calculated.

[0075] Finally, the emotional score loss I of the music attribute vector X is calculated, as defined in the formula below.

[0076]

[0077] c ik x is the k-th element of the vector in the musical emotion space i.k is the k-th element of the music attribute vector X, L is the number of music emotion spaces, m is the dimension of the music emotion space vector, γ i is the guide weight of the music emotion space i.

[0078] Specifically, the specific application of the above-mentioned symbolic music generation method based on deep neural network with emotional condition is as follows:

[0079] First, a VAE model is trained to encode the piano roll data into a latent space, and then a diffusion model with a DiT architecture is trained on the piano roll data. The diffusion model is trained using dataset-based conditions: MAESTRO and Pop (pop1k7 and pop909), following the classifier-free guidance with a dropout rate of 0.1. The diffusion model is trained for 1.2M steps, and 1000-step DDPM is used as the default sampling method unless otherwise specified. All experiments are performed on an NVIDIA V100 GPU.

[0080] Then, SCG random control guidance is used to solve the non-differentiable problem of emotional condition in the process of reverse diffusion. Starting from the 750th step of reverse diffusion, each step does not calculate the gradient but first performs 16 samplings, then predicts the denoising results of the samplings, calculates the regular loss of these results, selects the completely denoised result with the smallest loss, and generates the sampling of the result as the next denoising result. Starting from the 120th step of reverse diffusion, each step does not calculate the gradient but first performs 16 samplings, then predicts the denoising results of the samplings, calculates other regular losses of these results, and then decodes using the pre-trained VAE network, calculates the regular loss of the non-differentiable function with emotional condition, and selects the denoised result with the smallest total loss, and generates the sampling of the result as the next denoising result.

[0081] Finally, the reverse diffusion obtains symbolic music that meets the emotional condition and other conditions, as shown in Figure 4 .

[0082] The embodiment has been verified by a large number of experiments, and the effectiveness of the network model and the rule guidance of the emotional black box is verified. The symbolic music generated by the method achieves good results in traditional evaluation methods and achieves good results in the music emotion evaluation index MAV. It is also verified that MAV is more in line with human emotional judgment and more convenient than traditional evaluation methods.

[0083] Specifically:

[0084] 1) The method of the present application can flexibly (add new conditions without training the network model) generate symbolic music that meets the conditions according to various music emotional conditions and other music conditions; refer to Table 6.

[0085] 2) The generated symbolic music has high quality and authenticity; refer to Table 1 and Table 2.

[0086] 3) The subjective bias of music emotion labels is eliminated; refer to Table 3, Table 4 and Table 5.

[0087] 4) The learning ability of the network model for valence features is improved; refer to Table 4 and Table 5.

[0088] 5) An objective evaluation index for music emotion of deep neural network is proposed, which fills the gap of music emotion objective analysis and provides convenience for music emotion related tasks. The index is more consistent with human emotion cognition and eliminates subjective bias; refer to Table 4, Table 5 and Table 6.

[0089] The method of the present application is compared with the methods of EMOGEN, EOMPIA and TRANSFORMER-GAN, and objective standards such as pitch range (distance between highest point and lowest point), number of pitch categories and polyphonic pitch and other indicators related to music quality are selected. For the symbolic music generated in each emotion category, 20*5, a total of 100 samples are generated, of which 20 are emotionless categories. The samples are evaluated according to the above characteristics, and then the average value of the results is obtained. The overlapping area between the attribute set and the data set is taken as AOA. The expected value is the AOA calculated by taking part of the data in the data set as a set and the other part of the data as a set. Other methods such as EMOPIA, emogen and trans-gan also calculate AOA. The results are shown in Table 1.

[0090] Table 1

[0091] EOM IA EMOGEN TRANSFORMER-GAN The method of the invention Expected value AOA 0.903 0.890 0.892 0.919 0.949

[0092] The results of the present application have higher AOA compared with other methods such as EMOPIA, emogen and trans-gan. This shows that the data generated by the method of the present application is more consistent with the original data.

[0093] There are many subjective evaluations of music quality. The present application is compared with the methods of EMOGEN, EOMPIA and TRANSFORMER-GAN, and the authenticity index is selected. The authenticity score is 1-5 points. For the symbolic music generated in each emotion category, 20*5, a total of 100 samples are generated, of which 20 are emotionless categories. Other methods have 20 music samples for each method. I also extracted 40 samples from the EMOPIALAKH POP909 POP1K7 and 4 other data sets. A total of 200 samples are extracted to score the authenticity by 50 people. The score results are shown in Table 2.

[0094] Table 2

[0095] EOM IA EMOGEN TRANSFORMER-GAN The method of the invention Authenticity 3.31 3.51 3.45 3.49 Dataset 4.26 4.26 4.12 4.03 Ratio 0.78 0.82 0.84 0.87

[0096] eompia and emogen are more inclined to popular music, and trans-gan and the method of the present application add a large amount of MIDI data of classical music competition. The subjective authenticity of different data sets is also investigated, and the score of the real data set cannot reach more than 4.5, and the score of popular music is generally higher than that of classical music. If the ratio of the authenticity score of the generated music sample and the data set is calculated, the method of the present application is higher than other methods, which indicates that the generated music more highly simulates the data set.

[0097] Table 3

[0098] EOM IA EMOGEN TRANSFORMER-GAN The method of the invention Accuracy rate (*) 0.418 — — — Accuracy rate (** 0.433 0.650 — — Accuracy rate 0.448 0.673 0.642 0.703

[0099] In the subjective evaluation of music emotion of the present application, the participants need to give specific valence and arousal scores, which are mapped to different emotion spaces, so as to reduce the influence of subjective bias. The scores of valence and arousal are both 1-5. The samples and participants are the same as above. The score results are shown in Table 3. The accuracy rate (*) is the highest emotion accuracy rate of the emotion-conditioned music generation result obtained by the subjective evaluation survey of the emopia author, the accuracy rate (**) is the highest emotion accuracy rate of the emotion-conditioned music generation result obtained by the subjective evaluation survey of the emogen author, and the accuracy rate is the subjective evaluation survey result of the present application. Due to subjective bias and other reasons, the results of different survey groups will have some differences, and the accuracy rate (*) and the accuracy rate (**) are for reference of the results of the present application. The emotion accuracy rate of the emotion-conditioned music generation result of the present application is slightly higher than the results of emogen and trans-gan. The accuracy rate of the survey group for emotion judgment is also slightly higher than the results of other methods, which is within a reasonable range.

[0100] The objective evaluation of music emotion is usually using a music emotion classifier. The emopia trained a linear transformer emotion classifier, which achieved good results as the objective evaluation method of most methods afterwards. The present application proposes a music emotion objective evaluation index MAV, and the beneficial effects of MAV are shown in the table.

[0101] Table 4

[0102] EOM IA MAV 4Q 0.684 0.731 A 0.822 0.900 V 0.833 0.870

[0103] The objective evaluation method of the present application is more objective than the method of emopia, and the accuracy of valence is obviously improved, which to some extent reduces the ambiguity of valence.

[0104] In Table 5, the results of the present application are higher than those of other methods EMOPIA, emogen, trans-gan, L-trans and MAV. Among them, the accuracy on MAV is higher because the black box loss guidance is used in the reverse diffusion process, which is similar to MAV calculation. It is worth noting that other eompia and emogen methods achieve better results on MAV because MAV is more suitable for the characteristics of the data itself than the L-trans classifier. The validation experiments of the two methods on the data set prove this phenomenon. MAV is more accurate for the emotional classification of the data set, and it reduces the influence of subjective bias and is more accurate for classic music emotion clips. Moreover, the MAV method gives a new emotional space classification for the ambiguous part in the 4Q or VA model, such as data that is both Q1 and Q2.

[0105] Table 5

[0106] EOM IA EMOGEN TRANSFORMER-GAN The method of the invention 4Q (L-trans) 0.418 0.715 0.703 0.740 A (L-trans) 0.690 0.892 0.880 0.885 V (L-trans) 0.583 0.761 0.779 0.800 4Q (MAV) 0.548 0.723 0.691 0.793 A (MAV) 0.701 0.871 0.859 0.932 V (MAV) 0.623 0.770 0.781 0.891

[0107] Previous work can only generate 4Q emotional conditions. The method of the present application can generate any emotional condition. Taking 8Q as an example, the ambiguous emotions between Q1 and Q2, Q2 and Q3, Q3 and Q4, and Q4 and Q1 are divided into Q12, Q23, Q34 and Q41. The accuracy of the generated results is calculated using the MAV index, as shown in Tables 3-6, 8Q (strict) indicates that the generated result strictly belongs to a certain emotional category, and 8Q (lenient) indicates that the generated result falls within adjacent emotional space. The experimental results show that the accuracy of 8Q (lenient) is higher, and the corresponding music can be effectively generated in subtle emotions, but the accuracy of 8Q (strict) is lower, and the accuracy of the subtle emotions of the music still needs to be improved.

[0108] Table 6

[0109] MAV 8Q (strict) 0.447 8Q (lenient) 0.859

[0110] Preferably, the embodiments of the present application also provide a specific implementation of an electronic device capable of implementing all the steps of the above-mentioned embodiment of the method for generating symbolic music with emotional conditions. The electronic device specifically includes the following contents:

[0111] a processor, a memory, a communications interface, and a bus;

[0112] The processor, the memory, and the communications interface communicate with each other through the bus; the communications interface is used to realize the information transmission between the server-side device, the metering device, and the user-side device, and other related devices.

[0113] The processor is configured to invoke a computer program in the memory, and the processor implements all steps in the method for generating symbolic music with emotional conditions in the above embodiments when executing the computer program.

[0114] Embodiments of the present application also provide a computer readable storage medium capable of implementing all steps in the method for generating symbolic music with emotional conditions in the above embodiments, and the computer readable storage medium stores a computer program, and the computer program implements all steps in the method for generating symbolic music with emotional conditions in the above embodiments when executed by a processor.

[0115] Each of the above-described embodiments in the present specification is described in a progressive manner, and the same or similar parts among the embodiments can be mutually referred to. Each of the embodiments focuses on the difference from other embodiments. In particular, for the hardware + program type embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0116] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.

[0117] Although the present application provides method operation steps as embodiments or flowcharts, more or less operation steps can be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is only one of the many execution orders of the steps, and does not represent the only execution order. When the device or client product in practice is executed, the method order shown in the embodiments or the drawings can be executed in sequence or in parallel (for example, in the environment of parallel processor or multi-thread processing).

[0118] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can be embodied in the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk memory, CD-ROM, optical memory, etc.) having computer usable program code contained therein.

[0119] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 The functions of a flow or multiple flows and / or a block or multiple blocks in conjunction with the disclosed aspects can be implemented on practitioners' computers in an interactive mode or in a batch mode. Figure 1

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow Figure 1 The functions of a flow or multiple flows and / or a block or multiple blocks in conjunction with the disclosed aspects can be implemented on practitioners' computers in an interactive mode or in a batch mode. Figure 1

[0121] The present application is not limited to the embodiments described above. The above description of specific embodiments is intended to describe and illustrate the technical solutions of the present application, and the specific embodiments described above are merely illustrative and are not restrictive. Those skilled in the art can make many forms of specific changes under the guidance of the present application without departing from the purpose of the present application and the scope of protection of the claims, and these all belong to the protection scope of the present application.​​

Claims

1. A method for generating symbolic music with emotional conditions, characterized in that, include: S1. Prepare the dataset, which includes a MIDI dataset and a piano roll dataset. The piano roll dataset is obtained by preprocessing the MIDI dataset. Both the MIDI dataset and the piano roll dataset contain symbolic music data covering classical and popular music. S2. Train a VAE network to encode the piano roll data into the latent space; S3. Pre-trained diffusion model; The diffusion model adopts the DiT architecture and is trained on piano roll data with a time series length of 10.24 s after segmentation. S4. Use stochastic control-guided SCG to perform forward evaluation of non-differentiable rule functions, and work with diffusion models in a plug-and-play manner to achieve non-differentiable rule guidance for diffusion models; S5. Set up the Music Emotion Evaluation Index (MAV). Construct a multi-dimensional emotion space using several music attributes that can influence music emotion. Each multi-dimensional emotion space contains several multi-dimensional music emotion spaces. Each type of music emotion is considered as a music emotion space. Calculate which music emotion space a music segment belongs to to obtain its emotion score. Map the emotion score to a two-dimensional VA model to obtain an objective evaluation index. S6. Set a non-differentiable rule function with emotional conditions to guide the diffusion model in generating music segments that meet the emotional conditions during the back diffusion process; the non-differentiable rule function with emotional conditions consists of three parts: First, after the backdiffusion reaches step T, sentiment guidance is added to the backdiffusion process of the diffusion model. The prediction results generated by the diffusion model at each step are decoded into midi data by the trained VAE network, where the prediction results generated by the diffusion model at each step are piano roll data; the number of steps T ranges from 100 to 150 to balance training time and sentiment guidance. Then, jSymbolic is used to obtain various music attributes of the MIDI data, and the music attribute vector is calculated. ; Finally, calculate the music attribute vector. Emotional score loss The formula is defined as follows: ; It is a space for musical emotions The vector of the first One element, It is a music attribute vector The One element, It is the quantity of emotional space in music. It is the dimension of the musical emotional space vector. It is a space for musical emotions The guiding weight.

2. The method for generating symbolic music with emotional conditions according to claim 1, characterized in that, The datasets include the MAESTRO dataset, the Pop1k7 dataset, and the Pop909 dataset. The MIDI dataset is preprocessed to obtain the piano roll dataset. Specifically, the MIDI data is first cut into 10.24-second segments, and then the pitch, note start, and pedal are extracted from the MIDI data and pieced together to form the piano roll data.

3. The method for generating symbolic music with emotional conditions according to claim 1, characterized in that, In step S3, to generate music segments of arbitrary length, the scoring function of each piano roll data is aggregated using the scoring function aggregation mechanism of DiffCollage.

4. The method for generating symbolic music with emotional conditions according to claim 1, characterized in that, In step S4, SCG stochastic control is introduced at the t-step of the back diffusion to calculate... Posterior mean, then random sampling The posterior mean is obtained from n samples arrive Estimate the denoising result for each noisy sample. arrive Calculate the sample that minimizes the X0 loss, and use the corresponding sample as the result of this denoising step. .

5. The method for generating symbolic music with emotional conditions according to claim 1, characterized in that, In step S5, a point representing a general emotional attribute is determined for each multidimensional plane using a clustering algorithm. The similarity between the attributes of the music segment and these general emotional attributes is calculated to obtain the Music Emotion Evaluation Index (MAV). The formula for calculating the Music Emotion Evaluation Index (MAV) is as follows: ; ; ; in It is a musical attribute vector of a musical fragment with a certain emotional tag. It is a clustering of musical emotions. It is a vector of the musical emotion space i. It refers to the number of musical fragments. It is a space for musical emotions The vector of the first One element, It is the dimension of the musical emotional space vector. It is the first of the music fragment attribute vectors One element, The scores are based on different musical emotional spaces. It is the quantity of emotional space in music. It is the mean. It is the standard deviation. It is a feature vector of the musical emotional space. It is the weight of the emotional space in music.

6. The method for generating symbolic music with emotional conditions according to claim 1 or 5, characterized in that, In step S5, jSymbolic is first used to select music attributes highly correlated with musical emotion to form a multidimensional emotion space; the music attribute vectors of music segments in the MIDI dataset are calculated and then clustered to obtain different music emotion spaces; the center value of the music emotion space is obtained by calculating the value of the cluster center and represents the general emotion. The generation process of symbolic music is separated from the emotion label to avoid the subjective bias of the emotion label and better control the generation of emotional music. The generation process of symbolic music is also the process of back diffusion; to calculate the music emotion evaluation index MAV of a music segment, its music attribute vector needs to be calculated first. Then the music attribute vector Scores in different musical emotional spaces are mapped onto a two-dimensional VA model to obtain objective evaluation indicators.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the symbolic music generation method with emotional conditions as described in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the symbolic music generation method with emotional conditions as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method, system, and medium for affective music recommendation and composition

    CA3169171A1

  • Speech synthesis model training method, electronic equipment and storage medium

    CN115762464A