Music generation and diffusion method based on emotion guidance
By using technical means such as VAE-Diffusion framework and emotional feature encoder in Tibetan music generation, the problems of insufficient emotional expression, inefficiency and insufficient contextual consistency in music generation are solved, and a higher quality and consistent music generation effect is achieved.
Patent Information
- Application Number
- CN202510195088.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has problems such as insufficient emotional expression, low efficiency in processing high-dimensional feature, and insufficient consistency of music context in Tibetan music generation.
The music generation diffusion method based on the VAE-Diffusion framework is adopted to extract the potential characteristics of the sound source data through a variational autoencoder and model it during the diffusion process. Introduce emotional feature encoder, Token Drop strategy and Self-Conditioning mechanism to improve the quality and consistency of music generation.
It effectively improves the emotional expression ability, processing efficiency and context consistency in Tibetan music generation, and the generated music is more in line with specific emotional needs and has higher diversity and continuity.
Smart Images

Figure CN120032610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of music generation, and in particular to a music generation diffusion method based on emotion guidance. Background Art
[0002] In recent years, artificial intelligence technology has made significant progress in the field of music creation, but research on automatic generation of Tibetan music is still relatively scarce. Existing research faces three main challenges in Tibetan music generation: lack of expression ability of specific emotions, inefficient processing of high-dimensional features, and insufficient consistency of music context.
[0003] Deep learning has made significant progress in the field of music generation, greatly improving the efficiency and expressiveness of music creation. According to different generation methods, the current models in the field of music generation are mainly divided into two categories: Transformer-based autoregressive models and Diffusion-based diffusion models.
[0004] Autoregressive models usually capture the temporal dependencies of music by gradually predicting the next element in the sequence to generate continuous notes or audio signals. This type of model performs well in processing music sequence generation. Regressive models are represented by WaveNet, which can generate short music clips by modeling scalar quantized waveform samples. However, due to the autoregressive method of sample-by-sample generation, its sampling efficiency is low. To improve efficiency, researchers usually encode waveform samples into low-temporal resolution discrete potential representations (Tokens). These encoders (such as VQ-VAE and its variants) are trained by combining perceptual adversarial loss to support autoregressive Transformer modeling Tokens sequences, thereby significantly improving generation efficiency. Among such models, JukeBox is a representative framework that can generate music with specific emotions, styles, and instrumental characteristics from lyrics text, promoting breakthrough progress in music generation technology. In addition, MusicTransformer improves the ability to model long-term dependencies by introducing relative position encoding, and can generate music with complex chord structures. MuseNet is based on a multi-layer LSTM and Transformer architecture and supports multi-track music generation, covering a variety of styles from classical to popular. However, such models mainly focus on sound source separation, making it difficult to generate creative new music content, and may face problems of inefficiency and error accumulation when generating long sequences.
[0005] In contrast, diffusion models are particularly outstanding in music generation, especially in modeling complex data distributions and generating whole music clips. Diffusion models generate data by gradually adding noise and de-noising, and have strong distribution learning capabilities. For example, DiffSound uses the Mel VQ-VAE encoder to generate discrete intermediate representations, and models the Tokens sequence through a discrete diffusion model, which significantly improves the generation efficiency and detail expression. Some methods further use continuous latent variables in the spectral domain or waveform domain as intermediate representations of diffusion. StableAudio2 combines waveform domain VAE and uses diffusion to model its latent variables to generate the entire song. Although diffusion models can generate complete mixed music clips, most methods cannot effectively separate separate sound sources. The ideal music generation method should be able to generate and separate separate sound sources at the same time, such as controlling the volume ratio of piano and drums, making the generated music more interpretable and controllable. To this end, some studies have turned to multi-track modeling methods, such as directly generating music notes or MIDI representations and using a synthesizer to decode them into a single waveform; or by modeling multi-track music tracks, such as StemGen using a masked language model to model Encodec tokens to generate a single instrument sound source. There are also bass accompaniments based on mixed sound sources generated through latent diffusion models, and background accompaniments generated based on human voice sources. MSDM proposes to model four instrument sound sources simultaneously on the waveform domain diffusion model, and GMSDI is extended to text conditional generation on this basis to support a wider range of music data sets. In addition, there are also literatures that believe that music is composed of multiple closely related audio tracks, and propose a multi-source diffusion model MusicLDM, which can handle music generation and sound source separation tasks simultaneously under a unified framework, providing a new direction for achieving highly controllable and creative music generation. StableAudio2 not only achieves efficient generation of the entire song by diffusion modeling the latent variables of waveform domain VAE, but also provides a more complete solution for music generation technology, further expanding its application scenarios.
[0006] Although the above methods have made significant progress, there are still many limitations when directly used for training to generate Tibetan music, which are particularly evident in the following three aspects. First, the emotional expression of music is insufficient. For example, when generating a piece of Tibetan music, the existing model cannot accurately capture its unique emotional characteristics, resulting in the generated music emotions not matching the theme. Second, redundant features affect the generation efficiency. When generating long music clips, the model often needs to process high-dimensional and redundant feature data. Low-contribution or even irrelevant tokens not only increase the computational cost, but may also introduce noise, affecting the final generation quality. Finally, there is a lack of contextual consistency. When dealing with multi-instrument concertos, existing models often cannot effectively utilize previously generated tracks (such as piano or flute), resulting in the subsequent generated instruments (such as drums or bass) lacking coordination with the previous melody. Summary of the invention
[0007] In order to solve the problems existing in the prior art, the purpose of the present invention is to provide a music generation diffusion method based on emotion guidance. The present invention explores the tasks of simultaneously processing music generation and sound source separation, in order to achieve high-quality music generation with strong interpretability and controllability.
[0008] To achieve the above object, the technical solution adopted by the present invention is a music generation diffusion method based on emotion guidance, comprising the following steps:
[0009] Step 1: Compress music audio by training a variational autoencoder (VAE) to extract latent features and use a diffusion model to model latent variables.
[0010] Step 2: Complete the diffusion process of music based on emotion guidance: embed the emotion guidance model to generate music with specific emotions; randomly selected tokens are discarded to improve efficiency; the latent variables generated by the previous diffusion process are used as conditional input to enhance the consistency of the generated results.
[0011] As a further improvement of the present invention, the step 1 specifically comprises the following steps:
[0012] Step 1.1: The variational autoencoder VAE is used to represent the music source S∈R containing N samples in the waveform domain N Compressed into a compact and continuous latent space while ensuring that the reconstruction result is perceptually indistinguishable from the original sound source; given an input signal S, the encoder maps it to the posterior distribution: in, is the potential posterior mean, ∑ z (S) is the posterior covariance matrix, D is the time domain downsampling factor, and C is the latent space dimension;
[0013] Step 1.2: After encoding, sampling and input it into the decoder to reconstruct the signal S; using the posterior mean z s =μ z (S) as a potential representation;
[0014] Step 1.3: Based on the potential diffusion model, in the forward diffusion process, the original sample x 0 After T steps of adding noise step by step, a series of noisy samples x are generated. 1 ,x 2 ,...,x T ; At each time step t, sample x t The conditional probability distribution of is determined by the sample x at the previous moment t-1 Determine, its mathematical form is: Among them, β 1 ,…,β t ,…,β T is a predefined noise scheduling parameter;
[0015] Step 1.4: According to the properties of Gaussian distribution, we can deduce: in α t =1-β t ; By sampling And use the reparameterization technique to get the sample
[0016] Step 1.5: Reverse generation process from pure noise sample x T Start by gradually denoising and reconstructing x T-1 ,x T-2 ,…,x 0 , and finally obtain realistic samples; the reverse process is defined as the conditional probability distribution p θ (x t-1 |x t ), learned through a neural network, is used to approximate q(x t-1 |x t ,x 0 );
[0017] Step 1.6: To learn p θ (x t-1 |x t ), the training model output ∈ θ (x t ,t) to restore the generated x t The noise ∈ added when ; the loss function of the training diffusion model is:
[0018] Step 1.7: During inference, given x t and predicted noise, from p θ (x t-1 |x t )sampling: in
[0019] As a further improvement of the present invention, the step 2 specifically comprises the following steps:
[0020] Step 2.1, introduce the emotion feature encoder, embed the music emotion features into the diffusion model through the cross attention mechanism, and guide the diffusion model to generate music clips that meet specific emotions;
[0021] Step 2.2: Improve the Token Drop strategy so that it randomly drops some tokens during training.
[0022] Step 2.3: Propose a Self-Conditioning mechanism, which uses the previous generation results of the diffusion model as conditional input to provide contextual information for subsequent generation, thereby ensuring the consistency of music melody and emotion.
[0023] As a further improvement of the present invention, the step 2.1 is specifically as follows:
[0024] Assume that the potential representation of multiple input music sources is where z i is the potential representation of the i-th music clip, K is the number of clips; the music emotion information is generated by the emotion feature encoder, which maps the emotion description to the feature matrix E∈R in the latent space M×C , where M is the number of sentiment features; during the generation process, a cross-attention mechanism is used to combine sentiment information with the latent representation: where A∈R K×M is the attention weight, It is the potential representation after integrating emotional information.
[0025] As a further improvement of the present invention, the step 2.2 is specifically as follows:
[0026] The potential representation z t Divide into small blocks of size p×p, each small block is flattened into a vector to form a token, the total number of tokens is Subsequently, the tokens are reshaped into a matrix where d = c × p 2 Represents the dimension of a single token;
[0027] Based on the dynamic masking mechanism, the token matrix u is selectively masked: 1) the masking ratio ρ is defined to determine the number of tokens ρN that need to be masked; 2) the mask matrix By randomly or based on a specific weight, some tokens are selected for masking, M[i]=1 indicates masked, and M[i]=0 indicates unmasked; 3) The masked token is represented as: u′=M⊙u+(1-M)⊙0, where ⊙ represents an element-by-element product operation;
[0028] The encoder of the diffusion model focuses on the processing of the unmasked token u′ and generates the feature representation q; the side interpolator Int(·) is introduced to recover the masked token and fill the masked area by interpolation. The formula is: k = (1-M)·q+M·Int(q), where Int(q) estimates the masked token according to the encoder output q; the interpolated token k is embedded into the input decoder in combination with the position to recover the complete potential representation And restore high-resolution audio data through variational autoencoder VAE.
[0029] As a further improvement of the present invention, the step 2.3 is specifically as follows:
[0030] Using Self-Conditioning to introduce the estimate of the previous time step in the denoising network Enable the network to refer to historical information to improve the prediction of the current time step; in Self-Conditioning, modify the estimate of the denoising network to: in is the estimate of the previous time step t+1, and x is t and Splicing;
[0031] During the training phase, the Self-Conditioning input is set to zero, that is, Compute a preliminary estimate: in is to represent x based only on the current noise t and the estimated results at time step t;
[0032] When estimating with Self-Conditioning, a preliminary estimate is obtained in the first forward propagation After that, by stopping the gradient operation, Used as Self-Conditioning input for the second forward propagation: The denoising network is then optimized using the output of the two forward passes to accurately estimate x 0 ;
[0033] In the diffusion process, according to the time schedule σ(t) = t, the forward diffusion process is defined as: in
[0034] Sampling by solving the inverse ODE process The fraction By Neural Network approximation and trained with score matching loss.
[0035] As a further improvement of the present invention, the music is Tibetan music, which gradually solves the problems of lack of expression ability of specific emotions, low efficiency of high-dimensional feature processing, and insufficient consistency of music context in Tibetan music generation.
[0036] To solve the above problems, this algorithm proposes an emotion-guided diffusion method based on the VAE-Diffusion framework, which uses variational autoencoders to extract key potential features of the sound source data and model them in the diffusion process.
[0037] The beneficial effects of the present invention are:
[0038] Existing research faces three main challenges in Tibetan music generation: lack of expression ability of specific emotions, low efficiency of high-dimensional feature processing, and insufficient consistency of music context; to solve the above problems, the present invention proposes a diffusion model based on emotion guidance for Tibetan music generation method. The algorithm is based on the Latent Diffusion framework, which compresses music audio to extract latent features by training a shared variational autoencoder (VAE), and uses a diffusion model to model latent variables. In the entire diffusion process, the present invention combines the following three innovations to improve the quality and consistency of music generation. First, an emotion feature encoder is introduced to embed the music emotion features into the diffusion model through a cross-attention mechanism, guiding the diffusion model to generate music clips that meet specific emotions, thereby better meeting the emotional needs of Tibetan music. Second, the Token Drop strategy is improved to randomly discard some tokens during the training process, enhance the robustness of the model to missing information, improve the diversity and continuity of generated music, and effectively filter redundant information to reduce computational costs. Third, a Self-Conditioning mechanism is proposed, which uses the model's previous generation results as conditional input to provide contextual information for subsequent generation, thereby ensuring the consistency of musical melody and emotion, especially improving coordination in multi-instrument concerto. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flow chart of an embodiment of the present invention;
[0040] Figure 2 A schematic diagram showing characteristics of audio generated by different models in an embodiment of the present invention;
[0041] Figure 3 This is a schematic diagram of the Token Drop strategy in an embodiment of the present invention;
[0042] Figure 4 This is a schematic diagram of the effect of Token Drop on training efficiency in an embodiment of the present invention;
[0043] Figure 5 Schematic diagram of comparative analysis of frequency spectrum characteristics of music generated by different models and real music in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0045] Example
[0046] like Figure 1As shown, a music generation diffusion method based on emotion guidance includes:
[0047] Step 1: The whole method is mainly divided into two steps: VAE encodes and decodes the original music to extract latent features; and completes the diffusion process of music based on emotion guidance.
[0048] Step 2: The VAE module is used to represent the music source S∈R containing N samples in the waveform domain N Compressed into a compact and continuous latent space, while ensuring that the reconstruction result is perceptually indistinguishable from the original source. Given an input signal S, the encoder maps it to a posterior distribution: in, is the potential posterior mean, ∑ z (S) is the posterior covariance matrix, D = 320 is the time domain downsampling factor, and C = 80 is the latent space dimension.
[0049] Step 3: After encoding, sampling And input it into the decoder to reconstruct the signal S. To further simplify the extraction of potential features, this method directly uses the posterior mean z s =μ z (S) as a potential representation.
[0050] Step 4: Model the music generation process in the latent space. This method is based on the latent diffusion model. The diffusion model is a type of probabilistic generative model that learns the mapping relationship between noise and data by iteratively refining the noise samples. Specifically, in the forward diffusion process, the original sample x 0 After T steps of adding noise step by step, a series of noisy samples x are generated. 1 ,x 2 ,...,x T At each time step t, the sample x t The conditional probability distribution of is determined by the sample x at the previous moment t-1 Determine, its mathematical form is: Among them, β 1 ,…,β t ,…,β T are predefined noise scheduling parameters.
[0051] Step 5: According to the properties of Gaussian distribution, it can be deduced that: in α t =1-β t By sampling And using the reparameterization technique, we can get the sample
[0052] Step 6: The reverse generation process is from pure noise samples x T Start by gradually denoising and reconstructing x T-1 ,x T-2 ,…,x 0 , and finally obtain realistic samples. The reverse process is defined as the conditional probability distribution p θ (x t-1 |x t ), learned through a neural network, is used to approximate q(x t-1 |x t ,x 0 ).
[0053] Step 7: To learn p θ (x t-1 |x t ), we only need to train the model output ∈ θ (x t ,t) to restore the generated x t The noise ∈ added when . The loss function of the training diffusion model is:
[0054] Step 8: At inference time, given x t and predicted noise, can be obtained from p by the following formula θ (x t-1 |x t )sampling: in
[0055] In the diffusion process, three strategies are integrated: embedding music emotional conditions to guide the diffusion process, TokenDrop strategy to improve efficiency, and Self-Conditioning strategy to enhance the consistency of generated results.
[0056] Step 9: Further introduce music emotion information to guide the generation of latent variables. Suppose the latent representation of multiple input music sources is where z i is the potential representation of the ith music clip, and K is the number of clips. Music emotion information is generated through an emotion encoder, which maps the emotion description to a feature matrix E∈R in the latent space M×C , where M is the number of sentiment features. During the generation process, a cross-attention mechanism is used to combine sentiment information with the latent representation:
[0057] where A∈R K×M is the attention weight, It is the potential representation after integrating emotional information.
[0058] Step 10: For the TokenDrop strategy, Figure 3 As shown, the potential representation z t It is divided into small blocks of size p×p, each of which is flattened into a vector to form a token, and the total number of tokens is These tokens are then reshaped into the matrix where d = c × p 2 Represents the dimension of a single token.
[0059] Step 11: To reduce redundant calculations, this paper designs a dynamic masking mechanism to selectively mask the token matrix u. The specific implementation is as follows: 1) Define the masking ratio ρ to determine the number of tokens that need to be masked ρN. 2) Construct the mask matrix By randomly or based on a specific weight, some tokens are selected for masking, M[i] = 1 means masked, and M[i] = 0 means unmasked. 3) The masked token is represented as: u′ = M⊙u+(1-M)⊙0, where ⊙ represents the element-by-element product operation. At this point, only the unmasked tokens are retained to participate in the subsequent diffusion modeling, which significantly reduces the computational overhead.
[0060] Step 12: The encoder of the diffusion model focuses on the processing of the unmasked token u′ and generates the feature representation q. To recover the masked token, this paper introduces a side interpolator Int(·) to fill the masked area by interpolation. The formula is: k = (1-M)·q + M·Int(q), where Int(q) estimates the masked token based on the encoder output q. The interpolated token k is embedded into the decoder in combination with the position to recover the complete potential representation And restore high-resolution audio data through the VAE decoder.
[0061] Step 13: In standard diffusion sampling, at each time step t, the denoising network When using only x t As input, generate a pair x 0 Estimates The standard diffusion model generates a pair x independently at each time step t 0 Estimates Failure to use the estimate of the previous time step to gradually optimize the result. Self-Conditioning introduces the estimate of the previous time step into the denoising network. This enables the network to refer to historical information to improve the prediction of the current time step. In Self-Conditioning, the estimate of the denoising network is modified as follows: in is the estimated result of the previous time step t+1. In order to integrate Self-Conditioning in the network input, it is usually done by converting x t and Splicing.
[0062] Step 14: During the training phase, in order to approximate the inference behavior of Self-Conditioning while maintaining computational efficiency, a preliminary estimate (without Self-Conditioning) is made. Set the Self-Conditioning input to zero, i.e. Compute a preliminary estimate: in is to represent x based only on the current noise t and the estimated result at time step t.
[0063] Step 15: Estimation with Self-Conditioning. Get a preliminary estimate in the first forward propagation After that, by stopping the gradient (stop-gradient) operation, Used as Self-Conditioning input for the second forward propagation: The denoising network is then optimized using the output of the two forward passes to accurately estimate x 0 .
[0064] Step 16: During the diffusion process, according to the time schedule σ(t)=t, the forward diffusion process is defined as: in
[0065] Step 17: Sampling by solving the ODE in reverse The fraction By Neural Network approximation and trained with score matching loss.
[0066] The present embodiment is further described below by experiments:
[0067] 1. Objective indicators:
[0068] This embodiment evaluates the performance of multiple audio generation models, as shown in Table 1, including DAC, AudioGen, Encodec, MSDM, MSLDM, and the models EDMM (TD-50%) and EDMM (TD-0%) proposed in this embodiment, and quantitatively evaluates the generated audio samples through three key indicators: Fréchet Audio Distance (FAD), Jensen-Shannon Divergence (JSD), and Number of Distinct Bins (NDB). Among them, FAD is used to measure the distribution distance between the generated audio and the real audio, and the smaller the value, the higher the similarity; JSD is used to measure the similarity of the feature distribution of the generated audio and the real audio, and the value range is between [0,1]. The smaller the value, the more similar the two groups of distributions are; NDB is used to evaluate the diversity of the generated data, which is measured by calculating the number of different "boxes" distributed in the feature space of the generated audio. The larger the NDB value, the higher the diversity of the generated audio.
[0069] This embodiment adopts a two-stage training strategy to optimize the quality of Tibetan music generation. The first-stage model EDMM (TD-50%) is trained with a token drop rate of 50% and a self-conditioning rate of 80%; the second-stage model EDMM (TD-0%) maintains the same self-conditioning rate (80%), but completely cancels the token drop mechanism (i.e., the token drop rate is 0%), so that the model can be trained with complete token information.
[0070]
[0071]
[0072] Table 1
[0073] 2. Visual analysis:
[0074] Through Figure 2 The visualization of the music waveform shown demonstrates the characteristics of audio generated by different models. By comparing the real audio GT with the present embodiment, it can be clearly observed that the two have a high degree of similarity in waveform, which indicates that the generated audio of the present embodiment excels in retaining the characteristics of real music.
[0075] 3. Tibetan Music Generation Human Evaluation:
[0076] In order to evaluate the performance of different audio generation models in terms of auditory quality, the experiment selected 50 generated audio results, as shown in Table 2, and scored them from four dimensions through subjective evaluation by human listeners: Intelligence, Naturalness, Quality, and Synchronization. The dimensional evaluation results can reflect the listeners' overall feeling about the audio generation effect.
[0077]
[0078] Table 2
[0079] 4. Experimental data set: The Tibetan music data set used in this experiment contains samples of various music audio formats, which are processed by the Wiener denoising algorithm to improve the sound quality and effectively remove background noise. All audio samples are uniformly set to a sampling frequency of 24KHz to ensure the consistency of the input data. The data set is divided into a training set and a test set, where the training set accounts for 80% of the total data (about 4000 samples) and the test set accounts for 20% (about 1000 samples), ensuring that samples of different music styles and performers are evenly distributed.
[0080] 5. Experimental settings: In the process of emotional information extraction, the emotion recognition model is used to extract low-level features such as Mel-frequency cepstral coefficients (MFCC) and rhythm features from the audio. The audio samples are manually annotated and emotional labels including happiness, sadness, anger and calmness are constructed. The emotion recognition model is trained based on these annotated data to predict the emotional labels of unlabeled audio samples. In the experimental setting of the diffusion model, a score-based diffusion model is used in combination with SourceVAE to extract the potential representation of the audio signal. During the training process, the audio samples in the training set are used for model optimization, and the loss function includes Mel reconstruction loss, feature matching loss, adversarial loss and KL divergence loss. In terms of parameter settings, the learning rate is set to 0.001, the number of diffusion steps is 15000, and the noise standard deviation range is set to [0.01,3] to ensure the convergence and generation effect of the model. During the testing phase, the generated audio samples were quantitatively evaluated using the Fréchet Audio Distance (FAD), Jensen-Shannon Divergence (JSD) and Number of Distinct Bins (NDB) indicators. Among them, FAD is used to measure the distribution distance between the generated audio and the real audio. The smaller the value, the higher the similarity. JSD is used to measure the similarity of the feature distribution of the generated audio and the real audio. The value range is between [0,1]. The smaller the value, the more similar the two groups of distributions are. NDB is used to evaluate the diversity of the generated data. It is measured by calculating the number of different "boxes" in which the generated audio is distributed in the feature space. The larger the NDB value, the higher the diversity of the generated audio.
[0081] Experimental results:
[0082] (1) The comparative experiment is shown in Table 3:
[0083]
[0084] Table 3 (2) The impact of Token Drop and Self-Conditioning ratio is shown in Table 4:
[0085]
[0086] Table 4(3) The ablation experiments of different modules are shown in Table 5:
[0087]
[0088] Table 5
[0089] (4) The impact of Token Drop on training efficiency Figure 4 shown.
[0090] (5) Comparative analysis of the spectral characteristics of music generated by different models and real music. Figure 5 shown.
[0091] (6) Human verification is shown in Table 6:
[0092]
[0093] Table 6
[0094] The above-mentioned embodiments only express the specific implementation of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A music generation diffusion method based on emotion guidance, characterized in that: The following steps are involved: Step 1: Compress music audio by training a variational autoencoder (VAE) to extract latent features and use a diffusion model to model latent variables. Step 2: Complete the diffusion process of music based on emotion guidance: embed the emotion guidance model to generate music with specific emotions; randomly selected tokens are discarded to improve efficiency; the latent variables generated by the previous diffusion process are used as conditional input to enhance the consistency of the generated results.
2. The method for generating and diffusing music based on emotion guidance according to claim 1, characterized in that: The step 1 specifically comprises the following steps: Step 1.1: The variational autoencoder VAE is used to represent the music source S∈R containing N samples in the waveform domain N Compressed into a compact and continuous latent space while ensuring that the reconstruction result is perceptually indistinguishable from the original sound source; given an input signal S, the encoder maps it to the posterior distribution: in, is the latent posterior mean, ∑ z (S) is the posterior covariance matrix, D is the time domain downsampling factor, and C is the latent space dimension; Step 1.2: After encoding, sampling and input it into the decoder to reconstruct the signal S; using the posterior mean z s =μ z (S) as a potential representation; Step 1.3: Based on the potential diffusion model, in the forward diffusion process, the original sample x0 is gradually added with noise in T steps to generate a series of noisy samples x1, x2, ..., x T ; At each time step t, sample x t The conditional probability distribution of is determined by the sample x at the previous moment t-1 Determine, its mathematical form is: Among them, β1,…,β t ,…,β T is a predefined noise scheduling parameter; Step 1.4: According to the properties of Gaussian distribution, we can deduce: in α t =1-β t ; By sampling And use the reparameterization technique to get the sample Step 1.5: Reverse generation process from pure noise sample x T Start by gradually denoising and reconstructing x T-1 ,x T-2 ,…,x0, and finally obtain realistic samples; the reverse process is defined as the conditional probability distribution p θ (x t-1 |x t ), learned through a neural network, is used to approximate q(x t-1 |x t ,x0); Step 1.6: To learn p θ (x t-1 |x t ), the training model output ∈ θ (x t ,t) to restore the generated x t The noise ∈ added when ; the loss function of the training diffusion model is: Step 1.7: During inference, given x t and predicted noise, from p θ (x t-1 |x t )sampling: in 3. The method for generating and diffusing music based on emotion guidance according to claim 2, characterized in that: The step 2 specifically includes the following steps: Step 2.1, introduce the emotion feature encoder, embed the music emotion features into the diffusion model through the cross attention mechanism, and guide the diffusion model to generate music clips that meet specific emotions; Step 2.2: Improve the Token Drop strategy so that it randomly drops some tokens during training. Step 2.3: Propose a Self-Conditioning mechanism, which uses the previous generation results of the diffusion model as conditional input to provide contextual information for subsequent generation, thereby ensuring the consistency of music melody and emotion.
4. The method for generating and diffusing music based on emotion guidance according to claim 3 is characterized in that: The step 2.1 is as follows: Assume that the potential representation of multiple input music sources is where z i is the potential representation of the i-th music clip, K is the number of clips; the music emotion information is generated by the emotion feature encoder, which maps the emotion description to the feature matrix E∈R in the latent space M×C , where M is the number of sentiment features; during the generation process, a cross-attention mechanism is used to combine sentiment information with the latent representation: where A∈R K×M is the attention weight, It is the potential representation after integrating emotional information.
5. The method for generating and diffusing music based on emotion guidance according to claim 4, characterized in that: The step 2.2 is specifically as follows: The potential representation z t Divide into small blocks of size p×p, each small block is flattened into a vector to form a token, the total number of tokens is Subsequently, the tokens are reshaped into a matrix where d = c × p 2 Represents the dimension of a single token; Based on the dynamic masking mechanism, the token matrix u is selectively masked: 1) the masking ratio ρ is defined to determine the number of tokens ρN that need to be masked; 2) the mask matrix By randomly or based on a specific weight, some tokens are selected for masking, M[i]=1 indicates masked, and M[i]=0 indicates unmasked; 3) The masked token is represented as: u'=M⊙u+(1-M)⊙0, where ⊙ represents an element-by-element product operation; The encoder of the diffusion model focuses on the processing of the unmasked token u' and generates the feature representation q; the side interpolator Int(·) is introduced to recover the masked token and fill the masked area by interpolation. The formula is: k = (1-M)·q+M·Int(q), where Int(q) estimates the masked token according to the encoder output q; the interpolated token k is embedded into the input decoder in combination with the position to recover the complete potential representation And restore high-resolution audio data through variational autoencoder VAE.
6. The method for generating and diffusing music based on emotion guidance according to claim 5, characterized in that: The step 2.3 is as follows: Using Self-Conditioning to introduce the estimate of the previous time step in the denoising network Enable the network to refer to historical information to improve the prediction of the current time step; in Self-Conditioning, modify the estimate of the denoising network to: in is the estimate of the previous time step t+1, and x is t and Splicing; During the training phase, the Self-Conditioning input is set to zero, that is, Compute a preliminary estimate: in is to represent x based only on the current noise t and the estimated results at time step t; When estimating with Self-Conditioning, a preliminary estimate is obtained in the first forward propagation After that, by stopping the gradient operation, Used as Self-Conditioning input for the second forward propagation: The denoising network is then optimized using the output of the two forward passes so that it can accurately estimate x0; In the diffusion process, according to the time schedule σ(t) = t, the forward diffusion process is defined as: in Sampling by solving the inverse ODE process The fraction By Neural Network approximation and trained with score matching loss.
7. The method for generating and diffusing music based on emotion guidance according to any one of claims 1 to 6, characterized in that: The music is Tibetan music, which gradually solves the problems of lack of expression ability of specific emotions, low efficiency of high-dimensional feature processing, and insufficient consistency of music context in Tibetan music generation.