A fast text-guided sound effect generation method and system based on segmented rectification
By using the distillation learning method of segmented rectification in the sound effect generation model, the student diffusion model is trained. The trajectory of the model in each time period is straight, solving the problem of slow generation of existing sound effect generation models and achieving fast and high-quality sound effect generation.
Patent Information
- Application Number
- CN202510260365.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing sound effect generation model is slow to generate and difficult to deploy to scenarios with high performance requirements.
The rapid text-guided sound effect generation method based on segmented rectification is adopted. The teacher's diffusion model is pre-trained and the student's diffusion model is trained using distillation learning. The differential equation trajectory of the student's diffusion model in each time period is a linear trajectory, and the noise Mel spectrogram characteristics of the sampling time step are calculated using linear interpolation.
The student diffusion model can generate sound effects of the same quality with a generation step much lower than the teacher diffusion model, significantly improving the speed and quality of sound effects generation.
Smart Images

Figure CN119811365B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound effect generation, and in particular to a fast text-guided sound effect generation method and system based on piecewise rectification. Background Art
[0002] At present, automated sound effect generation has attracted wide attention and become an important field of artificial intelligence. This field aims to enable users to achieve automated sound effect generation by inputting descriptive text about a sound effect, which plays an important role in the development of film dubbing, game dubbing, short video production, and virtual reality technology.
[0003] Existing technologies can be divided into two categories: 1) sound effect generation models based on autoregressive models; 2) sound effect generation models based on diffusion models.
[0004] The autoregressive sound effect generation model encodes the sound effect waveform into a discrete audio latent variable sequence using a vector quantization variational autoencoder and generates discrete latent variable tokens one by one, achieving good performance. However, due to the limitation of its model structure, the model needs to be run once for generating each discrete latent variable. For a discrete latent variable sequence of length T, the model needs to be run T times, consuming a large amount of time and making it difficult to be deployed in scenarios with high performance requirements. Although existing technologies use a vector quantization variational autoencoder to compress waveform data into discrete latent codes to reduce the length of the audio data sequence and thus reduce the number of times the autoregressive model needs to be run, the compression ability for data length is limited. A too high compression ratio will correspond to a lower reconstruction quality and reduce the quality of the audio generated by the autoregressive model. Existing sound effect generation models based on diffusion models also need to run the model multiple times for multi-step denoising to achieve high sound effect generation quality. Although existing technologies can take a subsequence of the time step sequence of the diffusion model to achieve generation with fewer steps, since the inverse diffusion step of the diffusion model is an ordinary differential equation (ODE) solver (ODESolver) based on the diffusion function, due to the limitation of the mathematical form of the inverse solver, the ordinary differential equation trajectory is a curved trajectory. When the number of sampling times is too small, it is impossible to avoid the further expansion of errors, and the reduction of steps is limited, resulting in a slow generation speed of the sound effect generation model, which also limits the deployment of the sound effect generation model in scenarios with high performance requirements. Summary of the Invention
[0005] To overcome the problem of slow generation speed of existing sound effect generation models, the present invention provides a fast text-guided sound effect generation method and system based on piecewise rectification. A pre-trained teacher diffusion model is used to distill a student diffusion model with the same structure as the teacher diffusion model by piecewise rectification objective. The student diffusion model can generate high-quality sound effects with much fewer reverse diffusion steps than the teacher diffusion model, so as to achieve fast text-guided sound effect generation.
[0006] To achieve the above object, the specific technical solution adopted by the present invention is as follows:
[0007] In the first aspect, the present invention proposes a fast text-guided sound effect generation method based on piecewise rectification, including:
[0008] Obtain a training set composed of sound effect waveform data and its description text, and pre-train a diffusion model as the teacher diffusion model; during pre-training, the teacher diffusion model uses the text features of the description text as a guide to gradually predict the noise for the noisy Mel spectrogram features based on the sound effect waveform data;
[0009] Use the teacher diffusion model to train a student diffusion model by distillation learning. Divide the time steps into several time periods based on a predefined number of segments, and the differential equation trajectory of the student diffusion model is a straight-line trajectory within each segment;
[0010] In each training batch of the distillation learning stage, randomly sample time steps within a random time period, calculate the noisy Mel spectrogram features corresponding to the start and end time steps of the teacher diffusion model within the selected random time period, and use linear interpolation to calculate the noisy Mel spectrogram features corresponding to the sampled time steps; the student diffusion model uses the text features of the description text as a guide to predict the noise for the noisy Mel spectrogram features corresponding to the sampled time steps, and calculates the loss with the inter-segment denoising noise of the teacher diffusion model to update the student diffusion model;
[0011] In the sound effect generation stage, the user provides the description text, initializes the noise, and the student diffusion model uses the text features of the description text as a guide to run the reverse diffusion denoising process segment by segment and generate the final sound effect.
[0012] Furthermore, the teacher diffusion model and the student diffusion model are diffusion models with the same structure.
[0013] Furthermore, the teacher diffusion model and the student diffusion model include:
[0014] A sine encoding layer for sine encoding the time steps to obtain the diffusion model time step features;
[0015] A first fully connected layer for projecting the diffusion model time step features to generate time step embedding features;
[0016] A second fully connected layer for projecting text features to generate text embedding features;
[0017] A one-dimensional convolutional layer for projecting the noisy Mel spectrogram features to generate noisy Mel spectrogram embedding features;
[0018] N stacked Transformer blocks for extracting encoded features from the concatenation result of the time step embedding features, text embedding features, and noisy Mel spectrogram embedding features;
[0019] A group normalization and one-dimensional convolutional layer that takes the encoded features as input to generate an output sequence;
[0020] A predicted noise output layer for cropping out the part corresponding to the sound effect waveform data from the output sequence as the predicted noise.
[0021] Furthermore, in the pre-training stage of the teacher diffusion model, the diffusion model is updated with the loss between the predicted noise and the actual noise of the diffusion model.
[0022] Furthermore, the distillation learning stage of the student diffusion model includes:
[0023] Initializing the student diffusion model parameters with the teacher diffusion model parameters and dividing the total number of time steps T into several time periods;
[0024] Sampling a training batch of data from the dataset D, where the dataset D consists of the Mel spectrogram features and text features corresponding to the training set; performing the following steps a to d for each training batch until the student diffusion model converges:
[0025] Step a, randomly select one of the time periods , and randomly sample time steps within the selected time period , and calculate the time steps , corresponding to the start and end time steps of the selected time period , of the teacher diffusion model;
[0026] Step b, run the diffusion noise addition process based on the Mel spectrogram features and random noise to obtain the noisy Mel spectrogram features of the teacher diffusion model at the time step ; then gradually run the inverse diffusion denoising process of the teacher diffusion model based on the noisy Mel spectrogram features and text features until obtaining the noisy Mel spectrogram features of the teacher diffusion model at the time step ;
[0027] Step c, calculate the sampling time step by using linear interpolation The corresponding noisy Mel spectrogram features :
[0028] Step d, the student diffusion model uses the text features as a guide to predict the noise for the noisy Mel spectrogram features corresponding to the sampling time step and calculate the loss function to update the student diffusion model.
[0029] Furthermore, in step a, the calculation formulas for the time steps , are as follows:
[0030] ;
[0031] ;
[0032] wherein, represents rounding down.
[0033] Furthermore, in step c, the formula for the linear interpolation method is as follows:
[0034] .
[0035] Furthermore, in step d, the calculation formula for the loss function is as follows:
[0036] ;
[0037] wherein, represents the loss of the student diffusion model, represents the parameters of the student diffusion model, represents the student diffusion model, represents the mean coefficient of the Gaussian distribution, represents the variance coefficient of the Gaussian distribution, represents the norm, is the inter-segment denoising noise of the teacher diffusion model.
[0038] Furthermore, the sound effect generation stage of the described student diffusion model includes:
[0039] Extract the text features of the description text provided by the user, and randomly sample Gaussian noise;
[0040] Divide the total number of time steps into several time periods, and traverse each time period in the order from the last to the first. The traversal process is as follows: guided by the text features of the description text, gradually run the denoising process of the inverse diffusion of the student diffusion model within each time period, and finally generate the Mel spectrogram features without noise.
[0041] Use the pre-trained decoder to decode the Mel spectrogram features without noise to obtain the Mel spectrogram, and then use the pre-trained vocoder to convert the Mel spectrogram into an audio waveform.
[0042] In a second aspect, the present invention provides a fast text-guided sound effect generation system based on piecewise rectification for implementing the above-mentioned fast text-guided sound effect generation method based on piecewise rectification.
[0043] The beneficial effects of the present invention are as follows:
[0044] The present invention first pre-trains a teacher diffusion model to achieve the purpose of gradually predicting noise for the noisy Mel spectrogram features based on sound effect waveform data, and then uses the teacher diffusion model to train a student diffusion model using distillation learning. Based on the curved ordinary differential equation trajectory of the teacher diffusion model, the student diffusion model divides the trajectory into multiple segments and performs distillation learning with a linear rectification objective within each segment. The noisy Mel spectrogram features corresponding to the sampling time steps in the student diffusion model are calculated by linear interpolation. Therefore, the student diffusion model can learn the ordinary differential equation trajectory of a straight line within the segment. Since the expected ordinary differential equation trajectory within each segment becomes a straight line trajectory, the student diffusion model can use a larger sampling step size without causing excessive errors, and can generate sound effects of the same quality with far fewer generation steps than the teacher diffusion model, overcoming the problem of slow generation speed of existing sound effect generation models and significantly improving the sound effect quality of the diffusion model with low number of steps. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a schematic diagram of a fast text-guided sound effect generation method based on piecewise rectification shown in an embodiment of the present invention.
[0046] Figure 2 is a schematic diagram of the diffusion process and the inverse diffusion process shown in an embodiment of the present invention.
[0047] Figure 3 is a schematic diagram of the structure of the diffusion model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following further elaborates and explains the present invention in conjunction with specific embodiments. The embodiments are only examples of the present disclosure content and do not delimit the scope of limitation. Without conflict, the technical features of each embodiment of the present invention can be combined accordingly.
[0049] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0050] The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all steps. For example, some steps can be further decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0051] Figure 1 Shown is a schematic diagram of a method for generating fast text-guided sound effects based on piecewise rectification. In the training stage, a sound effect and its description text are input to train a diffusion model to achieve text-guided sound effect generation; the diffusion model obtained from the above training is used as the teacher diffusion model, and a student diffusion model with the same structure as the teacher diffusion model is distilled with piecewise rectification objective. In the inference stage, a description text of a sound effect and Gaussian sampled noise are input, and the student diffusion model gradually removes noise in segments to generate the final sound effect. This student diffusion model can generate high-quality sound effects with far fewer reverse diffusion steps than the teacher diffusion model.
[0052] Specifically, a method for generating fast text-guided sound effects based on piecewise rectification proposed by the present invention mainly includes the following steps:
[0053] Step 1, obtain the sound effect waveform data and the description text data of the sound effect waveform data as the training set; preprocess the sound effect waveform data to obtain the sound effect mel spectrogram data.
[0054] Step 2, extract the mel spectrogram features of the sound effect mel spectrogram data, and extract the text features of the description text.
[0055] Step 3, randomly sample the time step t, add noise to the mel spectrogram features using the diffusion process to obtain the noisy mel spectrogram features.
[0056] Step 4, use the sampled time step, text features, and noisy mel spectrogram features as the input of the diffusion model, and the diffusion model predicts the noise at the sampled time step. Use the mean square error loss between the actual noise added to the mel spectrogram features at the sampled time step and the predicted noise of the diffusion model as the loss function, and use the trained diffusion model as the teacher diffusion model;
[0057] Step 5: Define a student diffusion model with the same structure as the teacher diffusion model, initialize the student diffusion model with the parameters of the teacher diffusion model, and train the student diffusion model in a piecewise rectified manner using the training set and the teacher diffusion model; the student diffusion model can generate sound effects of the same quality with far fewer generation steps than the teacher diffusion model, achieving fast text-guided sound effect generation.
[0058] Step 6: The user provides a description text and extracts text features. The student diffusion model samples noise from a Gaussian distribution and performs an inverse diffusion process conditioned on the text feature input for the sampling result to generate Mel spectrogram features; decode the Mel spectrogram features to obtain a Mel spectrogram, and then convert it into an audio waveform through a pre-trained vocoder to obtain the final sound effect.
[0059] In the above step 1, the process of preprocessing the sound effect waveform data is as follows: Resample the sound effect waveform data, and convert the sound effect waveform data into sound effect Mel spectrogram data through short-time Fourier transform and Mel filters. In this embodiment, for the sound effect waveform-description text data pair (a, y), where is the audio waveform. For an audio waveform resampled to 16,000 Hz, if the duration of the audio waveform is n seconds, then . represents the description text corresponding to the sound effect waveform. Convert the sound effect waveform data into sound effect Mel spectrogram data through short-time Fourier transform and Mel filters in sequence , where is the number of Mel bins, is the number of Mel frames, is much smaller than .
[0060] In the above step 2, an optional implementation process for obtaining Mel spectrogram features and text features is as follows:
[0061] For the sound effect Mel spectrogram , use a pre-trained Mel spectrogram encoder to encode and obtain Mel spectrogram features:
[0062]
[0063] where E is a pre-trained Mel spectrogram encoder, , are the downsampling rates of the number of Mel bins and the number of Mel frames respectively, is the Mel spectrogram feature. In this embodiment, the Mel spectrogram encoder adopts the encoder part in a variational autoencoder. The structure and pre-training process of the variational autoencoder are well-known common knowledge in the art and will not be elaborated here.
[0064] For the description text , use a pre-trained text encoder to Convert to text features In this embodiment, the pre-trained text encoder is implemented using the existing Flan-T5-Large model, which will not be elaborated here.
[0065] In step 3 above, as Figure 2 shown, the process of adding noise to the Mel spectrogram features using the diffusion process is as follows:
[0066] Taking the diffusion process as an example, it is expressed as:
[0067]
[0068] According to the properties of the Gaussian distribution, it can be derived that:
[0069]
[0070] Suppose there are T time steps, sample the time step t from the uniform distribution of {1, 2, …, T}, and according to the predefined time step parameter , add noise to the Mel spectrogram features in the way of , where represents the noise-added Mel spectrogram features, is the noise sampled from the Gaussian distribution with a mean of 0 and a variance of , and this sampled noise has the same shape as the Mel spectrogram features , is the identity matrix; represents the product of the first t time step parameters, represents the i-th time step parameter.
[0071] In step 4 above, the inverse diffusion process is implemented using the diffusion model shown in Figure 3 , and is gradually denoised and restored . The diffusion model includes a sine encoding layer, a first fully connected layer, a second fully connected layer, and a one-dimensional convolutional layer; the sine encoding layer is used to perform sine encoding on the sampled time step t to obtain the diffusion model time step features; the learnable first fully connected layer, second fully connected layer, and one-dimensional convolutional layer are used to project the diffusion model time step features, text features, and noise-added Mel spectrogram features into the feature space of the diffusion model respectively to obtain the time step embedding features, text embedding features, and noise-added Mel spectrogram embedding features. After the three embedding features are connected, they are input into the Transformer block to extract features. The result after stacking N Transformer blocks is then passed through group normalization and one-dimensional convolution to obtain the output sequence, and the part with the same length as the number of Mel frames is cropped from the output sequence as the predicted noise.
[0072] In a specific implementation of the present invention, an optional implementation process for predicting the noise at the sampling time step t using a diffusion model is as follows:
[0073] 4.1) First, encode the time step embedding feature, text embedding feature, and noisy Mel spectrogram embedding feature:
[0074]
[0075]
[0076]
[0077]
[0078] Among them, represents two learnable first fully connected layers and second fully connected layers, represents sine encoding, represents a learnable one-dimensional convolutional layer, represents feature concatenation; represents the randomly sampled time step, represents the text feature extracted by the text encoder, represents the noisy Mel spectrogram feature; , , represent the time step embedding feature, text embedding feature, and noisy Mel spectrogram embedding feature respectively; h represents the total concatenated feature, and d is the dimension of the feature space of the Transformer block in the diffusion model.
[0079] 4.2) Input the total concatenated feature into the Transformer block, process it through N Transformer blocks in sequence, and then pass the generated result through group normalization and one-dimensional convolution to obtain the output sequence.
[0080] 4.3) Extract the predicted noise from the output sequence.
[0081] During the operation of the diffusion model, the channel dimension of the output sequence is the same as the dimension of the Mel spectrogram feature, and the length of the time dimension of the output result is the same as the time dimension of the total concatenated feature ( ). To make the length of the time dimension the same as the time dimension of the Mel spectrogram feature, crop the part of the output result with a length of at the back as the final predicted noise, that is, .
[0082] 4.4) Use the mean square error loss between the actual noise added to the Mel spectrogram feature at time step t and the predicted noise of the diffusion model as the loss function, and the loss function is expressed as follows:
[0083]
[0084] in, represents the diffusion model;
[0085] The parameters obtained by training are The diffusion model of is used as the teacher diffusion model.
[0086] In step 5 above, an optional implementation of training the student diffusion model is as follows:
[0087] 5.1) Define the following parameters:
[0088] Define the data set D, which consists of the Mel-spectrogram features corresponding to the training set data and text features composition.
[0089] The teacher diffusion model is denoted as , the student diffusion model is denoted as , the time step parameter of the teacher diffusion model is recorded as , the reverse diffusion sampling step of the teacher diffusion model is recorded as The number of segments is recorded as ,create Time periods:
[0090]
[0091] in, =0, .
[0092] 5.2) Initialize the parameters of the student diffusion model:
[0093] Using teacher diffusion model parameters Initialize the parameters of the Student Diffusion Model : .
[0094] 5.3) From the dataset Sample a mini-batch of text features and Mel-spectrogram features ,in Text features representing the descriptive text of the sound effect, Represents a mel-spectrogram feature.
[0095] 5.4) Randomly sample a number k from {1,…,K}, and select Randomly sample time steps from a uniform distribution , calculate the teacher diffusion model and , The corresponding time step is:
[0096]
[0097]
[0098] Among them, is the total number of time steps of the teacher diffusion model, represents rounding down.
[0099] 5.5) Calculate in the time period and , and obtain through interpolation;
[0100] Specifically, first run the diffusion and noise addition process to obtain the noise-added Mel spectrogram features corresponding to steps:
[0101]
[0102] After that, based on , run the reverse diffusion steps of the teacher diffusion model until is obtained, and the process is:
[0103] Denote , , and update in each step:
[0104]
[0105]
[0106] Until , at this time is .
[0107] After obtaining , denote:
[0108]
[0109]
[0110] Among them, represents the mean coefficient of the Gaussian distribution, represents the variance coefficient of the Gaussian distribution;
[0111] Finally, the to be solved is obtained by linear interpolation of the teacher diffusion model's steps and the noise-added features of steps:
[0112]
[0113] 5.6) Calculate the loss function of the student diffusion model as follows:
[0114]
[0115] Calculate the gradient with respect to the loss function for this batch, and use the optimizer to update the parameters of the student diffusion model for updating.
[0116] 5.7) Return to step 5.3), continuously traverse the dataset to update the parameters of the student diffusion model , until the loss function converges; since the parameters of the student diffusion model are initialized with the pre-trained teacher diffusion model parameters, the loss function of the student diffusion model can converge quickly.
[0117] Since the input of the student diffusion model is obtained by linear interpolation between and , the student diffusion model can learn the ordinary differential equation trajectory of the straight line within the segment, and can use a larger step size than the teacher diffusion model for sampling within each segment without large numerical errors, thereby greatly reducing the sampling steps and improving the sound effect generation speed. The student diffusion model can generate sound effects of the same quality with far fewer generation steps than the teacher diffusion model, thereby realizing a fast text-guided sound effect generation system.
[0118] In the above step 6, an optional implementation process for generating sound effects in a text-guided manner using the student diffusion model is as follows:
[0119] 6.1) Define the following parameters:
[0120] The student diffusion model is denoted as , the number of segments is denoted as K, the number of sampling steps within each segment is denoted as m, and the sound effect description text is ; Sample noise from Gaussian noise.
[0121] 6.2) Use the student diffusion model to implement the following reverse diffusion process:
[0122] Set , , , where represents the generation step size of the student diffusion model;
[0123] Let , , the segment position and corresponding teacher diffusion model time step cumulative multiplication parameters are and , the mean coefficient of the Gaussian distribution corresponding to the DDIM sampling of the teacher diffusion model can be calculated. and the variance coefficient :
[0124]
[0125]
[0126] When , according to the sampling method of DDIM, denote the mean coefficient of the Gaussian distribution linearly interpolated at the position of randomly sampled in the segment as:
[0127]
[0128] Denote the variance coefficient of the Gaussian distribution after linear interpolation as:
[0129]
[0130] The end position of the segment predicted by the student model is:
[0131]
[0132] Then the direction vector of the Euler solver for each step of the student diffusion model is:
[0133]
[0134] Therefore, at each step, run the Euler solver to update and t:
[0135]
[0136]
[0137] Until ; when , the student diffusion model runs to the end of the segment, and at this time, update the segment:
[0138]
[0139]
[0140] Start running the sampling of the next segment.
[0141] Until ; at this time, the student diffusion model successfully denoises the Gaussian noise, and the at this time is the clean Mel spectrogram feature, which is denoted as 。
[0142] 6.3) Use a decoder to process the Mel spectrogram features Decode to obtain the Mel spectrogram, and then the pre-trained vocoder can convert the Mel spectrogram into an audio waveform. In this embodiment, the Mel spectrogram decoder uses the decoder part in the variational autoencoder. The structure and pre-training process of the variational autoencoder are well-known common knowledge in the art and will not be elaborated here.
[0143] The above method will be applied to the following embodiments to demonstrate the technical effects of the present invention. The specific steps in the embodiments will not be elaborated.
[0144] The present invention evaluates the audio generation metrics on the test set of the AudioCaps audio-effect-description text dataset. Compare the generation metrics of the TANGO audio generation model, the teacher diffusion model, and the student diffusion model. The results are shown in Table 1. Among them, IS and FAD are audio quality evaluation metrics, and KL and CLAP are speech-text alignment metrics. The larger the IS and CLAP, the better; the smaller the FAD and KL, the better.
[0145] Table 1: Test results of the AudioCaps audio dataset of the present invention
[0146]
[0147] As can be seen from Table 1, when generating at a low number of steps of 4 steps, the student diffusion model has a significant advantage compared to the teacher diffusion model and the TANGO audio generation model; after increasing the number of steps, the three models gradually become equal. The method proposed by the present invention can generate high-quality audio with a much lower number of generation steps than the comparison models, realizing fast text-guided audio generation.
[0148] Based on the same inventive concept, in this embodiment, a fast text-guided audio generation system based on piecewise rectification is also provided. This system is used to implement the above embodiment. Terms such as "module" and "unit" used below can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.
[0149] In this embodiment, a fast text-guided audio generation system based on piecewise rectification includes:
[0150] A teacher diffusion model pre-training module, which is used to obtain a training set composed of audio waveform data and its description text, and pre-train a diffusion model as the teacher diffusion model; during pre-training, the teacher diffusion model is guided by the text features of the description text and gradually predicts the noise for the noisy Mel spectrogram features based on the audio waveform data;
[0151] A distillation learning module based on a rectified linear unit target, which is used to train a student diffusion model by distillation learning using a teacher diffusion model. The differential equation trajectory of the student diffusion model is a straight-line trajectory, and the time steps are divided into several time periods based on the straight-line trajectory; in each training batch of the distillation learning stage, time steps are randomly sampled within a random time period, the noisy mel spectrogram features corresponding to the start and end time steps within the selected random time period of the teacher diffusion model are calculated, and the noisy mel spectrogram features corresponding to the sampled time steps are calculated by linear interpolation; the student diffusion model is guided by the text features describing the text, predicts the noise for the noisy mel spectrogram features corresponding to the sampled time steps, and calculates the loss with the inter-segment denoising noise of the teacher diffusion model to update the student diffusion model;
[0152] A sound effect generation module based on the student diffusion model, which initializes the noise and uses the student diffusion model to guide the inverse diffusion denoising process segment by segment with the text features of the description text provided by the user to generate the final sound effect.
[0153] In a specific implementation of the present invention, it may further include: a text encoder, which is used to convert the description text into text features;
[0154] A mel spectrogram variational autoencoder, whose encoder part is used to extract the mel spectrogram features of the sound effect waveform data, and the decoder part is used to reconstruct the mel spectrogram features into a mel spectrogram;
[0155] A vocoder, which is used to convert the mel spectrogram into an audio waveform.
[0156] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment, and the implementation methods of the remaining modules are not elaborated here. The system embodiments described above are only illustrative, where the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention scheme. Those of ordinary skill in the art can understand and implement it without creative work.
[0157] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0158] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A method for generating fast text-guided sound effects based on segmented rectification, characterized in that: include: A training set consisting of sound effect waveform data and its description text is obtained, and a diffusion model is pre-trained as a teacher diffusion model; during pre-training, the teacher diffusion model is guided by the text features of the description text to gradually predict noise based on the noisy Mel-spectrogram features of the sound effect waveform data; Using the teacher diffusion model to train a student diffusion model using distillation learning, the time step is divided into a number of time periods based on a predefined number of segments, and the differential equation trajectory of the student diffusion model is a straight line trajectory in each segment; Both the teacher diffusion model and the student diffusion model are diffusion models based on Transformer blocks; the distillation learning stage of the student diffusion model includes: Initialize the student diffusion model parameters using the teacher diffusion model parameters and divide the total time step T into several time periods; Sample a training batch of data from a dataset D, where the dataset D consists of the Mel-spectrogram features corresponding to the training set and text features Composition; perform the following steps a to d in each training batch until the student diffusion model converges: Step a: Randomly select one of the time periods , and randomly sample time steps within the selected time period , calculate the teacher diffusion model and the first and last time steps of the selected time period , The corresponding time step , ; Step b, based on Mel-spectrogram features and random noise to run the diffusion and noise process, and obtain the teacher diffusion model at time step Noise-added Mel-spectrogram features ; Based on the noisy Mel spectrum features and text features Step by step, run the inverse diffusion denoising process of the teacher diffusion model until the teacher diffusion model is obtained at time step Noise-added Mel-spectrogram features ; Step c: use linear interpolation to calculate the sampling time step Corresponding noisy Mel-spectrogram features : Step d, student diffusion model based on text features To guide, the sampling time step Corresponding noisy Mel-spectrogram features Predict noise and calculate the loss function to update the student diffusion model; In each training batch of the distillation learning phase, randomly sample time steps in a random time period, calculate the noisy Mel-spectrogram features corresponding to the first and last time steps in the selected random time period of the teacher diffusion model, and use linear interpolation to calculate the noisy Mel-spectrogram features corresponding to the sampling time step; the student diffusion model is guided by the text features describing the text, predicts the noise of the noisy Mel-spectrogram features corresponding to the sampling time step, and calculates the loss with the inter-segment denoising noise of the teacher diffusion model to update the student diffusion model; In the sound effect generation stage, the user provides a description text and initializes the noise. The student diffusion model is guided by the text features of the description text, runs the inverse diffusion denoising process section by section, and generates the final sound effect.
2. The method for generating fast text-guided sound effects based on segmented rectification according to claim 1, characterized in that: The teacher diffusion model and the student diffusion model are diffusion models with the same structure.
3. The method for generating fast text-guided sound effects based on segmented rectification according to claim 2, characterized in that: The teacher diffusion model and the student diffusion model include: The sinusoidal coding layer is used to perform sinusoidal coding on the time step to obtain the time step characteristics of the diffusion model; The first fully connected layer is used to project the diffusion model time step features to generate time step embedding features; The second fully connected layer is used to project text features to generate text embedding features; A one-dimensional convolutional layer for projecting the noisy Mel-spectrogram features to generate noisy Mel-spectrogram embedding features; N stacked Transformer blocks, which are used to extract encoding features from the concatenation of time-step embedding features, text embedding features, and noisy Mel-spectrogram embedding features; Group normalization and 1D convolutional layers, which take the encoded features as input and generate output sequences; The prediction noise output layer is used to cut out the part corresponding to the sound effect waveform data from the output sequence as the prediction noise.
4. The method for generating fast text-guided sound effects based on segmented rectification according to claim 1 or 3, characterized in that: During the pre-training phase of the teacher diffusion model, the diffusion model is updated with the loss between the predicted noise of the diffusion model and the actual noise.
5. The method for generating fast text-guided sound effects based on segmented rectification according to claim 1, characterized in that: In step a, the time step , The calculation formula is as follows: ; ; in, Indicates rounding down.
6. The method for generating fast text-guided sound effects based on segmented rectification according to claim 1, characterized in that: In step c, the formula for linear interpolation is as follows: 。 7. The method for generating fast text-guided sound effects based on segmented rectification according to claim 1, characterized in that: In step a and step d, the loss function is calculated as follows: ; in, represents the student diffusion model loss, represents the parameters of the Student diffusion model, represents the student diffusion model, represents the mean coefficient of the Gaussian distribution, represents the variance coefficient of the Gaussian distribution, represents the norm, is the inter-segment denoising noise of the teacher diffusion model.
8. The method for generating fast text-guided sound effects based on segmented rectification according to claim 3, characterized in that: The sound effect generation stage of the student diffusion model includes: Extract text features of the description text provided by the user, and,randomly sample Gaussian noise; The total time step is divided into several time periods, and each time period is traversed in order from the last time period to the first time period. The traversal process is as follows: guided by the text features describing the text, the inverse diffusion denoising process of the student diffusion model is gradually run in each time period, and finally a noise-free Mel spectrum feature is generated; The pre-trained decoder is used to decode the noise-free Mel-spectrogram features to obtain a Mel-spectrogram, and then the pre-trained vocoder is used to convert the Mel-spectrogram into an audio waveform.
9. A fast text-guided sound effect generation system based on segmented rectification, used to implement the method of claim 1; characterized in that: The system comprises: A teacher diffusion model pre-training module is used to obtain a training set consisting of sound effect waveform data and its description text, and pre-train a diffusion model as a teacher diffusion model; during pre-training, the teacher diffusion model is guided by the text features of the description text to gradually predict the noise based on the noisy Mel-spectrogram features of the sound effect waveform data; A distillation learning module based on a linear rectification target is used to train a student diffusion model using distillation learning using a teacher diffusion model, and divide the time step into several time periods based on a predefined number of segments, and the differential equation trajectory of the student diffusion model is a straight line trajectory in each segment; in each training batch of the distillation learning stage, randomly sample time steps in a random time period, calculate the noisy Mel spectrum features corresponding to the first and last time steps of the teacher diffusion model in the selected random time period, and use linear interpolation to calculate the noisy Mel spectrum features corresponding to the sampled time steps; the student diffusion model is guided by the text features describing the text, predicts the noise of the noisy Mel spectrum features corresponding to the sampling time steps, and calculates the loss with the inter-segment denoising noise of the teacher diffusion model to update the student diffusion model; The sound effect generation module based on the student diffusion model initializes the noise, uses the student diffusion model to guide the text features of the description text provided by the user, runs the inverse diffusion denoising process section by section and generates the final sound effect.
Citation Information
Patent Citations
Picture generation acceleration method for large text graph diffusion model
CN118229817A
Speech code conversion speech synthesis method based on multi-task denoising diffusion implicit model
CN118762686A