Diffusion model speech enhancement method based on progressive jump connection
By introducing progressive jump connections and optimization prediction goals into the diffusion model, the problem of conditional information loss in the diffusion model is solved, efficient speech enhancement in complex noise environments is achieved, and speech quality and generation consistency is improved.
Patent Information
- Application Number
- CN202510404705.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing speech enhancement technology based on diffusion model has problems inconsistent with the generated speech and conditional signals, lack of stability and consistency when utilizing conditional information, and conditional information is gradually lost when the network depth increases, affecting the generation effect.
The progressive jump connection strategy and optimization prediction target are adopted. By introducing jump connections in the diffusion model, the condition information is directly input to each residual layer, and the condition information weight of each layer is adjusted through the jump connection coefficient to establish an effective connection between the model output and the condition information, so as to avoid the loss of condition information with the increase of depth.
It improves the speech enhancement stability and generation consistency of the model in complex noise environments, enhances the speech quality and noise suppression effect, and improves the generalization ability of the model and the accuracy of the generation results.
Smart Images

Figure CN120260600A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of artificial intelligence, and in particular relates to a diffusion model speech enhancement method based on progressive jump connections. Background Art
[0002] Speech enhancement (SE) refers to the processing of noisy speech data to reduce noise and other distortions to improve the quality and intelligibility of speech signals. This technology has important practical value in many fields and can improve the quality of voice communication, enhance the effectiveness of voice control systems, and improve the accuracy and efficiency of automatic speech recognition systems. However, speech enhancement technology still faces many challenges in practical applications, especially in how to use noisy speech as conditional information to help generate speech.
[0003] In recent years, diffusion models have been introduced into speech enhancement tasks due to their ability in conditional generation. They gradually restore clear speech by iteratively adding and removing Gaussian noise, and have shown good potential in improving the generalization ability of the model. However, existing methods based on diffusion models generally have the problem of semantic inconsistency between generated speech and conditional signals, resulting in a lack of stability and consistency in the enhancement results. In addition, diffusion models often ignore the guidance of conditional information during training and sampling, which further affects the enhancement effect. At present, many methods have been studied to solve the problem of conditional information utilization, but they usually have the following problems:
[0004] (1) Failure to input conditional information into the model through appropriate methods;
[0005] (2) Conditional information is only input at the initial stage of the model, lacking an effective fusion mechanism throughout the entire network structure, which makes it difficult for conditional information to continuously influence the generation process;
[0006] (3) Intermediate features rather than original conditional signals are often transmitted between the hidden layers of the model. As the network depth increases, the detailed features in the conditional speech are gradually lost, affecting the consistency and quality of the generated speech.
[0007] In summary, how to better utilize conditional information to generate perceived speech and avoid the collapse of conditional information in existing speech enhancement technologies based on diffusion models is a key issue that needs to be solved in the current field of speech enhancement. These issues have hindered the further development of diffusion models in speech enhancement tasks. Summary of the invention
[0008] The object of the present invention is to address the above-mentioned technical problems existing in the prior art, and propose a diffusion model speech enhancement method based on progressive skip connections, introducing a progressive skip connection strategy to mitigate the mutual information loss between the input and output that gradually occurs as the network depth increases. Additionally, the present invention also explicitly establishes the connection between the model output and the conditional information by using a prediction target with a higher utilization rate of conditional information.
[0009] The goal of single-channel speech enhancement is to improve the quality of speech signals contaminated by noise. Assume that the noisy signal z represents the mixture of clean speech c and noise γ. This relationship can be expressed by the following mathematical formula:
[0010] z = c + γ (1);
[0011] Therefore, the goal of the speech enhancement (SE) task is to recover the clean speech signal c from the noisy speech input z. The usual discriminative method is to train a model to directly learn the mapping from the noisy speech input z to the clean speech signal c. This not only relies on a large amount of data for training but also has unsatisfactory generalization ability for out-of-distribution data. To improve the generalization ability of the model, the diffusion-based speech enhancement model disrupts the clear speech by adding Gaussian noise with different intensities in the forward process, and then learns a scoring function to reverse this disruption and reconstruct the clear speech. In the forward process, the one-step noise addition formula can be expressed as:
[0012]
[0013] where c t represents all intermediate time-step data starting from the clean data, and also represents the speech signal gradually contaminated by noise during the diffusion process. η(t) represents the noise scale function (or signal-to-noise ratio), which controls the ratio between the clean signal c0 and the Gaussian noise during the noise addition process, and also represents the intensity of noise injection. ξ is a randomly sampled standard Gaussian noise. t represents the diffusion time step, which not only represents the current data's stage in the entire noise addition or denoising process but also represents the noise intensity contained in the current data. Therefore, according to this noise addition formula, the data will gradually be contaminated into a Gaussian noise that follows the standard normal distribution.
[0014] To use this noise addition process to generate data, a neural network μ θ is usually parameterized to predict the scoring function Δ ct logm t (c t |c0) at the current time step. From the formula, this represents the logarithmic gradient of the current data distribution at time t.
[0015] Specifically, the input of this neural network is (t, z, c t)The loss function LF to be minimized for the triple has the following objective:
[0016]
[0017] wherein, represents the expectation over all variables t, z, c0. z is the noisy speech, which is regarded as conditional information and used as the input of the model. μ θ (t, z, c t ) is the output of the model. The neural network can be any one of a U-shaped network with residual modules, a U-shaped network with time encoding, or a deep neural network with a Transformer architecture. In a specific implementation manner, the noise predictor is one of Diffusion Transformer, DiffWave, Latent Diffusion, etc. The diffusion model modeled according to the above definition can use the clean speech c0 and the noisy speech z to train a model using conditional noise, and gradually generate data starting from a standard normal distribution in the reverse process.
[0018] To better utilize the conditional information for training the model, the present invention finds that the Gaussian noise ξ and the noisy speech z are independent of each other. Therefore, using the conditional information z is of little help for predicting the Gaussian noise ξ. Specifically, by further deriving the scoring function, we can obtain:
[0019]
[0020] Obviously, according to the above formula, the main objective for the model to learn is the Gaussian noise ξ contained in c t . Therefore, using the conditional information z is of little help for predicting the Gaussian noise ξ that is independent of it.
[0021] Based on this, the present invention proposes a more effective prediction objective to explicitly establish the connection between the model and the conditional information. According to the discrete form of the forward process, it is found that by appropriate parameterization, predicting the Gaussian noise ξ is mathematically equivalent to predicting the clean speech c0:
[0022] ∥μ θ (t, z, c t ) - ξ∥ 2 = δ t ∥μ θ (t, z, c t ) - c0∥ 2 (5);
[0023] wherein, is a set of weight coefficients related to t. The present invention finds through experiments that ignoring the weight δ t helps to improve the training stability. Therefore, the final training objective is:
[0024]
[0025] By changing the prediction target of the model, the model can better utilize conditional information to learn and recover the clean speech signal. The training process designed in the present invention is as follows: First, a pair of paired samples are selected from the speech enhancement dataset, including the clean speech signal c0 and the corresponding noisy speech signal z. To ensure the consistency of the model input and training efficiency, preprocessing operations need to be performed on this pair of data, mainly including steps such as clipping, padding, and alignment, to ensure that c0 and z have the same length and valid region in the time dimension. Then, a diffusion time step t is randomly sampled from the preset time step range, and the clean speech signal c0 is added noise through a randomly sampled Gaussian noise according to formula (2) to obtain the data c at time step t t , which exactly simulates the forward noise addition process of the diffusion model. Then, the (t, z, c t ) triple is input into the model to attempt to recover the original clean speech signal c0, and finally, the model is optimized according to the loss function formula (6) proposed in the present invention.
[0026] In addition, the present invention captures that as the network depth increases, the features of the input itself are gradually lost in the model output. First, the neural network abstracts the input data through layer-by-layer non-linear transformations. The low-level detailed features (such as edges, textures, etc.) are gradually compressed into more semantic representations (such as object categories or overall structures) at higher levels. Although this abstraction process helps the model capture high-level information, it also causes the low-level detailed features of the input to be gradually weakened or even lost. Second, according to the information bottleneck theory, the network will actively compress the input information during training, only retaining the parts useful for the output task. This means that even if some input features are highly correlated with the input data, if they are irrelevant to the task objective, they will be discarded by the network, further exacerbating the loss of input features. Finally, in the absence of residual connections or skip structures, deep networks are also prone to the problem of gradient attenuation. As the network depth increases, the gradient gradually weakens during backpropagation, resulting in the difficulty of effectively retaining input features during propagation, and ultimately further reducing the correlation between the output data and the input features. Therefore, directly sending the conditional information z into each hidden layer through skip connections can effectively handle the phenomenon of conditional collapse and avoid the reduction of the priority of conditional information as the network depth increases.
[0027] Based on this, the present invention provides a diffusion model speech enhancement method based on progressive skip connections, and its steps are as follows:
[0028] (S1) Obtain time embeddings according to the set time steps;
[0029] (S2) Obtain the noisy speech embedding;
[0030] (S3) Using the noisy speech and randomly sampled standard Gaussian noise as inputs, sampling is performed according to the set time steps using a speech predictor to obtain clean speech; the speech predictor includes several sequentially arranged residual layers; project the conditional information obtained by adjusting the noisy speech embedding through the projection module and the skip connection coefficient to each residual layer.
[0031] In the above step (S1), first perform positional encoding on the set time steps, and then obtain the time embedding through processing by a linear layer and a swish function. Further, the time steps can be encoded by sine and cosine functions.
[0032] In the above step (S2), the noisy speech embedding with the same length as the standard Gaussian noise is obtained by performing operations such as clipping, padding, and alignment on the noisy speech.
[0033] In the above step (S3), for the time step t, the sampling steps of the speech predictor are as follows:
[0034] 1) Add noise to the output of the speech predictor at the previous time step to obtain the noisy speech c at time step t t , and the superposition of the noisy speech c t and the noisy speech z is used as the input of the speech predictor;
[0035] 2) Send the time embedding at time step t to each residual layer;
[0036] 3) Project the conditional information obtained by adjusting the noisy speech embedding through the projection module and the skip connection coefficient to each residual layer;
[0037] 4) The outputs of each residual layer are all superimposed on the output of the last residual layer.
[0038] In the above step (3), the speech predictor is any one of a U-shaped network with a residual module, a U-shaped network with time encoding, or a deep neural network with a Transfomer architecture. In a specific implementation, the speech predictor is one of Diffusion Transformer, DiffWave, Latent Diffusion, etc.
[0039] Furthermore, the structures of all residual layers are the same. Each residual layer generates two outputs. One output is added to the output of the previous residual layer and then fed as input into the next residual layer, while the other output is accumulated by the skip connection and fed into the last residual layer as the final residual output. The speech predictor further includes a convolutional layer a before the first residual layer and an activation function (ReLU) and a convolutional layer b after the last residual layer.
[0040] Furthermore, the residual layer provided by the present invention is sequentially connected by a dilated convolutional layer, a gated activation function, and a convolutional layer c; and time embedding is superimposed on the input features of the residual layer, and conditional information is superimposed on the output of the dilated convolutional layer. The gated activation function includes a tanh activation function and a Sigmoid activation function arranged in parallel. The output of the dilated convolutional layer with superimposed conditional information is split into two parts according to the dimension and fed into the two activation functions respectively. Finally, the two activated results are dot-multiplied and then split according to the dimension again to obtain the two outputs of this residual layer, which are used as the input of the next residual layer and the residual for accumulation respectively.
[0041] Furthermore, the projection module includes a first convolutional layer, a ReLu function, a second convolutional layer, and a GLU function arranged in sequence.
[0042] To further process the changes between the residual layers in the speech predictor, the present invention proposes a post-training progressive skip connection strategy.
[0043] Specifically, the present invention sets a skip connection coefficient for each residual layer to control the intensity of the conditional information received by each residual layer from the skip connection, so as to adapt to the unique changes between layers. To maintain stable gradient flow and avoid overfitting caused by the model relying too much on the skip connection, these skip connection coefficients are fixed at 1 during the training of the speech predictor, while they can be adjusted during the inference sampling, such as amplifying the skip connection to compensate for the gradually decreasing mutual information as the network depth increases.
[0044] After extracting a small amount of data from the validation dataset for experiments, the present invention empirically designs a linearly amplified skip connection coefficient, which is represented by the function:
[0045]
[0046] where k . represents the skip connection coefficient of the nth residual layer, n = 1,..., L, L represents the number of residual layers, and k ) , k / are hyperparameters. After experimental verification, this set of hyperparameters can effectively improve the model's ability to utilize conditional information during inference, thereby improving the quality of the denoised speech.
[0047] In the verification sampling stage, the present invention adopts a fast sampling method, and high-quality denoised speech can be generated only by using two-step sampling.
[0048] Specifically, in the sampling stage, the present invention makes full use of conditional information. Instead of adopting the traditional sampling starting from random Gaussian noise, two time steps t and t - r (50 ≥ t, T is the total number of time steps, and r, t - r > 0) are first selected for accelerated sampling. At this time, the above step (S3) includes the following steps:
[0049] (S31) For time step t, the noisy speech z is subjected to noise addition processing to obtain the noisy speech z t , which is expressed as:
[0050]
[0051] (S32) Using the noisy speech z and the noisy speech z t as inputs, a speech predictor is used to perform the first-step sampling to obtain the denoised speech , which is expressed as:
[0052]
[0053] (S33) For time step t - r, the weighted sum of the denoised speech and the noisy speech is subjected to noise addition processing to obtain the noisy speech c t-4 ; here, the average value of the denoised speech and the noisy speech can be subjected to noise addition processing, which is expressed as:
[0054]
[0055] (S34) Using the noisy speech z and the noisy speech c t-4 as inputs, a speech predictor is used to perform the second-step sampling to obtain the clean speech , which is expressed as:
[0056]
[0057] After two-step sampling like this is the clean speech finally predicted by this solution. Where ξ and ξ′ are two independent Gaussian noises. Through this method, the present invention can achieve fast sampling while using conditional information to generate excellent denoised speech.
[0058] Therefore, the present invention first establishes the correlation between the model output and the conditional information by modifying the prediction target, changing the Gaussian noise content existing in the prediction data to clean speech with higher predicted mutual information. Subsequently, a progressive skip connection strategy is proposed to enhance the model's utilization of conditional information. Specifically, skip connections are added between the residual layers in the speech prediction period of the original diffusion model, and the conditional information is explicitly projected and input into each residual layer. Additionally, to address the problem that the mutual information between the input and output gradually decreases as the network depth increases, a set of skip connection coefficients is designed for each hidden layer during model inference to guide the weight of the conditional information input to this layer through the skip connection. This can effectively maintain the transmission and update of conditional information in each layer, avoiding the situation where conditional information gradually disappears or is weakened as the network depth increases.
[0059] Compared with the prior art, the diffusion model speech enhancement method based on progressive skip connections provided by the present invention has the following beneficial effects:
[0060] (1) The present invention inputs the noisy speech as conditional information into each residual layer in the speech prediction period through skip connections, avoiding the reduction of the priority of conditional information as the network depth increases and during the training process, thus effectively solving the problem of conditional collapse in the diffusion model speech enhancement method;
[0061] (2) The design of the skip coefficients in the training and sampling stages of the present invention not only ensures the stability of the training process but also can ensure the ability to flexibly adjust the utilization of conditional information in the speech prediction period during the sampling process without additional training, effectively improving the stability of speech enhancement and enhancing the generalization ability.
[0062] (3) Different from the existing speech enhancement and denoising methods based on diffusion models that usually directly predict noise (it is difficult to utilize the conditional information in the model input using this prediction target), the present invention uses clear speech as the prediction target during training. This design not only proves to be equivalent to predicting Gaussian noise mathematically but also enables the model to learn the features of clean speech more through the features of noisy speech; this can not only help the model understand the relationship between clear speech and noise, thereby effectively improving the accuracy of speech enhancement but also enable the model to better utilize conditional information during the denoising process, improving the consistency and quality of the generated results. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a schematic block diagram of the diffusion model speech enhancement method based on progressive skip connections provided by an embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of the residual layer structure;
[0065] Figure 3 Schematic diagram of the projection module structure
[0066] Term Explanation
[0067] DDPM is the abbreviation of Denoising Diffusion Probabiblistic Model, which represents "Denoising Diffusion Probability Model", a learning framework for generative models; for the mathematical principle derivation and implementation method, please refer to the reference [Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. In NIPS, 2020]
[0068] The standard Wiener process is an important stochastic process in probability theory, also known as Brownian motion, which has some important mathematical properties. For example, it has the Markov property and its increments follow a normal distribution Specific Embodiments
[0069] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention
[0070] Embodiment
[0071] The speech predictor used in this embodiment is as Figures 1 to 2 shown, which includes a convolutional layer a, a number of residual layers, an activation function (ReLu), and a convolutional layer b arranged in sequence. Each residual layer generates two outputs. One output is superimposed on the output of the previous residual layer and then sent as input to the next residual layer, and the other output is accumulated by skip connection to the last residual layer as the residual output. The structures of the residual layers are all the same, which are composed of a dilated convolutional layer, a gated activation function, and a convolutional layer c connected in sequence. The gated activation function includes a tanh activation function and a Sigmoid activation function arranged in parallel, and the outputs of the two activation functions are multiplied to obtain the output of this residual layer. The convolutional layer a, the convolutional layer b, and the convolutional layer c are all one-dimensional convolutional layers, and the kernel size is 1×1
[0072] For the time step t, the speech predictor takes the superposition of the noisy speech and the speech data c t at time step t as input; the speech data c tIt can be the noisy speech after adding noise to the output of the speech predictor at the previous time step. And time embedding is superimposed on the input features of the residual layer, and the conditional information after the noisy speech is adjusted by the projection module and the skip connection coefficient is superimposed on the output of the dilated convolutional layer.
[0073] The time embedding is obtained according to the following operations: First, the position encoding of the time step is performed through the sine and cosine functions, and then the time embedding is obtained through processing by a linear layer and a swish function.
[0074] The conditional information is obtained by processing the noisy speech through the projection module and adjusting the skip connection coefficient. The projection module includes a first convolutional layer, a ReLu function, a second convolutional layer, and a GLU function arranged in sequence. Both the first convolutional layer and the second convolutional layer are one-dimensional convolutions, and the kernel size is 1×1.
[0075] The skip connection coefficient is defined as:
[0076]
[0077] where k n represents the skip connection coefficient of the nth layer of the residual layer, n = 1,…,L, and L represents the number of residual layers.
[0078] The training objective of the above speech predictor is to minimize the following loss function:
[0079]
[0080] In order to maintain stable gradient flow and at the same time avoid overfitting caused by the model relying too much on skip connections, during the training stage of the speech predictor, the skip connection coefficients are all fixed to 1 to force the model to learn how to use conditional signals and hierarchical information; then in the validation sampling stage, the present invention adopts a set of empirically designed skip connection coefficients, which are obtained by summarizing through multiple groups of experiments on a small amount of data extracted from the validation set, and are determined by the current hidden layer depth, the total number of hidden layers, and the hyperparameters k L and k1.
[0081] Based on the above speech predictor, this embodiment provides a speech enhancement method for a diffusion model based on progressive skip connections, and its steps are as follows:
[0082] (S1) Obtain the time embedding according to the set time step;
[0083] (S2) Obtain the noisy speech embedding;
[0084] (S3) Use the noisy speech and randomly sampled standard Gaussian noise as inputs, and sample them using a speech predictor according to the set time steps to obtain clean speech; the speech predictor includes several sequentially arranged residual layers; project the noisy speech into the conditional information obtained by adjusting the projection module and skip connection coefficients to each residual layer.
[0085] Specifically, in the sampling stage, this embodiment adopts an accelerated sampling method, and selects two time steps t and t - r (50 ≥ T ≥ t, and r, t - r > 0) for accelerated sampling. At this time, the above step (S3) includes the following steps:
[0086] (S31) For time step t, perform noise addition processing on the noisy speech z to obtain the noisy speech z t , expressed as:
[0087]
[0088] (S32) Use the noisy speech z and the noisy speech z t as inputs, and use the speech predictor to perform the first sampling to obtain the denoised speech expressed as:
[0089]
[0090] (S33) For time step t - r, perform noise addition processing on the weighted sum of the denoised speech and the noisy speech to obtain the noisy speech c t-4 ; here, noise addition processing can be performed on the average value of the denoised speech and the noisy speech, expressed as:
[0091]
[0092] (S34) Use the noisy speech z and the noisy speech c t-4 as inputs, and use the speech predictor to perform the second sampling to obtain the clean speech expressed as:
[0093]
[0094] Next, the effectiveness of the diffusion model speech enhancement method based on progressive skip connections provided by the present invention will be verified in combination with specific data sets.
[0095] (I) Data Set
[0096] In this embodiment, the VoiceBank-DEMAND dataset is used for model training and testing. For this dataset, there are known data pairs of the speech to be enhanced and the clean speech, and the dataset has been divided into a training set and a test set. See the reference [Botinhao, Cassia Valentini, et al. Investigating RNN-based speech enhancement methods for noise-robust text-to-speech. In 9th ISCA speech synthesis workshop, 2016].
[0097] (2) Speech Predictor
[0098] In this embodiment, a deep neural network (a non-autoregressive feedforward dilated convolutional architecture neural network, DiffWave) is used as the speech predictor. For the specific implementation, see the reference [Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In ICLR, 2020].
[0099] (3) Metrics
[0100] To comprehensively evaluate the effectiveness of the speech enhancement algorithm, the present invention uses the Perceptual Evaluation of Speech Quality (PESQ), Short-Time Objective Intelligibility (STOI), the Mean Opinion Score prediction of speech signal distortion (CSIG), and the Mean Opinion Score prediction of background noise (CBAK) to evaluate the performance. The higher the score of each metric, the better the performance.
[0101] PESQ (Perceptual Evaluation of Speech Quality): This is an objective metric for evaluating speech quality, aiming to simulate the human ear's perception of speech quality. It generates a score by comparing the enhanced or synthesized speech with the original reference speech, usually ranging from -0.5 to 4.5. The higher the score, the better the speech quality.
[0102] STOI (Short-Time Objective Intelligibility): This metric is used to evaluate the intelligibility of speech, that is, the ease with which a listener can understand the speech content in a noisy environment. This metric calculates the correlation between the reference speech and the speech to be evaluated in the short-time spectrum, obtaining a score between 0 and 1. The higher the value, the higher the intelligibility. STOI is particularly suitable for noisy speech environments and is an important tool for measuring the effectiveness of speech enhancement systems in improving intelligibility.
[0103] CSIG (Prediction of Mean Opinion Score for Speech Signal Distortion): This metric is used to measure the degree of distortion between the enhanced speech and the reference clean signal, reflecting the overall quality of the signal. The scoring range of CSIG is from 1 to 5, where a higher score indicates better signal quality and lower distortion.
[0104] CBAK (Prediction of Mean Opinion Score for Background Noise): This metric is used to evaluate the background noise level in the speech signal. CBAK also uses a scoring range of 1 to 5, where a higher score indicates better suppression of background noise and clearer speech.
[0105] (IV) Explanation of the Remaining Methods in the Table
[0106] The content in parentheses after each method name represents the model prediction target of this method, where SP represents clean speech Speech (c0 mentioned in the previous text), and N represents the noise component Noise in the noisy speech (ξ mentioned in the previous text).
[0107] *: The metric reflected by the test set data without any processing.
[0108] DiffWave: A diffusion-based DiffWave speech enhancement model without any additional design. For the specific implementation, please refer to the literature
Kong, Zhifeng, et al. DiffWave: A Versatile Diffusion Model for Audio Synthesis. In ICLR
[0109] DiffWave+SC: Added skip connections to the DiffWave model, but did not adopt a progressive skip coefficient design.
[0110] SCSE: The diffusion model speech enhancement model based on progressive skip connections proposed in the present invention, which adopts progressive skip connections on the basis of the Diffwave model.
[0111] The results of the evaluation metrics obtained after evaluating all methods on the dataset are shown in Table 1. NFE (Number of Function Evaluations, that is, the number of times the noise predictor is called) is also used to evaluate the model efficiency. The optimal result for each metric is shown in bold.
[0112] Table 1: Effects of Each Method on Speech Enhancement in the VoiceBank-DEMAND Dataset
[0113]
[0114] According to the results in Table 1, first of all, the performance of all baseline models is significantly better than that of the Diffwave(N) model. This result indicates that each baseline model has significantly improved the speech enhancement effect by effectively utilizing the conditional information strategy. Compared with the Diffwave(N) model that only relies on the original input, these methods can fully play the role of additional conditional information during the enhancement process, thereby improving the quality and clarity of the speech. Secondly, the training objective of the model has changed from the original noise prediction to clean speech prediction, which has further brought a significant performance improvement. This change not only optimizes the training process of the model but also increases the mutual information between the conditional information and the prediction target, enabling the model to better utilize the noisy speech to recover the features of the original clean speech. In this way, the model can not only more accurately remove noise but also improve the quality of the enhanced speech while maintaining the naturalness of the speech. Finally, although skip connections are widely used in deep neural networks, directly adding them to the model does not bring the expected effect. The reason is that directly introducing skip connections often makes it difficult to effectively handle the differences between layers and the complex changes in conditional information. The progressive skip connection strategy of the present invention dynamically adjusts the weights of the skip connections, enabling the conditional information to be effectively utilized continuously as the network deepens, thereby effectively avoiding the phenomenon of conditional collapse. Through this design, the present invention can not only better cope with the challenges brought by the increase in network depth but also maintain the stability and accuracy of its speech enhancement effect under complex conditions.
[0115] In summary, the present invention proposes a diffusion model-based speech enhancement method with progressive skip connections, which combines the progressive skip connection strategy and the optimized prediction target, and significantly improves the speech enhancement effect of the diffusion model in complex noise environments. First of all, the present invention modifies the prediction target of the traditional diffusion model, changing from directly predicting Gaussian noise to predicting clean speech with higher mutual information. This adjustment enhances the connection between the model output and the conditional information and improves the effective utilization of the conditional information. Secondly, to solve the problem that the conditional information gradually disappears as the network depth increases, the present invention introduces a progressive skip connection strategy in the diffusion model. Between each layer of the model, the skip connection introduces the conditional information into each residual layer through a projection module, and by designing a set of skip connection coefficients, the weights of the conditional information input to each layer are adjusted, thereby effectively avoiding the phenomenon of conditional collapse and enhancing the model's dependence on the conditional information. Through this method, the model can maintain the effective utilization of the conditional information at each stage of the generation process, avoiding the problem that the conditional information is gradually weakened in the traditional diffusion model. The experimental results show that the speech enhancement method of the present invention is superior to the existing diffusion models in terms of speech quality, noise suppression, and enhancement effect, especially in complex noise environments, showing stronger robustness and higher speech enhancement efficiency. This method provides an innovative solution for efficient speech enhancement in complex noise environments.
[0116] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A voice enhancement method for diffusion models based on progressive skip connections, characterized in that, The steps are as follows: (S1) Obtain a time embedding according to a set time step; (S2) Obtain a noisy speech embedding; (S3) Using the noisy speech and randomly sampled standard Gaussian noise as inputs, sample according to a set time step using a speech predictor to obtain clean speech; the speech predictor includes a number of sequentially arranged residual layers; project the conditional information obtained by adjusting the noisy speech embedding through a projection module and a skip connection coefficient to each residual layer.
2. The method for voice enhancement of a diffusion model based on progressive skip connections according to claim 1, wherein, In step (S3), for time step t, the sampling steps of the speech predictor are as follows: 1) Add noise to the output of the speech predictor at the previous time step to obtain the noisy speech c at time step t t , and the superposition of the noisy speech c t and the noisy speech z is used as the input to the speech predictor; 2) Send the time embedding at time step t to each residual layer; 3) Project the conditional information obtained by adjusting the noisy speech embedding through a projection module and a skip connection coefficient to each residual layer; 4) The outputs of each residual layer are all added to the output of the last residual layer.
3. The method for voice enhancement of the diffusion model based on progressive skip connections according to claim 1 or 2, wherein The speech predictor is any one of a U-shaped network with a residual module, a U-shaped network with time encoding, or a deep neural network with a Transfomer architecture.
4. The method for voice enhancement of the diffusion model based on progressive skip connection according to claim 3, wherein The speech predictor is one of Diffusion Transformer, DiffWave, and Latent Diffusion.
5. The method for voice enhancement of a diffusion model based on progressive skip connections according to claim 3, wherein, The structures of each residual layer are the same, and are sequentially connected by a dilated convolutional layer, a gated activation function, and a convolutional layer; and a time embedding is superimposed on the input features of the residual layer, and conditional information is superimposed on the output of the dilated convolutional layer.
6. The method for voice enhancement of the diffusion model based on progressive skip connection according to claim 3, wherein The loss function used for training the speech predictor is: Among them, represents the expectation of the noisy speech z and the clean speech c0 at time step t, μ θ (t, z, c t ) represents the output of the speech predictor.
7. The method for voice enhancement of the diffusion model based on progressive skip connection according to claim 1 or 2, characterized in that The projection module includes a first convolutional layer, a ReLu function, a second convolutional layer, and a GLU function arranged in sequence.
8. The method for voice enhancement of the diffusion model based on progressive skip connection according to claim 1 or 2, characterized in that The skip connection coefficient is: where k n represents the skip connection coefficient of the n-th residual layer, n = 1, …, L, where L represents the number of residual layers, and k1, k L are hyperparameters.
9. The method for voice enhancement of a diffusion model based on progressive skip connections according to claim 1, wherein In the sampling stage, an accelerated sampling method is adopted, and two time steps t and t - r (T ≥ t, and r, t - r > 0) are selected for accelerated sampling; at this time, step (S3) includes the following steps: (S31) For time step t, the noisy speech z is processed by adding noise to obtain the noise-added speech z t , which is expressed as: where η(·) represents a noise scale function, and ξ represents randomly sampled standard Gaussian noise; (S32) Using the noisy speech z and the noise-added speech z t as inputs, perform the first-step sampling using a speech predictor to obtain the denoised speech which is expressed as: (S33) For time step t-r, add noise to the weighted sum of the denoised speech and the noisy speech to obtain the noisy speech c t-( ; (S34) Using the noisy voice z and the noise-added voice c t-( as inputs, perform the second sampling using a voice predictor to obtain the clean voice which is expressed as:
10. The method for voice enhancement of the diffusion model based on progressive skip connections according to claim 9, characterized in that, In step (S33), add noise to the denoised speech and the average value of the noisy speech, which is expressed as: where ξ′ represents randomly sampled standard Gaussian noise.