Speech enhancement method based on distributed enhancement diffusion model

By constructing a distribution-enhanced diffusion model, designing a generalized random differential equation perturbation kernel and introducing a decoupled noise shuffling strategy, the problem of lack of theoretical support for the mean interpolation strategy in the diffusion model is solved, and efficient speech recovery and quality improvement in complex noise environments are achieved.

CN120496565AActive Publication Date: 2025-08-15CHENGDU AVEN DIGITAL INFORMATION TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510841319.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-15
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The mean interpolation strategy in the existing diffusion model lacks unified understanding and theoretical support, and is difficult to regulate flexibly, limiting its promotion and application in different noise environments.

Method used

A speech enhancement method based on a distribution-enhanced diffusion model is constructed. By designing a generalized random differential equation perturbation kernel, combining a decoupled noise shuffling strategy, a new noisy speech sample is generated, the training distribution is expanded, and the model adaptability is improved.

Benefits of technology

It significantly improves the model's voice recovery ability in complex noise environments, realizes the theoretical construction of a unified interpolation framework, and ensures the feasibility of real-time or near-real-time voice enhancement in practice, improving voice quality and intelligibility indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496565A_ABST
    Figure CN120496565A_ABST
Patent Text Reader

Abstract

The invention discloses a speech enhancement method based on a distribution enhancement diffusion model, relates to the technical field of single-channel speech enhancement, and constructs a unified interpolation framework for speech enhancement under a diffusion model. The method is characterized in that a novel generalized stochastic differential equation disturbance kernel is designed to unify existing interpolation strategies and reveal the essential effect of the disturbance kernel as a distribution enhancement mechanism. According to the method, a general interpolation formula is provided based on mean interpolation, various existing variants are covered, the role of interpolation in data enhancement is determined theoretically, and the effectiveness and stability of voice processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of single-channel speech enhancement, and more particularly to a speech enhancement method based on a distribution enhancement diffusion model. Background Art

[0002] Speech enhancement (SE) aims to restore clear speech from noisy audio. It is widely used in scenarios such as voice communication, intelligent assistants, and speech recognition preprocessing, and is particularly well-suited for complex or non-stationary noise environments. Traditional methods rely on signal processing techniques such as frequency-domain filtering and spectral subtraction, but are limited in effectiveness when mixed with multiple noises. With the development of deep learning, modern speech enhancement methods continue to optimize speech quality, real-time performance, and adaptability, driving the rapid evolution of related technologies.

[0003] In recent years, diffusion models have been introduced into speech enhancement tasks due to their powerful generative capabilities. Their basic principle is to diffuse clean speech into noisy data by gradually adding Gaussian noise, and then use a neural network to learn the inverse diffusion process to achieve denoising. To improve training stability, mean interpolation strategies are widely adopted to construct an intermediate state between clean and noisy speech. Early methods were mostly based on discrete-time diffusion models. Subsequent research has extended this to a continuous-time diffusion framework, combining it with drift term optimization of stochastic differential equations, demonstrating stronger adaptability and generalization under complex noise conditions. Despite some progress, the role of mean interpolation in diffusion models still lacks a unified understanding and theoretical support. Design strategies often rely on experience and are difficult to flexibly adjust, limiting their application in different application scenarios.

[0004] Therefore, it is necessary to propose a speech enhancement method based on the distribution enhancement diffusion model to solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to solve the problem that the role of mean interpolation in diffusion models still lacks unified understanding and theoretical support, and the design strategy mostly relies on experience and is difficult to flexibly control, which limits its promotion in different application scenarios.

[0006] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0007] A speech enhancement method based on a distributed enhancement diffusion model comprises the following steps:

[0008] S1. Construct the relationship between the noisy speech signal c, the clean speech signal a, and the noise n:

[0009] c=a+n;

[0010] S2, based on the diffusion model, gradually adds Gaussian noise to the clean speech signal a through the forward process to generate a noisy speech signal at , whose forward process is expressed as a stochastic differential equation:

[0011] da t =f(a, t)dt+g(t)dw;

[0012] Where f(a, t) is the drift coefficient, g(t) is the diffusion coefficient, and dw is the standard Wiener process;

[0013] S3. In the reverse process, the scoring function is learned through the neural network To reconstruct the clean speech signal. Furthermore, the closed solution of the forward process at any time step t is:

[0014] a t =α t a0+σ t ε;

[0015] Among them, a t is the scaling factor of the clean signal a0, β t is the noise scaling function, ε is Gaussian noise.

[0016] Furthermore, the mean interpolation strategy is adopted to perform weighted averaging of the noisy speech signal c and the clean speech signal a. The forward generalized stochastic differential equation is expressed as:

[0017]

[0018] Among them, η t The time-dependent perturbation kernel coefficient, is the disturbance term.

[0019] Furthermore, the noisy sample of the mean interpolation strategy at any time step t is expressed as:

[0020] a t =α t [η t a0+(1-η t )c]+β t ε.

[0021] Furthermore, the mean interpolation is extended from the perspective of distribution enhancement, and the noise n and Gaussian noise ε are introduced into the conversion process. The expression is:

[0022]

[0023] Among them, κ t is the coefficient that changes with time.

[0024] Furthermore, the method further includes a distribution enhancement step:

[0025] Decouple the noise n from the noisy speech signal to obtain a pure noise signal;

[0026] Randomly shuffle the noise signal to generate perturbed noise ;

[0027] The noise after shuffling The clean speech a0 is recombined to generate a new noisy speech sample.

[0028] Furthermore, the forward process of the distribution enhancement is expressed as:

[0029]

[0030] in, is the noise signal after shuffling.

[0031] Furthermore, the generalized stochastic differential equation of the reverse process is expressed as:

[0032]

[0033] in, By neural network μ θ predict.

[0034] Furthermore, the method is applicable to single-channel speech enhancement tasks, and is used to recover a clean speech signal a from a noisy speech input c.

[0035] Furthermore, the distribution enhancement method is performed in the training phase, and the inference phase uses the reverse process of the standard diffusion model to generate the speech enhancement result.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. This invention, by raising mean interpolation to a diffusion kernel with distribution enhancement level and combining it with a decoupled noise shuffling strategy, not only constructs a unified and comprehensive interpolation framework in theory, but also significantly improves the model's speech recovery ability in complex noisy environments in practice.

[0038] 2. The method of the present invention also has significant advantages in practical application. Because noise shuffling only occurs during the training phase, the inference phase still uses the standard inverse SDE generation process, without the need for additional computational overhead or complex noise operations. This means that during deployment, it can be seamlessly compatible with existing diffusion model accelerated sampling technologies, ensuring the feasibility of real-time or near-real-time speech enhancement.

[0039] 3. This invention introduces a unified interpolation coefficient into both standard deviation exploding and variance preserving diffusion kernels, constructing a generalized SDE perturbation kernel that encompasses all existing mean interpolation variants. This innovation not only preserves the model's accurate simulation of the clean speech-noise mixing mechanism but also provides a flexible and controllable parameterization path for noise reconstruction. Experimental results demonstrate that the diffusion model using this method achieves or exceeds the current state-of-the-art across a number of mainstream speech quality and intelligibility metrics, including PESQ, STOI, CSIG, CBAK, and COVL. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a principle block diagram of the diffusion model speech enhancement method based on distribution enhancement provided by an embodiment of the present invention.

[0041] Figure 2 is the speech predictor used in this embodiment. DETAILED DESCRIPTION

[0042] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0043] SDE stands for Stochastic Differential Equation, which is used to construct continuous-time diffusion generative models. Compared to discrete diffusion models, SDE provides a more flexible modeling approach. For related theory and implementation, please refer to [Song Y et al., Score-based generative modeling through stochastic differential equations, ICLR 2021].

[0044] See also Figure 1 A speech enhancement method based on a distribution enhancement diffusion model is proposed. The present invention aims to construct a unified interpolation framework for speech enhancement under the diffusion model. The core of the method is to design a new generalized stochastic differential equation (SDE) perturbation kernel to unify existing interpolation strategies and reveal its essential role as a distribution enhancement mechanism. Based on mean interpolation, the method proposes a universal interpolation formula that covers a variety of existing variants. It theoretically clarifies the role of interpolation in data enhancement and improves the effectiveness and stability of speech processing.

[0045] Building on this foundation, the present invention further introduces a distribution enhancement method with noise shuffling. By recombining the decoupled noise components to generate new clean-noisy signal pairs, this method effectively expands the training distribution and improves model performance. The proposed SDE perturbation kernel not only unifies and generalizes interpolation strategies but also enhances the model's adaptability to complex noise, driving speech enhancement technology towards greater robustness and generalization.

[0046] The goal of single-channel speech enhancement is to improve the quality of a noise-contaminated speech signal. Assume that the noisy signal c represents a mixture of clean speech a and noise n. This relationship can be expressed mathematically as follows:

[0047] c=a+n

[0048] Therefore, the goal of the speech enhancement (SE) task is to recover a clean speech signal a from a noisy speech input c. Discriminative methods directly learn the mapping from noisy speech input c to clean speech signal a, which relies on large amounts of data and has limited generalization capabilities. To improve the generalization ability of the model, diffusion-based speech enhancement models corrupt the clear speech by adding varying degrees of Gaussian noise in the forward process, and then learn a scoring function to reverse this corruption and reconstruct the clear speech. The traditional forward process can be expressed as the following stochastic differential equation (SDE):

[0049] da t =f(a, t)dt+g(t)dw,

[0050] Where f is the drift coefficient, which controls the shift of the mean value during the diffusion process; g is the diffusion coefficient, which controls the amount of Gaussian white noise injected at each time step; and w is the standard Wiener process. The reverse process, i.e., solving the stochastic differential equation, can be expressed as:

[0051]

[0052] The scoring function for the current time step is Usually a neural network is parameterized by μ θ To predict.

[0053] The special property of the forward process of the above stochastic differential equation is that it has a closed-form solution at any time step t:

[0054] a t =α t a0+σ t ε,

[0055] Where a t It represents all the intermediate data from the clean data, and also represents the speech signal that is gradually contaminated by noise during the diffusion process. t β represents the scaling factor of the clean signal a0.t represents the noise scaling function, which controls the ratio between the clean signal a0 and the Gaussian noise during the noise addition process and also represents the intensity of the noise injection. t represents the diffusion time step, which represents the stage of the current data in the entire noise addition or denoising process.

[0056] In this paper, the mean interpolation method is a special form of the above standard process. By performing a weighted average between the clean speech and the noisy signal, the model can better utilize information when processing the data. This strategy is particularly suitable for the training process of diffusion models because it helps smooth the transition and reduces the impact of noise on model performance. Specifically, the forward stochastic differential equation can be rewritten as:

[0057]

[0058] In this equation, represents the disturbance term, ηt is a time-dependent function, and represents the perturbation kernel of the SDE. According to the above equation, the noisy sample at any time step t is:

[0059] a t =α t [η t a o +(1-η t )c]+β t ε.

[0060] By comparing the above formulas, we can see that mean interpolation incorporates the noise signal into the forward SDE by interpolating a and c. We study mean interpolation from the perspective of distribution enhancement. The above formula can be expressed in a more concise form:

[0061]

[0062] where κ t is a time-varying coefficient. The form of this equation is similar to distributional augmentation, in which speech noise n, Gaussian noise ε, and time step t all participate in the transformation process. Unlike traditional data augmentation methods, which apply simple transformations to individual samples, distributional augmentation introduces a series of transformations, enabling the model to learn the underlying distribution of the transformed data under specific conditions. This approach promotes better generalization by training the model on both individual variations and their broader distribution.

[0063] Under the unified mean interpolation framework, the present invention further proposes a distribution enhancement method, which generates more diverse training samples by decoupling clean speech from its corresponding noise and shuffling and recombining them. Specifically, the noise part extracted from the original noisy speech signal (that is, the clean speech is subtracted from the noisy speech) is first separated separately to obtain a pure noise signal that is irrelevant to the speech content. Subsequently, these noise samples are randomly shuffled to break their fixed correspondence with the original speech. Then, the shuffled noise is recombined with multiple different clean speech signals to generate a new batch of noisy speech data. In this process, there is no need to rely on the original speech-noise pairing relationship, which significantly expands the coverage of training samples. Compared with traditional methods, this scheme improves the diversity of training data and the adaptability of the model to complex noise without introducing additional noise sources, and has good practicality and generalizability. The forward process of this scheme can be expressed as:

[0064]

[0065] in Corresponding to n in the previous equation, it represents the decoupled and perturbed speech noise. This solution is implemented during the training phase and remains consistent with the original solution during the inference phase. Therefore, the present invention can generate samples through a standard inference process, as described above.

[0066] The speech predictor used in this embodiment is as follows Figure 2 As shown in the figure, its input and output are designed as follows: The speech predictor is based on a multi-resolution U-Net structure, which performs well in generation and segmentation tasks. The network processes complex inputs by treating the real and imaginary parts of complex numbers as two independent channels, respectively, to adapt to the network's processing requirements for real numbers. The input layer of the network uses a 3x3 Conv2D layer with a stride of 1. Feature extraction is performed in a contraction path through residual blocks. Each residual block contains a Conv2D layer, group normalization, upsampling or downsampling of a finite impulse response (FIR) filter, and a Swish activation function. Each upsampling layer consists of three residual blocks, and each downsampling layer consists of two residual blocks connected in series. The last block performs upsampling or downsampling. A global attention mechanism is introduced in the 16x16 resolution and bottleneck layer to better learn global dependencies in the feature map.

[0067] The speech predictor network receives a complex spectrogram containing the speech signal, with the real and imaginary components as separate channel inputs. The network outputs a complex spectrogram of the generated clean speech, containing the corresponding real and imaginary components. To make the model time-dependent, the network architecture incorporates information about the progress of the current diffusion process. A common approach is to use Fourier embedding to map the scalar time coordinate to an M-dimensional vector, which is then integrated into each residual block. The network also employs a progressive enhancement input mechanism, providing a downsampled version of each feature map in the contracting path, an approach that has shown good performance in high-resolution image generation. The downsampling operation shares weights for each resolution in the progressive enhancement, and the output of the expanding path follows the same progressive enhancement approach. This design enables the network to efficiently generate high-quality speech signals while maintaining the ability to process complex spectrograms.

[0068] Based on the above-mentioned speech predictor, this embodiment provides a speech enhancement method based on a distributed enhancement diffusion model, the steps of which are as follows:

[0069] (1) Obtain time embedding according to the set time step;

[0070] (2) Obtaining noisy speech embedding;

[0071] (3) Obtaining enhanced speech through a speech predictor;

[0072] The effectiveness of the speech enhancement method based on the diffusion model of distribution enhancement provided by the present invention is verified in conjunction with a specific data set.

[0073] (1) Dataset

[0074] This example uses the VoiceBank-DEMAND dataset for model training and testing. This dataset contains known pairs of speech to be enhanced and clean speech, and is divided into training and test sets. Reference [Botinhao, Cassia Valentini, et al. Investigating RNN-based speech enhancement methods for noise-robust text-to-speech. In 9th ISCA speech synthesis workshop, 2016].

[0075] (2) Speech Predictor

[0076] This embodiment uses a deep neural network (Noise Conditioned Score Network, NCSN++) as a speech predictor, and the specific implementation reference is [J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, "Speechenhancement and dereverberation with diffusion-based generative models," TASLP, vol. 31, pp. 2351-2364, 2023].

[0077] (3) Indicators

[0078] In order to comprehensively evaluate the effectiveness of the speech enhancement algorithm, the present invention uses the perceptual evaluation of speech quality (PESQ), short-term objective intelligibility (STOI), the mean opinion score prediction of speech signal distortion (CSIG), the mean opinion score prediction of background noise (CBAK), and the mean opinion score prediction of the overall effect (COVL). For each indicator, the higher the score, the better the performance.

[0079] PESQ (Perceptual Speech Quality): An objective method for measuring speech quality that simulates the human ear's perception of audio quality and scores the difference between a reference speech and the processed speech.

[0080] ESTOI (Extended Short-Time Objective Intelligibility): Used to assess the intelligibility of speech in noisy environments. It is calculated based on the correlation between the reference and target speech in the short-time spectrum and is an important indicator for measuring speech clarity.

[0081] CSIG (Mean Opinion Score Prediction of Speech Signal Distortion): used to reflect the overall fidelity of enhanced speech, focusing on evaluating the degree of signal distortion and reflecting the naturalness and clarity of speech.

[0082] CBAK (Mean Opinion Score Prediction of Background Noise): Used to measure the degree of residual background noise in speech. The scoring result can reflect the cleanliness of the speech and the background noise suppression effect.

[0083] COVL (Mean Opinion Score Prediction of Overall Effectiveness): Provides a comprehensive assessment of perceived audio quality, taking into account the effects of speech clarity and background noise.

[0084] (IV) Explanation of the remaining methods in the table

[0085] The contents in the brackets after the names of various methods represent two commonly used variance types of SDE, where E stands for Exploding (the noise scale β mentioned above). tgrows explosively over time, while the scaling of the clean signal α t unchanged), P stands for Preserving (the noise scale β mentioned above t It grows slightly over time and is scaled with the clean signal α t Keep the sum of squares equal to 1).

[0086] *: Indicators reflected by the test set data without any processing.

[0087] DIS: Discriminative speech enhancement model. For a detailed implementation, see [J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, "Speech enhancement and dereverberation with diffusion-based generative models," TASLP, vol. 31, pp. 2351-2364, 2023]

[0088] DM(E): Diffusion-based speech enhancement model without any additional design, using SDE with exploded variance.

[0089] DM(P): Diffusion-based speech enhancement model without any additional design, using variance-preserving SDE.

[0090] MIDM(E): A diffusion-based speech enhancement model with mean interpolation and SDE with exploded variance.

[0091] MIDM(P): A diffusion-based speech enhancement model that introduces mean interpolation and uses variance-preserving SDE.

[0092] Ours: The speech enhancement diffusion model based on distribution enhancement proposed in this paper adopts a decoupled noise shuffling strategy under the mean interpolation framework.

[0093] The evaluation metrics of all methods evaluated on the VoiceBank-DEMAND dataset are shown in Table 1. The best results for each metric are shown in bold.

[0094] Table 1

[0095]

[0096] Results Analysis: Table 1 reports the speech enhancement performance of the present invention and all baselines on the VoiceBank-DEMAND dataset. As shown in Table 1, the present invention achieves the best performance on all metrics on the VoiceBank-DEMAND dataset. The performance of the present invention and MIDM(P) is better than DM(P) (without any noise enhancement), and MIDM(E) is better than DM(E). This highlights that the interpolation method is essentially a data enhancement technique using paired noise signals. Introducing unpaired noise signals further improves performance, confirming the positive impact of the distribution enhancement method proposed in this paper.

[0097] In general, the present invention systematically reflects on and improves the mean interpolation strategy in the existing diffusion model, and reinterprets it as a special distribution-enhanced diffusion kernel, thereby providing a new theoretical support and practical path for speech enhancement tasks. By unifying the interpolation method in the perturbation kernel of the stochastic differential equation (SDE), the present invention proves that mean interpolation is not just a simple interpolation weighted operation, but a distribution enhancement method that actively introduces noise sample diversity and expands the distribution of training data through distribution transformation. Based on this framework, the present invention further designs and introduces a decoupled noise shuffling mechanism - that is, clean speech and unpaired noise are recombined to generate new speech-noise pairs, which significantly expands the coverage of the distribution seen by the model during the training phase. This shuffling method breaks the limitation of the one-to-one correspondence between noise and speech in traditional practices, enabling the model to learn to remain robust in more diverse and complex noise backgrounds, thereby effectively improving the adaptability to noise scenes and speech enhancement performance.

[0098] At the implementation level, this invention introduces a unified interpolation coefficient into the standard variance exploding and variance preserving diffusion kernels, constructing a generalized SDE perturbation kernel that covers all existing mean interpolation variants. This innovation not only preserves the model's accurate simulation of the clean speech-noise mixing mechanism but also provides a flexible and controllable parameterization path for noise reconstruction. Experimental results show that the diffusion model using the method of this invention reaches or exceeds the current state-of-the-art in multiple mainstream speech quality and intelligibility indicators, such as PESQ, STOI, CSIG, CBAK, and COVL.

[0099] In addition, the method of the present invention also has significant advantages in practical application. Since noise shuffling only occurs in the training phase, the inference phase still follows the standard inverse SDE generation process, without the need for additional computational overhead or complex noise operations. This means that it can be seamlessly compatible with existing diffusion model acceleration sampling technology during deployment, ensuring the feasibility of real-time or near-real-time speech enhancement. In summary, the present invention not only constructs a unified and comprehensive interpolation framework in theory by raising the mean interpolation to a diffusion kernel at the distribution enhancement level and combining it with a decoupled noise shuffling strategy, but also significantly improves the model's speech recovery ability in complex noise environments in practice, providing practical and feasible innovative solutions for various practical application scenarios, including remote communications, hearing aids, and intelligent human-computer interaction systems.

[0100] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. The scope of patent protection of the present invention shall be based on the claims. Any equivalent structural changes made using the contents of the description of the present invention shall also be included in the scope of protection of the present invention.

Claims

1. A speech enhancement method based on a distributed enhancement diffusion model, characterized by: The following steps are involved: S1. Construct the relationship between the noisy speech signal c, the clean speech signal a, and the noise n: C=a+nσ S2, based on the diffusion model, gradually adds Gaussian noise to the clean speech signal a through the forward process to generate a noisy speech signal a t , whose forward process is expressed as a generalized stochastic differential equation: in t =f(a,t)dt+g(t)dw; Where f(a, t) is the drift coefficient, g(t) is the diffusion coefficient, and dw is the standard Wiener process; S3. In the reverse process, the scoring function is learned through the neural network To reconstruct the clean speech signal.

2. The speech enhancement method based on the distributed enhancement diffusion model according to claim 1, characterized in that: The closed-form solution of the forward process at any time step t is: a t =a t a0+s t e; Among them, a t is the scaling factor of the clean signal a0, β t is the noise scaling function, ε is Gaussian noise.

3. The speech enhancement method based on the distributed enhancement diffusion model according to claim 2, characterized in that: The mean interpolation strategy is used to perform weighted averaging of the noisy speech signal c and the clean speech signal a. The forward generalized stochastic differential equation is expressed as: Among them, η t The time-dependent perturbation kernel coefficient, is the disturbance term.

4. The speech enhancement method based on the distributed enhancement diffusion model according to claim 1, characterized in that: The noisy sample of the mean interpolation strategy at any time step t is expressed as: a t =a t [or t a0+(1-η t )c]+β t Yes.

5. The speech enhancement method based on the distributed enhancement diffusion model according to any one of claims 1 to 4, characterized in that: From the perspective of distribution enhancement, mean interpolation is extended, and noise n and Gaussian noise ε are introduced into the conversion process. The expression is: Among them, κ t is the coefficient that changes with time.

6. The speech enhancement method based on the distributed enhancement diffusion model according to claim 5, characterized in that: Further including the distribution enhancement step: Decouple the noise n from the noisy speech signal to obtain a pure noise signal; Randomly shuffle the noise signal to generate perturbed noise The noise after shuffling The clean speech signal a0 is recombined to generate a new noisy speech sample.

7. The speech enhancement method based on the distributed enhancement diffusion model according to claim 6, characterized in that: The forward process of distribution enhancement is expressed as: in, is the noise signal after shuffling.

8. The speech enhancement method based on the distributed enhancement diffusion model according to claim 7, characterized in that: The generalized stochastic differential equation of the reverse process is expressed as: in, Predicted by the neural network μθ.

9. The speech enhancement method based on the distributed enhancement diffusion model according to claim 1, characterized in that: The method is applicable to single-channel speech enhancement tasks and is used to recover a clean speech signal a from a noisy speech input c.

10. The speech enhancement method based on the distributed enhancement diffusion model according to claim 1, characterized in that: The distribution enhancement method is performed in the training phase, and the inference phase uses the reverse process of the standard diffusion model to generate speech enhancement results.

Citation Information

Patent Citations

  • Speech synthesis model training method, electronic equipment and storage medium

    CN115762464A

  • Voice conversion method and related equipment

    CN117995209A

  • Industrial noise scene speech enhancement method and system based on rapid sampling denoising diffusion model

    CN119339733A

  • Schrodinger bridge-based diffusion model speech enhancement method and system

    CN120032650A

  • Method and apparatus for improved estimation of non-stationary noise for speech enhancement

    EP1760696A2