Speech enhancement method and device based on time-frequency envelope guided denoising diffusion process
Through the time-frequency envelope-guided denoising diffusion process, a neural network is used to estimate the time-frequency characteristics and envelope modulation noise of the speech signal, which solves the performance limitation problem of speech enhancement in complex acoustic environments and achieves more targeted speech key frequency recovery and higher enhancement effect.
Patent Information
- Application Number
- CN202510861224.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
AI Technical Summary
Existing speech enhancement methods have limited performance in complex acoustic environments, fail to fully utilize the time-frequency characteristics of speech signals, and adopt an isotropic noise addition strategy without considering the importance differences of speech signals at different time-frequency positions, resulting in a lack of targeted recovery of key frequency components of speech and limited enhancement effect.
Through a denoising diffusion process guided by the time-frequency envelope, a neural network is used to estimate the amplitude spectrum of the clean speech signal, extract the time-frequency envelope features, construct envelope-modulated noise and perform anisotropic noise modulation, adjust the diffusion path, and use the diffusion model of the encoder-decoder architecture for training and reverse sampling to reconstruct the clean speech signal.
It improves the speech feature loss problem of the traditional diffusion model in high-noise environments, improves the speech enhancement performance in complex acoustic environments, and improves the naturalness and intelligibility of speech.
Smart Images

Figure CN120708639A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of speech enhancement technology, and more particularly, to a speech enhancement method and apparatus based on a time-frequency envelope guided denoising diffusion process. Background Art
[0002] The goal of speech enhancement is to recover a pure speech signal from a noisy one. This is crucial in many applications, such as improving speech recognition accuracy in noisy environments or improving speech quality in communication systems. However, the performance of speech enhancement systems is significantly limited in complex acoustic environments, as noise often significantly obscures key speech features. Traditional speech enhancement techniques rely primarily on signal processing methods such as spectral subtraction and Wiener filtering, or statistical model-based methods such as the Kalman filter. However, these methods often make strong assumptions about the statistical properties of noise and speech, such as noise stationarity or short-term stationarity of speech. In practical applications, these assumptions often fail, especially in complex acoustic environments and under extremely low signal-to-noise ratio conditions, severely limiting their performance. In recent years, deep learning technology has made significant progress in the field of speech enhancement. Methods based on generative models such as generative adversarial networks (GANs) and variational autoencoders (VAEs) are capable of learning the complex distributional characteristics of speech. However, these methods still face challenges in training stability and generation quality. While the recently emerging diffusion model has demonstrated excellent performance in image generation, its direct application to speech enhancement suffers from high computational overhead and a lack of speech-specific prior guidance. Existing diffusion model speech enhancement methods primarily employ two strategies: directly modeling the sample points using isotropic diffusion in the time domain, and modeling the complex spectrum using isotropic diffusion in the frequency domain. However, these methods suffer from two major limitations: first, they fail to fully utilize the time-frequency characteristics of the speech signal to guide the diffusion process, resulting in a lack of specificity in recovering key speech frequency components; second, they employ an isotropic noise addition strategy, failing to consider the varying importance of the speech signal at different time-frequency locations.
[0003] Therefore, there is an urgent need for a speech enhancement method that can significantly improve the naturalness and intelligibility of the enhanced speech. Summary of the Invention
[0004] This application describes a speech enhancement method and device based on a time-frequency envelope guided denoising diffusion process, which can solve the above technical problems.
[0005] According to a first aspect, a method for speech enhancement based on a time-frequency envelope guided denoising diffusion process is provided, the method comprising:
[0006] Based on the acquired noisy speech signal, an estimate of the amplitude spectrum of the clean speech signal in the noisy speech signal is obtained through a neural network; based on the estimate of the amplitude spectrum of the clean speech signal, the time-frequency envelope of the clean speech signal is extracted; based on the time-frequency envelope characteristics of the clean speech signal, envelope-modulated noise of the clean speech signal is constructed; the envelope-modulated noise is obtained by performing various anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope characteristics of the clean speech signal; based on the envelope-modulated noise of the clean speech signal, various anisotropic adjustments are made to the diffusion path of the diffusion process, and an estimate of the complex spectrum of the clean speech signal is obtained through a diffusion model.
[0007] Based on the above further embodiment, it also includes: obtaining the time-frequency envelope, including: performing a logarithmic transformation on the estimation of the amplitude spectrum of the clean speech signal to obtain the logarithmic amplitude spectrum of the clean speech signal; obtaining the real-valued cepstral coefficients of the logarithmic amplitude spectrum based on the logarithmic amplitude spectrum; extracting the vocal tract cepstrum of the real-valued cepstral coefficients of the logarithmic amplitude spectrum through a predetermined mask; and reconstructing the frequency domain envelope of the clean speech signal based on the vocal tract cepstrum to obtain the time-frequency envelope of the clean speech signal.
[0008] Based on the above further embodiment, obtaining an estimate of the amplitude spectrum of a clean speech signal in a noisy speech signal includes: framing and windowing the noisy speech signal, taking N sampling points as a frame, and then windowing each frame; performing a Fourier transform on each frame to obtain its complex spectrum representation. Using a noise reduction discriminant neural network, using the complex spectrum of each frame as input, a rough estimate of the amplitude spectrum of the clean signal is obtained.
[0009] Based on the above further embodiment, N is any positive integer between 256 and 512; the discriminative front-stage network adopts an encoder-decoder architecture, including a convolutional coding layer, a temporal modeling layer and a decoding and reconstruction layer, inputs the complex spectrum of noisy speech, and estimates the amplitude spectrum of clean speech.
[0010] Based on the above further embodiment, N is 510.
[0011] Based on the above further embodiment, obtaining the modulation masking matrix of the time-frequency envelope includes:
[0012] Preprocess and logarithmically transform the complex spectrum of the noisy speech signal to obtain the logarithmic amplitude spectrum;
[0013] X log (k,f)=log((X m (k,f))+∈)
[0014] Where ∈ is a small constant, f is the frequency index of the short-time Fourier transform, and k is the time index of the short-time Fourier transform.
[0015] To X log Perform an inverse Fourier transform to obtain real-valued cepstral coefficients:
[0016] C(k,q)=Re[IFFT(X log (k,f 1…F ))]
[0017] Where q is the cepstrum index.
[0018] Use the low-frequency cepstrum lifting method to extract the channel envelope cepstrum C through the predefined masks M[q] and W[q] vocal [q,k].
[0019]
[0020] Among them, n co is the cepstrum truncation parameter, and N is the Fourier transform length.
[0021] Reconstruct the frequency domain envelope and perform time domain smoothing to obtain the time-frequency envelope:
[0022]
[0023] Among them: E[f,k] represents the frequency domain envelope, Re represents the operation of taking the real part of the complex number, exp is the natural base symbol, represents one-dimensional convolution smoothing, E smooth Represents the time-frequency envelope obtained after time smoothing.
[0024] Based on the above further embodiment, the cepstrum truncation parameter n co The value range is 20 to 40, and the convolution kernel length l of the convolution smoothing is 2 to 9. The convolution kernel value is 1 / l.
[0025] Based on the above further embodiment, n co The value is 29.
[0026] Based on the above further embodiment, the length l of the convolution kernel along the time domain is 3.
[0027] Based on the above further embodiment, the envelope modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope of the clean speech signal; including: generating standard white Gaussian noise with the same frequency dimension as the time-frequency envelope; and calculating a modulation masking matrix of the time-frequency envelope, where the modulation masking matrix of the time-frequency envelope is the product of a preset modulation intensity coefficient and the time-frequency envelope; and using the modulation masking matrix of the time-frequency envelope to perform anisotropic noise modulation on the standard white Gaussian noise with the same dimension as the time-frequency envelope.
[0028] Based on the above further embodiment, a standard Gaussian white noise with the same spectrum dimension is generated:
[0029]
[0030] in, represents the complex normal distribution, Represents a complex field.
[0031] Calculate the envelope modulation masking matrix, denoted as e m :
[0032] e m =E smooth ×β
[0033] Where β is the modulation intensity coefficient;
[0034] Perform anisotropic noise modulation on standard white noise:
[0035] n m =e m ·n
[0036] where · represents point-by-point multiplication.
[0037] Based on the above further embodiment, the modulation intensity coefficient ranges from 0.1 to 2.0.
[0038] Based on the above further embodiment, the modulation intensity coefficient is set to 0.3.
[0039] Based on the above further embodiment, the diffusion model is trained based on the noise samples generated by the forward diffusion process. The training process includes: constructing the marginal distribution of the diffusion state at any time t:
[0040]
[0041] Among them, the intermediate state of the diffusion process is x t , the complex spectrum of the clean speech signal is x0, and the complex spectrum of the original noisy speech signal is y, is the diffusion scheduling parameter corresponding to time step t; diag(e m) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; T is the total number of time steps in the diffusion process;
[0042] For each time step, the state x at each time step t is sampled from the marginal distribution t Together with the original noisy speech signal complex spectrum y, it is input into the diffusion model f to estimate the initial state complex spectrum x0. The mean square error between the initial state complex spectrum estimated by the diffusion model and the corresponding true complex spectrum is calculated:
[0043]
[0044] Among them, the intermediate state of the diffusion process is x t , the original noisy speech signal complex spectrum is y, the clean speech signal complex spectrum is x0, and the envelope modulation masking matrix is e m ,f(x t ,y,e m ,t) represents the predicted value of the diffusion model f at each time step t;
[0045] The gradient descent method is used to update the diffusion model parameters.
[0046] Based on the above further embodiment, reasoning the diffusion model based on the reverse diffusion process includes:
[0047] The starting point of the reverse diffusion process is to sample x from the prior distribution according to the mean and variance of the prior distribution. T ,
[0048]
[0049] Where y is the complex spectrum of the noisy speech signal, is the diffusion scheduling parameter corresponding to time step T, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix, Indicates that the random variable z follows a complex Gaussian distribution, the mean of which is the zero vector 0, and the covariance matrix is the identity matrix I;
[0050] Perform inverse sampling of envelope modulation:
[0051]
[0052] in, is the initial state distribution estimated by the diffusion neural network, Control the state evolution of the reverse process, where is the diffusion scheduling parameter corresponding to time step t, Indicates that the random variable z obeys a complex Gaussian distribution, the mean of which is the zero vector 0 and the covariance matrix is the identity matrix I. The intermediate state of the diffusion process is x t ;
[0053] The estimation of the complex spectrum of the clean speech signal is obtained by combining the initial state information estimated at each step of the noise reduction diffusion model.
[0054] Based on the above further embodiment, obtaining an estimate of the complex spectrum of the clean speech signal through a diffusion model, the method further includes:
[0055] Based on the estimation of the complex spectrum of the clean speech signal, an estimated value of the time domain speech signal of the clean speech is obtained by reconstructing through inverse Fourier transform.
[0056] Based on the above further embodiment, the noise scheduling of the diffusion process adopts the exponential rule p is any real number between 0.01 and 2.
[0057] Based on the above further embodiment, p is set to 0.3.
[0058] Based on the above further embodiment, the diffusion step number T ranges from 3 to 10.
[0059] Based on the above further embodiment, the diffusion step number T is set to 6.
[0060] Based on the above further embodiment, the first time step of the diffusion process Take any real number between 0.001 and 0.009.
[0061] Based on the above further embodiment, the first time step of the diffusion process Take 0.003.
[0062] Based on the above further embodiment, the last time step of the diffusion process Take any real number between 0.9 and 1.
[0063] Based on the above further embodiment, the last time step of the diffusion process Take 0.999.
[0064] Based on the above further embodiment, the diffusion neural network adopts an encoder-decoder architecture and has a time-conditional embedding function, which can adjust the network behavior according to the diffusion time step information to achieve time-adaptive denoising prediction. According to a second aspect, an anisotropic diffusion speech enhancement device based on time-frequency envelope modulation is provided, the device comprising: an amplitude spectrum estimation module for obtaining an estimate of the amplitude spectrum of a clean speech signal in the noisy speech signal through a neural network based on the acquired noisy speech signal;
[0065] A time-frequency envelope extraction module, configured to extract the time-frequency envelope of the clean speech signal based on an estimation of the amplitude spectrum of the clean speech signal;
[0066] an anisotropic noise modulation module, configured to construct envelope-modulated noise of the clean speech signal based on the time-frequency envelope characteristics of the clean speech signal; the envelope-modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope characteristics of the clean speech;
[0067] The envelope-guided diffusion module is used to perform anisotropic adjustments on the diffusion path of the diffusion process based on the envelope-modulated noise of the clean speech signal, and obtain the complex spectrum of the clean speech signal through a diffusion model.
[0068] Based on the above further embodiment, a diffusion model training module is further included to construct the marginal distribution of the diffusion state at any time t:
[0069]
[0070] Among them, the complex spectrum of the pure speech signal is x0, the complex spectrum of the original noisy speech signal is y, and the intermediate state of the diffusion process is x t ,, the diffusion scheduling parameter corresponding to time step t is Indicates x under x0,y conditions t The conditional probability distribution of Represents x t Follow the mean The variance is The complex normal distribution, diag(e m ) means the diagonal element is e m vector, a diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix;
[0071] For each time step, the state x at each time step t is sampled from the conditional distribution tTogether with the original noisy speech signal complex spectrum y, it is input into the diffusion model f to estimate the initial state complex spectrum x0. The mean square error between the initial state complex spectrum estimated by the diffusion model and the corresponding true complex spectrum is calculated: T is the total number of time steps in the diffusion process; the loss function formula is as follows:
[0072]
[0073] Among them, the intermediate state of the diffusion process is x t , the original noisy speech signal complex spectrum is y, the clean speech signal complex spectrum is x0, f(x t ,y,e m ,t) represents the predicted value of the diffusion model f at each time step t;
[0074] The gradient descent method is used to update the diffusion model parameters.
[0075] The reverse sampling reconstruction module is used to sample from the prior distribution as the starting point of the reverse diffusion process. The starting point of the diffusion process is: sampling x from the prior distribution according to the mean and variance of the prior distribution T ,
[0076]
[0077] Where y is the complex spectrum of the noisy speech signal, is the diffusion scheduling parameter corresponding to time step T, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix, Indicates that the random variable z follows a complex Gaussian distribution, the mean of which is the zero vector 0, and the covariance matrix is the identity matrix I;
[0078] Perform inverse sampling of envelope modulation:
[0079]
[0080] in, is the initial state distribution estimated by the diffusion neural network, Controls the state evolution of the reverse process, where α t is the diffusion scheduling parameter corresponding to time step t, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; Indicates that the random variable z obeys a complex Gaussian distribution, the mean of which is the zero vector 0 and the covariance matrix is the identity matrix I. The intermediate state of the diffusion process is x t ;
[0081] The estimation of the complex spectrum of the clean speech signal is obtained by combining the initial state information estimated at each step of the noise reduction diffusion model.
[0082] According to a third aspect, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by one or more processors, the speech enhancement method based on the time-frequency envelope guided denoising diffusion process as described in any of the above technical solutions is implemented.
[0083] According to a fourth aspect, an electronic device is provided, comprising a memory and one or more processors, wherein a computer program is stored on the memory, and when the computer program is executed by the one or more processors, a speech enhancement method based on a time-frequency envelope guided denoising diffusion process as described in any one of the above technical solutions is implemented.
[0084] The systems and methods described in the embodiments of this specification leverage the time-frequency characteristics of speech signals to guide the diffusion process, resulting in a more targeted recovery of key speech frequency components. Furthermore, an anisotropic noise addition strategy takes into account the varying importance of speech signals at different time-frequency locations. Consequently, these methods address the issues of traditional diffusion models, which suffer from loss of speech features and limited enhancement in high-noise environments. Furthermore, they improve the performance of diffusion-based speech enhancement in complex acoustic environments, providing a new and effective solution for speech signal processing in these environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0086] Figure 1 A schematic diagram illustrating the structure of a speech enhancement method based on a time-frequency envelope guided denoising diffusion process provided by an embodiment of this specification is shown;
[0087] Figure 2 A flow chart of a speech enhancement method based on a time-frequency envelope guided denoising diffusion process provided in an embodiment of this specification is shown;
[0088] Figure 3 A schematic diagram of a diffusion model training process provided by an embodiment of this specification is shown;
[0089] Figure 4 A schematic diagram of a process for performing reasoning using a diffusion model provided in an embodiment of this specification is shown;
[0090] Figure 5 A structural schematic diagram of a speech enhancement device based on a time-frequency envelope guided denoising diffusion process provided by an embodiment of this specification is shown. DETAILED DESCRIPTION
[0091] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0092] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0093] In the description of the embodiments of the present application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of the present application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0094] In the description of the embodiments of this application, the term "and / or" is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more.
[0095] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0096] Generative Adversarial Network (GAN): A deep learning model consisting of a generator and a discriminator. Through adversarial training between the two, the generator can generate realistic new data samples.
[0097] Variational Autoencoder (VAE): A deep generative model based on probabilistic modeling that combines the ideas of traditional autoencoders with Bayesian inference methods to learn the latent representations of data and generate new data samples from these latent representations.
[0098] The goal of speech enhancement is to recover a pure speech signal from a noisy one. This is crucial in many applications, such as improving speech recognition accuracy in noisy environments or improving speech quality in communication systems. However, the performance of speech enhancement systems is significantly limited in complex acoustic environments, as noise often significantly obscures key speech features. Traditional speech enhancement techniques rely primarily on signal processing methods such as spectral subtraction and Wiener filtering, or statistical model-based methods such as the Kalman filter. However, these methods often make strong assumptions about the statistical properties of noise and speech, such as noise stationarity or short-term stationarity of speech. In practical applications, these assumptions often fail, especially in complex acoustic environments and under extremely low signal-to-noise ratio conditions, severely limiting their performance. In recent years, deep learning technology has made significant progress in the field of speech enhancement. Methods based on generative models such as generative adversarial networks (GANs) and variational autoencoders (VAEs) are capable of learning the complex distributional characteristics of speech. However, these methods still face challenges in training stability and generation quality. While the recently emerging diffusion model has demonstrated excellent performance in image generation, its direct application to speech enhancement suffers from high computational overhead and a lack of speech-specific prior guidance. Existing diffusion model speech enhancement methods primarily employ two strategies: directly modeling the sample points using isotropic diffusion in the time domain, and modeling the complex spectrum using isotropic diffusion in the frequency domain. However, these methods suffer from two major limitations: first, they fail to fully utilize the time-frequency characteristics of the speech signal to guide the diffusion process, resulting in a lack of specificity in recovering key speech frequency components; second, they employ an isotropic noise addition strategy, failing to consider the varying importance of the speech signal at different time-frequency locations.
[0099] In order to solve the above problems, Figure 1 As shown, the present invention proposes a speech enhancement method based on a time-frequency envelope-guided denoising diffusion process. Specifically, the method includes: first, performing a short-time Fourier transform on a noisy speech signal in a complex acoustic environment to obtain a complex spectral representation of the noisy speech signal. Based on this information, a neural network is used to roughly estimate the amplitude spectrum of the clean speech in the noisy speech signal. For example, the neural network can be a pre-stage discriminant neural network.
[0100] For example, a noisy speech signal is framed and windowed, with N = 510 sampling points as a frame. Each frame is then windowed, using a Hanning window as the windowing function. Each frame is then Fourier transformed to obtain its complex spectrum representation, denoted as y. A noise reduction discriminant neural network is used, which takes y as input and obtains a rough estimate of the amplitude spectrum of the clean speech signal, denoted as X. m The discriminative front-end network can, for example, adopt an encoder-decoder architecture, including a convolutional coding layer, a temporal modeling layer, and a decoding and reconstruction layer. It inputs the complex spectrum of the noisy speech signal and estimates the amplitude spectrum of the corresponding clean speech signal.
[0101] Then, based on the roughly estimated amplitude spectrum of the clean speech signal, the cepstrum analysis method is used to obtain the time-frequency envelope features that represent the key clues of speech, including:
[0102] The roughly estimated amplitude spectrum is preprocessed and logarithmically transformed to obtain the logarithmic amplitude spectrum:
[0103] X log (k,f)=log((X m (k,f))+∈)
[0104] Where ∈ is a small constant, f is the frequency index of the short-time Fourier transform, and k is the time index of the short-time Fourier transform.
[0105] To X log Perform an inverse Fourier transform to obtain real-valued cepstral coefficients:
[0106] C(k,q)=Re[IFFT(X log (k,f 1…F ))]
[0107] Where q is the cepstrum index.
[0108] Use the low-frequency cepstrum lifting method to extract the channel envelope cepstrum c through the predefined masks M[q] and W[q] vocal [q,k].
[0109]
[0110] Among them, n co is the cepstrum truncation parameter, and N is the Fourier transform length. For example, n co =29.
[0111] Reconstruct the frequency domain envelope and perform time domain smoothing to obtain the time-frequency envelope:
[0112]
[0113] Among them: E[f,k] represents the frequency domain envelope, Re represents the operation of taking the real part of the complex number, exp is the natural base symbol, represents one-dimensional convolution smoothing, E smooth Represents the time-frequency envelope obtained after time smoothing. For example, in one-dimensional convolution smoothing, the convolution kernel length l along the time domain is 3; the convolution kernel value is 1 / 3.
[0114] Afterwards, the standard Gaussian white noise is modulated using the time-frequency envelope features to generate anisotropic noise with the target speech structure characteristics, including:
[0115] Generate standard white Gaussian noise with the same dimensions as the spectrum:
[0116]
[0117] in, represents the complex normal distribution, Represents a complex field.
[0118] Calculate the envelope modulation masking matrix, denoted as e m :
[0119] e m =E smooth ×β
[0120] Wherein, β is the modulation intensity coefficient; for example, β=0.3.
[0121] Perform anisotropic noise modulation on standard white noise:
[0122] n m =e m ·n
[0123] where · represents point-by-point multiplication.
[0124] After acquiring envelope-modulated anisotropic noise, the denoising diffusion process is constructed using this envelope-modulated noise. During the forward and backward diffusion processes, the diffusion path is adjusted by the envelope-modulated noise to ensure the effective preservation and enhancement of key speech cues in complex acoustic environments, including:
[0125] Construct the forward diffusion process:
[0126]
[0127] The intermediate state of the diffusion process is The complex spectrum characteristics of the original noisy speech signal obtained by short-time Fourier transform are: The complex spectrum characteristics of the pure speech signal obtained by short-time Fourier transform are Where K and F represent the time and frequency dimensions after short-time Fourier transform, respectively. Unless otherwise specified, the complex spectra mentioned in the embodiments of this application refer to the complex spectra generated by this method and have the same data structure defined by the time and frequency dimensions K and F. t |x t-1 ,x0,y) represents x t-1 ,x0,y t The conditional probability distribution of Represents x t Follow the mean x t-1 +α t (y-x0), with variance α t diag(e m ) 2 The complex normal distribution, diag(e m ) means the diagonal element is e m vector, a diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; α t Controls the state evolution of the t-th step forward process; T is the total number of time steps in the forward diffusion process;
[0128] The corresponding reverse process is:
[0129]
[0130] in, Control the state evolution of the reverse process, is the diffusion scheduling parameter at time step t. For example, when using the exponential scheduling rule, Where p is any real number between 0.01 and 2, with a typical value of 0.3; The diffusion step number T ranges from 3 to 10, with a typical value of 6. The first time step of the diffusion process Take any real number between 0.001 and 0.009, with a typical value of 0.003; the last time step Take any real number between 0.9 and 1, with a typical value of 0.999. represents the noise intensity of the inverse process at time step t.
[0131] At each time step of the forward diffusion process, the initial state complex spectrum is estimated using the diffusion model based on the state of the denoising diffusion process at that time step and the complex spectrum of the noisy speech signal. The mean square error between the estimated initial state complex spectrum and the corresponding true complex spectrum is calculated, and the diffusion model parameters are optimized using the gradient descent method to train the denoising diffusion model, including:
[0132] According to the constructed forward diffusion process, the marginal distribution of the diffusion state at any time t is obtained when the initial diffusion state x0 and the noisy speech y are known:
[0133]
[0134] Sample the state x at each time step t from the above marginal distribution t and the original noisy speech signal complex spectrum y, modulation masking matrix e m Together, input the diffusion model f and estimate the initial state x0; where, is the diffusion scheduling parameter corresponding to time step t, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; T is the total number of time steps in the diffusion process. The mean square error between the initial state complex spectrum estimated by the diffusion model and the corresponding true complex spectrum is calculated:
[0135]
[0136] Among them, the intermediate state of the diffusion process is x t , the original noisy speech signal complex spectrum is y, the clean speech signal complex spectrum is x0, e m is the modulation masking matrix, f(x t ,y,e m ,t) represents the predicted value of the diffusion model f at each time step t.
[0137] Gradient descent is used to update the diffusion model parameters. For example, the diffusion neural network uses an encoder-decoder architecture with a time-conditional embedding function, which can adjust the network behavior according to the diffusion time step information and achieve time-adaptive denoising prediction.
[0138] After the diffusion model is trained, the prior distribution is sampled as the starting point for the reverse diffusion process. By gradually performing reverse sampling and combining the initial state information estimated at each step by the denoising diffusion model, an estimate of the complex spectrum of the clean speech signal is obtained. The time domain speech signal is reconstructed through an inverse short-time Fourier transform to achieve high-quality speech enhancement, including:
[0139] Sample x from the prior distribution according to its mean and variance T ,
[0140]
[0141] Where y is the complex spectrum of the noisy speech signal, is the diffusion scheduling parameter corresponding to time step T, diag(e m ) means the diagonal element is em Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix, Indicates that the random variable z follows a complex Gaussian distribution with a mean of zero vector 0 and a covariance matrix of the identity matrix I.
[0142] Perform inverse sampling of envelope modulation:
[0143]
[0144] in, is the initial state distribution estimated by the diffusion neural network, Control the state evolution of the reverse process, where is the diffusion scheduling parameter corresponding to time step t, Indicates that the random variable z obeys a complex Gaussian distribution, the mean of which is the zero vector 0 and the covariance matrix is the identity matrix I. The intermediate state of the diffusion process is x t .
[0145] The final estimate of the complex spectrum of the clean speech signal is obtained by combining the initial state information estimated at each step of the noise reduction diffusion model.
[0146] The above-mentioned method provided in the embodiments of this specification fully utilizes the time-frequency characteristics of the speech signal to guide the diffusion process, making it more targeted in recovering the key frequency components of the speech. An anisotropic noise addition strategy is adopted, taking into account the differences in the importance of the speech signal at different time-frequency positions. Therefore, the embodiments of this specification can effectively improve the problems of traditional diffusion models in high-noise environments, such as loss of speech features and limited enhancement effects, and improve the performance of speech enhancement based on diffusion models in complex acoustic environments, providing a new and effective solution for speech signal processing in complex acoustic environments.
[0147] The following combination Figure 2 The present invention introduces a speech enhancement method based on the time-frequency envelope guided denoising diffusion process in detail. Figure 1 and Figure 2 The method is described, and specifically comprises the following steps:
[0148] Step 210: Based on the acquired noisy speech signal, obtain an estimate of the amplitude spectrum of the clean speech signal in the noisy speech signal through a neural network.
[0149] In one possible implementation, a short-time Fourier transform is performed on the noisy speech signal to obtain a complex spectrum representation of the noisy speech signal as a complex spectrum feature. Based on the complex spectrum feature of the noisy speech signal, a pre-stage discriminant neural network is used to roughly estimate the amplitude spectrum of the clean speech.
[0150] For example, the noisy speech signal is framed and windowed, with N sampling points as a frame signal. Each frame signal is then windowed using a Hanning window. Each frame signal is then Fourier transformed to obtain its complex spectrum representation, denoted as y. A noise reduction discriminant neural network is used, which takes y as input and obtains a rough estimate of the amplitude spectrum of the clean speech signal, denoted as X. m .
[0151] For example, the number of sampling points N is any positive integer between 256 and 512, and N can be 510.
[0152] For example, the discriminative front-end network adopts an encoder-decoder architecture, including a convolutional coding layer, a temporal modeling layer, and a decoding and reconstruction layer. It inputs the complex spectrum of a noisy speech signal and estimates the amplitude spectrum of the corresponding clean speech signal.
[0153] Step 220: Extract the time-frequency envelope of the clean speech signal based on the estimation of the amplitude spectrum of the clean speech signal.
[0154] In one possible implementation, based on a roughly estimated pure speech amplitude spectrum, a cepstrum analysis method is used to obtain time-frequency envelope features representing key speech clues.
[0155] For example, by Figure 1 As shown, the amplitude spectrum of the pure signal is estimated X m Perform logarithmic transformation, inverse Fourier transform, low-frequency cepstrum truncation, Fourier reconstruction, exponential operation and time domain smoothing to obtain the pure signal amplitude spectrum estimate X m The time-frequency envelope of .
[0156] For example, the low-frequency cepstrum lifting method is used to extract the channel envelope cepstrum C through the predefined masks M[q] and W[q] vocal [q,k].
[0157]
[0158] For example, the cepstrum truncation parameter n co The value range is 20 to 40, for example, the value is 29; the convolution kernel length l of the one-dimensional convolution smoothing along the time domain is 2 to 9, for example, the value is 3; the value of the convolution kernel is 1 / l.
[0159] Step 230: Construct envelope-modulated noise of the clean speech signal based on the time-frequency envelope characteristics of the clean speech signal; the envelope-modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope characteristics of the clean speech signal.
[0160] In one possible implementation, a standard Gaussian white noise with the same dimension as the complex spectrum of the original noisy speech signal is generated:
[0161]
[0162] in, represents the complex normal distribution, represents a complex number. k represents the time dimension index of the complex spectrum of the noisy speech signal, and f represents the frequency dimension index of the complex spectrum of the noisy speech signal. The shape of the standard Gaussian white noise is the same as the shape of the complex spectrum of the noisy speech signal.
[0163] Calculate the envelope modulation masking matrix, denoted as e m :
[0164] e m =E smooth ×β
[0165] Where β is the modulation intensity coefficient;
[0166] Perform anisotropic noise modulation on standard white noise:
[0167] n m =e m ·n
[0168] where · represents point-by-point multiplication.
[0169] For example, the modulation intensity coefficient β ranges from 0.1 to 2.0, preferably 0.3.
[0170] Step 240: Based on the envelope modulated noise of the clean speech signal, perform anisotropic adjustments on the diffusion path of the diffusion process to obtain a clean speech signal through a diffusion model.
[0171] Specifically, the method further includes steps 310 to 330 for training the diffusion model:
[0172] Step 310: Construct the marginal distribution of the diffusion state at any time t.
[0173] For example, the marginal distribution is:
[0174]
[0175] in, is the diffusion scheduling parameter corresponding to time step t; diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix, and T is the total number of time steps in the diffusion process.
[0176] Step 320: For each time step, sample the state x at each time step t from the marginal distribution. t .
[0177] Step 330: The state of each time step t and the original noisy speech complex spectrum y and the envelope masking matrix e are combined. m , input diffusion model f.
[0178] In one possible implementation, the diffusion model f adopts an encoder-decoder architecture with time-conditional embedding capabilities, which can adjust the network behavior according to the diffusion time step information to achieve time-adaptive denoising prediction.
[0179] Step 340 : Estimate the initial state complex spectrum x0, calculate the mean square error between the initial state complex spectrum value estimated by the diffusion model and the corresponding true complex spectrum, and update the diffusion model parameters using the gradient descent method.
[0180] In one possible implementation, the loss function is formulated as follows:
[0181]
[0182] Among them, the intermediate state of the diffusion process is x t , the original noisy speech signal complex spectrum is y, the clean speech signal complex spectrum is x0, and the envelope masking matrix is e m ,f(x t ,y,e m ,t) represents the predicted value of the diffusion model f at each time step t.
[0183] Specifically, the method further includes steps 410 to 430 for performing reasoning using the trained diffusion model:
[0184] Step 410: Sampling from the prior distribution as the starting point of the reverse diffusion process,
[0185] In one possible implementation, x is sampled from the prior distribution according to its mean and variance T ,
[0186]
[0187] Where y is the complex spectrum of the noisy speech signal, where y is the complex spectrum of the noisy speech signal, is the diffusion scheduling parameter corresponding to time step T, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix, Indicates that the random variable z follows a complex Gaussian distribution with a mean of zero vector 0 and a covariance matrix of the identity matrix I.
[0188] Step 410: Perform inverse sampling of envelope modulation:
[0189] In one possible implementation, inverse sampling of the envelope modulation is performed:
[0190]
[0191] in, is the initial state distribution estimated by the diffusion neural network, Control the state evolution of the reverse process, where is the diffusion scheduling parameter corresponding to time step t, Indicates that the random variable z obeys a complex Gaussian distribution, the mean of which is the zero vector 0 and the covariance matrix is the identity matrix I. The intermediate state of the diffusion process is x t .
[0192] Step 410: Combine the initial state information estimated by the noise reduction diffusion model at each time step to obtain a final estimate of the complex spectrum of the clean speech signal.
[0193] In the above method provided in the embodiment of this specification, the time-frequency characteristics of the speech signal are fully utilized to guide the diffusion process, which is more targeted in recovering the key frequency components of the speech; an anisotropic noise addition strategy is adopted, taking into account the importance differences of the speech signal at different time-frequency positions.
[0194] The following combination Figure 5 The present invention introduces a speech enhancement device based on a time-frequency envelope guided denoising diffusion process, which includes:
[0195] An amplitude spectrum estimation module is used to obtain an estimate of the amplitude spectrum of a clean speech signal in the noisy speech signal through a neural network based on the obtained noisy speech signal;
[0196] A time-frequency envelope extraction module, configured to extract the time-frequency envelope of the clean speech based on the estimation of the amplitude spectrum of the clean speech;
[0197] an anisotropic noise modulation module, configured to construct envelope-modulated noise of the clean speech signal based on the time-frequency envelope characteristics of the clean speech signal; the envelope-modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope characteristics of the clean speech signal;
[0198] The envelope-guided diffusion module is used to perform anisotropic adjustments on the diffusion path of the diffusion process based on the envelope-modulated noise of the clean speech signal, and obtain an estimate of the complex spectrum of the clean speech signal through a diffusion model.
[0199] Based on the above further embodiment, a diffusion model training module is further included to construct the marginal distribution of the diffusion state at any time t:
[0200]
[0201] in, is the diffusion scheduling parameter corresponding to time step t; diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; and T is the total number of time steps in the diffusion process.
[0202] For each time step, the state x at each time step t is sampled from the conditional distribution t and the original noisy speech signal complex spectrum y, envelope masking matrix e m Together, input the diffusion model f, estimate the initial state spectrum x0, and calculate the mean square error between the initial state complex spectrum estimated by the diffusion model and the corresponding true complex spectrum:
[0203]
[0204] The gradient descent method is used to update the diffusion model parameters.
[0205] The reverse sampling reconstruction module is used to sample from the prior distribution as the starting point of the reverse diffusion process. The starting point of the diffusion process is: sampling x from the prior distribution according to the mean and variance of the prior distribution T ,
[0206]
[0207] Perform inverse sampling of envelope modulation:
[0208]
[0209] in, is the initial state distribution estimated by the diffusion neural network,
[0210] The estimation of the complex spectrum of the clean speech signal is obtained by combining the initial state information estimated at each step of the noise reduction diffusion model.
[0211] The present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by one or more processors, the anisotropic diffusion speech enhancement method based on time-frequency envelope modulation as described in any of the above technical solutions is implemented.
[0212] In addition, an electronic device is provided, comprising a memory and one or more processors, wherein a computer program is stored on the memory, and when the computer program is executed by the one or more processors, the anisotropic diffusion speech enhancement method based on time-frequency envelope modulation as described in any of the above technical solutions is implemented.
[0213] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.
Claims
1. A speech enhancement method based on a time-frequency envelope guided denoising diffusion process, comprising: Based on the acquired noisy speech signal, obtaining an estimate of the amplitude spectrum of a clean speech signal in the noisy speech signal through a neural network; Extracting the time-frequency envelope of the clean speech signal based on the estimation of the amplitude spectrum of the clean speech signal; Constructing envelope-modulated noise of the clean speech signal based on the time-frequency envelope of the clean speech signal; The envelope modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope characteristics of the pure speech signal; Based on the envelope modulation noise of the clean speech signal, anisotropic adjustment is performed on the diffusion path of the diffusion process, and an estimation of the complex spectrum of the clean speech signal is obtained through a diffusion model.
2. The method according to claim 1, characterized in that The acquisition of the time-frequency envelope includes: Performing a logarithmic transformation on the estimate of the amplitude spectrum of the clean speech signal to obtain a logarithmic amplitude spectrum of the clean speech signal; Based on the logarithmic magnitude spectrum, obtaining real-valued cepstral coefficients of the logarithmic magnitude spectrum; Extracting the vocal channel cepstrum of the real-valued cepstrum coefficients of the logarithmic magnitude spectrum through a predetermined mask; Based on the vocal channel cepstrum, the frequency domain envelope of the clean speech signal is reconstructed to obtain the time-frequency envelope of the clean speech signal.
3. The method according to any one of claims 1-2, characterized in that The envelope modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope of the pure speech signal; including: Generate standard Gaussian white noise with the same dimension as the time-frequency envelope; as well as, Calculating a modulation masking matrix of the time-frequency envelope, where the modulation masking matrix of the time-frequency envelope is a product of a preset modulation intensity coefficient and the time-frequency envelope; Anisotropic noise modulation is performed on the standard Gaussian white noise having the same dimension as the time-frequency envelope using a modulation masking matrix of the time-frequency envelope.
4. The method according to claim 1, wherein Based on the envelope modulation noise of the clean speech signal, anisotropically adjusting the diffusion path of the diffusion process, and obtaining the complex spectrum of the clean speech signal through the diffusion model, including: The forward diffusion process is constructed as follows: Among them, the complex spectrum of the pure speech signal is The intermediate state of the diffusion process is The complex spectrum of the original noisy speech signal is K and F are the time and frequency dimensions after short-time Fourier transform, represents the complex field; q(x t |x t-1 ,x0,y) represents x t-1 ,x0,y t The conditional probability distribution of Represents x t Follow the mean x t-1 +α t (y-x0), with variance α t diag(e m ) 2 The complex normal distribution, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m ∈R K×F is the envelope modulation masking matrix; α t Controls the state evolution of the t-th step forward process; T is the total number of time steps in the forward diffusion process; The corresponding reverse process is constructed as follows: Among them, the intermediate state of the diffusion process is x t , the complex spectrum of the clean speech signal is x0, and the complex spectrum of the original noisy speech signal is y, Control the state evolution of the reverse process, where is the diffusion scheduling parameter corresponding to time step t, represents the noise intensity of the inverse process at time step t.
5. The method according to claim 4, characterized in that The diffusion model is trained based on the noise samples generated by the forward diffusion process. The training process includes: Construct the marginal distribution of the diffusion state at any time t: Among them, the intermediate state of the diffusion process is x t , the complex spectrum of the clean speech signal is x0, and the complex spectrum of the original noisy speech signal is y, is the diffusion scheduling parameter corresponding to time step t, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; T is the total number of time steps in the diffusion process; For each time step, the state x at each time step t is sampled from the marginal distribution t Together with the original noisy speech signal complex spectrum y, it is input into the diffusion model f to estimate the initial state complex spectrum x0. The mean square error between the initial state complex spectrum estimated by the diffusion model and the corresponding true complex spectrum is calculated: Among them, the intermediate state of the diffusion process is x t , the original noisy speech signal complex spectrum is y, the clean speech signal complex spectrum is x0, and the envelope modulation masking matrix is e m ,f(x t ,y,e m ,t) represents the predicted value of the diffusion model f at each time step t; The gradient descent method is used to update the diffusion model parameters.
6. The method according to claim 4, characterized in that The reasoning of the diffusion model based on the reverse diffusion process includes: The starting point of the reverse diffusion process is to sample x from the prior distribution according to the mean and variance of the prior distribution. T , Where y is the complex spectrum of the noisy speech signal, is the diffusion scheduling parameter corresponding to time step t, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; Indicates that the random variable z follows a complex Gaussian distribution, the mean of which is the zero vector 0, and the covariance matrix is the identity matrix I; Perform inverse sampling of envelope modulation: in, is the initial state distribution estimated by the diffusion neural network, Control the state evolution of the reverse process, where is the diffusion scheduling parameter corresponding to time step t, represents the noise intensity of the inverse process at time step t; Indicates that the random variable z obeys a complex Gaussian distribution, the mean of which is the zero vector 0 and the covariance matrix is the identity matrix I. The intermediate state of the diffusion process is x t ; The estimation of the complex spectrum of the clean speech signal is obtained by combining the initial state information estimated at each step of the noise reduction diffusion model.
7. The method according to claim 1, characterized in that Obtaining an estimate of the complex spectrum of the clean speech signal through a diffusion model, the method further comprising: Based on the estimation of the complex spectrum of the clean speech signal, an estimated value of the time domain speech signal of the clean speech is obtained by reconstructing through inverse Fourier transform.
8. An anisotropic diffusion speech enhancement device based on time-frequency envelope modulation, characterized in that: The device comprises: An amplitude spectrum estimation module is used to obtain an estimate of the amplitude spectrum of a clean speech signal in the noisy speech signal through a neural network based on the obtained noisy speech signal; A time-frequency envelope extraction module, configured to extract the time-frequency envelope of the clean speech signal based on an estimation of the amplitude spectrum of the clean speech signal; an anisotropic noise modulation module, configured to construct envelope-modulated noise of the clean speech signal based on the time-frequency envelope characteristics of the clean speech signal; the envelope-modulated noise is obtained by performing anisotropic noise modulation on Gaussian white noise based on the time-frequency envelope characteristics of the clean speech signal; The envelope-guided diffusion module is used to perform anisotropic adjustments on the diffusion path of the diffusion process based on the envelope-modulated noise of the clean speech signal, and obtain the complex spectrum of the clean speech signal through a diffusion model.
9. The method according to claim 8, characterized in that The device further comprises: Diffusion model training module, used to construct the marginal distribution of the diffusion state at any time t: Among them, the complex spectrum of the pure speech signal is x0, the complex spectrum of the original noisy speech signal is y, and the intermediate state of the diffusion process is x t , the diffusion scheduling parameter corresponding to time step t is q(x t |x0,y) means x under the condition x0,y t The conditional probability distribution of Represents x t Follow the mean The variance is The complex normal distribution, diag(e m ) means the diagonal element is e m vector, a diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix; For each time step, the state x at each time step t is sampled from the conditional distribution t Together with the original noisy speech signal complex spectrum y, it is input into the diffusion model f to estimate the initial state complex spectrum x0. The mean square error between the initial state complex spectrum estimated by the diffusion model and the corresponding true complex spectrum is calculated: T is the total number of time steps in the diffusion process; the loss function formula is as follows: Among them, the intermediate state of the diffusion process is x t , the original noisy speech signal complex spectrum is y, the clean speech signal complex spectrum is x0, f(x t ,y,e m ,t) represents the predicted value of the diffusion model f at each time step t; The gradient descent method is used to update the diffusion model parameters. The reverse sampling reconstruction module is used to sample from the prior distribution as the starting point of the reverse diffusion process. The starting point of the diffusion process is: sampling x from the prior distribution according to the mean and variance of the prior distribution T , Where y is the complex spectrum of the noisy speech signal, is the diffusion scheduling parameter corresponding to time step t, diag(e m ) means the diagonal element is e m Flattened vector, diagonal matrix with all other elements set to 0, e m is the envelope modulation masking matrix, Indicates that the random variable z follows a complex Gaussian distribution, the mean of which is the zero vector 0, and the covariance matrix is the identity matrix I; Perform inverse sampling of envelope modulation: in, is the initial state distribution estimated by the diffusion neural network, Control the state evolution of the reverse process, where is the diffusion scheduling parameter at time step t, represents the noise intensity of the inverse process at time step t; Indicates that the random variable z obeys a complex Gaussian distribution, the mean of which is the zero vector 0 and the covariance matrix is the identity matrix I. The intermediate state of the diffusion process is x t ; The estimation of the complex spectrum of the clean speech signal is obtained by combining the initial state information estimated at each step of the noise reduction diffusion model.
10. A computer storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method for speech enhancement based on the time-frequency envelope guided denoising diffusion process as claimed in any one of claims 1 to 7 is implemented.