Speech enhancement method and system based on conditional stream matching and vocoder

By introducing a speech enhancement method of conditional stream matching and vocoder in the Meer spectrum domain, an end-to-end system is built, which solves the problems of noise robustness, speech distortion and computing efficiency of traditional speech enhancement methods, and achieves efficient and natural speech enhancement effects.

CN120526784APending Publication Date: 2025-08-22HANGZHOU DIANZI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510536941.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Traditional speech enhancement methods have shortcomings in noise robustness, speech distortion, computing efficiency and context information utilization, and the model structure makes it difficult to balance the relationship between processing speed and generation quality.

Method used

Using a speech enhancement method based on conditional stream matching and vocoder, a conditional stream matching method is introduced through the Meer spectrum domain to build an end-to-end speech enhancement system, combining the conditional stream matching noise reduction module and the neural network vocoder module to realize the complete processing flow from input noise speech to output high-quality speech signals.

Benefits of technology

It significantly improves the naturalness and clarity of speech, reduces the complexity of model calculations, and improves the efficiency and quality of speech enhancement, adapting to complex and diverse practical noise scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526784A_ABST
    Figure CN120526784A_ABST
Patent Text Reader

Abstract

The invention discloses a speech enhancement method and system based on conditional stream matching and a vocoder, and the method comprises the following steps: S1, constructing a Mel spectrum extraction module which is used for converting an input noisy speech into a noisy Mel spectrum; s2, constructing a condition flow matching noise reduction module which is used for processing the noisy Mel spectrum obtained in the step S1 and outputting an enhanced Mel spectrum; and S3, constructing a neural network vocoder module which is used for restoring the enhanced Mel spectrum obtained in the step S2 into a time domain voice waveform so as to obtain an enhanced voice signal. The speech enhancement method combining conditional stream matching and the vocoder is proposed for the first time, the conditional stream matching method is innovatively introduced into the Mel-frequency spectrum domain, an end-to-end speech enhancement system is constructed, and a complete processing flow from noise speech input to high-quality speech signal output is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech enhancement, and in particular to a speech enhancement method and system based on conditional stream matching and vocoder. Background Art

[0002] In daily life, voice is one of the most frequently encountered and used information. Due to noisy environments and various factors affecting network transmission, voice signals inevitably contain noise, such as car horns on the road or electrical noise caused by network signal issues during transmission. With the rapid development of intelligent networks and information transmission, and the increasing penetration of AI technology into our daily lives, the accuracy of voice information transmission is becoming increasingly important to our daily lives, and the optimization and improvement of voice enhancement models are becoming increasingly important.

[0003] Speech enhancement technology aims to improve speech quality through various algorithms. Noise reduction, the most important area of ​​speech enhancement, is used to improve speech damaged by noise. Research can be divided into single-channel and multi-channel areas, depending on the number of channels in the speech signal. Single-channel speech enhancement, as the foundation for other research, is particularly important.

[0004] Traditional speech enhancement architectures are primarily based on the application of Fourier transforms. A speech waveform in the time domain undergoes a Fourier transform to generate its amplitude and phase spectra in the time-frequency domain. After processing, an inverse Fourier transform is then used to restore the enhanced speech waveform. With the continuous advancement of digital signal processing theory, it has become an effective addition to traditional speech enhancement models, spawning numerous classic algorithms, such as subspace speech enhancement, Wiener filtering, and spectral subtraction. Most of these algorithms utilize Fourier transforms to process the transformed spectra of noise and non-noise components. For example, traditional spectral subtraction can separate the noise spectral components, while Wiener filtering applies a filter in the time domain to filter out noise.

[0005] Traditional speech enhancement methods (such as spectral subtraction and statistical models) as well as early deep learning approaches (such as mask learning and time-frequency mapping) face a fundamental limitation: they are essentially "corrective" approaches, making local adjustments to noisy speech rather than regenerating high-quality speech from a global distribution perspective. Furthermore, they require strictly aligned noisy and clean speech pairs for training. This makes most traditional speech enhancement methods insufficiently robust to noise, and they are optimized only for specific noises, resulting in significant performance degradation when faced with complex and diverse real-world noise. Furthermore, some methods over-process the speech signal for noise reduction, resulting in speech distortion that affects intelligibility and quality. While some complex algorithms are effective, they are computationally intensive, making them difficult to process in real time on resource-constrained devices. Furthermore, traditional methods often ignore the context of the speech signal and fail to fully understand speech semantics and prosody, which compromises the effectiveness of the enhancement.

[0006] At present, the existing technology still has at least the following technical defects:

[0007] On the one hand, traditional speech enhancement methods have poor noise robustness. Most are optimized for specific noise types, such as additive white Gaussian noise. However, real-world noise is complex and diverse, such as street traffic noise or the clamor of people in restaurants. When encountering noise types for which they have not been trained, their performance degrades significantly. Furthermore, many traditional methods focus primarily on the current speech frame, ignoring the contextual information of the speech signal. This prevents these methods from fully understanding the semantics and rhythm of speech when dealing with complex speech scenarios, compromising the enhancement effect.

[0008] On the other hand, traditional processing methods require making assumptions about the distribution of the audio signal. If these assumptions differ significantly from the actual situation, the enhanced speech will differ significantly from the pure speech, potentially leading to distortion and making it difficult to adapt to today's communication requirements. This has led to the development of a series of speech enhancement models that leverage the advantages of various deep learning networks. However, due to the complex relationships between various speech attributes, the design of neural network models struggles to balance quality and speed—that is, the quality of speech generation and the speed of model training and speech processing.

[0009] Therefore, in view of the defects of the existing technology, it is necessary to propose a technical solution to solve the technical problems existing in the existing technology. Summary of the Invention

[0010] Given this, a speech enhancement method and system based on conditional stream matching and a vocoder are essential. This paper proposes a speech enhancement method combining conditional stream matching and a vocoder for the first time, innovatively introducing conditional stream matching in the mel-spectrogram domain. This system constructs an end-to-end speech enhancement process, implementing a complete flow from noisy speech input to high-quality speech output. Compared with existing diffusion models and other speech enhancement methods, this method improves the naturalness and clarity of speech, demonstrating high practical value.

[0011] In order to solve the technical problems existing in the prior art, the technical solutions of the present invention are as follows:

[0012] A speech enhancement method based on conditional stream matching and vocoder comprises the following steps:

[0013] A speech enhancement method based on conditional stream matching and vocoder comprises the following steps:

[0014] Step S1: construct a Mel spectrum extraction module to convert the input noisy speech into a noisy Mel spectrum;

[0015] Step S2: Construct a conditional stream matching denoising module to process the noisy Mel spectrum obtained in step S1 and output an enhanced Mel spectrum;

[0016] Step S3: constructing a neural network vocoder module for restoring the enhanced Mel spectrum obtained in step S2 into a time-domain speech waveform, thereby obtaining an enhanced speech signal;

[0017] Among them, the conditional stream matching denoising module and the neural network vocoder module are obtained through pre-training.

[0018] As a further improvement, in step S2, the conditional flow matching denoising module is implemented using a conditional flow matching method based on a continuous-time diffusion model. By training a neural network to estimate the path direction vector field, an enhanced Mel-spectrogram can be iteratively generated from Gaussian noise during the inference phase. The method includes the following steps:

[0019] Step S21: Pre-training a conditional stream matching noise reduction module, wherein the path for recovering clean speech from a noisy state is learned by simulating a diffusion process and direction prediction;

[0020] Step S22: Perform inference using the conditional stream matching denoising module trained in step S21, wherein starting from random noise, the clean speech spectrum is gradually restored under the guidance of the noisy Mel spectrum condition.

[0021] As a further improvement, step S21 includes the following steps:

[0022] Obtain training data, including obtaining noisy speech and corresponding clean speech, and extracting their mel spectrograms as training targets;

[0023] Feature preparation, calculate the clean Mel spectrum x, sample noise z from the standard Gaussian distribution, randomly sample the diffusion time step t, and the diffusion time step t∈(0,1) adopts a uniform random distribution;

[0024] Intermediate state construction: construct the intermediate state y and target direction vector u in diffusion based on x, z, and t;

[0025] Time step embedding, the diffusion time step t is input into the sinusoidal position encoder and the multi-layer perceptron layer to generate a time embedding vector;

[0026] Deep decoder Decoder forward propagation, input y, conditional feature μ, time step t into the Decoder, predict the direction vector

[0027] Calculate the loss, based on the difference between the predicted direction and the target direction, calculate the conditional flow matching loss function;

[0028] Backpropagation and optimization, backpropagation gradients, using optimizers to update decoder parameters, clipping excessive gradients to prevent unstable training;

[0029] Learning rate scheduling, which uses a strategy of automatically reducing the learning rate based on performance stagnation, and dynamically adjusts the optimizer learning rate when the performance of the validation set stagnates;

[0030] Logging and evaluation: record loss values ​​and training status, and regularly evaluate performance indicators on the validation set;

[0031] Repeat the above process until the training is completed.

[0032] As a further improvement, step S22 includes the following steps:

[0033] S221: Generate a Gaussian random noise spectrum z with the same structure as the Mel spectrum μ of the noisy speech as the starting point of the path;

[0034] S222: Divide the time steps evenly into the sequence t0, t1, ..., t N , as discrete time points on the path, used for state update;

[0035] S223: At each time step t i , the current state y i , the noisy speech Mel spectrum μ and the time vector are fed into the decoder to obtain the predicted direction vector of the current position:

[0036]

[0037] in, Represents the path direction prediction vector of step i; y i is the current intermediate spectrum state; μ is the corresponding noisy Mel spectrum, which serves as the generation guide condition; t i is the current time step;

[0038] S224: Update the state using Euler method:

[0039]

[0040] Among them, y i+1 is the next state spectrum; Δt is the interval between adjacent time steps; The direction prediction value of the current position;

[0041] S225: After all time steps are updated, the last frame y N This is the enhanced speech spectrum.

[0042] As a further improvement, in step S3, a neural network vocoder module is constructed and pre-trained based on the HiFi-GAN structure, so as to perform inference using the trained neural network vocoder module.

[0043] As a further improvement, step S3 includes the following steps:

[0044] Step S31: performing feature transformation on the input enhanced Mel spectrum to obtain a high-dimensional initial feature map;

[0045] Step S32: upsampling the feature map layer by layer to restore it to a time length that matches the time-domain speech, wherein the frame-level spectrum is mapped to the time-domain length of the sampling level through the layer-by-layer upsampling operation;

[0046] Step S33: using a multi-scale residual structure to model speech details in each upsampling stage;

[0047] Step S34: Map the processed feature map into a time-domain speech waveform.

[0048] As a further improvement, in step S1 , the mel spectrum extraction module is implemented by using a spectrum conversion method based on Fourier analysis and a mel filter bank.

[0049] The present invention also discloses a speech enhancement system based on conditional stream matching and vocoder, comprising:

[0050] Mel spectrum extraction module, used to convert the input noisy speech into a noisy Mel spectrum;

[0051] A conditional stream matching denoising module is used to process the noisy Mel spectrum obtained in step S1 and output an enhanced Mel spectrum;

[0052] A neural network vocoder module is used to restore the enhanced Mel-spectrogram obtained in step S2 to a time-domain speech waveform, thereby obtaining an enhanced speech signal;

[0053] Among them, the conditional stream matching denoising module and the neural network vocoder module are obtained through pre-training.

[0054] As a further improvement, the conditional stream matching denoising module is implemented using a conditional stream matching method based on a continuous-time diffusion model. By training a neural network to estimate the path direction vector field, it is possible to iteratively generate an enhanced Mel-spectrogram from Gaussian noise during the inference phase.

[0055] The conditional stream matching denoising module is pre-trained, where it learns the path to recover clean speech from a noisy state by simulating the diffusion process and directional prediction. The trained conditional stream matching denoising module is used for inference, where it starts from random noise and gradually recovers the clean speech spectrum under the guidance of the noisy Mel spectrum condition.

[0056] As a further improvement, the training process of the conditional stream matching denoising module includes the following steps:

[0057] Obtain training data, including obtaining noisy speech and corresponding clean speech, and extracting their mel spectrograms as training targets;

[0058] Feature preparation, calculate the clean Mel spectrum x, sample noise z from the standard Gaussian distribution, randomly sample the diffusion time step t, and the diffusion time step t∈(0,1) adopts a uniform random distribution;

[0059] Intermediate state construction: construct the intermediate state y and target direction vector u in diffusion based on x, z, and t;

[0060] Time step embedding, the diffusion time step t is input into the sinusoidal position encoder and the multi-layer perceptron layer to generate a time embedding vector;

[0061] Deep decoder Decoder forward propagation, input y, conditional feature μ, time step t into the Decoder, predict the direction vector

[0062] Calculate the loss, based on the difference between the predicted direction and the target direction, calculate the conditional flow matching loss function;

[0063] Backpropagation and optimization, backpropagation gradients, using optimizers to update decoder parameters, clipping excessive gradients to prevent unstable training;

[0064] Learning rate scheduling, which uses a strategy of automatically reducing the learning rate based on performance stagnation, and dynamically adjusts the optimizer learning rate when the performance of the validation set stagnates;

[0065] Logging and evaluation: record loss values ​​and training status, and regularly evaluate performance indicators on the validation set;

[0066] Repeat the above process until the training is completed.

[0067] Compared with the prior art, the technical solution of the present invention has the following technical effects:

[0068] 1. This paper introduces a conditional stream matching framework into the main structure of speech enhancement. By modeling the conditional vector field, it improves the enhancement effect while significantly reducing the model's computational complexity. It also innovatively introduces a conditional stream matching method in the Mel-spectrogram domain, achieving more efficient speech enhancement processing.

[0069] 2 The present invention uses a decoder structure combined with time embedding to accurately model the conditional direction estimation in the diffusion path, thereby improving the intelligibility and naturalness of the enhanced speech.

[0070] 3. The present invention integrates an efficient neural network vocoder to directly restore the enhanced Mel spectrum to high-quality time-domain speech, improving the overall speech restoration quality and inference efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 The present invention is a flowchart of a speech enhancement method based on conditional stream matching and vocoder.

[0072] Figure 2 Schematic diagram of the algorithm framework of the speech enhancement method based on conditional stream matching and vocoder in the present invention.

[0073] Figure 3 This is a flowchart of the training phase of the conditional stream matching denoising module in the present invention.

[0074] Figure 4 This is a basic structural diagram of the conditional stream matching noise reduction module decoder in the present invention.

[0075] Figure 5 This is a flowchart of the inference phase of the conditional stream matching denoising module in the present invention.

[0076] Figure 6 Schematic diagram of the algorithm structure of the neural network vocoder module in the present invention.

[0077] Figure 7 Schematic diagram of the structure of the multi-scale residual block in the neural network vocoder module of the present invention.

[0078] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0079] The technical solution provided by the present invention will be further described below with reference to the accompanying drawings.

[0080] During their research, the applicant discovered that the existing technology has at least the following problems: first, traditional speech enhancement methods have problems with noise robustness, speech distortion, computational efficiency, and contextual information utilization; second, there is a problem with the balance between processing speed and generation quality due to different model structures.

[0081] In addition, among existing technical solutions, flow matching, as an emerging generative modeling method, has demonstrated good performance in many fields. Representative studies closely related to the present invention include: the FlowSE model [1] and the Flow Matching model [2] for speech enhancement, and the Matcha-TTS model [3] for speech synthesis.

[0082] FlowSE model: This model introduces the flow matching method into the speech enhancement task and constructs a probabilistic path based on noisy speech conditions. The core idea is to regard noisy speech as the starting point and clear speech as the end point, and establish a smooth path from noise to clean speech along the conditional vector field. The model models the dynamic evolution of the path through ordinary differential equations (ODEs) and uses the conditional mean and variance of the optimal transmission form (linear transition from noise to clarity) to generate intermediate samples. Conditional Flow Matching (CFM) loss is used for supervised learning during the training phase. Experiments show that the model can achieve the speech enhancement effect of 60 evaluations required by the traditional diffusion model with only 5 function evaluations (NFE=5).

[0083] Flow Matching Model: This model further expands the application of flow matching methods in speech enhancement and makes four key improvements to the original conditional flow matching architecture: the introduction of prior information (directly incorporating noisy speech into the probabilistic path construction), data prediction loss (directly predicting the target speech instead of vector field regression), deterministic reasoning (omitting the sampling process to achieve fixed output), and an early stopping mechanism (reducing the number of inference steps to avoid over-processing). These improvements, applied to both the training and inference stages, enable the model to not only achieve high-quality denoising within five steps but also achieve comparable performance to traditional multi-step models with a single inference step, achieving both efficiency and effectiveness.

[0084] Matcha-TTS model: This model is applied to speech synthesis tasks. It adopts a non-autoregressive structure and builds an acoustic modeling framework with continuous normalized flow (CNF) as the core. The model is trained using the optimal transmission condition flow matching (OT-CFM) method to construct a time-evolving vector field with a simple structure and stable path. Its decoder uses the U-Net architecture that combines 1D convolution and Transformer, with parallel computing capabilities and low memory overhead. Matcha-TTS can achieve high-fidelity speech synthesis with a relatively small number of function evaluation steps. It has the advantages of fast synthesis speed, high naturalness score, and no need for external alignment. It is one of the best performing stream matching TTS models currently available.

[0085] To address the technical problems existing in the prior art, the present invention proposes a speech enhancement method and system based on conditional stream matching and a vocoder. This is a speech enhancement algorithm based on the acoustic representation of speech signals. To address the first technical problem mentioned above, the present invention employs stream matching technology to transform the traditional "corrective" approach to speech enhancement into a "generative" approach, regenerating high-quality speech from a holistic distributional perspective. By learning the probability distribution transformation between noise and clean speech, the system adapts to various noise types and effectively enhances speech even when encountering untrained noise. Its efficient likelihood calculation capability enables faster training and inference, meeting real-time processing requirements. Combined with sequence modeling techniques, stream matching can also leverage speech context to accurately enhance speech in complex scenarios, significantly improving the enhancement effect – the core objective of the present invention. To address the second technical problem, the present invention employs Mel-spectrograms as an acoustic feature representation, reducing data dimensionality and computational complexity while preserving key speech information. This system also incorporates a high-fidelity adversarial neural network (HiFi-GAN) as a vocoder to achieve fast, high-quality time-domain speech reconstruction. This structure takes into account both the generation quality and inference efficiency of speech enhancement, solving the problem of achieving both speed and performance in traditional models.

[0086] See also Figure 1 , which is a flowchart of a speech enhancement method based on conditional stream matching and vocoder provided by the present invention, see Figure 2 , which is a basic framework diagram of the speech enhancement algorithm based on conditional stream matching and vocoder of the present invention, comprises the following steps:

[0087] Step S1: construct a Mel spectrum extraction module to convert the input noisy speech into a noisy Mel spectrum;

[0088] Step S2: Construct a conditional stream matching denoising module to process the noisy Mel spectrum obtained in step S1 and output an enhanced Mel spectrum;

[0089] Step S3: constructing a neural network vocoder module for restoring the enhanced Mel spectrum obtained in step S2 into a time-domain speech waveform, thereby obtaining an enhanced speech signal;

[0090] Among them, the conditional stream matching denoising module and the neural network vocoder module are obtained through pre-training.

[0091] In step S1, the Mel spectrum extraction module is implemented using a spectrum conversion method based on Fourier analysis and Mel filter bank. The specific steps are as follows:

[0092] Step S11: Calculate the Mel filter and window function according to the sampling rate, FFT window size, number of Mel bands and frequency range of the input signal to convert the time domain signal into a Mel spectrum.

[0093] Step S12: Convert the time domain signal into frequency domain representation through short-time Fourier transform to obtain spectrum information. The formula is:

[0094]

[0095] Where X(f,t) is the value of the spectrum at frequency f and time t, x[n] is the nth sample of the time domain signal, w[tn] is the window function, N is the FFT window size, and f is the frequency.

[0096] Step S13: Calculate the amplitude of the spectrum by taking the square root of the sum of the real and imaginary parts of the complex spectrum to obtain the amplitude spectrum of each frequency point. The formula is:

[0097]

[0098] Where Re(X(f)) is the real part of the spectrum, Im(X(f)) is the imaginary part of the spectrum, and ∈ is a minimum value used to avoid numerical overflow.

[0099] Step S14: The amplitude information of the spectrum is converted into a Mel spectrum through a Mel filter. The Mel filter maps the spectrum to the Mel scale, reflecting the energy distribution of the signal in different Mel frequency bands. The formula is:

[0100]

[0101] Among them, M(f m ) is the amplitude of the mth frequency band in the Mel spectrum, X(f) is the amplitude of the STFT spectrum, H(f m ,f) is the response of the Mel filter at frequency f.

[0102] Step S15: All time frames M(f m) value set to obtain a complete Mel spectrum matrix, which is used as the feature of the speech signal for subsequent processing or model input.

[0103] Among them, step S11 can be further described as follows:

[0104] S111: Convert the linear frequency axis to the Mel frequency axis using the frequency conversion formula:

[0105]

[0106] Where m is the frequency on the Mel scale and f is the linear frequency.

[0107] S112: Divide the frequency range into a number of Mel frequency bands, and calculate the interval between the frequency bands according to the number of Mel frequency bands.

[0108] S113: Design a Mel filter. Design a filter for each Mel frequency band according to the position and width of the Mel frequency band.

[0109] S114: Calculate the Hanning window. This window function is often used to reduce spectrum leakage in STFT calculation. The formula of the window function is:

[0110]

[0111] Where w[n] is the nth value of the window function, and N is the length of the window function.

[0112] In step S2, the Conditional Flow Matching (CFM) denoising module is implemented based on the Conditional Flow Matching (CFM) method of the continuous-time diffusion model. This module simulates the process of gradually recovering the clean Mel-spectrogram of the noisy speech, starting from a random noise state. This module estimates the path direction vector field by training a neural network, allowing it to iteratively generate the enhanced Mel-spectrogram from Gaussian noise during the inference phase. The specific steps are as follows:

[0113] Step S21: In the training phase, a conditional stream matching noise reduction module is pre-trained, wherein a path for recovering clean speech from a noisy state is learned by simulating a diffusion process and direction prediction;

[0114] Step S22: In the inference phase, the conditional stream matching denoising module trained in step S21 is used for inference, wherein starting from the random noise, the clean speech spectrum is gradually restored under the guidance of the noisy Mel spectrum condition.

[0115] As a further improvement, the specific training process of step S21 is shown in the following table:

[0116] Table 1 Conditional Flow Matching (CFM) denoising module training flow chart

[0117]

[0118]

[0119] See also Figure 3 , which is a flowchart of the training process of step S21, step S21 can be further described as follows:

[0120] S211: Construct a diffusion state. The noisy speech mel spectrum obtained in step S1 is used as the target spectrum x. Then, a random sample of a time step t∈(0,1) is taken, and a random noise vector z with the same structure as x is sampled from a standard normal distribution. An intermediate state y in the diffusion process is constructed according to the following formula:

[0121] y=(1-(1-σ min )·t)·z+t·x

[0122] Where y represents the current diffusion state, that is, the intermediate Mel spectrum state of the simulated speech at time t; x represents the Mel spectrum corresponding to the clean speech, which is used as the training target; z is the noise spectrum sampled from the standard normal distribution, with the same structure as x; t∈(0,1) is the diffusion time step randomly sampled on the time axis, σ min is the minimum diffuse noise proportional coefficient, which controls the noise intensity when t is close to 0.

[0123] S212: The intermediate state mel-spectrogram y constructed in step S211 is concatenated with its corresponding noisy speech mel-spectrogram (denoted as conditional feature μ) along the channel direction. This serves as the main input to the decoder, modeling the difference between the current state and the target spectrum. Simultaneously, the time position t of the diffusion process is encoded as a time vector using a specific function, indicating the network's current stage in the recovery path.

[0124] S213: Use the decoder structure to perform feature modeling and path direction prediction on the input. The decoder uses a conditional modeling network based on a one-dimensional U-Net structure and consists of three main stages: downsampling stage (Down Blocks), intermediate processing stage (Mid Blocks), and upsampling stage (Up Blocks). The input includes the intermediate diffusion state y, the conditional feature μ (i.e., the noisy speech mel spectrum), and the time step t during the diffusion process. It is used to extract multi-level context information and output the path direction prediction vector at each time position.

[0125] S213 can be divided into three stages:

[0126] S2131: Multiple downsampling modules are used to gradually shorten the input Mel-spectrogram sequence in the temporal dimension, enabling the model to holistically understand speech information over a longer timeframe. Each downsampling module includes a residual block (ResNetBlock), which extracts new local variation features while preserving the original features and dynamically adjusts them in conjunction with the diffusion time step; a multi-head attention module (Transformer), which identifies correlations between different moments in the speech sequence; and a downsampling operation (Downsample), which reduces feature length, thereby reducing computational cost and improving modeling efficiency.

[0127] S2132: Mid Blocks: Deep modeling is continued on the compressed feature map, using multi-layer residual blocks (ResNetBlock) and attention modules (Transformer) to further understand the semantic and structural changes of speech. The speech state at different stages is adapted based on the current diffusion time step, helping the network to more accurately determine how the current speech segment should be restored to a clean spectrum.

[0128] S2133: Upsampling phase (Up Blocks). The feature map is gradually restored to the same time length as the original input. Each layer first merges the current feature with the corresponding feature saved in the downsampling phase, learns detailed compensation information through the residual block (ResNetBlock), then uses the attention module (Transformer) to strengthen the prediction of key time points, and finally expands the time axis through the deconvolution module (Upsample).

[0129] In step S213, the decoder outputs a direction vector at each time position. This vector indicates the direction in which the current Mel spectrum state should change under the guidance of the noisy speech condition (i.e., the noisy Mel spectrum) to gradually approach the target clean speech Mel spectrum, thereby completing the path guidance and spectrum enhancement recovery. Figure 4 , which is a basic structural diagram of the conditional stream matching noise reduction module decoder in the present invention.

[0130] S214: The output of the decoder in step S213 is a path direction vector, which indicates the direction in which the current intermediate state should be updated to gradually restore the Mel spectrum of the clean speech. The direction predicted by the decoder is recorded as The goal of training is to As close as possible to the true target direction vector u, which is defined as follows:

[0131] u=x-(1-σ min )·z

[0132] Where x represents the Mel spectrum corresponding to the clean speech; σmin is the minimum diffusion coefficient, which is used to control the noise weight; z is the noise tensor sampled from the standard Gaussian distribution; u is the true direction required to “return from the intermediate state y to the clean speech x”.

[0133] In order to train the decoder to predict the correct direction, the Conditional Flow Matching Loss function is used, which is calculated as follows:

[0134]

[0135] Where, represents the total number of frequency bands in the spectrum; mask is the mask tensor of the valid speech frame, which is 0 in the time frame without speech and 1 in the time frame with speech, used to shield the interference of invalid frames on the loss; represents the directional vector component of the f-th frequency band predicted by the decoder; u f Represents the true value of the target direction vector in the fth frequency band.

[0136] See also Figure 5 , which is a flowchart of the reasoning process of step S22, step S22 can be further described as follows:

[0137] S221: Generate a Gaussian random noise spectrum z with the same structure as the Mel spectrum μ of the noisy speech as the starting point of the path.

[0138] S222: Divide the time steps evenly into the sequence t0, t1, ..., t N , as discrete time points on the path, used for state update.

[0139] S223: At each time step t i , the current state y i , the noisy speech Mel spectrum μ and the time vector are fed into the decoder to obtain the predicted direction vector of the current position:

[0140]

[0141] in, Represents the path direction prediction vector of step i; y i is the current intermediate spectrum state; μ is the corresponding noisy Mel spectrum, which serves as the generation guide condition; t i is the current time step.

[0142] S224: Update the state using Euler method:

[0143]

[0144] Among them, y i+1 is the next state spectrum; Δt is the interval between adjacent time steps; The direction prediction value for the current position.

[0145] S225: After all time steps are updated, the last frame y N This is the enhanced speech spectrum, which will be used as the input of the vocoder in step S3 to synthesize the final time-domain speech waveform.

[0146] As a further improvement, in step S3, the neural network vocoder module is constructed based on the HiFi-GAN structure, which is an important part of the present invention. Figure 6 , which shows the overall architecture of the neural network vocoder module in the present invention.

[0147] This module takes the enhanced Mel spectrum as input and generates a clear and natural time-domain speech waveform through feature mapping, layer-by-layer upsampling, residual modeling, and waveform synthesis. The specific steps are as follows:

[0148] Step S31: Perform feature transformation on the input enhanced Mel spectrum to obtain a high-dimensional initial feature map to improve feature representation capabilities and provide richer context information and channel dimension support for subsequent upsampling and detail modeling.

[0149] Step S32: Upsample the feature maps layer by layer to restore them to a time length that matches the time-domain speech. This layer-by-layer upsampling operation maps the frame-level spectrum to the sampling-level time-domain length, establishing a structural foundation for generating a continuously playable speech signal.

[0150] Step S33: Use a multi-scale residual structure to model speech details during each upsampling stage. The multi-scale residual structure helps simulate short-term and long-term dependencies in speech and enhances the ability to restore important acoustic details such as formants and phoneme boundaries.

[0151] Step S34: Map the processed feature map into a time-domain speech waveform.

[0152] Step S31 can be further described as follows:

[0153] S311: Enhanced Mel spectrum Input the one-dimensional convolution layer to perform feature extraction and preliminary modeling in the time dimension.

[0154] S312: Convert the input channel from 80 dimensions to a higher channel number C to enhance the representation capability of the model and adapt to subsequent upsampling processing.

[0155] S313: Use the LeakyReLU activation function on the output feature map to enhance its nonlinear expression ability and provide a good foundation for subsequent modeling.

[0156] Step S32 can further be described as follows:

[0157] S321: Input the previous feature map into the i-th upsampling module and perform a one-dimensional deconvolution operation, which is expressed as:

[0158] y i =ConvTranspose1d(x,stride=s i ,kernel=k i )

[0159] Among them, y i Represents the feature map of the upsampling output of the i-th layer; x represents the input feature map of the current upsampling layer; stride = s i Indicates the upsampling ratio of the deconvolution layer; kernel = k i Indicates the convolution kernel size of this layer.

[0160] S322: All upsampling layers expand the time axis layer by layer according to the following formula:

[0161]

[0162] Where T is the number of frames of the input Mel spectrum; T′ is the sampling length of the output time domain speech; s i is the upsampling ratio of the i-th layer; L is the total number of upsampling layers.

[0163] S323: After each upsampling layer, the channel dimension of the feature map is kept unchanged, and only the time dimension is expanded to provide higher time resolution input for subsequent residual modeling.

[0164] See also Figure 7 , which is a multi-scale residual block structure diagram of step S33 in the present invention, can be specifically expressed as the following steps:

[0165] S331: Multiple parallel residual blocks are connected after each level of upsampling module. Each residual block contains multiple sub-paths for simulating speech changes at different scales.

[0166] S332: The output of the j-th branch of the i-th layer is expressed as:

[0167] y ij =LeakyReLU(Conv1d(x,k ij ,d ij ))

[0168] Among them, y ij represents the output feature of the jth convolution path of the i-th layer; x is the input feature map of the path; k ij is the convolution kernel size; d ijis the expansion rate, used to expand the temporal perceptual range of the convolution without increasing the number of parameters. LeakyReLU is a nonlinear activation function with a negative slope, used to improve the model's nonlinear modeling capabilities in speech waveform modeling and avoid the vanishing gradient problem.

[0169] S333: All sub-path outputs are fused in an average manner and added to the input to form a residual connection:

[0170]

[0171] Among them, y i is the final output of the residual block in the i-th layer; N is the number of branches (for example, 3 receptive fields); x is the input feature map retained by the residual connection.

[0172] S334: The fused output will be used for upsampling in the next stage or for final output, ensuring the fusion of information at different scales and optimizing the representation of speech details.

[0173] Step S34 can further be described as follows:

[0174] S341: Apply one-dimensional convolution to the final residual module output feature map, compress the channel dimension and map it into a single-channel time domain signal:

[0175] z(t)=Conv1d post (x)

[0176] Among them, z(t) represents the output value at time point t after convolution; Conv1d post represents the post-processing one-dimensional convolutional layer; x is the output of the final residual module.

[0177] S342: Use the hyperbolic tangent function activation for the convolution output to limit the output value to [-1, 1]:

[0178]

[0179] in, represents the final audio signal value at time point t; tanh is the hyperbolic tangent activation function, which is used to standardize the numerical range.

[0180] S343: The final time domain waveform It is playable enhanced speech that can be directly used for storage or subsequent speech processing tasks.

[0181] See also Figure 2 The present invention also discloses a speech enhancement system based on conditional stream matching and vocoder, comprising:

[0182] Mel spectrum extraction module, used to convert the input noisy speech into a noisy Mel spectrum;

[0183] A conditional stream matching denoising module is used to process the noisy Mel spectrum obtained in step S1 and output an enhanced Mel spectrum;

[0184] A neural network vocoder module is used to restore the enhanced Mel-spectrogram obtained in step S2 to a time-domain speech waveform, thereby obtaining an enhanced speech signal;

[0185] Among them, the conditional stream matching denoising module and the neural network vocoder module are obtained through pre-training.

[0186] In the above technical solution, the conditional stream matching denoising module is implemented using a conditional stream matching method based on a continuous-time diffusion model. By training a neural network to estimate the path direction vector field, it is possible to iteratively generate an enhanced Mel-spectrum from Gaussian noise during the inference phase.

[0187] The conditional stream matching denoising module is pre-trained, where it learns the path to recover clean speech from a noisy state by simulating the diffusion process and directional prediction. The trained conditional stream matching denoising module is used for inference, where it starts from random noise and gradually recovers the clean speech spectrum under the guidance of the noisy Mel spectrum condition.

[0188] Furthermore, the training process of the conditional stream matching denoising module includes the following steps:

[0189] Obtain training data, including obtaining noisy speech and corresponding clean speech, and extracting their mel spectrograms as training targets;

[0190] Feature preparation, calculate the clean Mel spectrum x, sample noise z from the standard Gaussian distribution, randomly sample the diffusion time step t, and the diffusion time step t∈(0,1) adopts a uniform random distribution;

[0191] Intermediate state construction: construct the intermediate state y and target direction vector u in diffusion based on x, z, and t;

[0192] Time step embedding, the diffusion time step t is input into the sinusoidal position encoder and the multi-layer perceptron layer to generate a time embedding vector;

[0193] Deep decoder Decoder forward propagation, input y, conditional feature μ, time step t into the Decoder, predict the direction vector

[0194] Calculate the loss, based on the difference between the predicted direction and the target direction, calculate the conditional flow matching loss function;

[0195] Backpropagation and optimization, backpropagation gradients, using optimizers to update decoder parameters, clipping excessive gradients to prevent unstable training;

[0196] Learning rate scheduling, which uses a strategy of automatically reducing the learning rate based on performance stagnation, and dynamically adjusts the optimizer learning rate when the performance of the validation set stagnates;

[0197] Logging and evaluation: record loss values ​​and training status, and regularly evaluate performance indicators on the validation set;

[0198] Repeat the above process until the training is completed.

[0199] The structures and specific implementation processes of the Mel spectrum extraction module, the conditional stream matching denoising module and the neural network vocoder module are as described above and will not be repeated here.

[0200] The above embodiments are only intended to help understand the method and core concept of the present invention. It should be noted that, without departing from the principles of the present invention, a number of improvements and modifications may be made to the present invention by those skilled in the art, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.

[0201] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech enhancement method based on conditional stream matching and vocoder, characterized in that: The following steps are involved: Step S1: construct a Mel spectrum extraction module to convert the input noisy speech into a noisy Mel spectrum; Step S2: Construct a conditional stream matching denoising module to process the noisy Mel spectrum obtained in step S1 and output an enhanced Mel spectrum; Step S3: constructing a neural network vocoder module for restoring the enhanced Mel spectrum obtained in step S2 into a time-domain speech waveform, thereby obtaining an enhanced speech signal; Among them, the conditional stream matching denoising module and the neural network vocoder module are obtained through pre-training.

2. The speech enhancement method based on conditional stream matching and vocoder according to claim 1, characterized in that: In step S2, the conditional flow matching denoising module is implemented using a conditional flow matching method based on a continuous-time diffusion model. The path direction vector field is estimated by training a neural network so that an enhanced Mel spectrum can be iteratively generated from Gaussian noise during the inference phase. The module includes the following steps: Step S21: Pre-training a conditional stream matching noise reduction module, wherein the path for recovering clean speech from a noisy state is learned by simulating a diffusion process and direction prediction; Step S22: Perform inference using the conditional stream matching denoising module trained in step S21, wherein starting from random noise, the clean speech spectrum is gradually restored under the guidance of the noisy Mel spectrum condition.

3. The speech enhancement method based on conditional stream matching and vocoder according to claim 2, characterized in that: Step S21 includes the following steps: Obtain training data, including obtaining noisy speech and corresponding clean speech, and extracting their mel spectrograms as training targets; Feature preparation, calculate the clean Mel spectrum x, sample noise z from the standard Gaussian distribution, randomly sample the diffusion time step t, and the diffusion time step t∈(0,1) adopts a uniform random distribution; Intermediate state construction: construct the intermediate state y and target direction vector u in diffusion based on x, z, and t; Time step embedding, the diffusion time step t is input into the sinusoidal position encoder and the multi-layer perceptron layer to generate a time embedding vector; Deep decoder Decoder forward propagation, input y, conditional feature μ, time step t into the Decoder, predict the direction vector Calculate the loss, based on the difference between the predicted direction and the target direction, calculate the conditional flow matching loss function; Backpropagation and optimization, backpropagation gradients, using optimizers to update decoder parameters, clipping excessive gradients to prevent unstable training; Learning rate scheduling, which uses a strategy of automatically reducing the learning rate based on performance stagnation, and dynamically adjusts the optimizer learning rate when the performance of the validation set stagnates; Logging and evaluation: record loss values ​​and training status, and regularly evaluate performance indicators on the validation set; Repeat the above process until the training is completed.

4. The method for speech enhancement based on conditional stream matching and vocoder according to claim 3, characterized in that: Step S22 includes the following steps: S221: Generate a Gaussian random noise spectrum z with the same structure as the Mel spectrum μ of the noisy speech as the starting point of the path; S222: Divide the time steps evenly into the sequence t0, t1, ..., t N , as discrete time points on the path, used for state update; S223: At each time step t i , the current state y i , the noisy speech Mel spectrum μ and the time vector are fed into the decoder to obtain the predicted direction vector of the current position: in, Represents the path direction prediction vector of step i; y i is the current intermediate spectrum state; μ is the corresponding noisy Mel spectrum, which serves as the generation guide condition; t i is the current time step; S224: Update the state using Euler method: Among them, y i+1 is the next state spectrum; Δt is the interval between adjacent time steps; The direction prediction value of the current position; S225: After all time steps are updated, the last frame y N This is the enhanced speech spectrum.

5. The method for speech enhancement based on conditional stream matching and vocoder according to claim 1, characterized in that: In step S3, a neural network vocoder module is constructed and pre-trained based on the HiFi-GAN structure, so as to perform inference using the trained neural network vocoder module.

6. The method for speech enhancement based on conditional stream matching and vocoder according to claim 5, characterized in that: Step S3 includes the following steps: Step S31: performing feature transformation on the input enhanced Mel spectrum to obtain a high-dimensional initial feature map; Step S32: upsampling the feature map layer by layer to restore it to a time length that matches the time-domain speech, wherein the frame-level spectrum is mapped to the time-domain length of the sampling level through the layer-by-layer upsampling operation; Step S33: using a multi-scale residual structure to model speech details in each upsampling stage; Step S34: Map the processed feature map into a time-domain speech waveform.

7. The method for speech enhancement based on conditional stream matching and vocoder according to claim 1, characterized in that: In step S1, the Mel spectrum extraction module is implemented by using a spectrum conversion method based on Fourier analysis and Mel filter bank.

8. A speech enhancement system based on conditional stream matching and vocoder, characterized in that: include: Mel spectrum extraction module, used to convert the input noisy speech into a noisy Mel spectrum; A conditional stream matching denoising module is used to process the noisy Mel spectrum obtained in step S1 and output an enhanced Mel spectrum; A neural network vocoder module is used to restore the enhanced Mel-spectrogram obtained in step S2 to a time-domain speech waveform, thereby obtaining an enhanced speech signal; Among them, the conditional stream matching denoising module and the neural network vocoder module are obtained through pre-training.

9. The speech enhancement system based on conditional stream matching and vocoder according to claim 8, characterized in that: The conditional stream matching denoising module is implemented using a conditional stream matching method based on a continuous-time diffusion model. By training a neural network to estimate the path direction vector field, it is possible to iteratively generate an enhanced Mel spectrum from Gaussian noise during the inference phase. The conditional stream matching denoising module is pre-trained, where it learns the path to recover clean speech from a noisy state by simulating the diffusion process and directional prediction. The trained conditional stream matching denoising module is used for inference, where it starts from random noise and gradually recovers the clean speech spectrum under the guidance of the noisy Mel spectrum condition.

10. The speech enhancement system based on conditional stream matching and vocoder according to claim 9, characterized in that: The training process of the conditional stream matching denoising module includes the following steps: Obtain training data, including obtaining noisy speech and corresponding clean speech, and extracting their mel spectrograms as training targets; Feature preparation, calculate the clean Mel spectrum x, sample noise z from the standard Gaussian distribution, randomly sample the diffusion time step t, and the diffusion time step t∈(0,1) adopts a uniform random distribution; Intermediate state construction: construct the intermediate state y and target direction vector u in diffusion based on x, z, and t; Time step embedding, the diffusion time step t is input into the sinusoidal position encoder and the multi-layer perceptron layer to generate a time embedding vector; Deep decoder Decoder forward propagation, input y, conditional feature μ, time step t into the Decoder, predict the direction vector Calculate the loss, based on the difference between the predicted direction and the target direction, calculate the conditional flow matching loss function; Backpropagation and optimization, backpropagation gradients, using optimizers to update decoder parameters, clipping excessive gradients to prevent unstable training; Learning rate scheduling, which uses a strategy of automatically reducing the learning rate based on performance stagnation, and dynamically adjusts the optimizer learning rate when the performance of the validation set stagnates; Logging and evaluation: record loss values ​​and training status, and regularly evaluate performance indicators on the validation set; Repeat the above process until the training is completed.

Citation Information

Cited By

  • Speech enhancement method and device based on stream matching, equipment and medium

    CN120913576A

  • Speech enhancement method and system of stream matching sample level adaptive path

    CN121237110A

  • A single-step speech enhancement method and system based on a drift model

    CN122435941A