Speech enhancement method and device based on conditional average flow, equipment and storage medium

By employing a conditional average flow-based speech enhancement method, which utilizes single-step inverse Eulerian update and average flow learning, the problem of low speech enhancement efficiency in existing technologies is solved, and efficient speech quality enhancement is achieved in noisy environments.

CN121838784APending Publication Date: 2026-04-10PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing generative speech enhancement methods are inefficient in noisy environments, making it difficult to improve speech enhancement efficiency while maintaining speech quality.

Method used

A speech enhancement method based on conditional average flow is adopted. Initial speech features are obtained through short-time Fourier transform, and a single-step inverse Euler update is performed using the conditional average flow model. The enhanced speech signal is obtained by combining inverse short-time Fourier transform, thus realizing single-step inverse Euler update and average flow learning.

Benefits of technology

It improves the efficiency of speech enhancement, ensures speech quality, and reduces the complexity and computational cost of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838784A_ABST
    Figure CN121838784A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech enhancement, and particularly discloses a speech enhancement method and device based on conditional average flow, equipment and a storage medium. When the original voice signal is received, performing short-time Fourier transform to obtain an initial voice feature; and performing single-step reverse Euler updating on the initial speech features based on a conditional average flow model to obtain enhanced speech features, and performing inverse short-time Fourier transform to obtain enhanced speech signals. According to the invention, average flow learning is carried out through the conditional average flow model, the overall displacement rate with the same dimension as the input feature is obtained, single-step reverse Euler updating is realized, accurate displacement indication from the noisy feature to the clean feature is obtained, the speech enhancement quality is ensured, multi-step integration is not needed, and the speech enhancement efficiency is improved. The method is applied to financial and medical voice interaction systems, such as telemarketing, telephone return visit, remote inquiry and the like, high-quality clear voice can be quickly obtained, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech enhancement technology, and in particular to a speech enhancement method, apparatus, device and storage medium based on conditional averaging. Background Technology

[0002] With the deepening of digital transformation, voice interaction technology has become a key infrastructure for improving service efficiency and optimizing user experience in the financial and healthcare industries. Voice interaction is widely used in financial services such as telephone sales, telephone banking, and telephone follow-ups, as well as in medical systems such as remote consultations, patient follow-ups, and medical education. Therefore, voice enhancement to remove noise and ensure clear speech is fundamental to smooth communication between both parties.

[0003] Speech enhancement aims to recover clear speech from noisy signals and is a crucial front-end technology in communication systems, automatic speech recognition, and human-computer interaction. Traditional discriminative methods (such as spectral masking and deep convolutional networks) perform well under normal conditions, but often exhibit oversmoothing and distortion in noisy environments, leading to degraded speech quality. In recent years, generative models (diffusion models, fractional models, and normalized flow models) have achieved significant advantages in perceptual quality by learning clean speech distributions and reversing noise processes. However, these generative methods generally rely on multi-step iterative sampling to solve ordinary differential equations, requiring extensive function evaluations, which reduces speech enhancement efficiency. Therefore, improving speech enhancement efficiency while maintaining speech quality has become a pressing issue. Summary of the Invention

[0004] This application provides a speech enhancement method, apparatus, device, and storage medium based on conditional averaging flow, to improve speech enhancement efficiency while ensuring speech quality.

[0005] In a first aspect, this application provides a speech enhancement method based on conditional averaging, the method comprising: Upon receiving the original speech signal, a short-time Fourier transform is performed on the original speech signal to obtain initial speech features; The initial speech features are updated using a single-step inverse Euler model based on a preset conditional average flow model to obtain enhanced speech features. The enhanced speech features are subjected to inverse short-time Fourier transform to obtain the enhanced speech signal.

[0006] Secondly, this application also provides a speech enhancement device based on conditional averaging, the device comprising: The initial speech feature acquisition module is used to perform a short-time Fourier transform on the received original speech signal to obtain initial speech features. The enhanced speech feature acquisition module is used to perform a one-step inverse Euler update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features; The enhanced speech signal acquisition module is used to perform inverse short-time Fourier transform on the enhanced speech features to obtain the enhanced speech signal.

[0007] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the speech enhancement method based on conditional averaging as described above.

[0008] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the conditional average stream-based speech enhancement method as described above.

[0009] This application discloses a speech enhancement method, apparatus, device, and storage medium based on conditional average flow. Upon receiving an original speech signal, a short-time Fourier transform is performed on the original speech signal to obtain initial speech features. A single-step inverse Euler update is then performed on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features. Finally, an inverse short-time Fourier transform is performed on the enhanced speech features to obtain the enhanced speech signal. This application uses a conditional average flow model for average flow learning to obtain an overall displacement rate with the same dimension as the input features, achieving a single-step inverse Euler update. This not only obtains accurate displacement indicators from noisy features to clean features, ensuring speech enhancement quality, but also eliminates the need for multi-step integration, improving speech enhancement efficiency. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a first schematic flowchart of a speech enhancement method based on conditional averaging provided in an embodiment of this application; Figure 2 This is a second schematic flowchart of a speech enhancement method based on conditional averaging provided in an embodiment of this application; Figure 3 This is a third schematic flowchart of a speech enhancement method based on conditional averaging provided in an embodiment of this application; Figure 4A schematic block diagram of a speech enhancement device based on conditional averaging provided for embodiments of this application; Figure 5 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0014] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0016] This application provides a speech enhancement method, apparatus, device, and storage medium based on conditional averaging flow. The conditional averaging flow-based speech enhancement method can be applied to a server. It learns the averaging flow through a conditional averaging flow model to obtain an overall displacement rate with the same dimension as the input features, achieving a single-step inverse Eulerian update. This not only obtains accurate displacement indicators from noisy features to clean features, ensuring speech enhancement quality, but also eliminates the need for multi-step integration, improving speech enhancement efficiency. The server can be a standalone server or a server cluster.

[0017] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0018] Please see Figure 1 , Figure 1This is a schematic flowchart illustrating a speech enhancement method based on conditional averaging flow, provided in an embodiment of this application. This conditional averaging flow-based speech enhancement method can be applied in a server to learn the averaging flow through a conditional averaging flow model, obtaining an overall displacement rate with the same dimension as the input features. It achieves a single-step inverse Eulerian update, not only obtaining accurate displacement indicators from noisy features to clean features, ensuring speech enhancement quality, but also eliminating the need for multi-step integration, thus improving speech enhancement efficiency.

[0019] like Figure 1 As shown, the speech enhancement method based on conditional average flow specifically includes steps S101 to S103.

[0020] S101. Upon receiving the original speech signal, perform a short-time Fourier transform on the original speech signal to obtain initial speech features; In one embodiment, the original speech signal is a noisy speech signal, which is a time-domain signal. The time-domain signal is converted into time-frequency domain features through Fourier transform.

[0021] Specifically, the original speech signal is segmented into overlapping short time frames, a Hamming window is superimposed on each frame, and then a fast Fourier transform is performed on each frame to obtain the time-frequency domain features of each frame. The time-frequency domain features of all frames are spliced ​​together in chronological order to obtain the time-frequency domain features, i.e., the initial speech features.

[0022] In another embodiment, the original speech signal is preprocessed before performing a short-time Fourier transform on the original speech signal. The preprocessing includes standardizing the sampling rate and removing silence segments to reduce invalid computations.

[0023] S102. Based on the preset conditional average flow model, perform a one-step inverse Euler update on the initial speech features to obtain enhanced speech features; In one embodiment, a trained conditional average flow model is used to predict the average velocity field of the initial speech features to obtain the average velocity field. Then, a one-step inverse Eulerian update is performed on the initial speech features based on the average velocity field to obtain the clean features corresponding to the initial speech features, i.e., the enhanced speech features.

[0024] The average velocity field is the overall displacement rate with the same dimension as the input features, representing the direction and magnitude of the movement of the noisy features (i.e., the initial speech features) towards the clean features.

[0025] Furthermore, before performing a single-step inverse Eulerian update on the initial speech features based on the preset conditional average flow model to obtain enhanced speech features, the method further includes: performing Gaussian Fourier processing on the preset time parameters to obtain high-dimensional time features; and fusing the high-dimensional time features with the initial speech features and inputting them into the conditional average flow model.

[0026] In one embodiment, a preset time parameter t∈[0,1] corresponds to the noise pollution level; t=0 represents clean speech, and t=1 represents noisy speech. Gaussian Fourier transform is applied to the time parameter to convert it into a high-dimensional time feature. Specifically, linear processing is performed on the time parameter to project it into a high-dimensional space, obtaining an intermediate vector. Then, using a preset frequency sine and a pre-selected function, sine and cosine transforms are applied to this intermediate vector, and the results of the sine and cosine transforms are concatenated to obtain the high-dimensional time feature.

[0027] In one embodiment, a deep learning-based broadcasting mechanism extends high-dimensional temporal features to the spatial dimension of speech features, and concatenates the extended temporal features with the initial speech features in the channel dimension to generate fused features. These fused features are then input into a conditional average flow model for subsequent velocity field prediction and speech enhancement.

[0028] Furthermore, the step of performing a one-step backward Eulerian update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features includes: predicting the average velocity field of the initial speech features based on the conditional average flow model to obtain an average velocity field; and performing a one-step backward Eulerian update on the initial speech features based on the average velocity field to obtain the enhanced speech features.

[0029] In one embodiment, after inputting the fused features corresponding to the initial speech features and the high-level time features into the conditional average flow model, the conditional average flow model predicts the average velocity field based on the initial speech features, intermediate state speech features, and time features to obtain the average velocity field.

[0030] In a specific embodiment, the conditional average flow model adopts an architecture of NCSN++U-Net with added self-attention. The encoder extracts noise patterns in the time-frequency domain through convolutional layers and gradually compresses the spatial dimension to capture global features. The decoder restores the spatial dimension through transposed convolutional layers and finally outputs an average velocity field with the same shape as the initial speech features.

[0031] In one embodiment, after obtaining the average velocity field, the average velocity field is substituted into the inverse Euler formula, and the inverse Euler formula is applied to perform a one-step transformation on the initial speech features to obtain enhanced speech features. The inverse Euler formula is:

[0032] in, This is the output of the conditional average flow model, i.e., the enhanced speech signal estimate; y represents the average velocity field; y represents conditional information, which serves as a guide or condition for the generation process, and is usually the initial speech features corresponding to noisy speech. This represents the start time of the reverse process; This is the termination time of the reverse process, i.e., the time corresponding to the target state that the signal is expected to reach; The initial speech features to be enhanced.

[0033] S103. Perform inverse short-time Fourier transform on the enhanced speech features to obtain the enhanced speech signal.

[0034] In one embodiment, the enhanced time-frequency domain features are restored to the time-domain speech signal to complete the final enhancement. Specifically, an inverse short-time Fourier transform is performed on the time-frequency domain features of each frame to obtain the time-domain features. A Hamming window is added to the time-frequency domain features of each frame, and the features are spliced ​​in frame order to obtain continuous time-domain features, i.e., enhanced speech features.

[0035] In one embodiment, the parameters of the inverse short-time Fourier transform are consistent with those of the short-time Fourier transform, ensuring correct phase and amplitude restoration.

[0036] In the above embodiments, average flow learning is performed through a conditional average flow model to directly predict displacement within a finite interval, i.e., average velocity field prediction, to obtain the overall displacement rate with the same dimension as the input features. Then, a single-step reverse Euler update is performed based on the average velocity field. This not only obtains accurate indications of the shift from noisy features to clean features, ensuring the quality of speech enhancement, but also eliminates the need for multi-step integration, thus improving the efficiency of speech enhancement.

[0037] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a conditional average flow-based speech enhancement method provided in an embodiment of this application. This conditional average flow-based speech enhancement method can be applied in a server to train a pre-trained model using speech training samples to obtain a conditional average flow model. It eliminates the need for knowledge distillation or correction back-processing based on a teacher model, reducing the complexity of model training and improving training efficiency.

[0038] like Figure 2 As shown, the speech enhancement method based on conditional averaging flow includes steps S201 to S203 before step S101.

[0039] S201. Obtain training speech sample pairs, wherein each training speech sample pair includes a noisy speech feature and a clean speech feature; In one embodiment, clean speech can be extracted from a publicly available speech dataset that stores clear speech segments. Noisy speech signals can be speech segments obtained by adding various noises (such as street noise, office voices, white noise, etc.) and possible distortions to the extracted clean speech signals.

[0040] In one embodiment, a short-time Fourier transform is performed on the obtained clean speech signal and the corresponding noisy speech signal to obtain clean speech features and noisy speech features.

[0041] S202. Based on the training speech sample pairs, construct a bilinear conditional probability path, and perform random sampling based on the bilinear conditional probability path to obtain intermediate state speech features. In one embodiment, a continuous transition path from clean speech features to noisy speech features, i.e., a bilinear conditional probability path, is constructed based on training speech sample pairs. Then, specific points are sampled from the bilinear conditional path as intermediate state speech features.

[0042] In a specific embodiment, the bilinear conditional probability path is as follows:

[0043] in, Let be the mean vector of the speech signal at time t; Let be the standard deviation vector of the speech signal at time t; y represents clean speech features; y represents noisy speech features; t is a time parameter, ranging from [0,1], controlling the transition from clean to noisy speech. When t=0, ... , representing the clean speech endpoint, at t=1 , representing the noisy speech endpoint; These are the preset minimum and maximum standard deviation hyperparameters.

[0044] Further, the step of randomly sampling based on the bilinear conditional probability path to obtain intermediate state speech features includes: calculating the mean and variance corresponding to the random sampling time points based on the bilinear conditional path, and randomly sampling noise from a preset standard normal distribution; obtaining the intermediate state speech features based on the mean, the variance, and the noise.

[0045] In one embodiment, intermediate states are sampled from a bilinear path as input samples for model training.

[0046] Specifically, a time t is randomly selected from the interval [0,1], and the mean μ(t) and variance σ(t) are calculated based on t. A random noise vector is then sampled, with each dimension independently derived from a standard normal distribution. By sampling, the intermediate state speech features obtained at time t are obtained based on the mean, variance, and random noise vector. .

[0047] In the above embodiments, the introduction of a random noise vector z increases the diversity and robustness of the training data.

[0048] S203. Based on the intermediate state speech features and the noisy speech features, perform model training to obtain the conditional average flow model.

[0049] In one embodiment, a conditional average flow model is trained based on intermediate state speech features and noisy speech features to learn the average velocity field from noisy speech to clean speech, thereby achieving single-step inference.

[0050] Specifically, the intermediate state speech features are concatenated with the noisy speech features. At the same time, the time parameter t is transformed into a high-dimensional feature through Gaussian Fourier embedding. The intermediate state speech features, noisy speech features, and high-dimensional time features are then fused and input into the model architecture to be trained.

[0051] In one embodiment, the pre-trained model adopts an architecture that combines NCSN++U-Net and a self-attention module. NCSN++U-Net is an encoder-decoder structure used to capture local spectral features, while the self-attention module is used to enhance long-range dependency modeling and improve feature representation in complex noisy scenarios.

[0052] In one embodiment, the optimization strategy employs the Adam optimizer and EMA (Exponential Moving Average) to update weights. Distributed training is used, supporting multi-GPU parallel training to improve training efficiency. A course learning strategy is also employed to gradually increase the weights of the mean branch, mixing r=t samples to ensure boundary consistency (t=0 corresponds to clean speech endpoints, t=1 corresponds to noisy speech endpoints).

[0053] In the above embodiments, the pre-trained model is trained using voice training samples to obtain the conditional average flow model. This eliminates the need for knowledge distillation or correction back-processing based on the teacher model, reducing the complexity of model training and improving training efficiency.

[0054] Furthermore, such as Figure 3 As shown, the step of training the model based on the intermediate state speech features and the noisy speech features to obtain the conditional average flow model specifically includes steps S301 to S204.

[0055] S301. Based on the pre-trained model, the average velocity field is predicted for the intermediate state speech features and the noisy speech features to obtain the predicted average velocity field. In one embodiment, noisy speech features serve as conditional information, guiding the model's denoising direction. A pre-trained model performs multi-scale feature extraction on the input intermediate-state speech features and noisy speech features, identifying noise at different frequencies and time locations, capturing the correlation between noise in the intermediate-state speech features and time t, and outputting the corresponding predicted average velocity field.

[0056] S302. Based on the intermediate state speech features and the noisy speech features, perform average flow identity reasoning to obtain the target average velocity field; In one embodiment, the pre-trained model learns the Mean Flow identity by capturing the correlation between noise in intermediate state speech features and time t, and derives the local target mean velocity field (i.e., the true velocity field corresponding to the intermediate state speech features).

[0057] The target average velocity field represents the true direction and magnitude of the movement of intermediate-state speech features toward clean speech features at time t. It reflects the distribution of noise in the time-frequency domain and is the target of model learning.

[0058] Further, the step of inferring the average flow identity based on the intermediate state speech features and the noisy speech features to obtain the target average velocity field includes: calculating the instantaneous velocity based on the intermediate state speech features to obtain the instantaneous velocity target; obtaining the gradient parameters of the pre-trained model and the preset average flow identity; and inferring the average flow identity based on the instantaneous velocity target, the gradient parameters, and the noisy speech features to obtain the target average velocity field.

[0059] In one embodiment, the average velocity field u is the integral average of the instantaneous velocity field v over the time interval [r, t], reflecting the overall motion trend of the signal within this interval. The average velocity field function is expressed as:

[0060] in, The average velocity field function; This is the instantaneous velocity field function; y represents the intermediate state speech features at time t; r is the starting time point, t is the current time point; y represents the conditional information (noisy speech features). Let be the integral variable, which varies within the interval [r, t].

[0061] When r = t, it degenerates into instantaneous velocity. At this point, differentiating (tr)u yields the average flow identity:

[0062] Since the true average velocity field corresponding to the intermediate state speech features is unknown, the above average flow identity is therefore invalid. It is unknown, and is predicted using the current model. To approximate the true average velocity field, therefore, using replace According to the chain rule, It can be broken down into:

[0063] in, instantaneous velocity , It is the model output. For input The gradient; Output for the model The partial derivative with respect to time t. Therefore, after transforming the mean flow identity, the target mean velocity field can be obtained, and the function formula is:

[0064] in, The target value for training; is the instantaneous velocity target; c is the hyperparameter for stable training, usually set to 0.5; The average velocity field predicted by the model; and It reflects the local change trend predicted by the model.

[0065] S303. Based on the preset loss function, the predicted average velocity field, and the target average velocity field, obtain the loss function value; In one embodiment, the model is optimized by measuring the difference between the predicted average velocity field and the target average velocity field using a loss function, where the loss function is:

[0066] in, The stop gradient operator indicates that... It is treated as a constant, and its gradient is not calculated during backpropagation; It is the expectation operator (averaging over all random variables).

[0067] Understandably, the loss function value represents the degree of inaccuracy of the current model's prediction. The larger the loss function value, the worse the model's prediction result; the smaller the loss function value, the more accurate the model's prediction result.

[0068] S304. When the loss function value is less than a preset loss threshold, the conditional average flow model is obtained.

[0069] In one embodiment, the model parameters are optimized through iterative training to improve the predicted values. Approaching the theoretical target value When the loss function value is less than a preset loss threshold, the trained model at this point is used as a conditional average flow model.

[0070] In the above embodiments, the average flow identity is used, combined with the instantaneous path velocity and the current gradient information of the model, to obtain the target average velocity field, so as to guide the model to predict the local change trend and accelerate the model training convergence process.

[0071] Please see Figure 4 , Figure 4 This application provides a schematic block diagram of a conditional averaged flow-based speech enhancement device, which is used to perform the aforementioned conditional averaged flow-based speech enhancement method. The conditional averaged flow-based speech enhancement device can be configured on a server.

[0072] like Figure 4 As shown, the conditional averaged stream-based speech enhancement device 400 includes: The initial speech feature acquisition module 401 is used to perform a short-time Fourier transform on the original speech signal when the original speech signal is received to obtain the initial speech features. The enhanced speech feature acquisition module 402 is used to perform a one-step inverse Euler update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features; The enhanced speech signal acquisition module 403 is used to perform inverse short-time Fourier transform on the enhanced speech features to obtain the enhanced speech signal.

[0073] Furthermore, the enhanced speech feature acquisition module 401 includes: The average velocity field acquisition unit is used to predict the average velocity field of the initial speech features based on the conditional average flow model, and obtain the average velocity field. An enhanced speech feature acquisition unit is used to perform a single-step inverse Euler update on the initial speech features based on the average velocity field to obtain the enhanced speech features.

[0074] Furthermore, the speech enhancement device 400 based on conditional averaging also includes a model training module, which includes: The sample pair acquisition submodule is used to acquire training speech sample pairs, wherein each training speech sample pair includes a noisy speech feature and a clean speech feature; The random sampling submodule is used to construct a bilinear conditional probability path based on the training speech sample pairs, and to perform random sampling based on the bilinear conditional probability path to obtain intermediate state speech features. The model training submodule is used to train the model based on the intermediate state speech features and the noisy speech features to obtain the conditional average flow model.

[0075] Furthermore, the model training submodule includes: The unit for predicting the average velocity field is used to predict the average velocity field based on the intermediate state speech features and the noisy speech features using a pre-trained model, thereby obtaining the predicted average velocity field. The target average velocity field acquisition unit is used to perform average flow identity inference based on the intermediate state speech features and the noisy speech features to obtain the target average velocity field. The loss function calculation unit is used to obtain the loss function value based on the preset loss function, the predicted average velocity field, and the target average velocity field; The conditional average flow model acquisition unit is used to obtain the conditional average flow model when the loss function value is less than a preset loss threshold.

[0076] Furthermore, the target average velocity field acquisition unit includes: The instantaneous velocity target acquisition subunit is used to calculate the instantaneous velocity based on the intermediate state speech features to obtain the instantaneous velocity target; The target average velocity field acquisition sub-unit is used to obtain the gradient parameters of the pre-trained model and the preset average flow identity. Based on the instantaneous velocity target, the gradient parameters and the noisy speech features, the average flow identity is inferred to obtain the target average velocity field.

[0077] Furthermore, the random sampling submodule includes: The random sampling unit is used to calculate the mean and variance of the random sampling time points based on the bilinear conditional path, and to randomly sample noise from a preset standard normal distribution. The speech feature acquisition unit is used to obtain the intermediate state speech features based on the mean, the variance, and the noise.

[0078] Furthermore, the speech enhancement device 400 based on conditional averaging also includes a feature input module, which includes: The time feature acquisition unit is used to perform Gaussian Fourier processing on preset time parameters to obtain high-dimensional time features; The feature fusion input unit is used to fuse the high-dimensional temporal features with the initial speech features and input them into the conditional average flow model.

[0079] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0080] The aforementioned device can be implemented as a computer program, which can be used in, for example... Figure 5 It runs on the computer device shown.

[0081] Please see Figure 5 , Figure 5This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.

[0082] See Figure 5 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0083] The non-volatile storage medium can store an operating system and a computer program. This computer program includes program instructions that, when executed, cause the processor to perform any conditional averaging-based speech enhancement method.

[0084] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0085] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When executed by a processor, the computer program can enable the processor to perform any speech enhancement method based on conditional averaging.

[0086] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0087] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0088] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Upon receiving the original speech signal, a short-time Fourier transform is performed on the original speech signal to obtain initial speech features; The initial speech features are updated using a single-step inverse Euler model based on a preset conditional average flow model to obtain enhanced speech features. The enhanced speech features are subjected to inverse short-time Fourier transform to obtain the enhanced speech signal.

[0089] In one embodiment, when the processor performs a one-step inverse Euler update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features, it is configured to: Based on the conditional average flow model, the average velocity field is predicted from the initial speech features to obtain the average velocity field. The enhanced speech features are obtained by performing a one-step inverse Euler update on the initial speech features based on the average velocity field.

[0090] In one embodiment, before performing a one-step inverse Euler update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features, the processor is further configured to: Obtain training speech sample pairs, wherein each training speech sample pair includes a noisy speech feature and a clean speech feature; Based on the training speech sample pairs, a bilinear conditional probability path is constructed, and random sampling is performed based on the bilinear conditional probability path to obtain intermediate state speech features. The conditional average flow model is obtained by training the model based on the intermediate state speech features and the noisy speech features.

[0091] In one embodiment, when the processor performs model training based on the intermediate state speech features and the noisy speech features to obtain the conditional average flow model, it is configured to: Based on the pre-trained model, the average velocity field is predicted for the intermediate state speech features and the noisy speech features to obtain the predicted average velocity field. Based on the intermediate state speech features and the noisy speech features, average flow identity inference is performed to obtain the target average velocity field. The loss function value is obtained based on the preset loss function, the predicted average velocity field, and the target average velocity field; When the loss function value is less than a preset loss threshold, the conditional average flow model is obtained.

[0092] In one embodiment, when the processor performs average flow identity inference based on the intermediate state speech features and the noisy speech features to obtain the target average velocity field, it is configured to: Instantaneous velocity is calculated based on the intermediate state speech features to obtain the instantaneous velocity target; The gradient parameters of the pre-trained model and the preset average flow identity are obtained. Based on the instantaneous velocity target, the gradient parameters and the noisy speech features, the average flow identity is inferred to obtain the target average velocity field.

[0093] In one embodiment, when the processor performs random sampling based on the bilinear conditional probability path to obtain intermediate state speech features, it is used to: The mean and variance of the random sampling time points are calculated based on the bilinear conditional path, and noise is randomly sampled from the preset standard normal distribution. The intermediate state speech features are obtained based on the mean, the variance, and the noise.

[0094] In one embodiment, before performing a one-step inverse Euler update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features, the processor is further configured to: Gaussian Fourier transform is applied to the preset time parameters to obtain high-dimensional time features; The high-dimensional temporal features are fused with the initial speech features and input into the conditional average flow model.

[0095] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the conditional average stream-based speech enhancement methods provided in the embodiments of this application.

[0096] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0097] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of speech enhancement based on conditional average flow, characterized by, The method comprises the following steps: Upon receiving an original speech signal, performing a short-time Fourier transform on the original speech signal to obtain initial speech features; Performing a single-step reverse Euler update on the initial speech features based on a preset conditional mean flow model to obtain enhanced speech features; Performing an inverse short-time Fourier transform on the enhanced speech features to obtain an enhanced speech signal.

2. The method of claim 1, wherein, The step of performing a single-step reverse Euler update on the initial speech features based on a preset conditional mean flow model to obtain enhanced speech features comprises: Performing an average velocity field prediction on the initial speech features based on the conditional mean flow model to obtain an average velocity field; Performing a single-step reverse Euler update on the initial speech features based on the average velocity field to obtain the enhanced speech features.

3. The method of claim 1, wherein the conditionally averaged flow-based speech enhancement method is characterized by, Before the step of performing a single-step reverse Euler update on the initial speech features based on a preset conditional mean flow model to obtain enhanced speech features, the method further comprises the following steps: Obtaining a training speech sample pair, wherein each set of the training speech sample pair comprises a noisy speech feature and a clean speech feature; Based on the training speech sample pair, constructing a bilinear conditional probability path and performing random sampling based on the bilinear conditional probability path to obtain an intermediate state speech feature; Based on the intermediate state speech feature and the noisy speech feature, performing model training to obtain the conditional mean flow model.

4. The method of speech enhancement based on the condition average flow according to claim 3, characterized in that, The step of performing model training based on the intermediate state speech feature and the noisy speech feature to obtain the conditional mean flow model comprises: Based on a pre-trained model, performing an average velocity field prediction on the intermediate state speech feature and the noisy speech feature to obtain a predicted average velocity field; Based on the intermediate state speech feature and the noisy speech feature, performing an average flow identity inference to obtain a target average velocity field; Based on a preset loss function, the predicted average velocity field, and the target average velocity field, obtaining a loss function value; When the loss function value is less than a preset loss threshold, obtaining the conditional mean flow model.

5. The method of speech enhancement based on the condition average flow according to claim 4, characterized in that, The step of performing an average flow identity inference based on the intermediate state speech feature and the noisy speech feature to obtain a target average velocity field comprises: Based on the intermediate state speech feature, performing an instantaneous velocity calculation to obtain an instantaneous velocity target; Obtaining gradient parameters of a pre-trained model and a preset average flow identity, based on the instantaneous velocity target, the gradient parameters, and the noisy speech feature, performing an inference on the average flow identity to obtain the target average velocity field.

6. The method of speech enhancement based on the condition average flow according to claim 3, characterized in that, The step of performing random sampling based on the bilinear conditional probability path to obtain an intermediate state speech feature comprises: Based on the bilinear conditional path, calculating a mean value and a variance corresponding to a random sampling time point, and randomly sampling a noise from a preset standard normal distribution; Based on the mean value, the variance, and the noise, obtaining the intermediate state speech feature.

7. The method of speech enhancement based on the condition average flow according to any one of claims 1 to 6, characterized in that, Before the step of performing a single-step reverse Euler update on the initial speech features based on a preset conditional mean flow model to obtain enhanced speech features, the method further comprises the following steps: Performing a Gaussian Fourier transform on a preset time parameter to obtain a high-dimensional time feature; input the high-dimensional time feature and the initial speech feature into the conditional average flow model; concatenate the high-dimensional time feature and the initial speech feature to generate the model input parameter.

8. A speech enhancement apparatus based on a conditional average flow, characterized by Comprise: An initial speech feature obtaining module configured to perform short-time Fourier transform on an original speech signal to obtain initial speech features when the original speech signal is received; An enhanced speech feature obtaining module configured to perform single-step reverse Euler update on the initial speech features based on a preset conditional average flow model to obtain enhanced speech features; An enhanced speech signal obtaining module configured to perform inverse short-time Fourier transform on the enhanced speech features to obtain an enhanced speech signal.

9. A computer device, comprising: The computer device comprises a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program and implement the conditional average flow-based speech enhancement method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program enables the processor to implement the conditional average flow-based speech enhancement method according to any one of claims 1 to 7 when executed by the processor.