Speech enhancement method and device based on hidden space schrodinger bridge, equipment and medium
By employing a latent space Schrödinger bridge-based speech enhancement method, the problem of poor speech reconstruction quality in high-noise and low-signal-to-noise ratio environments in existing technologies is solved. This method enables speech restoration in high-sampling-rate and multi-distortion scenarios, thereby improving the stability and applicability of the speech enhancement system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-09-01
- Publication Date
- 2026-07-21
AI Technical Summary
Existing speech enhancement technologies cannot effectively utilize degraded prior information in high-noise and low-signal-noise-ratio environments, resulting in poor speech reconstruction quality and an inability to handle various complex distortion scenarios, as well as insufficient generalization ability.
A speech enhancement method based on the latent space Schrödinger bridge is adopted. The power spectrum energy characteristics of the audio signal are preserved by the target encoder, and a smooth evolution path from degradation to clean is constructed in the latent space by the Schrödinger bridge interpolation process. Combined with scale-variable regularization and latent space alignment mechanism, speech reconstruction in multi-distortion scenarios is achieved.
It improves the quality and fidelity of speech reconstruction, enhances the stability and applicability of the system in various application scenarios, can handle complex forms of high sampling rates and multiple distortions, and improves the speech restoration capability under extreme conditions.
Smart Images

Figure CN121075349B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a speech enhancement method, apparatus, device and medium based on a latent space Schrödinger bridge. Background Technology
[0002] Degraded speech refers to speech signals whose quality deteriorates during acquisition, transmission, or storage due to external interference or system defects. Speech enhancement aims to recover high-quality, intelligible clean speech from degraded speech and is a core technology in fields such as communication, hearing aids, and remote conferencing. Traditional speech enhancement methods are mostly based on mapping generative models (e.g., deep regression networks) or spectral subtraction, directly estimating the clean signal from the time-domain waveform or frequency-domain spectrogram. This method may be effective under controllable conditions (e.g., known noise type), but its generalization ability significantly decreases when facing scenarios with unknown noise types, high sampling rates (e.g., 48kHz) audio, and multiple distortions simultaneously (e.g., noise + reverberation + clipping). In recent years, path-modeling-based generative methods (e.g., diffusion models, stream matching models, adversarial generative models) have achieved excellent performance in speech enhancement.
[0003] However, regardless of the speech enhancement method, most current methods (such as diffusion-based, stream matching, or mask generation models) often fail to reconstruct missing or severely distorted spectral information in high-noise, low-signal-to-noise-ratio environments due to a lack of effective utilization of degradation priors, resulting in a decline in the quality of the enhanced speech. Summary of the Invention
[0004] This application provides a speech enhancement method, apparatus, device, and medium based on a latent space Schrödinger bridge, which can effectively solve the problem in related technologies that it is impossible to achieve structured reconstruction of missing spectral information in high noise and low signal-to-noise ratio environments, and can effectively improve the speech enhancement effect.
[0005] This application provides a speech enhancement method based on a latent space Schrödinger bridge, comprising the following steps:
[0006] Acquire degraded speech for which speech enhancement is desired;
[0007] The degraded speech is encoded into a degraded latent variable in the latent space by a target encoder. The target encoder is used to converge different types of degraded speech in the latent space to a distribution close to the corresponding clean latent variable. The target encoder has the function of maintaining the energy characteristics of the power spectrum of the audio signal.
[0008] The degraded speech is encoded into a degraded latent variable in the latent space by a target encoder. The target encoder is used to control the latent space to present a natural equivariant response to the amplitude scaling of the degraded speech based on the equivariant characteristics obtained during training, and to maintain the consistency of the spectral energy distribution of the degraded speech in the original time domain and the latent space during the process of encoding the degraded speech into the degraded latent variable.
[0009] The degenerate latent variables are processed using a target generation network based on a Schrödinger bridge to obtain the clean latent variables corresponding to the degenerate latent variables;
[0010] The clean latent variables are decoded using a target decoder to obtain the clean speech corresponding to the degraded speech.
[0011] This application also provides a speech enhancement device based on a latent space Schrödinger bridge, comprising the following modules:
[0012] The acquisition module is used to acquire the degraded speech that is to be enhanced.
[0013] The encoding module is used to encode the degraded speech into degraded latent variables in the latent space through a target encoder. The target encoder is used to converge different types of degraded speech in the latent space to a distribution close to the corresponding clean latent variables. The target encoder has the function of maintaining the energy characteristics of the power spectrum of the audio signal.
[0014] The processing module is used to process the degenerate latent variables through a target generation network based on Schrödinger bridge to obtain the clean latent variables corresponding to the degenerate latent variables;
[0015] The decoding module is used to decode the clean latent variables through the target decoder to obtain the clean speech corresponding to the degraded speech.
[0016] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a speech enhancement method based on a latent space Schrödinger bridge as described above.
[0017] This application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a speech enhancement method based on a latent space Schrödinger bridge as described above.
[0018] This application also provides a computer program product, including a computer program that, when executed by a processor, implements a speech enhancement method based on a latent space Schrödinger bridge as described above.
[0019] The speech enhancement method based on the latent space Schrödinger bridge proposed in this application has at least the following technical effects:
[0020] First, by employing a target encoder that maintains the power spectrum energy characteristics of the audio signal (i.e., scale variability), the problem of unstable model performance caused by drastic changes in the volume or dynamic range of the input signal can be solved, ensuring that the system can always extract accurate speech features, avoiding the loss of details or misjudgment of features due to amplitude differences, and improving the stability and applicability of the system in real and varied application scenarios.
[0021] Secondly, by employing a target generation network that constructs and utilizes the Schrödinger bridge interpolation process in the latent space, the problem of poor speech reconstruction quality caused by the failure of traditional generative models to effectively utilize degraded prior information in high-noise and low-signal-to-noise ratio environments can be solved. Since the target generation network can fully utilize the information of clean-degraded speech pairs during training, learning and constructing the smoothest evolutionary path from the degraded state to the clean state, it can accurately interpolate and restore speech even when spectral information is severely lacking, significantly improving speech restoration capability and content fidelity under extreme conditions. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating a speech enhancement method based on a latent space Schrödinger bridge, as shown in one embodiment of this application.
[0024] Figure 2 This is a schematic diagram illustrating the principle of a training target encoder according to an embodiment of this application.
[0025] Figure 3 This is a schematic diagram illustrating the principle of speech enhancement based on a latent space Schrödinger bridge, as shown in one embodiment of this application.
[0026] Figure 4 This is a schematic diagram illustrating the implementation principle of a speech enhancement method based on a latent space Schrödinger bridge, as shown in one embodiment of this application.
[0027] Figure 5 This is a structural block diagram of a speech enhancement device based on a latent space Schrödinger bridge, as shown in one embodiment of this application. Figure 6 This is a schematic diagram of the physical structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] Based on model structure, existing speech enhancement techniques can be categorized into mapping-based (e.g., Conv-TasNet, Gesper, Miipher, VoiceFixer, FINALLY) and path-building models (e.g., SGMSE, SBSE, UNIVERSE++, SpeechFlow, ResembleEnhance, AnyEnhance). Based on generation capability, they can be divided into discriminative and generative models. Regardless of the classification, each existing method in current speech enhancement and restoration techniques has significant limitations.
[0030] Mapping methods, such as those based on a single feedforward neural network or a Generative Adversarial Network (GAN), primarily rely on a direct mapping from input to output. Techniques like Conv-TasNet, Miipher, and Gesper are trained and optimized only under specific distortion conditions (e.g., specific types of noise or reverberation), resulting in limited functionality and a tendency to overfit to a single dataset. While methods like Voicefixer and FINALLY emphasize general restoration, modeling various distortion conditions and training on more datasets, their inherent limitations as mapping methods lead to poor generalization ability. When faced with new environments or distortion types with different distributions than the training data, model performance tends to drop significantly. Furthermore, discriminative models lack the ability to model the speech generation process, often focusing more on signal-level restoration while neglecting the naturalness and auditory quality of the speech, thus resulting in insufficient performance in subjective perceived quality.
[0031] For path pattern building methods, the main ones include the popular diffusion models, Schrödinger bridges, masking modeling, and flow models. Although they have achieved significant improvements in recovery quality and naturalness, they also have key limitations. Some models (such as SGMSE, SBSE, and SpeechFlow) are still limited to 16kHz speech processing and lack support for high-fidelity speech (such as 48kHz and 44.1kHz), making them unsuitable for studio-quality audio or music-related scenarios. Among other models, AnyEnhance is a mask generation paradigm. This type of model requires discretization of continuous audio signals, which introduces additional discretization loss. Furthermore, this type of model performs mask completion during training but generates from blanks during sampling, resulting in a significant difference between training and inference, leading to unstable generation quality. SpeechFlow and ResembleEnhance are both flow matching generation paradigms. Compared to diffusion and bridge models, these models do not introduce random Gaussian noise into the path, resulting in insufficient exploration of the path space and affecting the performance ceiling. Furthermore, SpeechFlow requires individual fine-tuning for each subtask after unified pre-training, making it unable to handle various signal distortion scenarios. While ResembleEnhance can handle complex tasks, the generated speech has poor listening quality. UNIVERSE++ uses a diffusion model to generate speech signals from Gaussian noise. Compared to flow matching and bridge models, this method does not fully utilize the prior conditions of distorted speech, resulting in a longer generation path and additional generation errors and performance losses.
[0032] In summary, existing speech enhancement techniques often only train on a single or at most two types of speech corruption (such as additive noise or reverberation), failing to provide unified processing for complex multi-source distortions in the real world (such as noise + reverberation + clipping + bandwidth limitations). A few works have attempted to build more general restoration capabilities, but issues such as inter-task interference and insufficient generalization of prompts remain, and a truly all-around speech restoration model has not yet been achieved. Therefore, regardless of the approach, challenges remain in generalization, task coverage, and practical deployment adaptability.
[0033] By integrating all the problems existing in the above-mentioned existing speech enhancement technologies, we can roughly identify the following three main issues:
[0034] First, narrowband limitations and dataset dependence: Most advanced generative methods (such as diffusion, flow, or mask generation models) are only optimized for low sampling rate audio of 16 kHz and cannot handle high-fidelity (48 kHz) audio; at the same time, most of these methods only achieve good results on specific datasets (such as single noise type or single distortion type) and perform poorly on compound distortion or mixed noise in real-world scenarios.
[0035] Second, insufficient utilization of prior information: Prior information refers to the useful information remaining in the original degraded speech. In high-noise, low-signal-to-noise ratio environments, existing generative models often fail to reconstruct missing or severely distorted spectral information (i.e., they cannot recover the severely damaged parts well) due to a lack of effective utilization of degraded prior information, resulting in a decline in the quality of the enhanced speech.
[0036] Third, the semantic gap in the latent space is large: even if some methods attempt to apply the generation process to the latent space, they fail to fully consider the semantic and perceptual gap between clean and degenerate speech pairs in the encoder-decoder design. Therefore, the latent space representation is difficult to achieve a smooth mapping between the degenerate and clean domains.
[0037] To address the aforementioned shortcomings, this application proposes a VoiceLBM technology solution for large-scale, multi-distortion, high-fidelity (e.g., 48 kHz) speech recovery, aiming to overcome the limitations of existing methods and achieve the following objectives:
[0038] First, it supports general speech enhancement for high sampling rates and multi-distortion scenarios: an end-to-end generative framework capable of processing 48 kHz audio is designed, which can take into account a variety of common degradations (such as additive noise, reverberation, downsampling, clipping, dynamic equalization, etc.) and their combinations in the same model, effectively solving the first problem mentioned above.
[0039] Second, improve speech reconstruction capability under low signal-to-noise ratio conditions: by constructing the Schrödinger Bridge (SB) interpolation process in the latent space, the information between clean and degraded speech pairs is fully utilized to improve the speech reconstruction quality under severe degradation conditions, effectively solving the second problem mentioned above.
[0040] Third, narrow the semantic gap in the latent space and enhance the generative generalization ability: introduce scale-equivariant regularization and latent alignment mechanism to make the encoder more consistent in its representation of clean speech and degraded speech in the latent space, thereby achieving smooth transformation between multiple data distributions and effectively solving the third problem mentioned above.
[0041] Next, this application will describe in detail how it addresses the aforementioned problems one by one. The speech enhancement method based on the latent space Schrödinger bridge provided in this application is implemented by a speech enhancement system or any electronic device with data processing capabilities. The method of this application will be introduced below using a speech enhancement system as the implementation subject.
[0042] Figure 1 This is a flowchart illustrating a speech enhancement method based on a latent space Schrödinger bridge, as shown in one embodiment of this application. (Refer to...) Figure 1 The method of this application may include:
[0043] Step 101: Obtain the degraded speech for which speech enhancement is desired.
[0044] In this embodiment, degraded speech can be any speech that carries noise or is distorted. In other words, any speech that is contaminated can be considered degraded speech.
[0045] Step 102: The degraded speech is encoded into a degraded latent variable in the latent space by the target encoder. The target encoder is used to converge different types of degraded speech in the latent space to a distribution close to the corresponding clean latent variable. The target encoder has the function of maintaining the energy characteristics of the power spectrum of the audio signal.
[0046] In this embodiment, the latent space is an abstract multidimensional space learned by the model, used to represent the potential features and structure of data (such as speech). Input data is mapped into low-dimensional vectors (latent variables), and the core is to compress data and extract key information.
[0047] In this embodiment, the target encoder is obtained by fine-tuning the encoder in a pre-trained scale-equal variational autoencoder. Since the encoder in the scale-equal variational autoencoder has energy preservation capability, that is, it can maintain the energy characteristics of the power spectrum of the audio signal, the target encoder also has energy preservation capability.
[0048] Energy originates from the signal's amplitude. Signal energy refers to the sum of the squares of the amplitude values at all sampling points over a period of time. The greater the amplitude, the stronger the signal's energy.
[0049] In this embodiment, the target encoder can maintain the energy characteristics of the power spectrum of the audio signal and effectively decouple the content information and amplitude (volume) information of the speech during the feature extraction stage, thereby solving the problem of performance instability caused by changes in the volume of the input signal in traditional models. In practical applications, the energy dynamic range of speech signals is extremely large. Traditional models may lose details due to excessively low volume or produce distortion due to excessively high volume. However, in the solution of this application, the target encoder is robust to scale changes in the signal. Regardless of how the volume of the input speech changes, it can accurately extract its acoustic content features, enabling the speech enhancement system to always provide reliable processing results.
[0050] Step 103: Process the degenerate latent variables using a target generation network based on the Schrödinger bridge to obtain the clean latent variables corresponding to the degenerate latent variables.
[0051] In this embodiment, the Schrödinger bridge problem is formulated as: Given an initial probability distribution... Smoothly transform into a final probability distribution In a stochastic process, the goal is to find an optimal process that is closest to or least deviates from a known prior stochastic process. The solution of this optimal process can then provide the solution from... arrive The conditional probability distribution at each intermediate moment is used to fully characterize the entire evolutionary path.
[0052] This application can pre-train a target generation network based on the Schrödinger Bridge theory. After obtaining the degenerate latent variables through step 102, the degenerate latent variables are input into the target generation network to obtain the corresponding clean latent variables.
[0053] The training process of the target generation network will be described in detail later.
[0054] Step 104: Decode the clean latent variables using the target decoder to obtain the clean speech corresponding to the degraded speech.
[0055] Execute step 104, and decode the clean latent variables through the target decoder to obtain the clean speech after enhancing the input degraded speech.
[0056] The speech enhancement method based on the latent space Schrödinger bridge proposed in this application has at least the following technical effects:
[0057] First, by employing a target encoder that maintains the power spectrum energy characteristics of the audio signal (i.e., scale variability), the problem of unstable model performance caused by drastic changes in the volume or dynamic range of the input signal can be solved, ensuring that the system can always extract accurate speech features, avoiding the loss of details or misjudgment of features due to amplitude differences, and improving the stability and applicability of the system in real and varied application scenarios.
[0058] Secondly, by employing a target generation network that constructs and utilizes the Schrödinger bridge interpolation process in the latent space, the problem of poor speech reconstruction quality caused by the failure of traditional generative models to effectively utilize degraded prior information in high-noise and low-signal-to-noise ratio environments can be solved. Since the target generation network can fully utilize the information of clean-degraded speech pairs during training, learning and constructing the smoothest evolutionary path from the degraded state to the clean state, it can accurately interpolate and restore speech even when spectral information is severely lacking, significantly improving speech restoration capability and content fidelity under extreme conditions.
[0059] In one implementation, based on the above embodiments, the target encoder is trained through the following steps:
[0060] Step a1: Obtain the first clean speech sample, the first degraded speech sample corresponding to the first clean speech sample, the first encoder, and the second encoder. The parameters of the second encoder are initialized from those of the first encoder. The second encoder is an encoder with adjustable parameters.
[0061] In this embodiment, the first encoder is the encoder in a pre-trained scale-equal variational autoencoder (VAE).
[0062] The first clean speech sample can be like... Figure 2 In As shown, the first degraded speech sample can be as follows: Figure 2 In As shown. Figure 2 This is a schematic diagram illustrating the principle of a training target encoder according to an embodiment of this application. Figure 2 middle, The subsequent encoder For the first encoder, The subsequent encoder is the second encoder. Indicates decoder, This represents the result after decoding the first clean latent variable sample. This represents the result after decoding the first degenerate latent variable sample.
[0063] Step a2: Encode the first clean speech sample into a first clean latent variable sample in the latent space using the first encoder.
[0064] exist Figure 2 middle, for Latent variables in the latent space (first clean latent variable sample).
[0065] Step a3: Encode the first degraded speech sample into a first degraded latent variable sample in the latent space using the second encoder.
[0066] exist Figure 2 middle, for Latent variables in the latent space (first degenerate latent variable sample).
[0067] Step a4: Adjust the parameters of the second encoder with the goal of minimizing the latent space distance between the first degenerate latent variable sample and the first clean latent variable sample. The adjusted second encoder is the target encoder.
[0068] In step a4, the alignment loss function shown in formula (1) can be used. This represents the total difference between the first degraded latent variable sample and the first clean latent variable sample:
[0069] (1)
[0070] in, These are the weighting coefficients for each loss term. Mean squared error loss is used to measure the distance between latent variable vectors. Cosine similarity loss is used to measure the directional consistency of latent variables. The cross-reconstruction loss is used to ensure that latent variables can be reconstructed from each other. To counteract the loss, it is used to promote the alignment of the latent space distribution. The specific representation of each loss term can be set according to actual needs. Figure 2 middle, That is, alignment loss function .
[0071] In one implementation, Specifically, it can be as follows:
[0072]
[0073] in, This is the first clean latent variable sample. For the first latent variable of degradation, The dimension of the latent variable.
[0074] Specifically, it can be:
[0075]
[0076] Specifically, it can be:
[0077]
[0078] in, This is the first clean speech sample. This is the first degraded speech sample. For decoders.
[0079] Specifically, it can be:
[0080]
[0081] in, It serves as a discriminator, used to distinguish between the first clean latent variable samples and the first degenerate latent variable samples.
[0082] When performing step a4, the goal is to minimize the value of the alignment loss function in formula (1) and adjust the parameters of the second encoder until the set stopping condition is met, such as the value of the loss function being less than a preset threshold.
[0083] The adjusted second encoder is the target encoder. In this embodiment, a dual-encoder structure is used to encode clean and degraded speech samples separately. Training is performed using multiple alignment losses (MSE, CosSim, cross-reconstruction loss, and adversarial loss) to narrow the distance between their latent space representations, enhancing the compactness and geometric consistency of the latent space. Specifically, MSE ensures that the latent variable values are similar, preserving energy amplitude information; CosSim maintains consistent feature orientation, ensuring that patterns such as spectral shape remain unchanged; cross-reconstruction loss forces latent variables to contain mutually reconstructable semantic information, achieving mutual translation; and adversarial loss eliminates the differences between the two at the probability distribution level, promoting global alignment. This strategy maps the latent variables of clean and degraded speech to the same semantic space, providing high-quality latent space representations for subsequent speech enhancement and denoising tasks, effectively improving the model's ability to process degraded speech.
[0084] Of course, in actual implementation, other alignment loss functions besides formula (1) can also be used, and this application does not impose specific restrictions on them.
[0085] This application employs a dual-encoder structure to perform asymmetric processing on clean and degraded speech, achieving latent space alignment and addressing the poor generalization ability problem in general speech enhancement. Specifically, by freezing the parameters of an encoder processing clean speech (the first encoder), its output is used as a baseline reference latent variable. Another trainable degradation encoder (the second encoder) is specifically designed to handle degraded speech. Its optimization objective is to adjust the output degradation latent variables by aligning the loss function. In the hidden space and The distribution is aligned. This mechanism forces the degenerate encoder to learn the ability to strip away various distortion components and extract core acoustic content, thereby eliminating the semantic representation differences between clean speech and various degenerate versions. Therefore, the model no longer learns a one-to-one mapping from a specific degenerate type to a clean signal, but masters a universal mapping function from any degenerate domain to a single clean domain. This can fundamentally improve the model's generalization ability, enabling it to handle novel noise and complex distortions not included in the training set, and laying a unified semantic foundation for the subsequent generation process.
[0086] The first encoder used in training the target encoder will be described in detail below. In one implementation, based on the above embodiments, the first encoder is trained through the following steps:
[0087] Step b1: Obtain the original speech samples and the scale-equal variational autoencoder to be trained, which includes a third encoder and a decoder.
[0088] Step b2: Encode the original speech samples into latent variable samples in the latent space using a third encoder.
[0089] In this step, we assume the original speech sample is The third encoder uses This means that, in this case, latent variable samples can be used express.
[0090] Step b3: Scale the latent variable samples using a random scaling factor to obtain scaled latent variables, and scale the original speech samples to obtain the theoretical scaled signal.
[0091] In this embodiment, it is assumed that the random scaling factor is... , , It can be set according to actual needs; therefore, the scaling latent variable is... The theoretical scaling signal is .
[0092] Step b4: Decode the scaled latent variables using a decoder to obtain the reconstructed signal.
[0093] Assuming the decoder is Then the reconstructed signal is .
[0094] Step b5: With the goal of minimizing the spectral energy difference between the reconstructed signal and the theoretically scaled signal, the third encoder and decoder are jointly trained, and the trained third encoder becomes the first encoder.
[0095] In step b5, the spectral energy difference between the reconstructed signal and the scaled speech sample can be minimized by minimizing the value of the loss function shown in formula (2).
[0096] (2)
[0097] in, For standard reconstruction losses, This loss term is used to measure the input. and its reconstruction output The similarity between them. For scale-variable reconstruction loss, The loss term measures the scaled input. and its reconstruction output The similarity between them. This application uses KL divergence regularization loss. As a latent space regularization term, other forms can also be used as latent space regularization terms depending on actual needs. and It's a hyperparameter.
[0098] In one implementation, Specifically, it can be:
[0099]
[0100] in, For the original speech sample signal at time... Sampling, The reconstructed signal output by the decoder at time The sampled values; This represents the total length of the signal (number of sampling points).
[0101] Specifically, it can be:
[0102]
[0103] in, For the dimension of latent variables, and These are the mean and standard deviation of the latent variable distribution of the encoder output, respectively.
[0104] Specifically:
[0105]
[0106] in , For random scaling factor, Scaling the signal theoretically, For the latent variables of the original speech samples, For signal reconstruction.
[0107] In actual implementation, the goal can be to minimize the value of the loss function in formula (2) and jointly train the third encoder and decoder until the set stopping condition is met, such as the value of the loss function being less than a preset threshold.
[0108] The trained third encoder is the same as the first encoder mentioned earlier.
[0109] In this embodiment, by jointly optimizing the standard reconstruction loss and the scale-variable reconstruction loss, the first encoder can accurately restore all the details of the original speech signal and has good amplitude robustness, enabling it to generate stable and consistent latent variables for speech with the same content but different volumes. Furthermore, the introduction of a KL divergence regularization term ensures that the latent variables are continuously and smoothly distributed in the latent space, avoiding the collapse of the feature space. Through the synergy of the three loss terms, a latent space that accurately represents the speech content and is insensitive to scale changes can be constructed, laying a solid foundation for subsequent improvements in speech enhancement.
[0110] Of course, in actual implementation, other loss functions besides formula (2) can be used, and this application does not impose specific restrictions on them.
[0111] In one implementation, based on the above embodiments, the target generation network is trained through the following steps:
[0112] Step c1: Obtain the second clean speech sample and its corresponding second degraded speech sample.
[0113] Step c2: Encode the second degraded speech sample into a second degraded latent variable sample in the latent space using the target encoder;
[0114] Step c3: Encode the second clean speech sample into a second clean latent variable sample in the latent space using the first encoder.
[0115] Step c4: Based on the closed-form solution of the Schrödinger bridge, determine the optimal intermediate distribution between the distribution of the second degenerate latent variable sample and the distribution of the second clean latent variable sample.
[0116] The core idea of Schrödinger's bridge theory is to find an optimal random path between two probability distributions, such that the probability transition of the path satisfies the dynamic constraints. By deriving the intermediate probability distribution from the distorted distribution to the clean distribution through this theory, the smoothness and analyzability of the generation process can be ensured.
[0117] In practical implementation, the second degenerate latent variable sample will be used. Second clean latent variable sample Considering the Gaussian endpoints, define the initial distribution. and target distribution Next, based on the Schrödinger bridge theory, the intermediate time points are derived. Distribution . This represents the intermediate probability distribution during the transition from a degenerate state to a clean state, ensuring that each step on the path conforms to the optimal evolution logic.
[0118] Step c5: With the goal of minimizing the Euclidean distance (L2 distance) between the distribution of the latent variables of the initial generator network output and the corresponding distribution in the optimal intermediate distribution, train the initial generator network based on the optimal intermediate distribution.
[0119] Next, an initial generator network is constructed using structures such as U-Net or Transformer, with the second degenerate latent variable as the input. and time step Output intermediate latent variables When training the initial generator network, the goal is to minimize the intermediate latent variables of the initial generator network's output. The distribution of the optimal intermediate distribution and the corresponding distribution (i.e. The difference between the two is used as the target for training, so that the initial generative network learns from the differences between the two. To meet of The mapping relationship, the parameters are optimized by the loss function, and the loss function may include... The initial generator network is trained through backpropagation, enabling the generated... Approaching step by step Finally, clean speech is decoded. In summary, the initial training objective of the generator network is to learn the parameters. , making the initial generator network Output The distribution should be as close as possible This ensures that the generated latent variables conform to the evolutionary logic of the theoretical path.
[0120] Minimizing the difference between the distribution of intermediate latent variables in the initial generation network output and the corresponding distribution in the optimal intermediate distribution can be achieved by minimizing the value of the loss function as shown in formula (3).
[0121] (3)
[0122] These are intermediate latent variables used to initially generate the network output. The latent variable corresponding to the second clean speech sample. For latent variable dimensions.
[0123] Among them, Figure 3 In the Schrödinger bridge forward process, that is, the process of gradually injecting noise into the latent variables, This is the second clean speech sample. This is the second degraded speech sample. It is the upper limit of noise intensity. As the second clean latent variable sample, For the second latent variable of degradation, to The process of doing this is the training process. Figure 3 This is a schematic diagram illustrating the principle of speech enhancement based on a latent space Schrödinger bridge, as shown in one embodiment of this application.
[0124] Step c6: Determine the target generator network based on the initial generator network after training.
[0125] exist Figure 3 In this context, DiT stands for Target Generation Network.
[0126] In one implementation, based on the above embodiments, the target decoder is obtained through the following steps:
[0127] Step d1: Obtain the third clean speech sample and its corresponding third degraded speech sample.
[0128] Step d2: Determine the clean latent variable prediction value corresponding to the third degraded speech sample through the target generation network.
[0129] Step d3: The clean latent variable predictions are decoded using the decoder in the trained scale-equal variational autoencoder to obtain the reconstructed speech.
[0130] Step d4: With the goal of minimizing the difference between the reconstructed speech and the third clean speech sample in perceptual speech quality evaluation and the difference in unified temporal objective speech quality measurement, the target generation network and the decoder are jointly optimized, and the optimized decoder is the target decoder.
[0131] Steps d1-d4 involve assembling all the previously trained modules (including the first encoder, the target encoder, the target generation network, and the decoder) and performing a comprehensive joint optimization. During the optimization of the target generation network and the decoder, the parameters of the first encoder and the target encoder need to be frozen.
[0132] In step d4, the difference between the third clean speech sample and the reconstructed speech in perceptual speech quality evaluation and the difference in unified temporal objective speech quality measurement can be minimized by minimizing the value of the loss function shown in formula (4).
[0133] (4)
[0134] in, The loss function corresponding to formula (3) is... This represents the Perceptual Evaluation of Speech Quality (PESQ) loss. This represents the loss of the Unified Time-Domain Objective Speech Quality Measure (U-TMOS). These are the weighting coefficients.
[0135] In one implementation, Specifically, it can be:
[0136]
[0137] intermediate latent variables Decoded speech (reconstructed speech) Original, clean speech. The value can be selected according to actual needs (e.g., 4.5).
[0138] Specifically, it can be:
[0139]
[0140] The value can be selected according to actual needs (e.g., 5.0).
[0141] In this embodiment, the optimized decoder is the target decoder, and the optimized generator network is the target generator network.
[0142] In this embodiment, after the Schrödinger bridge model (i.e., the generator network) in the latent space and VAE pre-training are completed, a perceptual feedback mechanism is additionally introduced to incorporate human subjective evaluation of speech quality (such as PESQ and U-TMOS) into the joint optimization objective. This joint fine-tuning method not only takes into account the accuracy of signal reconstruction but also significantly improves the naturalness and intelligibility of the enhanced speech in a real listening environment, thus extending the final generation effect from technical indicators to a more subjective experience that is closer to human hearing.
[0143] In conjunction with the above embodiments, in one implementation, after the entire speech enhancement system has been trained, the degenerate latent variables are processed through a Schrödinger bridge-based target generation network to obtain the clean latent variables corresponding to the degenerate latent variables. This process may include:
[0144] Step 1: Determine the time step sequence. The first time step in the time step sequence corresponds to the state of the degenerate latent variable.
[0145] Among them, the time step sequence is usually a uniformly distributed time series. ,in Corresponding to the degenerate state, This corresponds to a clean state.
[0146] Step 2: For each time step in the time step sequence, determine the corresponding latent variables through the target generation network.
[0147] Perform this step from arrive Intermediate latent variables are generated sequentially according to time steps. Initial state = (correspond During the iteration process, for each time step... Input the current latent variable and time step The next latent variable is calculated through a target generation network. .
[0148] Step 3: Identify the latent variable corresponding to the last time step in the time step sequence as the clean latent variable corresponding to the degenerate latent variable.
[0149] The latent variable corresponding to the last time step is , That is, close to clean latent variables. Output (corresponding) ).
[0150] exist Figure 3 In the Schrödinger bridge reverse process, This indicates degraded speech that requires speech enhancement. This represents the degenerate latent variable obtained by encoding through the target encoder. This represents the latent variables obtained by the target generation network at each time step. This represents a clean latent variable. This represents the clean speech obtained through the target decoder.
[0151] In conjunction with the above embodiments, in one implementation, the first clean speech sample and its corresponding first degraded speech sample are obtained from a sample library, and any degraded speech sample in the sample library is obtained in the following manner:
[0152] Obtain any clean speech sample from the sample library;
[0153] Different degradation operation types are randomly combined according to their respective probabilities, and degradation operations are applied to any clean speech sample based on the combination results to obtain the corresponding degraded speech sample. The different degradation operation types include at least: additive noise, reverberation, downsampling, clipping, and dynamic equalization.
[0154] In addition, the second clean speech sample, the second degraded speech sample, the third clean speech sample, and the third degraded speech sample mentioned above were all obtained from the sample library.
[0155] In this embodiment, by randomly combining common distortion operations such as additive noise, reverberation, downsampling, shearing, and dynamic equalization with preset probabilities to construct a multi-condition composite distortion operator T, we can not only enrich the diversity of training samples and simulate complex and varied real-world distortion scenarios, allowing the model to be exposed to more comprehensive distortion situations and improving its generalization adaptability to different distortion types, but also enhance the robustness of the model, enabling it to better resist various distortion interferences in practical applications and stably output high-quality results under training with diverse distortion combinations.
[0156] Figure 4 This is a schematic diagram illustrating the implementation principle of a speech enhancement method based on a latent space Schrödinger bridge, as shown in one embodiment of this application. Figure 4 In the middle, in the upper left area Represents original, clean speech; Indicates encoder; Representing latent variables, i.e. The code representation in the latent space after compression by the encoder; Indicates decoder; This indicates the reconstructed speech; Indicates a random scaling factor; This indicates clean speech with randomly altered volume. express The corresponding latent variables; Indicates to The reconstructed speech; Data Space refers to the space where the original, high-dimensional speech signal resides; Latent Space refers to the latent space; Latent Space Regularization is a regularization method, such as scaling; Energy Scaling and Energy Preserving are also known as energy scaling and energy preserving. and For the loss function, see formula (2) above. Figure 4 upper right area Represents original, clean speech; This represents the general distortion operator, which accepts various clean... And subject it to various random combinations of contamination; Indicates the process The processed speech generates a wide variety of distorted audio. Figure 4 There are multiple , indicating from the same It can generate multiple different types of damaged versions; Diverse Degradations are diverse degraded samples; It is to produce pure voice Input to encoder Obtained; It is Input to encoder Obtained; is the latent space alignment loss function, i.e., formula (1); Diverse Priors is the diversified prior; Converged Priors is the unified target latent variable; Latent Space Prior Convergence is the latent space prior convergence; Diverse Reconstructions is the diversified reconstruction result; Data Space Prior Convergence is the data space prior convergence; ConvergedReconstructions is the converged reconstruction result. This refers to a specific degraded speech input. The final enhanced output of the entire system; It is to produce pure voice The reconstruction result obtained after inputting the encoder; Let be the loss function for the final joint optimization, i.e., formula (4). Figure 4 Below, Bridge Forward is the forward bridge, formula It describes a theoretically existing, optimal evolutionary path, which can be understood as starting from a degenerate latent variable. The distribution is smoothly and optimally transformed into clean latent variables. The ideal process of distribution. Bridge Reverse is the reverse bridge. This describes the actual learned inverse process used for generation and repair, given a degenerate latent variable. At that time, the target generation network uses this formula to calculate step by step, and finally obtains the clean latent variables. In the two formulas above, Represents an intermediate moment in the evolutionary process. Latent variables; Represents time; Representing latent variables In a very small time step The minute changes within; Represents tiny steps in time; This represents a tiny increment in a standard stochastic process; and These are predefined functions, representing the drift term and diffusion term of the basic process, respectively; and It is the core function in Schrödinger's bridge theory. The learning objective of a neural network is to fit one of these two functions or a term derived from them. Let be a fractional function, and ∇ be the gradient operator.
[0157] This application employs a scale-variable autoencoder (VAE). The network structure of its encoder and decoder is optimized for processing high-dimensional data streams of 48kHz audio in terms of layer depth, convolutional receptive field, and upsampling / downsampling strategies, ensuring the model has the fundamental ability to process full-bandwidth signals. Secondly, this application transfers the computationally intensive core generation algorithm—an iterative repair process based on a Schrödinger bridge—from the vast original signal domain to the low-dimensional, compact latent space of the VAE. This avoids the enormous computational overhead and modeling difficulties encountered in directly generating in the high-dimensional spectrum or time domain, making efficient processing of high-fidelity audio computationally possible. Finally, the entire model can be trained on a database containing a large number of high-quality 48kHz samples, ensuring it fully learns the complex patterns of high-frequency signals. Therefore, by combining a high-fidelity network design with an efficient latent space generation strategy, this application, while ensuring computational feasibility, enables the model to fully learn and reconstruct the rich high-frequency harmonics and detailed information contained in 48kHz audio, overcoming the technical bottleneck of traditional generation methods being limited to narrowband audio of 16kHz.
[0158] In summary, the solution adopted in this application to address the proposed problem can be briefly summarized as follows:
[0159] (1) Construction of a general distortion model: Common distortion operations such as additive noise, reverberation, downsampling, shearing and dynamic equalization are randomly combined according to preset probabilities to obtain a multi-condition composite distortion operator T. This method can effectively solve the problem in the previous part (1) that most of the methods only achieve good results on specific datasets (such as a single noise type or a single distortion type), and perform poorly on composite distortion or when multiple noises are mixed in real-world scenarios.
[0160] This application integrates various speech data during the training phase, including multiple publicly available high-quality speech libraries, as well as various ambient noise and real reverberation sources, along with different types of sound distortion techniques, to generate extremely diverse training pairs through random combinations. This multi-source, multi-condition data construction strategy exposes the model to various real-world degradation scenarios during training, thereby possessing more comprehensive adaptability and maintaining good performance under various recording devices, microphone environments, and ambient noise combinations.
[0161] (2) Latent space coding: Design a scalable and equivariant audio VAE to map the original waveform to a compact latent space representation, while ensuring that the latent space has good robustness to signal amplitude and bandwidth scaling, and giving the latent space smoother geometric properties that are more conducive to deep neural network learning.
[0162] To address the complex relationship between amplitude variations in speech signals and sound quality, this application introduces a scale-variable regularization strategy: during training, the input signal is randomly scaled in amplitude, requiring the encoder to still learn the corresponding representation in the latent space for the scaled signal compared to the original signal. Through this training method, the model maintains structural consistency in the latent space under different volume conditions, ensuring that high-energy low-frequency components still dominate the latent space representation, thereby guaranteeing stable recovery of spectral details in subsequent generation processes. This design enables the autoencoder to not only remain robust under amplitude variations but also provides a sound geometric foundation for generation operations within the latent space.
[0163] (3) Latent Space Schrödinger Bridge Modeling: In the latent space, the clean latent variables and the distorted latent variables are regarded as Gaussian approximation endpoints, respectively. Through Schrödinger bridge theory, the analytical intermediate distribution and the corresponding smooth evolution formula are derived, and a bridge-type generative network is designed based on this to minimize the reconstruction error. This method can effectively solve the problem in aspect (2) above, where existing generative models often fail to reconstruct missing or severely distorted spectral information in high-noise and low-signal-noise ratio environments due to the lack of effective utilization of degenerate priors, resulting in a decrease in speech quality after enhancement.
[0164] By compressing audio into the latent space for bridge model generation, this application significantly reduces the runtime computational load compared to traditional diffusion or streaming models that operate directly in the spectrum or time domain. The interpolation process only requires iteration in a low-dimensional space, thereby accelerating the generation speed. This architecture naturally supports multiple deployment modes, including server-side batch processing, real-time online meeting noise suppression, and accelerated inference on mobile devices. The overall system balances high efficiency with consistency and scalability across multiple scenarios.
[0165] (4) Latent Space Alignment Module: Introducing a latent space alignment strategy: A dual encoder structure is adopted to encode the same pair of clean / degraded speech separately, and multiple alignment losses (MSE, CosSim, reconstruction loss, adversarial loss) are used to bring the latent space representations of degraded and clean speech closer together, enhancing the compactness and geometric consistency of the latent space. Even though some methods attempt to apply the generation process to the latent space, they have not fully considered the semantic and perceptual gap between clean and degraded speech pairs in the encoder-decoder design. Therefore, the latent space representation is difficult to achieve a smooth mapping between the degraded and clean domains, which can solve the third problem mentioned above.
[0166] This application employs a dual-encoder structure for clean speech-degraded speech pairs: one encoder processes the original clean speech signal, and the other processes the corresponding degraded speech signal. The two encoders share most parameters during training, but the encoder for the degraded speech signal undergoes additional fine-tuning using a dedicated alignment loss. The alignment loss includes reducing the distance between the two in the latent space representation space and maintaining the similarity in perceptual quality between their outputs in the data reconstruction space. Through these two constraints, the latent space representation corresponding to the degraded speech can more closely approximate the latent space representation corresponding to the clean speech, thus laying a semantic foundation for the subsequent random interpolation process.
[0167] (5) High-fidelity decoding and human feedback fine-tuning: After the Xueqiao model is trained, the decoder and Xueqiao model are fine-tuned together, and human feedback loss such as PESQ / U-TMOS is added to further optimize the subjective perception quality.
[0168] After the latent space bridge model and autoencoder are pre-trained, a perceptual feedback mechanism is introduced to incorporate human subjective evaluation of speech quality (such as expert scores or differentiable perceptual loss) into the joint optimization objective. This joint fine-tuning approach not only takes into account the accuracy of signal reconstruction but also significantly improves the naturalness and intelligibility of the enhanced speech in real-world listening environments, extending the final generated effect from technical metrics to a more subjective experience that closely resembles human hearing.
[0169] (6) System training process: Step 1: Pre-train scalable and equivariant VAE; Step 2: Fine-tune the distortion audio encoder based on alignment loss; Step 3: Construct the latent space Schrödinger bridge and train the generative network; Step 4: Fine-tune the output by jointly decoding and bridge model.
[0170] This application has achieved several significant effects and advantages in the field of general speech enhancement, as specifically reflected below:
[0171] First, a significant improvement in sound quality and intelligibility.
[0172] On multiple publicly available benchmark datasets, the enhancement performance of this application generally outperforms existing state-of-the-art models. Through joint fine-tuning using latent space interpolation and perceptual feedback, the subjective naturalness and intelligibility of the enhanced speech are significantly improved. User reviews show that whether it's everyday conversation, recordings of noisy streets, or heavily reverberated meeting recordings, clearer and more coherent speech contours can be heard without noticeable artificiality or musical noise artifacts. Comparative experiments demonstrate that, under the same noise and distortion intensity, the speech quality generated by this application surpasses traditional generative or discriminative methods, restoring more details and maintaining higher fidelity in key vowel parts of the human voice.
[0173] Second, robustness and generalization ability in various degradation scenarios.
[0174] Because the training phase incorporates various single or combined degradations such as additive noise, reverberation, bandwidth reduction, and signal distortion, this application exhibits excellent adaptability to unseen noise types or novel distortion combinations. Experiments demonstrate that even in the "live recording + microphone distortion" scenario presented in the test set, this application achieves high-quality denoising and dereverberation, while traditional methods often experience significant performance degradation under these mixed degradation conditions. The multi-source data construction strategy ensures that the model has used similar audio types at various sampling rates and in various scenarios, making this application applicable not only to standard benchmarks but also achieving equally excellent denoising results in different application scenarios such as telephone calls, live streaming, and short videos.
[0175] Third, detail restoration at high-fidelity sampling rates.
[0176] Traditional methods are often limited to low sampling rate environments (e.g., 16kHz), while this application is designed for high-fidelity audio (e.g., 48kHz), preserving more detail in the high-frequency range and avoiding suppression of high-frequency information. Experimental spectrum visualization shows that in regions where vocal consonants overlap with noise, this application preserves more vocal harmonics, rather than smoothing these areas as noise. This results in a more natural and spatial auditory experience.
[0177] Fourth, computational efficiency and deployment flexibility.
[0178] By concentrating the generation operations in the latent space, the computational load on the model during inference is significantly reduced. Compared to gradual diffusion in the time or frequency domains, this application only requires a few forward network operations on the lower-dimensional latent space representation, resulting in a significant speed improvement.
[0179] Fifth, theoretical interpretability and model stability.
[0180] The Schrödinger Bridge framework adopted in this application has a sound theoretical foundation. It is equivalent to finding the most suitable stochastic evolution path in the latent space, explaining from a probabilistic perspective why a natural transition can occur between noisy and clean spaces. Latent space alignment and scale equivariance mechanisms further ensure the semantic and perceptual coherence of this transition process. In actual training, latent space alignment reduces the distance between different domains, making it easier for the generative model to converge quickly; scale equivariance regularization avoids overly distorted or unreasonable distributions in the latent space, thereby improving training stability and reducing the probability of pattern collapse or speech distortion artifacts.
[0181] Sixth, application value and promotion prospects.
[0182] This application can be widely used in real-time call noise reduction (e.g., video conferencing, telephone customer service), post-production audio restoration (e.g., novel recording and refurbishment, old recording restoration), live streaming noise reduction, and smart home voice interaction noise suppression.
[0183] Specifically, the speech enhancement system of this application is applicable to the following fields: It can be integrated into terminal devices such as smartphones, tablets, and smart speakers to achieve real-time noise reduction and speech restoration during calls, recordings, and video conferencing, improving user experience. It can also be embedded in wearable devices such as hearing aids and headphones to help people with hearing impairments hear clearly in noisy environments. In communication and video conferencing, it is suitable for scenarios such as VoIP, online education, and telemedicine, effectively suppressing speech distortion caused by network packet loss, echo, and reverberation through cloud or edge-side inference. It can seamlessly integrate with mainstream conferencing software and customer service systems to improve the clarity and naturalness of remote communication. In content creation and post-processing, it can be used as a plugin or independent tool in the post-production workflow of film, radio, and podcasts to remove background noise, repair recording defects, and improve audio quality. It supports high sampling rates of 48 kHz and above to meet the high-fidelity requirements of professional recording and mixing.
[0184] In summary, through optimization and innovation of several key technical aspects, this application has achieved remarkable results in many aspects, including versatility, strong multi-degradation adaptability, high-fidelity reconstruction, robustness under low signal-to-noise ratio, computational efficiency, model stability, and practical application value, far exceeding the existing technical level and possessing extremely high research and application promotion value.
[0185] The speech enhancement device based on the latent space Schrödinger bridge provided in this application is described below. The speech enhancement device based on the latent space Schrödinger bridge described below can be referred to in correspondence with the speech enhancement method based on the latent space Schrödinger bridge described above.
[0186] Figure 5 This is a structural block diagram of a speech enhancement device based on a latent space Schrödinger bridge, as illustrated in one embodiment of this application. (Refer to...) Figure 5 The speech enhancement device 500 based on the latent space Schrödinger bridge of this application may include:
[0187] The acquisition module 501 is used to acquire the degraded speech that is to be enhanced.
[0188] The encoding module 502 is used to encode the degraded speech into degraded latent variables in the latent space through a target encoder. The target encoder is used to converge different types of degraded speech in the latent space to a distribution close to the corresponding clean latent variables. The target encoder has the function of maintaining the energy characteristics of the power spectrum of the audio signal.
[0189] Processing module 503 is used to process the degenerate latent variable through a target generation network based on Schrödinger bridge to obtain the clean latent variable corresponding to the degenerate latent variable;
[0190] The decoding module 504 is used to decode the clean latent variable through the target decoder to obtain the clean speech corresponding to the degraded speech.
[0191] According to the speech enhancement device 500 based on the latent space Schrödinger bridge provided in this application, the processing module 503 includes:
[0192] The first determining submodule is used to determine a time step sequence, wherein the first time step in the time step sequence corresponds to the state of the degradation latent variable;
[0193] The second determination submodule is used to determine the corresponding latent variables for each time step in the time step sequence through the target generation network.
[0194] The third determining submodule is used to determine the latent variable corresponding to the last time step in the time step sequence as the clean latent variable corresponding to the degenerate latent variable.
[0195] Figure 6 This is a schematic diagram of the physical structure of an electronic device according to an embodiment of this application, as shown below. Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions from the memory 630 to execute a speech enhancement method based on a latent space Schrödinger bridge.
[0196] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0197] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute a speech enhancement method based on a latent space Schrödinger bridge provided by the above methods.
[0198] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform a speech enhancement method based on a latent space Schrödinger bridge provided by the above methods.
[0199] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0200] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A speech enhancement method based on a latent space Schrödinger bridge, characterized in that, include: Acquire degraded speech for which speech enhancement is desired; The degraded speech is encoded into degraded latent variables in the latent space using a target encoder. The target encoder is used to converge different types of degraded speech in the latent space to a distribution close to the corresponding clean latent variables. The target encoder has the function of preserving the energy characteristics of the power spectrum of the audio signal. The target encoder is trained through the following steps: acquiring a first clean speech sample and its corresponding first degraded speech sample; encoding the first clean speech sample into a first clean latent variable sample in the latent space using a first encoder in a pre-trained scale-equal variational autoencoder; encoding the first degraded speech sample into a first degraded latent variable sample in the latent space using a second encoder to be trained, wherein the parameters of the second encoder are initialized from those of the first encoder; adjusting the parameters of the second encoder with the goal of minimizing the latent space distance between the first degraded latent variable sample and the first clean latent variable sample, and the adjusted second encoder is the target encoder. The degenerate latent variables are processed using a target generation network based on a Schrödinger bridge to obtain the clean latent variables corresponding to the degenerate latent variables; The clean latent variables are decoded using a target decoder to obtain the clean speech corresponding to the degraded speech.
2. The speech enhancement method based on the latent space Schrödinger bridge according to claim 1, characterized in that, The first encoder was trained through the following steps: Obtain raw speech samples; The original speech samples are encoded into latent variable samples in the latent space by the third encoder in the scale-equal variational autoencoder to be trained. The latent variable samples are scaled using a random scaling factor to obtain scaled latent variables, and the original speech samples are scaled to obtain theoretical scaled signals. The scaled latent variable is decoded by the decoder in the scale-equal variational autoencoder to be trained to obtain the reconstructed signal; With the goal of minimizing the spectral energy difference between the reconstructed signal and the theoretically scaled signal, the third encoder and the decoder are jointly trained, and the trained third encoder becomes the first encoder.
3. The speech enhancement method based on the latent space Schrödinger bridge according to claim 2, characterized in that, The target generation network is trained through the following steps: Obtain the second clean speech sample and its corresponding second degraded speech sample; The second degraded speech sample is encoded into a second degraded latent variable sample in the latent space by the target encoder. The second clean speech sample is encoded into a second clean latent variable sample in the latent space using the first encoder. Based on the closed-form solution of the Schrödinger bridge, the optimal intermediate distribution between the distribution of the second degenerate latent variable sample and the distribution of the second clean latent variable sample is determined; The initial generator network is trained based on the optimal intermediate distribution with the objective of minimizing the Euclidean distance between the distribution of the latent variables of the initial generator network output and the corresponding distribution in the optimal intermediate distribution. The target generator network is determined based on the trained initial generator network.
4. The speech enhancement method based on the latent space Schrödinger bridge according to claim 3, characterized in that, The target decoder is obtained through the following steps: Obtain the third clean speech sample and its corresponding third degraded speech sample; The clean latent variable prediction value corresponding to the third degraded speech sample is determined through the target generation network. The clean latent variable prediction value is decoded by the decoder in the trained scale-variable variational autoencoder to obtain the reconstructed speech. With the goal of minimizing the difference between the reconstructed speech and the third clean speech sample in perceptual speech quality evaluation and the difference in unified temporal objective speech quality measurement, the target generation network and the decoder are jointly optimized, and the optimized decoder is the target decoder.
5. The speech enhancement method based on the latent space Schrödinger bridge according to any one of claims 1-4, characterized in that, The process of processing the degenerate latent variables using a target generation network based on a Schrödinger bridge to obtain the clean latent variables corresponding to the degenerate latent variables includes: Determine a time step sequence, wherein the first time step in the time step sequence corresponds to the state of the degradation latent variable; For each time step in the time step sequence, the corresponding latent variable is determined by the target generation network; The latent variable corresponding to the last time step in the time step sequence is determined as the clean latent variable corresponding to the degenerate latent variable.
6. The speech enhancement method based on the latent space Schrödinger bridge according to any one of claims 1-4, characterized in that, The first clean speech sample and its corresponding first degraded speech sample are obtained from a sample library, and any degraded speech sample in the sample library is obtained in the following way: Obtain any clean speech sample from the sample library; According to the probability corresponding to each of the different degradation operation types, different degradation operation types are randomly combined, and degradation operations are applied to any clean speech sample according to the combination results to obtain the corresponding degraded speech sample. The different degradation operation types include at least: additive noise, reverberation, downsampling, clipping, and dynamic equalization.
7. A speech enhancement device based on a latent space Schrödinger bridge, characterized in that, include: The acquisition module is used to acquire the degraded speech that is to be enhanced. An encoding module is used to encode the degraded speech into degraded latent variables in the latent space using a target encoder. The target encoder is used to converge different types of degraded speech in the latent space to a distribution close to the corresponding clean latent variables. The target encoder has the function of preserving the energy characteristics of the power spectrum of the audio signal. The target encoder is trained through the following steps: acquiring a first clean speech sample and its corresponding first degraded speech sample; encoding the first clean speech sample into a first clean latent variable sample in the latent space using a first encoder in a pre-trained scale-equal variational autoencoder; encoding the first degraded speech sample into a first degraded latent variable sample in the latent space using a second encoder to be trained, wherein the parameters of the second encoder are initialized from those of the first encoder; adjusting the parameters of the second encoder with the goal of minimizing the latent space distance between the first degraded latent variable sample and the first clean latent variable sample, and the adjusted second encoder is the target encoder. The processing module is used to process the degenerate latent variables through a target generation network based on Schrödinger bridge to obtain the clean latent variables corresponding to the degenerate latent variables; The decoding module is used to decode the clean latent variables through the target decoder to obtain the clean speech corresponding to the degraded speech.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements a speech enhancement method based on a latent space Schrödinger bridge as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a speech enhancement method based on a latent space Schrödinger bridge as described in any one of claims 1 to 6.