Single-channel speech enhancement method and device based on mean inversion Schrodinger bridge
By using a single-channel speech enhancement method based on the mean-inverted Schrödinger bridge, the generation process is guided by posterior mean information, which solves the problems of high computational cost and insufficient generalization ability of hybrid methods and achieves more efficient speech enhancement results.
Patent Information
- Application Number
- CN202411900877.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-12-23
AI Technical Summary
Existing hybrid speech enhancement methods are computationally expensive and lack generalization ability in complex acoustic scenarios.
A single-channel speech enhancement method based on the mean-inverted Schrödinger bridge is adopted. By constructing discriminant and fractional models, the generation process is guided by posterior mean information, which reduces the number of neural network calls and improves the ability to retain high-frequency information and the accuracy of the generation process.
Significantly reduces computational costs, improves speech enhancement performance and generalization ability, and generates higher quality, cleaner speech.
Smart Images

Figure CN119763597B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech enhancement in intelligent speech technology, in particular to a single-channel speech enhancement method and device based on mean inversion Schrodinger bridge. BACKGROUND
[0002] Speech Enhancement (SE) technology aims to estimate the clean waveform from the noisy waveform, improve the naturalness and intelligibility of speech, and is widely used in speech recognition systems, online meetings, smart homes and other fields.
[0003] Generally speaking, speech enhancement methods can be divided into discriminative and generative methods. The discriminative method models the deterministic mapping between noisy speech and clean speech, which can effectively eliminate noise in speech waveform, but may cause speech distortion and over-denoising problems. The generative method implicitly or explicitly learns the latent data distribution of clean speech, which is more natural and more understandable, and has stronger generalization ability to various noise patterns. In particular, the generative method based on diffusion model regards the speech enhancement task as a conditional generation process or a probability distribution transmission process, which effectively improves the perceptual quality of the enhanced speech, but the reverse generation process requires multiple iterations, which has high computational cost and cannot meet the real-time speech processing requirements.
[0004] The latest research combines the advantages of the above two methods and develops a hybrid speech enhancement method, that is, a discriminative model and a diffusion model are cascaded. The discriminative model of this method is used to output the initial predicted waveform of the clean speech, and then the diffusion model is used to repair the blurring and distortion of the initial prediction, which reduces the number of iterations of the diffusion model to a certain extent and produces clean speech with higher perceptual quality. However, the current hybrid method still has some problems. On the one hand, the reverse generation process of the diffusion model still requires at least 20 neural network calls, and the computational cost is still high. On the other hand, due to the limitation of the generalization ability of the pre-discriminative model, the generalization ability is poor in complex acoustic scenes. SUMMARY
[0005] In view of the deficiencies in the prior art, the present application provides a single-channel speech enhancement method and device based on mean inversion Schrodinger bridge.
[0006] The present application achieves the above technical purpose by the following technical means.
[0007] The single-channel speech enhancement method based on mean inversion Schrodinger bridge comprises the following steps:
[0008] Mixing clean speech samples and noise speech samples according to different signal-to-noise ratios to obtain noisy speech samples, and forming a speech sample pair set composed of clean speech samples and noisy speech samples;
[0009] preprocessing the speech samples in the set of speech sample pairs, and then performing Fourier transform to obtain a pair of spectral complex number matrices;
[0010] constructing a discriminant model, taking the spectral complex number matrix of the noisy speech as input and taking the estimated spectral complex number matrix of the clean speech as output; adjusting parameters of the discriminant model based on the difference between the estimated spectral complex number matrix of the clean speech and the real spectral complex number matrix of the clean speech, and taking the discriminant model corresponding to the optimal parameters as the target discriminant model when the discriminant model meets the preset training requirement;
[0011] constructing a score model, parameterizing the inverse optimal shift score of the mean inversion Schrodinger bridge by using the score model to estimate the inverse optimal shift score in the inverse generation process; adjusting parameters of the score model based on the difference between the estimated inverse optimal shift score and the real inverse optimal shift score, and taking the score model corresponding to the optimal parameters as the target score model when the score model meets the preset training requirement;
[0012] given the noisy speech, parameterizing the inverse optimal shift score of the mean inversion Schrodinger bridge by using the trained target score model, performing the inverse generation process, and generating the clean speech.
[0013] Further, the process of obtaining the pair of spectral complex number matrices is as follows: for any speech sample pair (x, y), randomly intercepting sub-segments of x and y, and counting the maximum value of the sampling values in y, normalizing x and y respectively by using the maximum value to obtain x' and y', and obtaining the spectral complex number matrices X 0 and Y 0 corresponding to x' and y' by Fourier transform, and then performing amplitude compression on X 0 and Y 0 to obtain the pair of spectral complex number matrices (X, Y).
[0014] Further, the training process of the discriminant model is as follows:
[0015] inputting the spectral complex number matrix Y of the noisy speech into the discriminant model D θ , and outputting the estimated value
[0016] constructing a reconstruction loss function based on the clean speech where L(·) represents the difference between the two; taking the minimization of the loss function as the training target, performing gradient back propagation to update the parameters θ;
[0017] stopping the training when the preset training requirement is met; selecting the optimal parameters θ *The corresponding discriminative model is taken as the target discriminative model.
[0018] Further, the mean reversal Schrödinger bridge process is obtained as follows:
[0019] For a spectral complex matrix pair (X, Y), let the posterior mean of X conditioned on Y be denoted as When Y is given, the variable is from a conditional posterior distribution It is observed that the target discriminative model approximates the conditional posterior distribution, i.e.
[0020] The mean reversal process is denoted as at a specific time The sample obeys an edge Gaussian distribution where denotes a uniform distribution on the interval [0, 1],
[0021] A Schrödinger bridge process is defined to bridge the noisy speech distribution and the clean speech distribution, and the state sample of the Schrödinger bridge process at time t is denoted as X t The sample X t obeys an edge Gaussian distribution q(X t |X0, X1), and X0 = X, X1 = Y; at time t, the posterior mean of X t is denoted as The edge Gaussian distribution of the Schrödinger bridge is derived as follows:
[0022]
[0023] where is a point estimate of
[0024] The Schrödinger bridge process with the mean reversal process is called a mean reversal Schrödinger bridge, which replaces the boundary condition X0 = X of the Schrödinger bridge at time t with
[0025] Further, the mean reversal process is a Schrödinger bridge process bridging the posterior mean distribution and the clean speech distribution.
[0026] Further, the edge Gaussian distribution of the mean reversal process at time t is
[0027]
[0028] where denotes the mean of the edge distribution of the mean reversal process at time t, denotes the variance of the edge distribution of the mean-reverting process at time t, α t and σ t is the noise schedule of the mean-reverting process, signal-to-noise ratio
[0029] Further, the mean-reverting Schrödinger bridge is described by the following two stochastic differential equations:
[0030]
[0031] where f is the drift coefficient, g is the diffusion coefficient, W and are standard Wiener processes, Ψ t and is the Schrödinger factor, fractional and represent the forward optimal drift fraction, the backward optimal drift fraction;
[0032] At time t, the edge Gaussian distribution of the mean-reverting Schrödinger bridge is:
[0033]
[0034] where denotes the mean of the edge distribution of the mean-reverting Schrödinger bridge at time t, denotes the variance of the edge distribution of the mean-reverting Schrödinger bridge at time t, and is the noise schedule of the mean-reverting Schrödinger bridge, determined by f and g in the above stochastic differential equations, signal-to-noise ratio
[0035] Further, the training process of the fractional model is:
[0036] The spectral complex matrix Y of the noisy speech is input into the target discriminant model Output
[0037] Randomly sample time t According to the edge Gaussian distribution of the mean-reverting process at time t Calculate the mean of the distribution to obtain the point estimate at time t Next, according to the edge Gaussian distribution of the mean-reverting Schrödinger bridge at time t Sample to obtain the sample at time t where is random noise; finally, input the time t, sample X t into the fractional model Output At this time, the real backward optimal drift fraction is The estimated reverse optimal offset score is
[0038] According to the difference between the estimated reverse optimal offset score and the true reverse optimal offset score, the score model parameters are adjusted, and the difference is reflected in the difference between X0 and ; Specifically: define the data prediction loss function Minimize the loss function As the training target, perform gradient back propagation to update the parameters
[0039] When the preset training requirement is met, stop training; select the optimal parameters in the training process The corresponding score model as the target score model.
[0040] Further, the reverse generation process is specifically:
[0041] The time t j , the sample is input into the target score model to obtain
[0042] The reverse optimal offset score is estimated by Solve the reverse stochastic differential equation Calculate
[0043] Let j = j-1, if j > 0, repeat the above process, if j = 0, stop the reverse generation, the sample As the spectral complex matrix of the pure speech after noise speech enhancement, after inverse amplitude compression, inverse Fourier transform and inverse normalization, the waveform of the pure speech is finally obtained.
[0044] A single-channel speech enhancement device based on mean reversal Schrödinger bridge, comprising:
[0045] A sample pair construction module is constructed, which is composed of a pure speech sample and a noisy speech sample to form a speech sample pair set;
[0046] A sample processing module pre-processes and Fourier transforms the speech samples in the speech sample pair set to obtain a spectral complex matrix pair;
[0047] A discriminant model construction module constructs a discriminant model, estimates the spectral complex matrix of the pure speech, adjusts the parameters of the model based on the difference between the estimated spectral complex matrix of the pure speech and the true spectral complex matrix of the pure speech, and when the model meets the preset training requirement, the discriminant model corresponding to the optimal parameters is selected as the target discriminant model;
[0048] The score model construction module constructs a score model, estimates a reverse optimal offset score in a reverse generation process, adjusts parameters of the model based on a difference between the estimated reverse optimal offset score and a true reverse optimal offset score, and takes the score model corresponding to optimal parameters as a target score model when the model meets preset training requirements;
[0049] The pure speech generation module parameterizes the reverse optimal offset score of the mean inversion Schrodinger bridge using the trained target score model for a given noisy speech, performs a reverse generation process, and generates pure speech.
[0050] Compared with the prior art, the present application has the following beneficial effects:
[0051] (1) The present application proposes a single-channel speech enhancement framework based on a mean inversion Schrodinger bridge, defines the speech enhancement task as an optimal transmission problem between a noisy speech distribution and a pure speech distribution, solves the corresponding Schrodinger bridge problem using the diffusion model idea, and realizes pure speech generation. Compared with the existing single-channel speech enhancement method based on the diffusion model on the public data set, the present application has more advanced performance, and the number of neural network calls is significantly reduced, and the calculation cost is lower.
[0052] (2) The present application improves the traditional Schrodinger bridge and proposes a mean inversion Schrodinger bridge. The posterior mean (the posterior mean of the pure speech conditioned on the noisy speech) information is introduced as a generation process guide through a discriminative model. Unlike existing hybrid methods, the present application only calculates the reverse optimal offset score based on the approximate sampling of the edge Gaussian distribution by the discriminative model during the training of the score model, thereby avoiding the problem of limited generalization ability caused by the pre-discriminative model in the hybrid method. Specifically, the present application designs a mean inversion process to simulate the conversion process from the posterior mean to the pure speech. Then, by deriving the approximate form of the Gaussian edge distribution, the mean inversion process is introduced into the Schrodinger bridge as a guide for the reverse generation process. In the early stage of the reverse generation process, the mean inversion process is biased towards the posterior mean sample, and the mean inversion Schrodinger bridge process sample is located outside the data manifold. Under the guidance of the posterior mean, the complexity of the neural network estimating the reverse optimal offset score can be reduced, and the ability to preserve high-frequency information of the speech can be improved. In the later stage of the generation process, the mean inversion process gradually deviates towards the pure speech sample, and the mean inversion Schrodinger bridge process sample is located inside the data manifold, which is converted into a generation process of the pure speech sample, gradually restores the low-frequency information of the speech, and finally achieves the optimal conversion of the noisy speech and the pure speech under the guidance of the posterior mean, realizing the speech enhancement effect. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0054] Figure 1 The single-channel speech enhancement flowchart based on the mean inversion Schrödinger bridge described in the present application. DETAILED DESCRIPTION
[0055] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the protection scope of the present application.
[0056] As shown in Figure 1 , the single-channel speech enhancement method based on the mean inversion Schrödinger bridge provided by the present embodiment comprises the following steps:
[0057] S1, constructing a sample pair: mixing a pure speech sample and a noise speech sample according to different signal-to-noise ratios to obtain a noisy speech sample, and grouping the pure speech sample and the noisy speech sample to form a speech sample pair set.
[0058] In the present embodiment, in the computer, the VoiceBank+DEMAND speech enhancement benchmark data is obtained from the network, and the pre-constructed speech sample pair set D xy = {(x (i) ,y (i) )|x (i) ∈D x ,y (i) ∈D y ,i=1,2,...,n} is directly obtained using the VoiceBank+DEMAND speech enhancement benchmark data set, wherein (x (i) ,y (i) ) is the i-th group of pure speech and noisy speech sample pair, x (i) and y (i) represent the original waveform of the pure speech and the original waveform of the noisy speech respectively, D x = {x (1) ,x (2) ,...,x (n)} is a pure speech set, and D y = {y (1) ,y (2) ,...,y (n)} is a set of noisy speech, i is a sample index, and n is the total number of sample pairs. The clean speech in the dataset comes from the VoiceBank corpus, and the noisy speech comes from the DEMAND corpus, and is mixed according to signal-to-noise ratios of 0db, 5db, 10db and 15db.
[0059] The above sample pair construction method is not limited to using pre-constructed sample pairs of public datasets, and can also use clean speech samples from clean speech datasets such as VoiceBank, TIMIT, WJS0, mix noise speech from noise speech datasets such as DEMAND, WHAM!, and construct speech sample pairs according to a specific signal-to-noise ratio.
[0060] S2, speech preprocessing: after preprocessing the speech samples in the set of speech sample pairs, the corresponding spectral complex number matrix pair is obtained by Fourier transform.
[0061] In this embodiment, when any speech sample pair (x, y) ∈ D xy is preprocessed, x and y are first randomly truncated, and the maximum value y max of the sampling value in y is counted, and the maximum value is used to normalize x and y to obtain x' = x / y max , y' = y / y max ; then, Fourier transform is performed to obtain the corresponding spectral complex number matrix X 0 , Y 0 of x' and y'; then, amplitude compression is performed on X 0 and Y 0 to obtain the final spectral complex number matrix pair (X, Y), wherein where |·| represents the modulus, ∠(·) represents the phase angle, β s is a scaling factor, and β m is an exponential factor. In this embodiment, β s = 0.33 and β m = 0.5. In this embodiment, the original waveform x and the spectral complex number matrix X can represent clean speech, and the original waveform y and the spectral complex number matrix Y can represent noisy speech.
[0062] S3, discriminant model training: a discriminant model is constructed for estimating the posterior mean of clean speech conditioned on noisy speech. The model takes the spectral complex number matrix of noisy speech as input and the estimated spectral complex number matrix of clean speech as output. The parameters of the discriminant model are adjusted according to the difference between the estimated spectral complex number matrix of clean speech and the true spectral complex number matrix of clean speech. When the discriminant model meets the preset training requirements, the discriminant model corresponding to the optimal parameters is used as the target discriminant model.
[0063] In this embodiment, first construct a discriminative model D θ , train the model according to the following process:
[0064] (1) Select a sample (x, y) ∈ D xy , pre-process according to S2 to obtain a spectral complex number matrix pair (X, Y);
[0065] (2) Input the spectral complex number matrix Y of the noisy speech into D θ , output the estimated value of the spectral complex number matrix of the pure speech
[0066] (3) Based on the pure speech, construct a reconstruction loss function Where L(·) represents the difference between the two, which can usually use the Mean-Square Error (MSE); the training goal is to minimize the loss function , gradient back propagation is performed to update the parameters θ;
[0067] (4) Repeat (1) to (3) until the training requirements such as the maximum number of iterations are met, stop training; select the optimal parameters θ * Corresponding discriminative model as the target discriminative model; the target discriminative model is only used for training the score model in S4, and is not used for generating pure speech in S5.
[0068] S4, score model training: construct a score model for matching the inverse optimal offset score of the mean inversion Schrödinger bridge. The inverse optimal offset score of the mean inversion Schrödinger bridge is parameterized using the score model to estimate the inverse optimal offset score in the inverse generation process. According to the difference between the estimated inverse optimal offset score and the true inverse optimal offset score, the parameters of the score model are adjusted. When the preset training requirements are met, the score model corresponding to the optimal parameters is used as the target score model.
[0069] In this embodiment, the mean inversion Schrödinger bridge is a special Schrödinger bridge that improves the estimation ability of high-frequency information of pure speech by introducing posterior mean information, and reduces the difficulty of score matching during training.
[0070] Specifically, for a specific spectral complex number matrix pair (X, Y), first assume that the posterior mean of X given Y is a variable When Y is given, the variable can be observed from a conditional posterior distribution In this embodiment, the target discriminative model trained by S3 is used to approximate the conditional posterior distribution, that is,
[0071] Next, a Schrödinger's bridge process bridging the posterior mean distribution and the clean speech distribution is defined, called the mean inversion process. The mean inversion process is then applied at a specific time... The state sample is represented as sample Follows a marginal Gaussian distribution in This represents a uniform distribution on the interval [0,1]. At time t, the marginal Gaussian distribution of the mean reversal process can be specifically represented as:
[0072]
[0073] in, This represents the mean of the marginal distribution during the mean reversal process at time t. Then α represents the variance of the marginal distribution during the mean reversal process at time t. t and σ t It is noise scheduling in the mean-inversion process, signal-to-noise ratio In the early stages of the reverse generation process, the mean inversion process is biased towards the posterior mean sample; while in the later stages, the mean inversion process gradually shifts towards the clean speech sample, achieving the optimal conversion between the posterior mean and the clean speech.
[0074] Finally, a Schrödinger bridge process is defined to bridge the noisy speech distribution and the clean speech distribution, and the state sample of this Schrödinger bridge process at time t is represented as X. t Sample X t Follows a marginal Gaussian distribution q(X) t Given a region |X0,X1), and X0=X, X1=Y. This Schrödinger bridge can be described by the following two stochastic differential equations:
[0075]
[0076] Where f is the offset coefficient, g is the diffusion coefficient, and W and It is a standard Wiener process, Ψ t and It is the Schrödinger factor, a fraction. and Represents the forward optimal offset score and the reverse optimal offset score.
[0077] To introduce the mean reversal process into the Schrödinger bridge process, at time t, assume X... t by As a conditional expression, the approximate marginal Gaussian distribution of the Schrödinger bridge is derived as follows:
[0078]
[0079] in, is point estimate, which can be generally represented by the distribution mean. Thus, the Schrödinger bridge process with the mean inversion process is called the mean inversion Schrödinger bridge, which replaces the boundary condition X0=X of the traditional Schrödinger bridge at time t with X can be approximately sampled from the marginal Gaussian distribution t At time t, the marginal Gaussian distribution of the mean inversion Schrödinger bridge can be specifically represented as:
[0080]
[0081] wherein, denotes the mean of the marginal distribution of the mean inversion Schrödinger bridge at time t, denotes the variance of the marginal distribution of the mean inversion Schrödinger bridge at time t, and are the noise schedules of the mean inversion Schrödinger bridge, which are determined by f and g in the above stochastic differential equation respectively (the determination process is prior art), and the signal-to-noise ratio
[0082] The reverse generation process of the mean inversion Schrödinger bridge refers to the process of transforming from time t to 1 to 0, corresponding to the sample X t from the noisy speech to the pure speech. The mean inversion Schrödinger bridge is characterized in that: due to the approximate replacement of the boundary condition, at the initial stage of the reverse generation process, the mean inversion process is biased towards the posterior mean sample, and the mean inversion Schrödinger bridge process takes the posterior mean as a condition to maximize the restoration of high-frequency information of the pure speech from the noisy speech; and at the later stage, the mean inversion process is gradually biased towards the pure speech sample, and the mean inversion Schrödinger bridge process gradually restores the low-frequency information, and finally realizes the optimal conversion of the noisy speech and the pure speech under the guidance of the posterior mean.
[0083] In the fractional model training process, because the boundary conditions X0 and X1 are known, the reverse generation process of the mean inversion Schrödinger bridge is also known, and the reverse optimal offset fraction at time t can be represented as The reverse optimal offset fraction characterizes the optimal path of the sample from the noisy speech distribution to the pure speech distribution, and indicates the direction of the gradual transition from the noisy speech sample to the pure speech sample. In the actual speech enhancement process, only the boundary condition X1 is known, and the reverse optimal offset fraction at time t can be parameterized by constructing a fractional model The training target of the fractional model is to ensure that the estimated reverse optimal offset fraction can accurately guide the sample to evolve from the noisy state to the pure state, thereby realizing the effect of speech enhancement.
[0084] In this embodiment, when training the fractional model, a fractional model The model usually adopts a U-Net structure with a residual connection. In the use of fractional models Inverse optimal shift fraction parameterized by mean inversion Schrödinger bridge At this time, the fractional prediction can usually be performed by using the model to directly match the real fraction, or indirectly match the real fraction by using prediction noise, prediction data, etc. The training method of the embodiment for indirectly matching the real fraction by using the prediction data is described as follows, that is, Predicting X0 to estimate the inverse optimal shift fraction, denoted as
[0085] Specifically, the fractional model is trained according to the following process:
[0086] (1) Select a sample (x, y) ∈ D xy , and perform pre-processing according to S2 to obtain a spectral complex matrix pair (X, Y);
[0087] (2) Input the spectral complex matrix Y of the noisy speech into the target discriminant model output by S3 training
[0088] (3) Randomly sample time t According to the marginal Gaussian distribution of the above mean inversion process at time t Calculate the mean of the distribution to obtain the point estimate at time t Next, according to the marginal Gaussian distribution of the mean inversion Schrödinger bridge at time t Sample the sample at time t where is random noise; finally, input the time t and the sample X t into the fractional model output At this time, the real inverse optimal shift fraction is The estimated inverse optimal shift fraction is
[0089] (4) Adjust the fractional model parameters according to the difference between the estimated inverse optimal shift fraction and the real inverse optimal shift fraction (the difference between X0 and ); define the data prediction loss function Minimizing the loss function can improve the accuracy of the estimated inverse optimal shift fraction, where L(·) represents the difference between the two, and the mean square error can usually be used; the training target is to minimize the loss function , gradient back propagation is performed to update the parameters
[0090] (5) Repeat (1) to (4) until the training requirement such as the maximum number of iterations is met, and stop the training. Select the optimal parameters in the training process The corresponding score model is taken as the target score model.
[0091] S5, pure speech generation: Given noisy speech, the trained target score model is used to parameterize the inverse optimal offset score of the mean inversion Schrödinger bridge, and the inverse generation process is performed to generate pure speech. The initial state of the inverse generation process is the noisy speech, and the final state obtained after multiple iterations is the estimated pure speech.
[0092] In this embodiment, given the noisy speech y tgt , according to the preprocessing method of S2, the noisy speech is sequentially normalized, Fourier transformed, and amplitude compressed without intercepting subsegments to obtain the corresponding spectral complex matrix Y tgt . Let the time list of inverse generation be [t0, t1,..., t i ,..., t N ], where 0<t i <1, t0=0.03, t N =1, i=0,1,…,N, and N is the number of iterations. Let j=N, and the initial state The inverse generation is performed according to the following process:
[0093] (1) Input the sample j at time t to the target score model to obtain
[0094] (2) Estimate the inverse optimal offset score by Solve the inverse stochastic differential equation Calculate A simple solution is that the sample j-1 at time t obeys the posterior probability distribution of the Schrödinger bridge Under the condition that the sample is known, use the predicted of the score model to replace X0, and sample from
[0095] (3) Let j=j-1, if j>0, repeat (1) to (2), and if j=0, stop the inverse generation; the sample at time t0 is taken as the spectral complex matrix of the pure speech after noisy speech enhancement, and after inverse amplitude compression, inverse Fourier transform, and inverse normalization operations, the waveform of the pure speech is finally obtained.
[0096] The reverse generation process can directly sample step by step through the posterior probability distribution in (2) above to simulate the mean-reversing Schrodinger bridge, because the score model is trained by indirectly matching the real score by using the prediction data method. If the score model is trained by directly matching the real score or indirectly matching the real score by using the prediction noise method, the reverse generation process is similar to the above process, which is not described here.
[0097] In this embodiment, the discriminant model and the score model are trained using the VoiceBank+DEMAND and TIMIT+WHAM! data sets, respectively, to achieve speech enhancement. VoiceBank+DEMAND is a public speech enhancement benchmark data set, and the sample pairs have been pre-constructed. The TIMIT+WHAM! data set is generated by mixing clean speech from TIMIT with noise speech from WHAM! at a random signal-to-noise ratio to construct sample pairs according to the steps of S1; wherein the signal-to-noise ratio is uniformly sampled between -6db and 14db.
[0098] In this embodiment, the speech enhancement performance is compared based on the VoiceBank+DEMAND data set and three mainstream single-channel speech enhancement methods based on diffusion models, CDiffuSE, SGMSE+, and StoRM. The results are shown in Table 1.
[0099] Table 1 shows the speech enhancement performance comparison results of three mainstream methods on the VoiceBank+DEMAND data set
[0100]
[0101] PESQ is an objective speech quality perception evaluation index, and the larger the value is, the better; ESTOI is an objective short-time speech intelligibility evaluation index, and the larger the value is, the better; SI-SDR is used to evaluate the degree of speech distortion, and the larger the value is, the better; DNSMOS is a subjective speech quality perception evaluation index, and the larger the value is, the better.
[0102] As shown in Table 1, the method of this embodiment has improved in PESQ, ESTOI, SI-SDR, and DNSMOS compared with other methods, proving the advanced performance of the method. At the same time, the method only calls the neural network 3 times in the reverse generation process, and the calculation cost is significantly reduced.
[0103] In this embodiment, the generalization is compared based on the TIMIT+WHAM! data set and the VoiceBank+DEMAND data set, and two mainstream single-channel speech enhancement methods based on diffusion models, SGMSE+ and StoRM. The results are shown in Table 2.
[0104] Table 2: Generalization comparison results of two mainstream methods on TIMIT+WHAM! dataset
[0105]
[0106] where WER is the word error rate of the recognized text on the downstream speech recognition task, and the lower the value is, the better. Domain matched means that the method is trained on the TIMIT+WHAM! dataset and validated on the test set of the TIMIT+WHAM! dataset. Domain mismatched means that the method is trained on the VoiceBank+DEMAND dataset and validated on the test set of the TIMIT+WHAM! dataset.
[0107] As shown in Table 2, compared with other methods, the method of the present embodiment achieves the best performance in both domain matching and domain mismatching, proving the generalization ability of the method.
[0108] In the present embodiment, an ablation experiment is performed based on the VoiceBank+DEMAND dataset, and the results are shown in Table 3.
[0109] Table 3: Ablation experiment results on the VoiceBank+DEMAND dataset
[0110]
[0111] The second to fifth rows are all improvements on the traditional Schrödinger bridge method. In the methods of the second to fourth rows, the discriminative model exists in both the training process of the score model and the reverse generation process. The method of the fifth row, i.e. the method of the present embodiment, only exists in the training process of the score model.
[0112] As shown in Table 3, the method of the present embodiment achieves the highest objective speech perceptual quality (PESQ) and intelligibility (ESTOI) without the assistance of the discriminative model in the reverse generation process, and the remaining performance is also improved compared with the traditional Schrödinger bridge method.
[0113] Based on the same inventive concept, the present embodiment also provides a single-channel speech enhancement device based on the mean inversion Schrödinger bridge, comprising:
[0114] The sample pair construction module is composed of a pure speech sample and a noisy speech sample to form a set of speech sample pairs;
[0115] The sample processing module pre-processes and Fourier transforms the speech samples in the set of speech sample pairs to obtain a set of spectral complex number matrices;
[0116] The discriminant model construction module constructs a discriminant model, takes the spectral complex number matrix of the noisy speech as input, takes the estimated spectral complex number matrix of the clean speech as output, adjusts parameters of the discriminant model based on a difference between the estimated spectral complex number matrix of the clean speech and the real spectral complex number matrix of the clean speech, and takes the discriminant model corresponding to the optimal parameters as the target discriminant model when the discriminant model meets preset training requirements.
[0117] The score model construction module constructs a score model, parameterizes the inverse optimal shift score of the mean inversion Schrodinger bridge, estimates the inverse optimal shift score in the inverse generation process, adjusts parameters of the score model based on a difference between the estimated inverse optimal shift score and the real inverse optimal shift score, and takes the score model corresponding to the optimal parameters as the target score model when the score model meets preset training requirements.
[0118] The clean speech generation module parameterizes the inverse optimal shift score of the mean inversion Schrodinger bridge by using the trained target score model for a given noisy speech, performs the inverse generation process, and generates clean speech.
[0119] It should be noted that the single-channel speech enhancement device based on the mean inversion Schrodinger bridge provided in the above embodiments should be illustrated by the division of the above functional modules when performing speech enhancement and generating clean speech, and the above functions can be completed by different functional modules according to needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the above-described functions. In addition, the single-channel speech enhancement device based on the mean inversion Schrodinger bridge and the single-channel speech enhancement method based on the mean inversion Schrodinger bridge provided in the above embodiments belong to the same concept, and the specific implementation process and effects are described in detail in the single-channel speech enhancement method based on the mean inversion Schrodinger bridge, which will not be repeated here.
[0120] Based on the same inventive concept, the present embodiment further provides a computing device including a memory and one or more processors, the memory storing executable code, and the one or more processors executing the executable code to implement the above-described single-channel speech enhancement method based on the mean inversion Schrodinger bridge.
[0121] The computing device provided by the embodiment comprises, in addition to the processor and the memory, internal bus, network interface, memory and other hardware required by the business. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to realize the single-channel speech enhancement method based on the mean inversion Schrodinger bridge described in S1-S5. Of course, in addition to the software implementation, the present application does not exclude other implementation manners, such as logic device or the combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but also can be hardware or logic device.
[0122] Based on the same inventive concept, the embodiment further provides a computer readable storage medium having a program stored thereon, and the program is executed by a processor to realize the single-channel speech enhancement method based on the mean inversion Schrodinger bridge.
[0123] In the embodiment, the computer readable medium includes permanent and non-permanent, removable and non-removable media, which can realize information storage by any method or technology. The information can be computer readable instructions, data structures, program modules or other data.
[0124] The embodiments are preferred embodiments of the present application, but the present application is not limited to the above embodiments, and any obvious improvement, replacement or modification made by those skilled in the art without departing from the essential content of the present application shall fall within the protection scope of the present application.
Claims
1. A single-channel speech enhancement method based on mean inversion Schrdinger bridge, characterized in that: a clean speech sample and a noisy speech sample are mixed at different signal-to-noise ratios to obtain a noisy speech sample, and a set of speech sample pairs is formed by the clean speech sample and the noisy speech sample; the speech samples in the set of speech sample pairs are preprocessed, and then Fourier transform is performed to obtain a set of spectral complex number matrices; a discriminant model is constructed, with the spectral complex number matrix of the noisy speech as the input and the estimated spectral complex number matrix of the clean speech as the output; parameters of the discriminant model are adjusted based on the difference between the estimated spectral complex number matrix of the clean speech and the actual spectral complex number matrix of the clean speech, and when the discriminant model meets the preset training requirements, the discriminant model corresponding to the optimal parameters is taken as the target discriminant model; a score model is constructed, and the score model is used to parameterize the inverse optimal offset score of the mean inversion Schrdinger bridge to estimate the inverse optimal offset score in the inverse generation process; parameters of the score model are adjusted based on the difference between the estimated inverse optimal offset score and the actual inverse optimal offset score, and when the score model meets the preset training requirements, the score model corresponding to the optimal parameters is taken as the target score model; given a noisy speech, the inverse optimal offset score of the mean inversion Schrdinger bridge is parameterized by using the trained target score model, and the inverse generation process is performed to generate a clean speech.
2. The single-channel speech enhancement method based on mean inversion Sjohdahl bridge according to claim 1, characterized in that, The obtaining process of the spectrum complex matrix pair is as follows: for any speech sample pair (x, y), sub-segments of x and y are randomly intercepted, the maximum value of the sampling values in y is counted, the maximum value is used for normalization of x and y respectively to obtain x' and y', Fourier transform is performed to obtain the spectrum complex matrix X corresponding to x' and y' 0 , Y 0 corresponding to y', amplitude compression is performed on X 0 and Y 0 again to obtain the spectrum complex matrix pair (X, Y), x represents the original waveform of pure speech, and y represents the original waveform of noisy speech.
3. The single-channel speech enhancement method based on mean inversion Sjohdahl bridge according to claim 2, characterized in that, The training process of the discriminant model is: inputting a spectral complex matrix Y of noisy speech into a discriminant model D θ outputting an estimate of a spectral complex matrix of clean speech Reconstruction loss function is constructed based on clean speech where L(·) represents the difference between the two; to minimize the loss function For the training target, gradient back propagation is performed to update the parameters θ; stop training when the preset training requirements are met; Selecting the optimal parameter θ in the training process * The corresponding discriminant model is taken as the target discriminant model.
4. The single-channel speech enhancement method based on mean inversion Sjohdahl bridge according to claim 3, characterized in that, The acquisition process of the mean inversion Schrdinger bridge is: For a spectral complex matrix pair (X, Y), let the posterior mean of X given Y be a function of The variable From a conditional posterior distribution It is observed that the target discriminant model Approximating the conditional posterior distribution, i.e. The mean-reverting process is expressed at a specific time t as a state sample x subject to an edge Gaussian distribution where denotes a uniform distribution on the interval [0, 1], Define a Schrödinger bridge process that bridges a noisy speech distribution and a clean speech distribution, and represent the state sample of this Schrödinger bridge process at time t as X. t Sample X t Follows a marginal Gaussian distribution q(X) t |X0,X1), and X0=X, X1=Y; at time t, assume X t by As a conditional expression, the approximate marginal Gaussian distribution of the Schrödinger bridge is derived as follows: wherein is a point estimate of The Schrödinger bridge process that is introduced into the mean reverting process is called mean reverting Schrödinger bridge, which replaces the boundary condition X0= X at time t of the Schrödinger bridge by 5. The single-channel speech enhancement method based on mean inversion Sjohdahl bridge according to claim 4, characterized in that, The mean inversion process is a Schrdinger bridge process of bridging the posterior mean distribution and the clean speech distribution. 6.The single-channel speech enhancement method based on mean inversion Schrdinger bridge according to claim 4, characterized in that: At time t, the marginal Gaussian distribution of the mean inversion process is: wherein, denotes the mean of the edge distribution of the mean-reverting process at time t, denotes the variance of the edge distribution of the mean-reverting process at time t, a t and σ t is the noise schedule of the mean-reverting process, signal-to-noise ratio 7.The single-channel speech enhancement method based on mean inversion Schrdinger bridge according to claim 6, characterized in that: The mean inversion Schrdinger bridge is described by the following two stochastic differential equations: where f is the drift coefficient, g is the diffusion coefficient, W and is a standard Wiener process, Ψ t and is the Schrodinger factor, the fractional and represent the forward optimal drift fraction, the backward optimal drift fraction; At time t, the marginal Gaussian distribution of the mean inversion Schrdinger bridge is: wherein, denotes the mean of the mean-reverting Schrödinger bridge edge distribution at time t, denotes the variance of the mean-reverting Schrödinger bridge edge distribution at time t, and is the noise schedule of the mean-reverting Schrödinger bridge, determined by f and g in the stochastic differential equation above, the signal-to-noise ratio 8. The single-channel speech enhancement method based on mean inversion Sjohdahl bridge according to claim 7, characterized in that, The training process of the score model is: inputting a spectral complex matrix Y of noisy speech into a target discriminant model D θ* , outputting Randomly sampled time instant Edge Gaussian distribution of the mean-reverting process at time t Compute the mean of this distribution to get the point estimate at time t Next, the edge Gaussian distribution of the mean-reverting Schrödinger bridge at time t Sample to get a sample at time t where is random noise; Finally, at time t, sample X t Input score model Output The true reverse-optimal shift score at this time is The estimated reverse-optimal shift score is Adjust the score model parameters according to the difference between the estimated reverse optimal offset score and the true reverse optimal offset score, which is reflected in the difference between X0 and ; Specifically: define the data prediction loss function to minimize the loss function , Gradient back propagation is carried out to update the parameters stop training when the preset training requirements are met; Selecting optimal parameters during training The corresponding score model as the target score model.
9. The single-channel speech enhancement method based on mean inversion Sjohdahl bridge according to claim 8, characterized in that, The inverse generation process specifically includes: At time t j , sample input target score model, obtain By Estimating inverse optimal shift scores Solving inverse stochastic differential equations Computing Let j = j - 1, if j > 0, repeat the above process, if j = 0, stop the inverse generation, the sample at time t0 As the spectral complex matrix of the pure speech after the noisy speech enhancement, after the inverse amplitude compression, the inverse Fourier transform and the inverse normalization, the waveform of the pure speech is finally obtained.
10. A device for implementing the single channel speech enhancement method based on the mean inversion Sjohdahl bridge according to any one of claims 1 to 9, characterized in that, including: a sample pair module, which forms a set of speech sample pairs by a clean speech sample and a noisy speech sample; a sample processing module, which pre-processes the speech samples in the set of speech sample pairs and performs Fourier transform to obtain a set of spectral complex number matrices; a discriminant model construction module, which constructs a discriminant model, estimates the spectral complex number matrix of the clean speech, and adjusts the parameters of the model based on the difference between the estimated spectral complex number matrix of the clean speech and the actual spectral complex number matrix of the clean speech, and when the model meets the preset training requirements, the discriminant model corresponding to the optimal parameters is taken as the target discriminant model; The score model construction module constructs a score model, estimates a reverse optimal offset score in a reverse generation process, adjusts parameters of the model based on a difference between the estimated reverse optimal offset score and a true reverse optimal offset score, and takes a score model corresponding to optimal parameters as a target score model when the model meets preset training requirements. The pure speech generation module parameterizes the reverse optimal offset score of the mean inversion Schrödinger bridge by using the trained target score model for a given noisy speech, performs a reverse generation process, and generates pure speech.
Citation Information
Patent Citations
Speech enhancement method and system based on psychoacoustic domain weight loss function
CN113744749A
Speech enhancement algorithm based on improved phase spectrum compensation and full convolutional neural network
CN114242099A