Speech Enhancement Method and Device
By predicting and calculating the enhanced pseudo-real and pseudo-fiction spectra in the speech enhancement model, combined with simulated phase calculation formulas, the problem of lack of phase spectral enhancement in the existing technology is solved, and the speech enhancement effect with high quality and high signal-to-noise ratio is achieved.
Patent Information
- Application Number
- CN202310573048.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-05-17
AI Technical Summary
The existing speech enhancement model lacks enhancement to the phase spectrum, resulting in poor quality of the calculated enhanced speech waveform, low signal-to-noise ratio, and poor enhancement effect.
By obtaining the phase spectrum and amplitude spectrum of the noisy speech waveform, the enhanced pseudo-real and pseudo-fiction spectrum is predicted using the preset speech enhancement model, and the calculation is performed based on the simulated phase calculation formula to obtain the enhanced phase spectrum with the value range limited to the main value range, and a high-quality enhanced speech waveform is calculated based on the enhanced amplitude spectrum.
The quality and signal-to-noise ratio of the enhanced speech waveform are improved, the enhancement effect on the noisy speech waveform is improved, and the problem of unpredictable enhancement phase spectrum caused by phase winding characteristics is avoided.
Smart Images

Figure CN116386653B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech processing, and more specifically, to a speech enhancement method and device. Background Art
[0002] In real-life scenarios, the speech captured by devices is always inevitably interfered by noise, which greatly affects the practical applications of devices such as communication and hearing aids. Therefore, it is necessary to enhance the speech. Speech enhancement aims to recover a clean speech waveform from a speech waveform interfered by noise.
[0003] With the development of deep learning technology, usually a speech enhancement model is trained, and the trained speech enhancement model is used to enhance a noisy speech waveform to obtain an enhanced speech waveform. Existing speech enhancement models usually enhance the noisy amplitude spectrum of the noisy speech waveform, and then calculate the enhanced speech waveform based on the enhanced amplitude spectrum and the noisy phase spectrum. However, since the enhanced speech waveform is calculated based on the noisy phase spectrum and there is a lack of enhancement of the phase spectrum, the quality of the calculated enhanced speech waveform is poor, the signal-to-noise ratio is low, and the enhancement effect on the noisy speech waveform is poor. Summary of the Invention
[0004] In view of this, the present application provides a speech enhancement method and device, which are used to solve the problem that in the existing speech enhancement method, due to the lack of enhancement of the phase spectrum, the quality of the calculated enhanced speech waveform is poor, the signal-to-noise ratio is low, and the enhancement effect on the noisy speech waveform is poor.
[0005] To achieve the above object, the following solutions are proposed:
[0006] A speech enhancement method includes:
[0007] Obtain the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform;
[0008] Use a preset speech enhancement model to process the noisy phase spectrum and the noisy amplitude spectrum to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum;
[0009] The speech enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and predict an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, and calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with the value range limited within the principal value range;
[0010] The enhanced speech waveform corresponding to the noisy speech waveform is calculated based on the enhanced phase spectrum and the enhanced amplitude spectrum.
[0011] Preferably, after obtaining the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform, the method further includes:
[0012] Performing amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum;
[0013] The process of using a preset speech enhancement model to process the noisy phase spectrum and the noisy amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum includes:
[0014] Inputting the noisy phase spectrum and the compressed noisy amplitude spectrum into a preset speech enhancement model, so as to predict, using the speech enhancement model, a compressed enhanced amplitude spectrum mask corresponding to the compressed noisy amplitude spectrum, and multiplying the compressed enhanced amplitude spectrum mask point by point by the compressed noisy amplitude spectrum to obtain a compressed enhanced amplitude spectrum, and decompressing the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0015] Preferably, performing amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum includes:
[0016] Calculating the c-th power of the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum, where c is a preset compression factor;
[0017] The process of decompressing the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum includes:
[0018] Calculating the 1 / c-th power of the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0019] Preferably, the simulated phase calculation formula is:
[0020]
[0021] where p is the independent variable representing the enhanced pseudo-real part spectrum, q is the independent variable representing the enhanced pseudo-imaginary part spectrum, when p≥0, Sgn * (p)=1, when p<0, Sgn * (p)= -1, when q≥0, Sgn * (q)=1, when q<0, Sgn * (q)= -1.
[0022] Preferably, the training process of the speech enhancement model includes:
[0023] Obtain the training noisy phase spectrum and training noisy amplitude spectrum of the training noisy speech waveform, and the training clean phase spectrum and training clean amplitude spectrum of the training clean speech waveform corresponding to the training noisy speech waveform;
[0024] Input the training noisy phase spectrum and the noisy amplitude spectrum into the speech enhancement model, so as to use the speech enhancement model to predict the training enhanced pseudo real part spectrum and training enhanced pseudo imaginary part spectrum corresponding to the training noisy phase spectrum based on the input training noisy phase spectrum and the training noisy amplitude spectrum, and predict the training enhanced amplitude spectrum corresponding to the training noisy amplitude spectrum, and calculate the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the simulated phase calculation formula to obtain a training enhanced phase spectrum with the value range limited within the principal value range;
[0025] Calculate the amplitude loss based on the training clean amplitude spectrum and the training enhanced amplitude spectrum;
[0026] Calculate the phase loss based on the training clean phase spectrum and the training enhanced phase spectrum;
[0027] Train the speech enhancement model based on the target loss until the set training end condition is met, where the target loss includes the amplitude loss and the phase loss.
[0028] Preferably, the phase loss includes the instantaneous phase loss between the training clean phase spectrum and the training enhanced phase spectrum;
[0029] The calculation process of the instantaneous phase loss includes:
[0030] Calculate the first true distance of the instantaneous phase between the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function;
[0031] Calculate the instantaneous phase loss based on the first true distance.
[0032] Preferably, the phase loss further includes the group delay loss and instantaneous angular frequency loss between the training clean phase spectrum and the training enhanced phase spectrum;
[0033] The calculation process of the group delay loss includes:
[0034] Calculate the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis;
[0035] Calculate the second true distance of the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function;
[0036] Calculate the group delay loss based on the second true distance;
[0037] The calculation process of the instantaneous angular frequency loss includes:
[0038] Calculating the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis;
[0039] Based on the unwrapping function, calculating the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis;
[0040] Calculating the instantaneous angular frequency loss based on the third true distance.
[0041] Preferably, the unwrapping function is f AW (m), where f AW (m) = m - 2π·round(m / 2π), m is the independent variable, and round is the rounding function;
[0042] The calculating of the first true distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function includes:
[0043] Calculating the first true distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum X p is the training clean phase spectrum, is the training enhanced phase spectrum;
[0044] The calculating of the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function includes:
[0045] Calculating the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis Δ DF X p is the differential spectrum of the training clean phase spectrum along the frequency axis, is the differential spectrum of the training enhanced phase spectrum along the frequency axis;
[0046] The calculating of the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis based on the unwrapping function includes:
[0047] Calculating the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis Δ DT X p is the differential spectrum of the training clean phase spectrum along the time axis, It is the differential spectrum of the training enhanced phase spectrum along the time axis.
[0048] Preferably, calculating the instantaneous phase loss based on the first true distance includes:
[0049] Calculating the instantaneous phase loss based on the following formula
[0050] is a function for calculating the average value, ||||1 is the formula of the first norm, and t1 is the first true distance;
[0051] Calculating the group delay loss based on the second true distance includes:
[0052] Calculating the group delay loss based on the following formula
[0053] t2 is the second true distance;
[0054] Calculating the instantaneous angular frequency loss based on the third true distance includes:
[0055] Calculating the instantaneous angular frequency loss based on the following formula
[0056] t3 is the third true distance.
[0057] Preferably, before training the speech enhancement model based on the target loss until the set training end condition is met, it further includes:
[0058] Performing short-time Fourier transform on the training clean speech waveform to obtain the training clean short-time complex spectrum;
[0059] Reconstructing the training enhanced phase spectrum and the training enhanced amplitude spectrum to obtain the training enhanced short-time complex spectrum;
[0060] Calculating the short-time complex spectrum loss based on the training clean short-time complex spectrum and the training enhanced short-time complex spectrum;
[0061] The target loss further includes the short-time complex spectrum loss.
[0062] Preferably, after reconstructing the training enhanced phase spectrum and the training enhanced amplitude spectrum to obtain the training enhanced short-time complex spectrum, it further includes:
[0063] Performing inverse short-time Fourier transform on the training enhanced short-time complex spectrum to obtain the training enhanced speech waveform;
[0064] Calculate a waveform loss based on the training clean speech waveform and the training enhanced speech waveform;
[0065] The objective loss further includes the waveform loss.
[0066] Preferably, the speech enhancement model includes:
[0067] An encoder, a TS-Conformer module, a phase decoder, and an amplitude decoder;
[0068] The encoder is used to encode the training noisy phase spectrum and the training noisy amplitude spectrum to obtain high-dimensional time-frequency domain features;
[0069] The TS-Conformer module is used to process the high-dimensional time-frequency domain features to obtain processed high-dimensional time-frequency domain features;
[0070] The phase decoder is used to predict the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the processed high-dimensional time-frequency domain features, and calculate the training enhanced phase spectrum based on the simulated phase calculation formula for the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum;
[0071] The amplitude decoder is used to predict the training enhanced amplitude spectrum based on the processed high-dimensional time-frequency domain features.
[0072] Preferably, the encoder, the TS-Conformer module, the phase decoder, and the amplitude decoder are combined into a generator, and the speech enhancement model further includes a discriminator;
[0073] The discriminator is used to discriminate between the training clean amplitude spectrum and the training enhanced amplitude spectrum to obtain a discrimination index corresponding to the enhanced speech waveform;
[0074] Before training the speech enhancement model based on the objective loss until the set training end condition is satisfied, it further includes:
[0075] Calculate the discriminator loss;
[0076] Calculate an objective metric loss based on the discrimination index;
[0077] The objective loss further includes the objective metric loss and the discriminator loss;
[0078] Training the speech enhancement model based on the objective loss until the set training end condition is satisfied includes:
[0079] Train the discriminator based on the discriminator loss, and train the generator based on the amplitude loss, the phase loss, the short-time complex spectrum loss, the waveform loss, and the objective metric loss until the set training end condition is met.
[0080] Preferably, the discrimination metric is the discrimination normalized PESQ metric, and the discriminator is further configured to discriminate a pair of the training clean amplitude spectra to obtain the discrimination normalized PESQ metric corresponding to the training clean speech waveform;
[0081] Calculating the discriminator loss includes:
[0082] Calculate the discriminator loss based on the following formula
[0083]
[0084] where X m is the training clean amplitude spectrum, D(X m , X m ) is the discrimination normalized PESQ metric corresponding to the training clean speech waveform, is the training noisy amplitude spectrum, is the discrimination normalized PESQ metric corresponding to the enhanced speech waveform, Q PESQ is the true normalized PESQ metric corresponding to the enhanced speech waveform determined in advance;
[0085] Calculating the objective metric loss based on the discrimination metric includes:
[0086] Calculate the objective metric loss based on the following formula
[0087] is the function for calculating the average value, and ||||2 is the two-norm formula.
[0088] A speech enhancement device, comprising:
[0089] A noisy phase spectrum and noisy amplitude spectrum acquisition unit, configured to acquire a noisy phase spectrum and a noisy amplitude spectrum of a noisy speech waveform;
[0090] An enhanced phase spectrum and enhanced amplitude spectrum acquisition unit, configured to process the noisy phase spectrum and the noisy amplitude spectrum by using a preset speech enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum;
[0091] The speech enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and to predict an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, and to calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with a value range limited within the principal value range;
[0092] An enhanced speech waveform acquisition unit for calculating an enhanced speech waveform corresponding to the noisy speech waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum.
[0093] As can be seen from the above technical solution, the speech enhancement method provided by the embodiment of the present application acquires the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform, processes the noisy phase spectrum and the noisy amplitude spectrum by using a preset speech enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, calculates an enhanced speech waveform corresponding to the noisy speech waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum, and the speech enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and to predict an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, and to calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with a value range limited within the principal value range. It not only enhances the noisy amplitude spectrum of the noisy speech waveform, but also enhances the noisy phase spectrum of the noisy speech waveform, so that the enhanced speech waveform calculated according to the enhanced phase spectrum and the enhanced amplitude spectrum has high quality and high signal-to-noise ratio, improving the enhancement effect on the noisy speech waveform. Moreover, when the speech enhancement model predicts the enhanced phase spectrum, it does not directly enhance the noisy phase spectrum, but predicts the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum, and calculates the enhanced phase spectrum with a value range limited within the principal value range based on the preset simulated phase calculation formula, realizing the prediction of the enhanced phase spectrum, avoiding the problem that the enhanced phase spectrum cannot be predicted due to the winding characteristic of the phase, and further making the enhanced speech waveform calculated according to the enhanced phase spectrum and the enhanced amplitude spectrum have high quality and high signal-to-noise ratio, greatly improving the enhancement effect on the noisy speech waveform. Description of the Drawings
[0094] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0095] Figure 1 Flow chart of a voice enhancement method disclosed in an embodiment of the present application;
[0096] Figure 2 Schematic diagram of components of a voice enhancement model disclosed in an embodiment of the present application;
[0097] Figure 3 Schematic diagram of the training process of a voice enhancement model disclosed in an embodiment of the present application;
[0098] Figure 4 Schematic diagram of the structure of a voice enhancement device disclosed in an embodiment of the present application;
[0099] Figure 5 Hardware structure block diagram of a voice enhancement device disclosed in an embodiment of the present application. Specific implementation manners
[0100] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0101] The solution of the present application can be implemented based on a terminal with data processing capabilities, and the terminal can be a computer, a server, a cloud, etc.
[0102] An embodiment of the present application provides a voice enhancement solution. Next, the voice enhancement method of the present application will be described through the attached Figure 1 As shown in the following, the method may include: Figure 1 As shown in the following, the method may include:
[0103] Step S100, obtaining the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform.
[0104] Specifically, the noisy speech waveform is a speech waveform interfered by noise. The phase spectrum is a curve of phase changing with frequency, and the amplitude spectrum is a curve of amplitude changing with frequency. The phase spectrum and the amplitude spectrum of the speech waveform are important features of the speech waveform. Therefore, obtaining the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform is to enhance the noisy phase spectrum and the noisy amplitude spectrum, so as to enhance the noisy speech waveform.
[0105] Step S110: Process the noisy phase spectrum and the noisy amplitude spectrum by using a preset voice enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum. The voice enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum, and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with the value range limited within the principal value range.
[0106] Specifically, the phase spectrum and the amplitude spectrum of a voice waveform are important features of the voice waveform. Therefore, an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum and an enhanced phase spectrum corresponding to the noisy phase spectrum can be predicted by using a voice enhancement model, and then an enhanced voice waveform can be calculated based on the enhanced amplitude spectrum and the enhanced phase spectrum, which can make the enhanced voice waveform have high quality and high signal-to-noise ratio. Since the phase has a wrapping property, if the noisy phase spectrum is directly enhanced, the wrapped enhanced phase spectrum corresponding to the noisy phase spectrum cannot be predicted. Therefore, in the embodiment of the present application, the voice enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum, and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with the value range limited within the principal value range. Since the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum are predicted, and the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum are calculated based on a preset simulated phase calculation formula to obtain an enhanced phase spectrum with the value range limited within the principal value range, the finally obtained enhanced phase spectrum is a wrapped phase spectrum, which not only realizes the prediction of the enhanced amplitude spectrum but also realizes the prediction of the enhanced phase spectrum, avoiding the problem that the enhanced phase spectrum cannot be predicted due to the wrapping property of the phase. Among them, the principal value range is the value range threshold interval of the phase.
[0107] Step S120: Calculate an enhanced voice waveform corresponding to the noisy voice waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum.
[0108] Specifically, after processing the noisy phase spectrum and the noisy amplitude spectrum by using a preset voice enhancement model to obtain an enhanced phase spectrum and an enhanced amplitude spectrum, an enhanced voice waveform corresponding to the noisy voice waveform can be calculated according to the enhanced phase spectrum and the enhanced amplitude spectrum.
[0109] The voice enhancement method provided by the embodiment of the present application obtains the noisy phase spectrum and noisy amplitude spectrum of a noisy voice waveform, processes the noisy phase spectrum and noisy amplitude spectrum by using a preset voice enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, calculates an enhanced voice waveform corresponding to the noisy voice waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum. The voice enhancement model is configured to predict an enhanced pseudo-real part spectrum and an enhanced pseudo-imaginary part spectrum corresponding to the noisy phase spectrum based on the input noisy phase spectrum and noisy amplitude spectrum, and predict an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, and calculate the enhanced pseudo-real part spectrum and the enhanced pseudo-imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with a value range limited within the principal value range. It not only enhances the noisy amplitude spectrum of the noisy voice waveform, but also enhances the noisy phase spectrum of the noisy voice waveform, so that the enhanced voice waveform calculated according to the enhanced phase spectrum and the enhanced amplitude spectrum has high quality and high signal-to-noise ratio, improving the enhancement effect on the noisy voice waveform. Moreover, since when the voice enhancement model predicts the enhanced phase spectrum, it does not directly enhance the noisy phase spectrum, but predicts the enhanced pseudo-real part spectrum and the enhanced pseudo-imaginary part spectrum corresponding to the noisy phase spectrum, and calculates the enhanced phase spectrum with a value range limited within the principal value range based on the preset simulated phase calculation formula, realizing the prediction of the enhanced phase spectrum and avoiding the problem that the enhanced phase spectrum cannot be predicted due to the winding characteristic of the phase. This further makes the enhanced voice waveform calculated according to the enhanced phase spectrum and the enhanced amplitude spectrum have high quality and high signal-to-noise ratio, greatly improving the enhancement effect on the noisy voice waveform.
[0110] Optionally, considering that it is relatively difficult to directly predict the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, while it is relatively easy to predict the enhanced amplitude spectrum mask corresponding to the noisy amplitude spectrum, so the enhanced amplitude spectrum mask corresponding to the noisy amplitude spectrum can be predicted, and then the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum can be calculated based on the enhanced amplitude spectrum mask. Based on this, after the above step S100 obtains the noisy phase spectrum and noisy amplitude spectrum of the noisy voice waveform, the following steps may further be included:
[0111] Perform amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum.
[0112] Specifically, considering that most of the values of the amplitude spectrum mask are between 0 and 1, and a small part is between 0 and infinity. If the amplitude spectrum mask is restricted to between 0 and 1 when predicting the amplitude spectrum mask, for the small part of the amplitude spectrum mask between 0 and infinity, an accurate enhanced amplitude spectrum mask cannot be predicted. If the amplitude spectrum mask is not restricted when predicting the amplitude spectrum mask and is allowed to be between 0 and infinity, the range is very large and difficult to predict. Therefore, amplitude compression is performed on the noisy amplitude spectrum to obtain the compressed noisy amplitude spectrum. Amplitude compression of the noisy amplitude spectrum also compresses the corresponding amplitude spectrum mask, limiting the value of the amplitude spectrum mask to a small range, which is easy to predict.
[0113] Optionally, the c-th power of the noisy amplitude spectrum can be calculated to obtain the compressed noisy amplitude spectrum, where c is a preset compression factor, and an example is 0.3.
[0114] Based on this, the process of using the preset speech enhancement model to process the noisy phase spectrum and the noisy amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum may include:
[0115] Input the noisy phase spectrum and the compressed noisy amplitude spectrum into the preset speech enhancement model to use the speech enhancement model to predict the compressed enhanced amplitude spectrum mask corresponding to the compressed noisy amplitude spectrum, and multiply the compressed enhanced amplitude spectrum mask point by point by the compressed noisy amplitude spectrum to obtain the compressed enhanced amplitude spectrum, and decompress the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0116] Specifically, input the noisy phase spectrum and the compressed noisy amplitude spectrum into the preset speech enhancement model to use the speech enhancement model to predict the compressed enhanced amplitude spectrum mask corresponding to the compressed noisy amplitude spectrum. Multiply the compressed enhanced amplitude spectrum mask point by point by the compressed noisy amplitude spectrum to obtain the compressed enhanced amplitude spectrum, and then decompress the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0117] Optionally, if the compressed noisy amplitude spectrum is obtained by calculating the c-th power of the noisy amplitude spectrum, the 1 / c-th power of the compressed enhanced amplitude spectrum can be calculated to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0118] In the embodiments of the present application, the noisy magnitude spectrum is amplitude-compressed to obtain the compressed noisy magnitude spectrum. The noisy phase spectrum and the compressed noisy magnitude spectrum are input into a speech enhancement model to predict, using the speech enhancement model, a compressed enhanced magnitude spectrum mask corresponding to the compressed noisy magnitude spectrum, and the compressed enhanced magnitude spectrum mask is multiplied point by point with the compressed noisy magnitude spectrum to obtain a compressed enhanced magnitude spectrum. The compressed enhanced magnitude spectrum is decompressed to obtain the enhanced magnitude spectrum corresponding to the noisy magnitude spectrum. The enhanced magnitude spectrum is predicted by using the predicted magnitude mask method, and the noisy magnitude spectrum is amplitude-compressed to narrow the value range of the compressed enhanced magnitude spectrum mask, making it easier to predict and improving the efficiency of predicting the enhanced magnitude spectrum.
[0119] Optionally, the above simulation phase calculation formula can be:
[0120]
[0121] where p is the independent variable representing the enhanced pseudo real part spectrum, q is the independent variable representing the enhanced pseudo imaginary part spectrum. When p ≥ 0, Sgn * (p) = 1, when p < 0, Sgn * (p) = -1, when q ≥ 0, Sgn * (q) = 1, when q < 0, Sgn * (q) = -1.
[0122] Specifically, it can be seen from the above formula that the above formula simulates the calculation process from the real part and the imaginary part to the phase spectrum, and strictly limits the value range within the principal value interval (-π, π], realizing the prediction of the wrapped phase spectrum.
[0123] Optionally, the training process of the above speech enhancement model may include:
[0124] Obtain the training noisy phase spectrum, training noisy magnitude spectrum of the training noisy speech waveform, and the training clean phase spectrum, training clean magnitude spectrum of the training clean speech waveform corresponding to the training noisy speech waveform.
[0125] Specifically, if a speech enhancement model is to be trained, first, the training noisy phase spectrum, training noisy magnitude spectrum of the training noisy speech waveform, and the training clean phase spectrum, training clean magnitude spectrum of the training clean speech waveform corresponding to the training noisy speech waveform need to be obtained.
[0126] Input the training noisy phase spectrum and the noisy amplitude spectrum into the speech enhancement model, so as to use the speech enhancement model to predict the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum corresponding to the training noisy phase spectrum based on the input training noisy phase spectrum and the training noisy amplitude spectrum, and to predict the training enhanced amplitude spectrum corresponding to the training noisy amplitude spectrum, and calculate the training enhanced phase spectrum with the value range limited within the principal value range based on the simulated phase calculation formula for the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum.
[0127] Calculate the amplitude loss based on the training clean amplitude spectrum and the training enhanced amplitude spectrum.
[0128] Optionally, the amplitude loss can be calculated based on the following formula
[0129]
[0130] where is the function for calculating the average value, ||||2 is the formula for the two-norm, and X m is the training clean amplitude spectrum, is the training enhanced amplitude spectrum.
[0131] Calculate the phase loss based on the training clean phase spectrum and the training enhanced phase spectrum.
[0132] Specifically, in order to train the speech enhancement model, after obtaining the training enhanced phase spectrum corresponding to the training noisy phase spectrum and the training enhanced amplitude spectrum corresponding to the training noisy amplitude spectrum, calculate the amplitude loss based on the training clean amplitude spectrum and the training enhanced amplitude spectrum, and calculate the phase loss based on the training clean phase spectrum and the training enhanced phase spectrum.
[0133] Train the speech enhancement model based on the objective loss until the set training end condition is satisfied, where the objective loss includes the amplitude loss and the phase loss.
[0134] Specifically, the speech enhancement model can be trained based on the objective loss including the amplitude loss and the phase loss, and the relevant network parameters in the speech enhancement model can be adjusted to make the training enhanced phase spectrum predicted by the speech enhancement model the same as the training clean phase spectrum, and the training enhanced amplitude spectrum the same as the training clean amplitude spectrum.
[0135] Optionally, the above phase loss can include the following optional losses:
[0136] The instantaneous phase loss, group delay loss and instantaneous angular frequency loss between the training clean phase spectrum and the training enhanced phase spectrum.
[0137] In the embodiments of the present application, the phase loss includes three losses between the training clean phase spectrum and the training enhanced phase spectrum, namely, the instantaneous phase loss, the group delay loss, and the instantaneous angular frequency loss, so that the speech enhancement model finally trained based on the phase loss and the amplitude loss is more accurate.
[0138] Optionally, the calculation process of the above instantaneous phase loss may include:
[0139] Calculating a first true distance of the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function.
[0140] Specifically, due to the wrapping characteristic of the phase, both the training clean phase spectrum and the training enhanced phase spectrum are wrapped phase spectra, resulting in the fact that the absolute distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum is not their actual distance. If the absolute distance between the two is directly calculated, it will lead to the problem of error magnification caused by phase wrapping. Therefore, the embodiments of the present application preset an unwrapping function to calculate the first true distance of the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum based on the unwrapping function, so as to avoid the problem of error magnification caused by phase wrapping. The first true distance is the actual distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum.
[0141] Calculating the instantaneous phase loss based on the first true distance.
[0142] Specifically, after obtaining the first true distance of the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum, the instantaneous phase loss can be calculated based on the first true distance.
[0143] In the embodiments of the present application, the first true distance of the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum is calculated based on a preset unwrapping function, and then the instantaneous phase loss is calculated based on the first true distance. The first true distance is the actual distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum, which avoids the problem of error magnification caused by phase wrapping and makes the calculated instantaneous phase loss more accurate.
[0144] Optionally, the calculation process of the above group delay loss may include:
[0145] Calculating the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis.
[0146] Specifically, in order to ensure the continuity of the predicted enhanced phase spectrum along the frequency axis, the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis are calculated, so as to calculate the group delay loss of the training clean phase spectrum and the training enhanced phase spectrum based on the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis.
[0147] Calculate a second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function.
[0148] Specifically, the distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis can be calculated, and then the group delay loss is calculated based on the calculated distance. Considering that the phase has a wrapping characteristic, the absolute distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis is not the actual distance between the two. If the absolute distance between the two is directly calculated, the problem of error magnification will also occur. Therefore, the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis is calculated based on the unwrapping function, and the second true distance is the actual distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis.
[0149] Calculate the group delay loss based on the second true distance.
[0150] Specifically, after calculating the second true distance, the group delay loss is calculated based on the second true distance.
[0151] Optionally, the group delay loss can be calculated based on the following formula
[0152]
[0153] where Δ DF X p is the differential spectrum of the training clean phase spectrum along the frequency axis, is the differential spectrum of the training enhanced phase spectrum along the frequency axis, is the second true distance.
[0154] In the embodiments of the present application, the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis are calculated. The second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis is calculated based on the unwrapping function, and then the group delay loss is calculated based on the second true distance. The second true distance is the actual distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis, avoiding the problem of error magnification caused by phase wrapping and ensuring the continuity of the predicted enhanced phase spectrum along the frequency axis.
[0155] Optionally, the calculation process of the above instantaneous angular frequency loss may include:
[0156] Calculate the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis.
[0157] Specifically, in order to ensure the continuity of the predicted enhanced phase spectrum along the time axis, the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis are calculated, so as to calculate the instantaneous angular frequency loss between the training clean phase spectrum and the training enhanced phase spectrum based on the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis.
[0158] Calculate the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis based on the unwrapping function.
[0159] Specifically, the distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis can be calculated, and then the instantaneous angular frequency loss can be calculated based on the calculated distance. Considering the specific winding characteristics of the phase, the absolute distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis is not the actual distance between the two. If the absolute distance between the two is directly calculated, it will also lead to the problem of error expansion. Therefore, the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis is calculated based on the unwrapping function, and the third true distance is the actual distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis.
[0160] Calculate the instantaneous angular frequency loss based on the third true distance.
[0161] Specifically, after calculating the third true distance, calculate the instantaneous angular frequency loss based on the third true distance.
[0162] In the embodiments of the present application, the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis are calculated, the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis is calculated based on the unwrapping function, and then the instantaneous angular frequency loss is calculated based on the third true distance, avoiding the problem of error expansion caused by phase winding and ensuring the continuity of the predicted enhanced phase spectrum along the time axis.
[0163] Optionally, the above unwrapping function can be f AW (m), f AW (m) = m - 2π·round(m / 2π), where m is the independent variable and round is the rounding function.
[0164] Specifically, the independent variable m is the difference between the phases of two phase spectra. According to the winding property of the phase, when the absolute value |r| of the difference between the two phases is greater than π, the true distance between them is 2π - |r|; when the absolute value |r| of the difference between the two phases is not greater than π, the true distance between them is r. Therefore, this formula determines whether the true distance between the two phases is 2π - |m| or m by judging whether the absolute value of m is greater than π, thus achieving anti-winding.
[0165] Based on this, the process of calculating the first true distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum based on the preset anti-winding function may include:
[0166] Calculating the first true distance between the instantaneous phases of the training clean phase spectrum and the training enhanced phase spectrum X p is the training clean phase spectrum, is the training enhanced phase spectrum.
[0167] Optionally, the process of calculating the instantaneous phase loss based on the first true distance may include:
[0168] Calculating the instantaneous phase loss based on the following formula
[0169]
[0170] is the function for calculating the average value, ||||1 is the formula for the first norm, is the first true distance.
[0171] Optionally, the process of calculating the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis based on the anti-winding function may include:
[0172] Calculating the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis Δ DF X p is the differential spectrum of the training clean phase spectrum along the frequency axis, is the differential spectrum of the training enhanced phase spectrum along the frequency axis.
[0173] Optionally, the process of calculating the group delay loss based on the second true distance may include:
[0174] Calculating the group delay loss based on the following formula
[0175] t2 is the second true distance.
[0176] Optionally, the process of calculating the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis based on the anti-winding function may include:
[0177] Calculate the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis Δ DT X p is the differential spectrum of the training clean phase spectrum along the time axis, is the differential spectrum of the training enhanced phase spectrum along the time axis.
[0178] Optionally, the process of calculating the instantaneous phase loss based on the first true distance may include:
[0179] Calculating the instantaneous angular frequency loss based on the third true distance, including:
[0180] Calculating the instantaneous angular frequency loss based on the following formula
[0181] t3 is the third true distance.
[0182] In order to make the quality of the finally obtained enhanced speech waveform higher and the signal-to-noise ratio higher, in addition to the amplitude loss and phase loss, the above target loss may further include the loss between the training clean short-time complex spectrum corresponding to the training clean speech waveform and the training enhanced short-time complex spectrum corresponding to the training enhanced speech waveform, and the loss between the training clean speech waveform and the training enhanced speech waveform. Based on this, before training the speech enhancement model based on the target loss until the set training end condition is met, it may further include:
[0183] Perform short-time Fourier transform on the training clean speech waveform to obtain the training clean short-time complex spectrum.
[0184] Specifically, performing short-time Fourier transform on the training clean speech waveform can obtain the training clean short-time complex spectrum.
[0185] Reconstruct the training enhanced phase spectrum and the training enhanced amplitude spectrum to obtain the training enhanced short-time complex spectrum.
[0186] Specifically, reconstructing the training enhanced phase spectrum and the training enhanced amplitude spectrum can obtain the training enhanced short-time complex spectrum.
[0187] Perform inverse short-time Fourier transform on the training enhanced short-time complex spectrum to obtain the training enhanced speech waveform.
[0188] Specifically, performing an inverse short-time Fourier transform on the training enhanced short-time complex spectrum can obtain the training enhanced speech waveform.
[0189] Calculate the short-time complex spectrum loss based on the training clean short-time complex spectrum and the training enhanced short-time complex spectrum.
[0190] Specifically, after obtaining the training clean short-time complex spectrum and the training enhanced short-time complex spectrum, the short-time complex spectrum loss can be calculated based on the training clean short-time complex spectrum and the training enhanced short-time complex spectrum.
[0191] Optionally, the real and imaginary parts (X r , X i ) of the training clean short-time complex spectrum and the real and imaginary parts of the training enhanced short-time complex spectrum can be obtained. Then, calculate the short-time complex spectrum loss based on the following formula
[0192]
[0193] is the function for calculating the average value, and ||||2 is the formula for the two-norm.
[0194] Calculate the waveform loss based on the training clean speech waveform and the training enhanced speech waveform.
[0195] Specifically, the waveform loss can be calculated based on the training clean speech waveform and the training enhanced speech waveform.
[0196] Optionally, the waveform loss can be calculated x is the training clean speech waveform, is the enhanced speech waveform, and ||||1 is the formula for the one-norm.
[0197] In the embodiments of the present application, in addition to the amplitude loss and the phase loss, the above target loss further includes the short-time complex spectrum loss between the training clean short-time complex spectrum corresponding to the training clean speech waveform and the training enhanced short-time complex spectrum corresponding to the training enhanced speech waveform, and the waveform loss between the training clean speech waveform and the training enhanced speech waveform, so that the quality of the finally obtained enhanced speech waveform is higher and the signal-to-noise ratio is higher.
[0198] In some embodiments of the present application, the components of the above multi-speech enhancement model are introduced. Its components may include: an encoder, a TS-Conformer module, a phase decoder, and an amplitude decoder. The components of the speech enhancement model are introduced in turn below:
[0199] The encoder is used to encode the training noisy phase spectrum and the training noisy amplitude spectrum to obtain high-dimensional time-frequency domain features.
[0200] Specifically, considering that the time-frequency domain speech enhancement method can make the quality of the finally predicted enhanced speech waveform higher, the encoder encodes the training noisy phase spectrum and the training noisy amplitude spectrum to obtain high-dimensional time-frequency domain features.
[0201] Optionally, the training noisy amplitude spectrum and the training noisy phase spectrum can be concatenated, and the concatenated training noisy amplitude spectrum and the training noisy phase spectrum are input into the encoder, and the encoder encodes the concatenated training noisy amplitude spectrum and the training noisy phase spectrum to obtain high-dimensional time-frequency domain features.
[0202] Optionally, the training noisy amplitude spectrum can be first amplitude-compressed to obtain the training compressed noisy amplitude spectrum, and then the training compressed noisy amplitude spectrum and the training noisy phase spectrum are concatenated, and the concatenated training compressed noisy amplitude spectrum and the training noisy phase spectrum are input into the encoder, and the encoder encodes the concatenated training compressed noisy amplitude spectrum and the training noisy phase spectrum to obtain high-dimensional time-frequency domain features.
[0203] The TS-Conformer module is used to process the high-dimensional time-frequency domain features to obtain processed high-dimensional time-frequency domain features.
[0204] Specifically, the TS-Conformer module can be composed of N TS-Conformers (Two Stage Convolution-augmented Transformers). Processing the high-dimensional time-frequency domain features can capture local and global information of the high-dimensional time-frequency domain features. At the same time, in the time-frequency domain features, there is a correlation between the feature points along the time axis and the frequency axis. Capturing these correlations can better remove irrelevant noise features to achieve a better enhancement effect. Therefore, the high-dimensional time-frequency domain features can also be processed along the time axis and the frequency axis respectively to capture time and frequency dependencies, and obtain high-dimensional time-frequency domain features that capture local and global information and time and frequency dependencies.
[0205] The phase decoder is used to predict the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the processed high-dimensional time-frequency domain features, and calculate the training enhanced phase spectrum based on the simulated phase calculation formula.
[0206] The amplitude decoder is used to predict the training enhanced amplitude spectrum based on the processed high-dimensional time-frequency domain features.
[0207] Optionally, an amplitude decoder can be used to predict the training compressed enhanced amplitude spectrum mask corresponding to the training compressed noisy amplitude spectrum, multiply the training compressed enhanced amplitude spectrum mask point by point with the training compressed noisy amplitude spectrum to obtain the training compressed enhanced amplitude spectrum, and decompress the training compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0208] Based on this, refer to Figure 2 , in the embodiments of the present application, the components of the above-mentioned encoder, phase decoder, and amplitude decoder are introduced respectively:
[0209] The components of the encoder may include: a first two-dimensional convolutional layer, a first instance normalization layer, a first PReLU layer, a first dilated dense convolutional network, a second two-dimensional convolutional layer, a second instance normalization layer, and a second PReLU layer.
[0210] Among them, the first two-dimensional convolutional layer can be used to expand the dimensions of the concatenated training compressed noisy amplitude spectrum and the training noisy phase spectrum, and encode them into high-dimensional time-frequency domain features.
[0211] The first instance normalization layer can be used to normalize the high-dimensional time-frequency domain features to enable the model to converge and output the normalized high-dimensional time-frequency domain features.
[0212] The first PReLU layer can be used to activate the normalized high-dimensional time-frequency domain features, and by adding non-linear factors, solve the problem of linear inseparability, and output the activated high-dimensional time-frequency domain features.
[0213] The first dilated dense convolutional network can increase the receptive field of the model through multiple dilated convolutions, and avoid the problem of gradient explosion in the deep neural network through the dense connection of these convolutional layers. Input the activated high-dimensional time-frequency domain features into the first dilated dense convolutional network, and output the high-dimensional time-frequency domain features processed by the dilated dense neural network.
[0214] The second two-dimensional convolutional layer can be used to perform downsampling operations on the high-dimensional time-frequency domain features processed by the dilated dense neural network along the frequency axis, and output the high-dimensional time-frequency domain features with a low sampling rate to reduce the computational complexity and improve the model efficiency.
[0215] The second instance normalization layer can be used to normalize the high-dimensional time-frequency domain features with a low sampling rate and output the normalized high-dimensional time-frequency domain features with a low sampling rate.
[0216] The second PReLU layer can be used to activate the normalized high-dimensional time-frequency domain features with a low sampling rate and output the activated high-dimensional time-frequency domain features with a low sampling rate.
[0217] Based on this, N TS-Conformers can be used to process the activated high-dimensional time-frequency domain features with a low sampling rate to obtain the processed high-dimensional time-frequency domain features.
[0218] The components of the phase decoder may include: a second dilated dense convolutional network, a first two-dimensional transposed convolutional layer, a third instance normalization layer, a third PReLU layer, a third two-dimensional convolutional layer, a fourth two-dimensional convolutional layer, and a binary activation layer.
[0219] Among them, the second dilated dense convolutional network is used to process the high-dimensional time-frequency domain features processed by N TS-Conformers and output the high-dimensional time-frequency domain features that retain phase information.
[0220] The first two-dimensional transposed convolutional layer can be used to upsample the high-dimensional time-frequency domain features that retain phase information along the frequency axis to restore the frequency resolution of the time-frequency domain features downsampled by the second two-dimensional convolutional layer in the encoder, and output the processed high-dimensional time-frequency domain features that retain phase information.
[0221] The third instance normalization layer can be used to normalize the processed high-dimensional time-frequency domain features that retain phase information output by the first two-dimensional transposed convolutional layer and output the normalized high-dimensional time-frequency domain features that retain phase information.
[0222] The third PReLU layer can be used to activate the normalized high-dimensional time-frequency domain features that retain phase information and output the activated high-dimensional time-frequency domain features that retain phase information.
[0223] The third two-dimensional convolutional layer and the fourth two-dimensional convolutional layer can reduce the dimension of the activated high-dimensional time-frequency domain features that retain phase information by reducing the number of channels and output the training-enhanced pseudo real part spectrum and the training-enhanced pseudo imaginary part spectrum that retain phase information.
[0224] The binary activation layer can be used to calculate the training-enhanced phase spectrum from the training-enhanced pseudo real part spectrum and the training-enhanced pseudo imaginary part spectrum by simulating the phase calculation formula.
[0225] The components of the amplitude decoder may include: a third dilated dense convolutional network, a second two-dimensional transposed convolutional layer, a fourth instance normalization layer, a fourth PReLU layer, a fifth two-dimensional convolutional layer, an LSigmoid layer, and an enhanced amplitude spectrum calculation module.
[0226] The third dilated dense convolutional network is used to process the high-dimensional time-frequency domain features processed by N TS-Conformers and output the high-dimensional time-frequency domain features that retain amplitude information.
[0227] The second two-dimensional transposed convolutional layer can be used to upsample the high-dimensional time-frequency domain features that retain amplitude information along the frequency axis, so as to restore the frequency resolution of the time-frequency domain features downsampled by the second two-dimensional convolutional layer in the encoder, and output the processed high-dimensional time-frequency domain features that retain amplitude information.
[0228] The fourth instance normalization layer can be used to normalize the processed high-dimensional time-frequency domain features that retain amplitude information output by the second two-dimensional transposed convolutional layer, and output the normalized high-dimensional time-frequency domain features that retain amplitude information.
[0229] The fourth PReLU layer can be used to activate the normalized high-dimensional time-frequency domain features that retain amplitude information, and output the activated high-dimensional time-frequency domain features that retain amplitude information.
[0230] The fifth two-dimensional convolutional layer can be used to reduce the dimension of the high-dimensional time-frequency domain features that retain amplitude information by reducing the number of channels, and output the low-dimensional time-frequency domain features that retain amplitude information.
[0231] The LSigmoid layer can activate the low-dimensional time-frequency domain features that retain amplitude information, limit its value range to between 0 and 2, and output the bounded training compressed enhanced amplitude spectrum mask.
[0232] The enhanced amplitude spectrum calculation module can be used to multiply the training compressed enhanced amplitude spectrum mask point by point with the noisy amplitude spectrum to obtain the training compressed enhanced amplitude spectrum, and decompress the training compressed enhanced amplitude spectrum to obtain the training enhanced amplitude spectrum corresponding to the training noisy amplitude spectrum.
[0233] Optionally, considering that the above speech enhancement model can be trained in a generative adversarial training mode, based on this, the above encoder, the above TS-Conformer module, the above phase decoder and the above amplitude decoder can be combined into a generator, and the above speech enhancement model can further include a discriminator, and the discriminator can be used to make a judgment on the training clean amplitude spectrum and the training enhanced amplitude spectrum to obtain the judgment index corresponding to the enhanced speech waveform.
[0234] Optionally, the judgment index can be the judgment normalized PESQ (Perceptual Evaluation of Speech Quality) index, and the discriminator can also be used to make a judgment on a pair of training clean amplitude spectra to obtain the judgment normalized PESQ index corresponding to the training clean speech waveform.
[0235] Based on this, before training the above speech enhancement model based on the target loss until the set training end condition is met, it can further include:
[0236] Calculate the discriminator loss.
[0237] Specifically, the class generation adversarial training mode is a mode of training a model based on discriminator loss and generator loss. Therefore, before training the speech enhancement model based on the objective loss until the set training end condition is met, the discriminator loss is first calculated.
[0238] Optionally, the discriminator loss can be calculated based on the following formula
[0239]
[0240] where X m is the training clean magnitude spectrum, D(X m , X m ) is the decision-normalized PESQ metric corresponding to the training clean speech waveform, is the training noisy magnitude spectrum, is the decision-normalized PESQ metric corresponding to the enhanced speech waveform, Q PESQ is the true normalized PESQ metric corresponding to the enhanced speech waveform determined in advance.
[0241] Specifically, this formula defines the discriminator loss as the mean squared error loss between the decision-normalized PESQ metric corresponding to the training clean speech waveform obtained by the discriminator's decision on a pair of training clean magnitude spectra and the maximum value 1, and the mean squared error loss between the decision-normalized PESQ metric corresponding to the enhanced speech waveform obtained by the discriminator's decision on the training clean magnitude spectrum and the training enhanced magnitude spectrum and its true normalized PESQ metric.
[0242] Calculate the objective metric loss based on the decision metric.
[0243] Calculate the objective metric loss based on the following formula
[0244] is the function for calculating the average value, and ||||2 is the formula for the two-norm.
[0245] Specifically, this formula defines the objective metric loss as the mean squared error loss between the decision-normalized PESQ metric corresponding to the enhanced speech waveform obtained by the discriminator's decision on the training clean magnitude spectrum and the training enhanced magnitude spectrum and the maximum metric value 1.
[0246] Based on this, the objective loss can also include the objective metric loss and the discriminator loss.
[0247] The process of training the speech enhancement model based on the objective loss until the set training end condition is met can include:
[0248] Train the discriminator based on the discriminator loss, and train the generator based on the magnitude loss, the phase loss, the short-time complex spectrum loss, the waveform loss, and the objective metric loss until a set training end condition is met.
[0249] Specifically, since a generative adversarial-like mode is adopted to train the speech enhancement model, the discriminator is trained based on the discriminator loss, and the generator is trained based on the magnitude loss, the phase loss, the short-time complex spectrum loss, the waveform loss, and the objective metric loss until a set training end condition is met.
[0250] Optionally, the process of training the generator based on the magnitude loss, the phase loss, the short-time complex spectrum loss, the waveform loss, and the objective metric loss may include:
[0251] Calculate the generator loss based on the following formula
[0252]
[0253] where γ1, γ2, γ3, γ4, γ5 are preset hyperparameters, is the magnitude loss, is the phase loss, is the short-time complex spectrum loss, is the waveform loss, is the objective metric loss.
[0254] Based on the generator loss train the generator.
[0255] In the embodiments of the present application, adopting a generative adversarial-like training mode to train the speech enhancement model can improve the stability of the speech enhancement model and the accuracy of the enhanced phase spectrum and enhanced magnitude spectrum obtained by prediction.
[0256] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the process of training the above speech enhancement model. Next, the process of training the speech enhancement model will be introduced in combination with Figure 3 This process may include:
[0257] Obtain the training noisy speech waveform, perform short-time Fourier transform on the training noisy speech waveform to obtain the training noisy short-time complex spectrum, calculate the phase of the training noisy short-time complex spectrum to obtain the training noisy phase spectrum, calculate the amplitude of the training noisy short-time complex spectrum to obtain the training noisy amplitude spectrum, and perform amplitude compression on the training noisy phase spectrum to obtain the training compressed noisy amplitude spectrum. Perform short-time Fourier transform on the training clean speech waveform corresponding to the training noisy speech waveform to obtain the training clean short-time complex spectrum, calculate the phase of the training clean short-time complex spectrum to obtain the training clean phase spectrum, and calculate the amplitude of the training clean short-time complex spectrum to obtain the training clean amplitude spectrum.
[0258] Concatenate the training noisy phase spectrum and the training compressed noisy amplitude spectrum, and input the concatenated training noisy phase spectrum and the training compressed noisy amplitude spectrum into the encoder to encode the concatenated training noisy phase spectrum and the training compressed noisy amplitude spectrum using the encoder, and output high-dimensional time-frequency domain features. N TS-Conformers process the high-dimensional time-frequency domain features output by the encoder and output the processed high-dimensional time-frequency domain features. The phase decoder is used to predict the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the processed high-dimensional time-frequency domain features, calculate the training enhanced phase spectrum based on the simulated phase calculation formula, and the amplitude decoder predicts the training enhanced amplitude spectrum based on the processed high-dimensional time-frequency domain features.
[0259] Calculate the instantaneous phase loss, instantaneous angular frequency loss, and group delay loss based on the training clean phase spectrum and the training enhanced phase spectrum, calculate the amplitude loss based on the training clean amplitude spectrum and the training enhanced amplitude spectrum, reconstruct the training enhanced phase spectrum and the training enhanced amplitude spectrum to obtain the training enhanced short-time complex spectrum, take the real part of the training enhanced short-time complex spectrum to obtain the training enhanced real part, take the imaginary part of the training enhanced short-time complex spectrum to obtain the training enhanced imaginary part, take the real part of the training clean short-time complex spectrum to obtain the training clean real part, take the imaginary part of the training clean short-time complex spectrum to obtain the training clean imaginary part, calculate the real part loss between the training enhanced real part and the training clean real part, calculate the imaginary part loss between the training enhanced imaginary part and the training clean imaginary part, calculate the short-time complex spectrum loss based on the real part loss and the imaginary part loss, perform the inverse short-time Fourier transform on the training enhanced short-time complex spectrum to obtain the training enhanced speech waveform, calculate the waveform loss based on the training enhanced speech waveform and the training clean speech waveform, input the training clean amplitude spectrum and the training enhanced amplitude spectrum into the discriminator to obtain the discriminant normalized PESQ index corresponding to the training enhanced speech waveform, input a pair of training clean amplitude spectra into the discriminator to obtain the discriminant normalized PESQ index corresponding to the training clean speech waveform, calculate the discriminator loss based on the discriminant normalized PESQ index corresponding to the training enhanced speech waveform and the discriminant normalized PESQ index corresponding to the training clean speech waveform, and calculate the objective metric loss based on the discriminant normalized PESQ index corresponding to the training enhanced speech waveform.
[0260] Train the discriminator based on the discriminator loss, and train the generator based on the linear combination of the amplitude loss, phase loss, short-time complex spectrum loss, waveform loss, and objective metric loss until the set training end condition is met. The phase loss is the linear combination of the instantaneous phase loss, instantaneous angular frequency loss, and group delay loss. The generator includes an encoder, N TS-Conformers, a phase decoder, and an amplitude decoder.
[0261] Next, the speech enhancement device provided by the embodiments of the present application will be described. The speech enhancement device described below can be correspondingly referred to the speech enhancement method described above.
[0262] First, in combination with Figure 4 , the speech enhancement device will be introduced. As Figure 4 shown, the speech enhancement device may include:
[0263] A noisy phase spectrum and noisy amplitude spectrum acquisition unit 10, configured to acquire a noisy phase spectrum and a noisy amplitude spectrum of a noisy speech waveform;
[0264] An enhanced phase spectrum and enhanced amplitude spectrum acquisition unit 20, configured to process the noisy phase spectrum and the noisy amplitude spectrum by using a preset speech enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum;
[0265] The voice enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and to predict an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, and calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with the value range limited within the principal value range;
[0266] The enhanced speech waveform acquisition unit 30 is configured to calculate an enhanced speech waveform corresponding to the noisy speech waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum.
[0267] Optionally, the voice enhancement device may further include:
[0268] A compression unit configured to perform amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum.
[0269] Optionally, the process of the enhanced phase spectrum and enhanced amplitude spectrum acquisition unit using a preset voice enhancement model to process the noisy phase spectrum and the noisy amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum may include:
[0270] Input the noisy phase spectrum and the compressed noisy amplitude spectrum into a preset voice enhancement model to use the voice enhancement model to predict a compressed enhanced amplitude spectrum mask corresponding to the compressed noisy amplitude spectrum, and multiply the compressed enhanced amplitude spectrum mask by the compressed noisy amplitude spectrum point by point to obtain a compressed enhanced amplitude spectrum, and decompress the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0271] Optionally, the process of the compression unit performing amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum may include:
[0272] Calculate the c-th power of the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum, where c is a preset compression factor.
[0273] Optionally, the process of the enhanced phase spectrum and enhanced amplitude spectrum acquisition unit using the voice enhancement model to decompress the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum may include:
[0274] Calculate the 1 / c-th power of the compressed enhanced amplitude spectrum to obtain the enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
[0275] Optionally, the simulated phase calculation formula may be:
[0276]
[0277] where p is the independent variable representing the enhanced pseudo real part spectrum, q is the independent variable representing the enhanced pseudo imaginary part spectrum, when p ≥ 0, Sgn * (p) = 1, when p < 0, Sgn * (p) = -1, when q ≥ 0, Sgn * (q) = 1, when q < 0, Sgn * (q) = -1.
[0278] Optionally, the voice enhancement device may further include:
[0279] A model training unit for:
[0280] Obtain the training noisy phase spectrum and training noisy amplitude spectrum of the training noisy speech waveform, and the training clean phase spectrum and training clean amplitude spectrum of the training clean speech waveform corresponding to the training noisy speech waveform;
[0281] Input the training noisy phase spectrum and the noisy amplitude spectrum into the voice enhancement model, so as to use the voice enhancement model to predict the training enhanced pseudo real part spectrum and training enhanced pseudo imaginary part spectrum corresponding to the training noisy phase spectrum, and predict the training enhanced amplitude spectrum corresponding to the training noisy amplitude spectrum, and calculate the training enhanced phase spectrum with the value range limited within the principal value range based on the simulated phase calculation formula for the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum;
[0282] Calculate the amplitude loss based on the training clean amplitude spectrum and the training enhanced amplitude spectrum;
[0283] Calculate the phase loss based on the training clean phase spectrum and the training enhanced phase spectrum;
[0284] Train the voice enhancement model based on the target loss until the set training end condition is met, where the target loss includes the amplitude loss and the phase loss.
[0285] Optionally, the phase loss may include the instantaneous phase loss between the training clean phase spectrum and the training enhanced phase spectrum.
[0286] Based on this, the process of the model training unit calculating the phase loss based on the training clean phase spectrum and the training enhanced phase spectrum may include:
[0287] Calculate the first true distance of the instantaneous phase between the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function;
[0288] Calculate the instantaneous phase loss based on the first true distance.
[0289] Optionally, the phase loss may further include the group delay loss and the instantaneous angular frequency loss between the training clean phase spectrum and the training enhanced phase spectrum.
[0290] Based on this, the process of the model training unit calculating the phase loss based on the training clean phase spectrum and the training enhanced phase spectrum may further include:
[0291] Calculate the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis;
[0292] Calculate the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function;
[0293] Calculate the group delay loss based on the second true distance;
[0294] Calculate the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis;
[0295] Calculate the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis based on the unwrapping function;
[0296] Calculate the instantaneous angular frequency loss based on the third true distance.
[0297] Optionally, the unwrapping function may be f AW (m), f AW (m) = m - 2π·round(m / 2π), where m is the independent variable and round is the rounding function.
[0298] Optionally, the process of the model training unit calculating the first true distance of the instantaneous phase between the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function may include:
[0299] Calculate the first true distance of the instantaneous phase between the training clean phase spectrum and the training enhanced phase spectrum X p is the training clean phase spectrum, is the training enhanced phase spectrum.
[0300] Optionally, the process of the model training unit calculating the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function may include:
[0301] Calculate the second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis Δ DF X p is the differential spectrum of the training clean phase spectrum along the frequency axis, is the differential spectrum of the training enhanced phase spectrum along the frequency axis.
[0302] Optionally, the process by which the model training unit calculates the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis based on the unwrapping function may include:
[0303] Calculate the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis Δ DT X p is the differential spectrum of the training clean phase spectrum along the time axis, is the differential spectrum of the training enhanced phase spectrum along the time axis.
[0304] Optionally, the process by which the model training unit calculates the instantaneous phase loss based on the first true distance may include:
[0305] Calculate the instantaneous phase loss based on the following formula
[0306] is the function for calculating the average value, ||||1 is the formula for the first norm, and t1 is the first true distance.
[0307] Optionally, the process by which the model training unit calculates the group delay loss based on the second true distance may include:
[0308] Calculate the group delay loss based on the following formula
[0309] t2 is the second true distance.
[0310] Optionally, the process by which the model training unit calculates the instantaneous angular frequency loss based on the third true distance may include:
[0311] Calculate the instantaneous angular frequency loss based on the following formula
[0312] t3 is the third true distance.
[0313] Optionally, the model training unit may also be used for:
[0314] Perform a short-time Fourier transform on the training clean speech waveform to obtain a training clean short-time complex spectrum;
[0315] Reconstruct the training enhanced phase spectrum and the training enhanced amplitude spectrum to obtain a training enhanced short-time complex spectrum;
[0316] Calculate the short-time complex spectrum loss based on the training clean short-time complex spectrum and the training enhanced short-time complex spectrum.
[0317] Based on this, the target loss may further include the short-time complex spectrum loss.
[0318] Optionally, the model training unit may also be used for:
[0319] Perform an inverse short-time Fourier transform on the training enhanced short-time complex spectrum to obtain a training enhanced speech waveform;
[0320] Calculate the waveform loss based on the training clean speech waveform and the training enhanced speech waveform.
[0321] Based on this, the target loss may further include the waveform loss.
[0322] Optionally, the speech enhancement model may include:
[0323] An encoder, a TS-Conformer module, a phase decoder, and an amplitude decoder;
[0324] The encoder is used to encode the training noisy phase spectrum and the training noisy amplitude spectrum to obtain high-dimensional time-frequency domain features;
[0325] The TS-Conformer module is used to process the high-dimensional time-frequency domain features to obtain processed high-dimensional time-frequency domain features;
[0326] The phase decoder is used to predict the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the processed high-dimensional time-frequency domain features, and calculate the training enhanced phase spectrum based on the simulated phase calculation formula for the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum;
[0327] The amplitude decoder is used to predict the training enhanced amplitude spectrum based on the processed high-dimensional time-frequency domain features.
[0328] Optionally, the encoder, the TS-Conformer module, the phase decoder, and the amplitude decoder are combined into a generator, and the speech enhancement model may further include a discriminator.
[0329] The decider is used to make a decision on the training clean magnitude spectrum and the training enhanced magnitude spectrum, and obtain a decision metric corresponding to the enhanced speech waveform.
[0330] Optionally, the model training unit may also be used to:
[0331] Calculate the decider loss;
[0332] Calculate the objective metric loss based on the decision metric.
[0333] Based on this, the objective loss may further include the objective metric loss and the decider loss.
[0334] The process of the model training unit training the speech enhancement model based on the objective loss until a set training end condition is met may include:
[0335] Train the decider based on the decider loss, and train the generator based on the magnitude loss, the phase loss, the short-time complex spectrum loss, the waveform loss, and the objective metric loss until a set training end condition is met.
[0336] Optionally, the decision metric may be a decision normalized PESQ metric, and the decider is further used to make a decision on a pair of the training clean magnitude spectra to obtain the decision normalized PESQ metric corresponding to the training clean speech waveform.
[0337] Based on this, the process of the model training unit calculating the decider loss may include:
[0338] Calculate the decider loss based on the following formula
[0339]
[0340] where X m is the training clean magnitude spectrum, D(X m , X m ) is the decision normalized PESQ metric corresponding to the training clean speech waveform, is the training noisy magnitude spectrum, is the decision normalized PESQ metric t corresponding to the enhanced speech waveform, and Q PESQ is the true normalized PESQ metric corresponding to the enhanced speech waveform determined in advance;
[0341] The process of the model training unit calculating the objective metric loss based on the decision metric may include:
[0342] Calculate the objective metric loss based on the following formula
[0343] is a function for calculating the average value, and ||||2 is the formula for the two-norm.
[0344] The voice enhancement device provided by the embodiments of the present application can be applied to voice enhancement devices. Figure 5 shows a hardware structure block diagram of a voice enhancement device. Refer to Figure 5 , the hardware structure of the voice enhancement device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0345] In the embodiments of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4;
[0346] The processor 1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0347] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0348] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used to: implement each processing flow in the foregoing voice enhancement solution.
[0349] The embodiments of the present application also provide a storage medium, which can store a program suitable for being executed by a processor. The program is used to: implement each processing flow in the foregoing voice enhancement solution.
[0350] Finally, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0351] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other.
[0352] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice enhancement method, characterized in that, Including: Obtaining the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform; Processing the noisy phase spectrum and the noisy amplitude spectrum by using a preset speech enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum; The speech enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum based on the input noisy phase spectrum and the noisy amplitude spectrum, and predict an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum, and calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with the value range limited within the principal value range; Calculating an enhanced speech waveform corresponding to the noisy speech waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum.
2. The method according to claim 1, wherein After obtaining the noisy phase spectrum and the noisy amplitude spectrum of the noisy speech waveform, it further includes: Performing amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum; The process of processing the noisy phase spectrum and the noisy amplitude spectrum by using a preset speech enhancement model to obtain an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum includes: Inputting the noisy phase spectrum and the compressed noisy amplitude spectrum into a preset speech enhancement model to predict a compressed enhanced amplitude spectrum mask corresponding to the compressed noisy amplitude spectrum by using the speech enhancement model, and multiplying the compressed enhanced amplitude spectrum mask point by point by the compressed noisy amplitude spectrum to obtain a compressed enhanced amplitude spectrum, and decompressing the compressed enhanced amplitude spectrum to obtain an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
3. The method according to claim 2, wherein Performing amplitude compression on the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum includes: Calculating the c-th power of the noisy amplitude spectrum to obtain a compressed noisy amplitude spectrum, where c is a preset compression factor; The process of decompressing the compressed enhanced amplitude spectrum to obtain an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum includes: Calculating the 1 / c-th power of the compressed enhanced amplitude spectrum to obtain an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum.
4. The method according to claim 1, wherein The simulated phase calculation formula is: where p is the independent variable representing the enhanced pseudo real part spectrum, q is the independent variable representing the enhanced pseudo imaginary part spectrum, when p ≥ 0, Sgn * (p) = 1, when p < 0, Sgn * (p) = -1, when q ≥ 0, Sgn * (q) = 1, when q < 0, Sgn * (q) = -1.
5. The method according to claim 1, characterized in that The training process of the speech enhancement model includes: Obtaining a training noisy phase spectrum, a training noisy amplitude spectrum of a training noisy speech waveform, and a training clean phase spectrum and a training clean amplitude spectrum of a training clean speech waveform corresponding to the training noisy speech waveform; Inputting the training noisy phase spectrum and the noisy amplitude spectrum into the speech enhancement model to predict a training enhanced pseudo real part spectrum and a training enhanced pseudo imaginary part spectrum corresponding to the training noisy phase spectrum based on the input training noisy phase spectrum and the training noisy amplitude spectrum by using the speech enhancement model, and predicting a training enhanced amplitude spectrum corresponding to the training noisy amplitude spectrum, and calculating the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the simulated phase calculation formula to obtain a training enhanced phase spectrum with the value range limited within the principal value range; Calculating an amplitude loss based on the training clean amplitude spectrum and the training enhanced amplitude spectrum; Calculating a phase loss based on the training clean phase spectrum and the training enhanced phase spectrum; Train the speech enhancement model based on the target loss until a set training end condition is met, where the target loss includes the magnitude loss and the phase loss.
6. The method according to claim 5, wherein The phase loss includes the instantaneous phase loss between the training clean phase spectrum and the training enhanced phase spectrum; The calculation process of the instantaneous phase loss includes: Calculate the first true distance of the instantaneous phase between the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function; Calculate the instantaneous phase loss based on the first true distance.
7. The method according to claim 6, wherein The phase loss further includes the group delay loss and the instantaneous angular frequency loss between the training clean phase spectrum and the training enhanced phase spectrum; The calculation process of the group delay loss includes: Calculate the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis; Calculate the second true distance of the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function; Calculate the group delay loss based on the second true distance; The calculation process of the instantaneous angular frequency loss includes: Calculate the difference spectrum of the training clean phase spectrum along the time axis and the difference spectrum of the training enhanced phase spectrum along the time axis; Calculate the third true distance of the difference spectrum of the training clean phase spectrum along the time axis and the difference spectrum of the training enhanced phase spectrum along the time axis based on the unwrapping function; Calculate the instantaneous angular frequency loss based on the third true distance.
8. The method according to claim 7, characterized in that, The anti-winding function is f AW (m), where f AW (m) = m - 2π·round(m / 2π), m is the independent variable, and round is the rounding function; The step of calculating the first true distance of the instantaneous phase between the training clean phase spectrum and the training enhanced phase spectrum based on a preset unwrapping function includes: Calculate the first true distance of the instantaneous phase of the training clean phase spectrum and the training enhanced phase spectrum X p is the training clean phase spectrum, is the training enhanced phase spectrum; The step of calculating the second true distance of the difference spectrum of the training clean phase spectrum along the frequency axis and the difference spectrum of the training enhanced phase spectrum along the frequency axis based on the unwrapping function includes: Calculate a second true distance between the differential spectrum of the training clean phase spectrum along the frequency axis and the differential spectrum of the training enhanced phase spectrum along the frequency axis Δ DF X p is the differential spectrum of the training clean phase spectrum along the frequency axis, is the differential spectrum of the training enhanced phase spectrum along the frequency axis; The step of calculating the third true distance of the difference spectrum of the training clean phase spectrum along the time axis and the difference spectrum of the training enhanced phase spectrum along the time axis based on the unwrapping function includes: Calculate the third true distance between the differential spectrum of the training clean phase spectrum along the time axis and the differential spectrum of the training enhanced phase spectrum along the time axis Δ DT X p is the differential spectrum of the training clean phase spectrum along the time axis, is the differential spectrum of the training enhanced phase spectrum along the time axis.
9. The method according to claim 7, wherein The step of calculating the instantaneous phase loss based on the first true distance includes: Calculate the instantaneous phase loss based on the following formula is a function for calculating the average value, ||||1 is a formula for the first norm, and t1 is the first true distance; The step of calculating the group delay loss based on the second true distance includes: Calculate the group delay loss based on the following formula t2 is the second true distance; The step of calculating the instantaneous angular frequency loss based on the third true distance includes: Calculate the instantaneous angular frequency loss based on the following formula t3 is the third true distance.
10. The method according to claim 7, wherein Before training the speech enhancement model based on the target loss until a set training end condition is met, it further includes: Perform a short-time Fourier transform on the training clean speech waveform to obtain a training clean short-time complex spectrum; Reconstruct the training enhanced phase spectrum and the training enhanced magnitude spectrum to obtain a training enhanced short-time complex spectrum; Calculate the short-time complex spectrum loss based on the training clean short-time complex spectrum and the training enhanced short-time complex spectrum; The target loss further includes the short-time complex spectrum loss.
11. The method according to claim 10, wherein After reconstructing the training enhanced phase spectrum and the training enhanced magnitude spectrum to obtain a training enhanced short-time complex spectrum, it further includes: Perform an inverse short-time Fourier transform on the training enhanced short-time complex spectrum to obtain a training enhanced speech waveform; Calculate the waveform loss based on the training clean speech waveform and the training enhanced speech waveform; The target loss further includes the waveform loss.
12. The method according to claim 11, wherein The speech enhancement model includes: an encoder, a TS-Conformer module, a phase decoder, and an amplitude decoder; The encoder is used to encode the training noisy phase spectrum and the training noisy amplitude spectrum to obtain high-dimensional time-frequency domain features; The TS-Conformer module is used to process the high-dimensional time-frequency domain features to obtain processed high-dimensional time-frequency domain features; The phase decoder is used to predict the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the processed high-dimensional time-frequency domain features, and calculate the training enhanced phase spectrum by calculating the training enhanced pseudo real part spectrum and the training enhanced pseudo imaginary part spectrum based on the simulated phase calculation formula; The amplitude decoder is used to predict the training enhanced amplitude spectrum based on the processed high-dimensional time-frequency domain features.
13. The method according to claim 12, characterized in that The encoder, the TS-Conformer module, the phase decoder, and the amplitude decoder are combined into a generator, and the speech enhancement model further includes a discriminator; The discriminator is used to discriminate between the training clean amplitude spectrum and the training enhanced amplitude spectrum to obtain the discrimination index corresponding to the enhanced speech waveform; Before training the speech enhancement model based on the objective loss until the set training end condition is satisfied, it further includes: Calculating the discriminator loss; Calculating the objective metric loss based on the discrimination index; The objective loss further includes the objective metric loss and the discriminator loss; Training the speech enhancement model based on the objective loss until the set training end condition is satisfied, including: Training the discriminator based on the discriminator loss, and training the generator based on the amplitude loss, the phase loss, the short-time complex spectrum loss, the waveform loss, and the objective metric loss until the set training end condition is satisfied.
14. The method according to claim 13, characterized in that The discrimination index is the discrimination normalized PESQ index, and the discriminator is further used to discriminate between a pair of the training clean amplitude spectra to obtain the discrimination normalized PESQ index corresponding to the training clean speech waveform; The calculating the discriminator loss includes: Calculate the discriminator loss based on the following formula Among them, X m is the training clean magnitude spectrum, D(X m , X m ) is the decision-normalized PESQ metric corresponding to the training clean speech waveform, is the training noisy magnitude spectrum, is the decision-normalized PESQ metric corresponding to the enhanced speech waveform, Q PESQ is the true normalized PESQ metric corresponding to the enhanced speech waveform determined in advance; Calculating the objective metric loss based on the discrimination index includes: Calculate the objective metric loss based on the following formula is a function for calculating the average value, ||||2 is the formula for the two-norm.
15. A voice enhancement device, characterized in that, Includes: A noisy phase spectrum and noisy amplitude spectrum acquisition unit, configured to acquire a noisy phase spectrum and a noisy amplitude spectrum of a noisy speech waveform; An enhanced phase spectrum and enhanced amplitude spectrum acquisition unit, configured to process the noisy phase spectrum and the noisy amplitude spectrum by using a preset speech enhancement model to obtain an enhanced phase spectrum corresponding to the noisy phase spectrum and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum; The speech enhancement model is configured to predict an enhanced pseudo real part spectrum and an enhanced pseudo imaginary part spectrum corresponding to the noisy phase spectrum, and an enhanced amplitude spectrum corresponding to the noisy amplitude spectrum based on the input noisy phase spectrum and noisy amplitude spectrum, and calculate the enhanced pseudo real part spectrum and the enhanced pseudo imaginary part spectrum based on a preset simulated phase calculation formula to obtain an internal state representation of the enhanced phase spectrum with the value range limited within the principal value range; An enhanced speech waveform acquisition unit, configured to calculate an enhanced speech waveform corresponding to the noisy speech waveform according to the enhanced phase spectrum and the enhanced amplitude spectrum.