A single-channel speech enhancement system based on spectrogram compensation
By using a single-channel speech enhancement system based on spectral compensation and deep neural networks for pre-enhancement and spectral compensation, the speech distortion problem is solved, and a clear, intelligible, and high-quality speech enhancement effect is achieved in noisy environments.
Patent Information
- Application Number
- CN202111307973.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-11-05
AI Technical Summary
Existing deep learning-based speech enhancement methods suffer from speech distortion when dealing with noisy environments, resulting in poor perception quality and intelligibility of the enhanced speech.
A single-channel speech enhancement system based on spectrogram compensation is adopted, including a pre-enhancement module, a spectrogram compensation module, and a joint training module. Deep neural networks are used for pre-enhancement, the weight matrix of spectrogram compensation is estimated to fuse the pre-enhancement spectrogram with the original input spectrogram, and the speech quality is improved through a joint optimization module.
To obtain clear, intelligible, and better-quality speech in noisy environments, a pre-enhancement module removes most of the noise, a spectrogram compensation module recovers lost speech information, and a joint training module optimizes overall performance to improve speech enhancement.
Smart Images

Figure CN114038475B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of ranging, in particular to a single-channel speech enhancement system based on a spectrogram compensation. BACKGROUND
[0002] Speech as one of the main means of human communication, speech enhancement has always played an important role in speech signal processing. Speech enhancement refers to the technology of extracting useful speech signals from noise background when speech signals are disturbed or even submerged by various noises.
[0003] The actual speech interference can be divided into the following categories: ① Periodic noise, such as electrical interference, interference caused by rotating parts of the engine, etc. This kind of interference shows some discrete narrow frequency peaks; ② Impulse noise, such as noise interference caused by some electric sparks, discharge; ③ Wideband noise, which refers to Gaussian noise or white noise, which is characterized by wide frequency band, almost covering the entire speech frequency band; ④ Speech interference, such as picking up other people's speech in the microphone, or encountering a speech caused by a howling during transmission. The enhancement technology is different for dealing with the above various types of noise.
[0004] The goal of speech enhancement technology is to separate the target clean speech from the noisy environment and remove the background interference noise. When a speech contains background noise, it will seriously affect the performance of speech recognition, speaker recognition and hearing aid systems, so speech enhancement technology is particularly important.
[0005] In the development process of speech enhancement technology, the early research mainly adopts the methods based on spectral subtraction, Wiener filtering and statistics. However, these methods have very limited effect on non-stationary noise, so the application of these methods is also restricted. In recent years, with the development of computer technology, speech enhancement methods based on deep learning have developed a lot and received more and more attention.
[0006] The speech enhancement method based on deep learning uses a large number of pairs of noisy-clean speech data to train the speech enhancement model, and establishes the mapping relationship between the noisy speech feature parameters and the target clean speech signal feature parameters. In this way, for any input noisy speech signal, the noise-reduced speech signal can be output through the established enhancement model, so as to achieve the purpose of speech enhancement. The speech enhancement method based on deep learning modeling has many advantages over traditional methods, such as using the powerful modeling capability of deep learning to learn the mapping relationship between noisy speech and target speech signals. However, the biggest problem of speech enhancement is that the enhanced speech has distortion problem. Speech distortion will lose a lot of very important speech information, seriously affecting the perceptual quality and intelligibility of the enhanced speech, and restricting the performance of speech enhancement. SUMMARY
[0007] The technical problem solved by the present application is to provide a single-channel speech enhancement system based on spectrogram compensation to obtain clear, understandable and better quality speech in a noisy background environment.
[0008] To solve the above technical problems, the present application adopts the following technical solutions.
[0009] A single-channel speech enhancement system based on spectrogram compensation, comprising a pre-enhancement module, a spectrogram compensation module and a joint training module;
[0010] The pre-enhancement module is used to remove part of the interference signal in the speech.
[0011] The spectrogram compensation module is connected with the pre-enhancement module, and is used to obtain a weight matrix λ of spectrogram compensation, and fuse the pre-enhanced spectrogram and the original input spectrogram by using the weight matrix λ.
[0012] The joint training module is connected with the pre-enhancement module and the spectrogram compensation module, and is used to jointly train and optimize the pre-enhancement module and the spectrogram compensation module.
[0013] The single-channel speech enhancement system based on spectrogram compensation of the present application has the following structural features:
[0014] Preferably, the pre-enhancement module is a speech separation system trained by using a deep neural network.
[0015] Preferably, the output of the pre-enhancement module includes a pre-enhanced masking value
[0016] Preferably, the masking value is used to calculate the amplitude spectrum of the estimated target clean speech
[0017] Preferably, the spectrogram compensation module obtains a weight matrix λ by using the input generated by the pre-enhancement module.
[0018] Preferably, the final spectrogram after spectrogram compensation is calculated according to the weight matrix λ
[0019] Preferably, the final spectrogram after spectrogram compensation is used to calculate the enhanced speech signal in the time domain
[0020] Preferably, the input of the joint training module includes a pre-enhancement target function
[0021] Preferably, the input of the joint training module comprises a spectrogram compensation target function J SI-SNR .
[0022] Preferably, the total training target function J is calculated according to the pre-enhancement target function and the spectrogram compensation target function J SI-SNR The calculation formula of the total training target function J is:
[0023]
[0024] Wherein, a represents the weight of the pre-enhancement module and the spectrogram compensation module.
[0025] The present application has the following advantages:
[0026] The present application is a single-channel speech enhancement system based on spectrogram compensation, comprising a pre-enhancement module, a spectrogram compensation module and a joint training module; the pre-enhancement module is used to remove part of the interference signal in the speech; the spectrogram compensation module is connected with the pre-enhancement module and is used to obtain a weight matrix of spectrogram compensation, and the weight matrix is used to fuse the pre-enhanced spectrogram and the original input spectrogram; the joint training module is connected with the pre-enhancement module and the spectrogram compensation module and is used to jointly train and optimize the pre-enhancement module and the spectrogram compensation module.
[0027] The single-channel speech enhancement system based on spectrogram compensation has the following advantages:
[0028] (1) In the pre-enhancement module, the depth neural network is used to pre-enhance the speech containing noise to remove most of the background noise, so as to achieve the purpose of pre-enhancing the input speech signal;
[0029] (2) In the spectrogram compensation module, the weight matrix of spectrogram compensation is estimated first, and the pre-enhanced spectrogram and the original input spectrogram are fused by using the matrix, so as to realize the spectrogram compensation and further enhance the pre-enhanced speech, in order to solve the problem of speech distortion and loss of important speech information;
[0030] (3) In the joint training module, the pre-enhancement module and the spectrogram compensation module are jointly optimized, so as to improve the quality of the speech after the spectrum compensation while ensuring the pre-enhancement performance. Therefore, the separated speech is clearer, more understandable and better in sound quality than the method based on depth learning alone.
[0031] The single-channel speech enhancement system based on spectrogram compensation can maintain the enhanced speech with high sound quality, clear speech and intelligibility in a noisy background environment. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is a structural schematic diagram of a single-channel speech enhancement system based on a spectrum compensation of the present application;
[0033] Figure 2 is a structural schematic diagram of a pre-enhancement module in a single-channel speech enhancement system based on a spectrum compensation of the present application;
[0034] Figure 3 is a structural schematic diagram of a spectrum compensation module in a single-channel speech enhancement system based on a spectrum compensation of the present application;
[0035] Figure 4 is a structural schematic diagram of a joint training module in a single-channel speech enhancement system based on a spectrum compensation of the present application. DETAILED DESCRIPTION
[0036] The preferred embodiments of the present application will be described in detail with reference to the drawings, so that the objects, technical solutions and superiorities of the present application are more clearly understood, and the protection scope of the present application is more clearly defined. The present application will be further described in detail below with reference to the specific embodiments and with reference to the drawings.
[0037] It should be noted that, in the drawings or the description, similar or identical parts are denoted by the same reference numerals. In the drawings, the implementation is simplified or facilitated. Furthermore, the implementation not shown or described in the drawings is the form known to those skilled in the art. In addition, although this document can provide examples of parameters including specific values, it should be understood that the parameters do not necessarily equal the corresponding values, but can be approximately equal to the corresponding values within an acceptable error tolerance or design constraint.
[0038] As Figures 1-4 , a single-channel speech enhancement system based on a spectrum compensation of the present application includes a pre-enhancement module, a spectrum compensation module and a joint training module;
[0039] The pre-enhancement module is configured to remove part of the interference signals in the speech.
[0040] The spectrum compensation module is connected to the pre-enhancement module and configured to obtain a weight matrix λ of spectrum compensation, and fuse the pre-enhanced spectrum and the original input spectrum by using the weight matrix λ.
[0041] The joint training module is connected to the pre-enhancement module and the spectrum compensation module, and configured to jointly train and optimize the pre-enhancement module and the spectrum compensation module.
[0042] The pre-enhancement module is a speech separation system trained by using a deep neural network.
[0043] The output of the pre-enhancement module includes pre-enhanced masking values
[0044] Through the masking values The amplitude spectrum of the estimated target clean speech is calculated
[0045] First, the noisy speech is pre-enhanced by the pre-enhancement module to remove most of the background noise. Since the speech distortion will lose a lot of speech information, the spectrogram compensation module is used to compensate the spectrogram of the pre-enhanced speech and the original input speech. Finally, the joint optimization method is used to further improve the sound quality and intelligibility of speech enhancement.
[0046] The pre-enhancement module is used to remove most of the interference signals to play a pre-enhancement role and is trained by using a deep neural network. The output of the pre-enhancement module includes two parts: pre-enhanced masking values and the input of the spectrogram compensation module. Then, the amplitude spectrum of the estimated target clean speech is calculated by multiplying the amplitude spectrum of the original input speech and the pre-enhanced masking values The mean square error between the estimated amplitude spectrum and the true amplitude spectrum is calculated as the training objective function.
[0047] As Figure 2 is a structural diagram of the pre-enhancement module of the single-channel speech enhancement system based on spectrogram compensation. Figure 2 The pre-enhancement module is used to remove most of the interference signals to play a pre-enhancement role and is trained by using a deep neural network. The output of the pre-enhancement module includes two parts: pre-enhanced masking values and the input of the spectrogram compensation module h in , as shown in the following formula (1).
[0048]
[0049] Wherein, |Y(t, f)| represents the amplitude spectrum of the input noisy speech, t and f are the frame number and frequency block number of the input speech, respectively; f DNN (*) represents a mapping function based on a deep neural network. For convenience of description, we omit (t, f) in the following.
[0050] The pre-enhanced masking values can be obtained by multiplying the masking values with the amplitude spectrum |Y| of the original input speech to obtain the amplitude spectrum of the pre-enhanced speech as shown in the following formula (2).
[0051]
[0052] where ⊙ denotes the dot product.
[0053] For the pre-enhancement module, its training objective function is To calculate the mean square error between the pre-enhanced speech and the target clean speech magnitude spectrum, see the following formula (3).
[0054]
[0055] where TF denotes the number of time-frequency units, denotes the square Frobenius norm.
[0056] The speech spectrum compensation module obtains a weight matrix λ from the input generated by the pre-enhancement module.
[0057] According to the weight matrix λ, the final speech spectrum after compensation is obtained
[0058] According to the final speech spectrum after compensation The enhanced speech signal in the time domain is obtained
[0059] Based on the speech spectrum compensation module, the pre-enhancement module is connected, and is mainly used to solve the information loss problem caused by speech distortion of the pre-enhancement module. First, the input generated by the pre-enhancement module is used to estimate the weight matrix λ of speech spectrum compensation for each time-frequency unit; because the speech spectrum of the original input has no information loss, therefore, according to the weight matrix λ, the pre-enhanced speech features and the original input speech features are linearly weighted to realize speech spectrum compensation to recover the lost speech information due to speech distortion, further enhance the pre-enhanced speech, and improve the performance of speech enhancement.
[0060] The magnitude spectrum after speech spectrum compensation is used as the final enhanced feature. Then, the phase spectrum of the original input speech and the magnitude spectrum after speech spectrum compensation are inverse Fourier transformed to obtain the enhanced speech in the time domain. Finally, the scale invariant signal-to-noise ratio (SI-SNR) between the enhanced speech in the time domain and the target clean speech signal is calculated as the objective function of the module, and the SI-SNR is maximized.
[0061] Figure 3 is a structure diagram of a speech spectrum compensation module of a speech spectrum compensation-based single-channel speech enhancement system, which is connected with the pre-enhancement module, and is used to make up for the information loss problem caused by speech distortion. The speech spectrum compensation module first obtains h in from the deep neural network to obtain a deep representation h mend , see the following formula (4).
[0062] h mend = fDNN (h in ) (4)
[0063] Then, the deep representation h mend is subjected to a Sigmoid operation to obtain a weight matrix λ for the spectrum compensation, as shown in equation (5).
[0064]
[0065] where σ represents a Sigmoid activation function.
[0066] Taking λ as the weight matrix of the pre-enhanced spectrum and 1-λ as the weight matrix of the original input spectrum, the final spectrum after the spectrum compensation can be obtained by equation (6). as shown in equation (6).
[0067]
[0068] Finally, the enhanced spectral features are subjected to an inverse Fourier transform ISTFT together with the original noisy phase spectrum Φ y to obtain an enhanced speech signal in the time domain as shown in equation (7).
[0069]
[0070] For the training target of the spectrum compensation module, we directly define it on the time-domain speech signal, and take the scale-invariant signal-to-noise ratio (SI-SNR) as the objective function J SI-SNR , as shown in equations (8), (9) and (10).
[0071]
[0072] where x taget represents the target signal, x represents the target clean speech signal, represents the error signal, ||x|| 2 = <x, x> represents the energy of the signal.
[0073] The input of the joint training module includes a pre-enhancement objective function
[0074] The input of the joint training module includes a spectrum compensation objective function J SI-SNR .
[0075] According to the pre-enhancement objective function and the spectrum compensation objective function J SI-SNR , the calculation formula of the total training objective function J is:
[0076]
[0077] wherein, alpha represents the weight of the pre-enhancement module and the spectrogram compensation module.
[0078] The joint training module is used for jointly optimizing the respective modules, including the pre-enhancement module and the spectrogram compensation module.
[0079] Figure 4 The joint training module of the single-channel speech enhancement system based on spectrogram compensation is shown in the figure.
[0080] wherein, alpha represents the weight of the pre-enhancement module and the spectrogram compensation module.
[0081] In summary, the pre-enhancement module is used for pre-enhancing the input noisy speech to remove most of the noise signals. The final output of the whole speech enhancement system is shown in the figure.
[0082] First, a deep learning-based speech separation system is trained as a pre-enhancement module to pre-enhance the input noisy speech and remove most of the noise signals.
[0083] The spectrogram compensation module is connected with the pre-enhancement module and is used for obtaining a weight matrix of spectrogram compensation to perform spectrogram compensation on the pre-enhanced speech.
[0084] The joint training module is used for jointly training and optimizing the pre-enhancement module and the spectrogram compensation module.
[0085] The single-channel speech enhancement system based on spectrogram compensation has the following beneficial effects:
[0086] (1) In the pre-enhancement module, a deep neural network is used to pre-enhance the speech containing noise to remove most of the background noise, thereby achieving the purpose of pre-enhancing the input speech signal.
[0087] (2) In the spectrogram compensation module, a weight matrix of spectrogram compensation is first estimated to fuse the pre-enhanced spectrogram and the original input spectrogram, thereby achieving the function of spectrogram compensation and further enhancing the pre-enhanced speech.
[0088] (3) In the present application, in the joint training module, the joint optimization pre-enhancement module and the spectrogram compensation module are adopted, so that the quality of the speech after the spectrum compensation can be improved while the pre-enhancement performance is guaranteed.
[0089] The present application models the input noisy speech by using pre-enhancement and spectrogram compensation, so that the enhanced speech is more faithful, the perceptual quality and intelligibility are higher, and the performance of the speech enhancement system is improved.
[0090] In the single-channel speech enhancement system based on spectrogram compensation, a pre-enhancement module based on deep learning is constructed to pre-enhance the input noisy speech and remove most of the noise signals.
[0091] It will be obvious to a person skilled in the art that, as the application is not limited to the details of the foregoing exemplary embodiments but can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application, the application can be implemented in other concrete forms. Therefore, the embodiments should be considered exemplary and non-limiting, and the scope of the application is defined by the appended claims rather than the foregoing description, and all changes falling within the meaning and range of the equivalent elements of the claims are intended to be encompassed by the application. Any reference signs in the claims should not be considered limiting of the claims involved.
[0092] Furthermore, it should be understood that, although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the present specification is described in this way only for the sake of clarity, and a person skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by a person skilled in the art.
Claims
1. A single channel speech enhancement system based on spectrogram compensation, characterized in that, The pre-enhancement module, the spectrogram compensation module and the joint training module are included. The pre-enhancement module is configured to remove part of the interference signal in the voice. The spectrogram compensation module is connected with the pre-enhancement module and configured to obtain a weight matrix λ of spectrogram compensation, and fuse the pre-enhanced spectrogram and the original input spectrogram by using the weight matrix λ. The joint training module is connected with the pre-enhancement module and the spectrogram compensation module, and configured to jointly train and optimize the pre-enhancement module and the spectrogram compensation module.
2. The spectral basis compensation based single channel speech enhancement system according to claim 1, characterized in that, The pre-enhancement module is a voice separation system trained by using a deep neural network.
3. The spectral basis compensation based single channel speech enhancement system according to claim 1, wherein, The output of the pre-emphasis module includes pre-emphasized masking values 4. The spectral basis compensation based single channel speech enhancement system according to claim 3, wherein, by the masking values the amplitude spectrum of the estimated target clean speech is computed 5. The spectral basis compensation based single channel speech enhancement system according to claim 1, wherein, The spectrogram compensation module obtains the weight matrix λ by using the input generated by the pre-enhancement module.
6. The spectral basis compensation based single channel speech enhancement system according to claim 5, characterized in that, computing a final spectrum compensated spectrum from the weight matrix λ 7. The spectral basis compensation based single channel speech enhancement system according to claim 6, characterized in that, the compensated spectrogram the enhanced speech signal on the time domain 8. The spectral basis compensation based single channel speech enhancement system according to claim 1, wherein, The input of the joint training module includes a pre-enhanced objective function 9. The spectral basis compensation based single channel speech enhancement system according to claim 8, characterized in that, The input of the joint training module includes a spectrogram compensation objective function J SI-SNR .
10. The spectral basis compensation based single channel speech enhancement system according to claim 9, characterized in that, According to the pre-enhanced target function and the speech spectrum compensation target function J SI-SNR The calculation formula of the total training target function J is: Wherein, α represents the weight of the pre-enhancement module and the spectrogram compensation module.
Citation Information
Patent Citations
Speech enhancement method and device, electronic equipment and storage medium
CN112700786A
Speech enhancement model training method and device and speech enhancement method and device
CN113593594A