Single-channel Speech Enhancement Method Based on Masking Effect Using Complex Convolutional Recurrent Neural Network

Through the single-channel speech enhancement method of complex convolution recurrent neural network based on masking effect, speech is divided into sub-band parallel processing, combined with full-band feature modeling, speech amplitude and phase are directly enhanced, solving the problems of high complexity and limited effects in the existing technology, and achieving efficient speech enhancement effect.

CN114999510BActive Publication Date: 2025-07-11UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210467641.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-07-11
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

现有技术在语音增强中存在复杂度高且增强效果有限的问题,尤其在处理平稳或非平稳噪声时效果不佳。

Method used

The single-channel speech enhancement method of complex convolution recurrent neural network based on masking effect is adopted. By dividing single-channel speech into 22 subbands, the parallel complex convolution recurrent neural network model is processed, combined with full-band feature modeling, the amplitude and phase of the speech are directly enhanced, and the spectrum mode capture is captured using the human ear auditory masking effect.

Benefits of technology

In terms of enhancement effect and computing speed, the model size is reduced to 72% of the general neural network, while improving speech quality and intelligibility, achieving a PESQ effect of 97%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114999510B_ABST
    Figure CN114999510B_ABST
Patent Text Reader

Abstract

The present invention discloses a single-channel speech enhancement method based on a complex convolutional recurrent neural network using the masking effect, which includes: dividing the frequency dimension of the initial vector after Fourier transform into 22 sub-vectors with overlapping between adjacent sub-vectors closest to the critical band according to the frequency division method of the Bark band. The 22 sub-vectors represent 22 sub-bands and are fed into the corresponding 22 parallel complex convolutional recurrent network sub-band models; two full-band complex fully-connected layers are cascaded after the parallel sub-band models to obtain a complete complex convolutional recurrent neural network model based on the masking effect; using the complex ideal ratio mask (cIRM) as the training objective, and reconstructing the clean speech together with the cIRM and the original speech to be enhanced. The present invention can capture both the local spectral patterns within the sub-bands and the spectral patterns of the full band and the cross-dependencies between the sub-bands; making full use of the law of the auditory masking effect of the human ear, it has obvious advantages over general neural networks in terms of enhancement effect and computational speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech noise reduction, and particularly to a single-channel speech enhancement method based on a masking effect of a complex convolutional recurrent neural network. Background Art

[0002] Single-channel speech enhancement refers to eliminating or suppressing background noise under the condition of a single microphone to obtain higher speech quality and intelligibility. In recent years, deep learning methods have achieved excellent results in this regard. Especially in dealing with challenging scenarios such as non-stationary noise, deep learning methods are superior to traditional single-channel speech enhancement algorithms. Convolutional neural networks and recurrent neural networks are two widely used methods for speech enhancement. In 2019, [1] proposed an encoder-decoder network architecture for speech denoising, which is used to model the real and imaginary parts of the complex STFT spectrogram from the input mixture to the clean speech. A combination of multi-layer convolution and long short-term memory network (LSTM) is used to construct the encoder-decoder structure. The convolutional layer is responsible for extracting speech signal features, and the LSTM is responsible for capturing the temporal characteristics information before and after the speech. This article has very significant effects on improving speech quality and intelligibility, but the model size is large. Although both the amplitude and phase are improved compared with the traditional amplitude-only target, regarding the real and imaginary parts as two input channels and only using the same shared real-valued convolutional filter for convolution operation, this method is not restricted by the complex multiplication rule. In 2020, [2] proposed an improvement, using a complex network structure to replace the original traditional structure, making up for the defect that the original structure cannot simulate complex multiplication, and also shortening the depth of the convolutional layer. The volume of this model is reduced by half compared with [1], and the enhancement effect is also better than [1]. This model won the first place in the Real-Time Track (RT) of the 2020 DNS Challenge (Deep Noise Suppression Challenge, abbreviated as DNS). However, most of these methods are based on the full frequency band and cannot capture the frequency characteristics within the sub-band more accurately. To address this deficiency, the present invention proposes a parallel sub-band convolutional recurrent neural network structure based on the masking effect. It still uses a complex network structure to replace the traditional structure and proposes a combination of parallel sub-bands and the full frequency band. The sub-bands are divided according to the well-known critical band concept, and the sub-models correspond to the sub-bands one by one. 22 sub-band models capture the local spectral patterns within the corresponding critical sub-bands. After the parallel sub-bands, two fully connected layers are used to capture the spectral patterns of the full frequency band and the cross-dependencies between the sub-bands. This method makes full use of the law of the auditory masking effect of the human ear and enhances both the amplitude and phase of the speech at the same time. It has obvious advantages over general neural networks in terms of enhancement effect and calculation speed. Compared with [2], the model volume is only 72% of it, and the complexity is greatly reduced. However, in terms of the enhancement effect, the PESQ in this method reaches 97% of the method in [2]. In the case of needing to balance speech quality and computational cost at the same time, the present invention is superior to the method in [2].

[0003] Neural network-based speech enhancement methods can be divided into two major categories: time-frequency domain and time domain. Some time-domain methods directly learn the mapping from the original noisy signal waveform to the clean signal waveform to achieve end-to-end enhancement. These end-to-end speech enhancement methods avoid the step of manually extracting features and directly reconstruct the clean speech waveform. When there is enough training data, they generally perform better than methods that manually extract features. Time-frequency domain methods use time-frequency representations, and the log-magnitude power spectrum and mask are two common training objectives. Traditional speech enhancement methods usually enhance the magnitude spectrum of the signal and reconstruct the enhanced speech by combining the phase spectrum of the original speech. Some studies have shown that accurately estimating the phase spectrum can significantly improve the quality of the enhanced speech, so various phase enhancement algorithms have been proposed.

[0004] Under the interference of strong signals, nearby weak signals will be masked. Frequency components with higher loudness make the frequency components with lower loudness nearby less perceptible, and this phenomenon is called the masking effect. In 1940, Fletcher proposed the concept of the auditory critical band. He conducted a famous bandwidth experiment: when the bandwidth of the noise did not exceed the required frequency band range, the perceived loudness of the noise by the human ear did not change; when the bandwidth of the noise exceeded the width of the critical band, the noise would mask the pure tone and reduce the perceived sound. For this discovery, a special physiological unit - Bark was introduced, and 1 Bark corresponds to one critical band. The human hearing range is approximately 200 - 8000 Hz, so the full frequency band can be divided into sub-bands according to the Bark domain division standard, as Figure 1 shown.

[0005] [1] K.Tan, D.L.Wang, Complex spectral mapping with a convolutional recurrent network for monaural speech enhancement, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019.

[0006] [2] Y.Hu, Y.Liu, S.Lv, M.Xing, L.Xie, Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement (2020). Summary of the Invention

[0007] Problems to be solved by the present invention: Based on the problems of high complexity and limited speech enhancement effect existing in the prior art, the present invention provides a single-channel speech enhancement method based on a masked effect complex convolutional recurrent neural network, which has obvious advantages over general speech enhancement neural networks in terms of enhancement effect and model complexity, and can solve the problem of speech noise reduction under stationary or non-stationary noise interference.

[0008] Technical solution of the present invention: A single-channel speech enhancement method based on a masked effect complex convolutional recurrent neural network, comprising:

[0009] Step 1: Perform a discrete Fourier transform STFT with 1024 points on the single-channel original speech to be enhanced after framing and windowing to obtain an initial vector of length 513; then divide the initial vector of 513 points into 22 sub-vectors with overlapping adjacent sub-vectors according to the frequency division method of the Bark band. The 22 sub-vectors correspond to 22 sub-bands, and the overlapping points in the 22 overlapping sub-vectors are defined as error-tolerant points. The real part and the imaginary part of each sub-vector are used as two independent channels and fed into the corresponding 22 parallel complex convolutional recurrent neural network sub-band models; the input vector dimension of the sub-band model corresponds one-to-one with the length of the 22 sub-vectors; for the sake of simplicity of writing, hereinafter, the complex structure sub-band model is used to replace the complex convolutional recurrent neural network sub-band model. Concatenate the outputs of the complex structure sub-band models and input them into two full-band complex fully-connected layers for full-band feature modeling to obtain the speech to be enhanced as a complex ideal ratio mask (cIRM);

[0010] Step 2: Use the complex ideal ratio mask cIRM in Step 1 as the training target, reconstruct the frequency spectrum of the clean speech together with the cIRM and the original speech to be enhanced, and then perform an inverse Fourier transform ISTFT to finally obtain the enhanced speech.

[0011] The innovation points of the present invention are as follows: On the one hand, compared with most traditional neural networks that only enhance the speech amplitude spectrum, the present invention directly uses a complex network structure to enhance the complete frequency spectrum: according to the complex number operation rules, the real part and the imaginary part are used to simulate complex number operations to achieve the purpose of enhancing both the amplitude and the phase of the speech, improving the speech quality and intelligibility; on the other hand, the present invention divides the sub-bands and constructs sub-band models according to the law of the auditory masking effect of the human ear, and adopts a network structure combining sub-bands and full-bands, which can not only capture the local spectral patterns within the sub-bands, but also capture the spectral patterns of the full-band and the cross-dependence relationships between the sub-bands, and has obvious advantages over general neural networks in terms of enhancement effect and calculation speed.

[0012] In the said step 1: the frame length in the original speech to be enhanced after framing and windowing is 400, the frame shift is 100, the sampling rate of all audio is 16 kHz. After STFT transformation with 1024 points, two inputs, namely the real part and the imaginary part of the noisy signal, are obtained as follows:

[0013] Y(t,f) = S(t,f) + N(t,f)

[0014] Y = Y r + jY i

[0015] S = S r + jS i

[0016] Among them, Y(t,f) represents the single-channel original speech spectrum to be enhanced after STFT transformation, t represents the time dimension, f represents the frequency dimension, S(t,f) and N(t,f) represent the clean speech and background noise, Y and S represent the spectra of Y(t,f) and S(t,f), the subscripts r and i respectively represent the real part and the imaginary part of the spectrum, Y r and Y i are both 513-dimensional vectors. The frequency range of 0 - 8000 Hz is equally spaced and cut, and the corresponding frequency resolution is 8000(Hz) / 512 = 15.625(Hz); it is known that in the range of 0 - 8000 Hz, the critical band ranges of 22 standard Bark bands are: (20 - 100), (100 - 200), (200 - 300), (300 - 400), (400 - 510), (510 - 630), (630 - 770), (770 - 920), (920 - 1080), (1080 - 1270), (1270 - 1480), (1480 - 1720), (1720 - 2000), (2000 - 2320), (2320 - 2700), (2700 - 3150), (3150 - 3700), (3700 - 4400), (4400 - 5300), (5300 - 6400), (6400 - 7700), (7700 - 8000);

[0017] According to the above standards, 22 sub-bands are divided as follows: The ranges of the 22 sub-bands closest to the critical band are: (0 - 93.75), (93.75 - 203.125), (203.125 - 312.5), (296.875 - 406.25), (406.25 - 515.625), (515.625 - 625), (625 - 765.625), (765.625 - 921.875), (921.875 - 1078.125), (1078.125 - 1265.625), (1265.625 - 1484.375), (1484.375 - 1718.75), (1718.75 - 2000), (2000 - 2328.125), (2312.5 - 2703.125), (2703.125 - 3156.25), (3156.25 - 3703.125), (3703.125 - 4406.25), (4406.25 - 5296.875), (5296.875 - 6406.25), (6406.25 - 7703.125), (7703.125 - 8000), with the unit of Hz;

[0018] Under the condition that the frequency resolution is 15.625 (Hz), in order to improve the accuracy of the complex structure sub-band model, the concept of fault tolerance points is defined: In the above standard Bark band, there is a critical value between every two critical bands. For example, 100 Hz is the critical value between Bark1 and Bark2, and 200 Hz is the critical value between Bark1 and Bark2. When the frequency difference between the corresponding frequency of the point after STFT and the critical value is less than 8 Hz, this point is called a fault tolerance point; After calculation, there are 23 fault tolerance points in the entire frequency band, which are 93.75, 203.125, 296.875, 312.5, 406.25, 515.625, 625, 765.625, 921.875, 1078.125, 1265.625, 1484.375, 1718.75, 2000, 2312.5, 2328.125, 2703.125, 3156.25, 3703.125, 4406.25, 5296.875, 6406.25, 7703.125, with the unit of Hz; The fault tolerance points will be sent into both the left and right complex structure sub-band models for training at the same time; For these points, 23 groups of optimizable weights between 0 and 1 are introduced to calculate the outputs of 23 points. Each of the 23 groups has two weight coefficients and their sum is 1, as shown in the following formula:

[0019]

[0020] Among them, and respectively represent the outputs of the k-th fault-tolerant point in the i-th and j-th complex structure sub-band models, y k represents the actual output value of this point, and respectively represent the weights of the actual output value of this fault-tolerant point in the i-th and j-th sub-bands, and these weights are regarded as the complex structure sub-model parameters, and the optimal values are automatically obtained through gradient descent during training to ensure the reasonable output of the fault-tolerant point.

[0021] Abbreviate Y(t, f) as Y 2×513 , and divide it according to the above division method to obtain the inputs corresponding to 22 complex structure sub-band models where 2 represents the number of channels of the real part and the imaginary part, 513 represents the frequency dimension, and the superscript represents the number of the complex structure sub-band model to be input.

[0022] In step 1, the 22 complex structure sub-band models adopt the same encoding-decoding structure. There is a Long Short-Term Memory (LSTM) layer in the middle of the encoder and the decoder. One layer of transposed convolution is used as the encoder, and one layer of convolution is used as the decoder; and all structures in the model are complex network structures. These complex network structures simulate complex multiplication using real values to model the correlation between amplitude and phase.

[0023] The specific structure of the complex structure sub-band model is: one layer of complex transposed convolution + one layer of complex LSTM + one layer of complex convolution. The output of the complex convolution layer is obtained by combining four convolutions according to complex multiplication. The complex filter matrix W = W r + jW i is convolved with the complex vector X = X r + jX i where W r and W i are real matrices, and X r and X j are real vectors; the real part is used to simulate complex calculations. The vector X is convolved through the filter W to obtain:

[0024] F out =(X r *W r -X i *W i )+j(X r *W i +X i *W r )

[0025] where F outis the output of the complex convolutional layer. Similarly, there are also complex LSTM and complex fully connected layers, with the output being F out They are respectively defined as:

[0026] F rr = LSTM r (X r ); F ir = LSTM r (X i )

[0027] F ri = LSTM i (X r ); F ii = LSTM i (X i )

[0028] F out = (F rr - F ii ) + j(F ri + F ir )

[0029] F rr = Linear r (X r ); F ir = Linear r (X i )

[0030] F ri = Linear i (X r ); F ii = Linear i (X i )

[0031] F out = (F rr - F ii ) + j(F ri + F ir )

[0032] Among them, LSTM and Linear respectively represent the traditional LSTM and Linear neural network methods, and the subscripts r and i respectively represent the real part and the imaginary part of the corresponding network;

[0033] The real part and the imaginary part are two input channels, and the number of input and output channels of the transposed convolution and the convolutional layer is the same. Each complex structure sub-band model consists of three parts: the first part is the complex transposed convolutional layer + complex normalization + complex Relu activation function, the second part is the complex LSTM layer + complex fully connected layer + complex normalization + complex Relu activation function, and the third part is the complex convolutional layer;

[0034] After obtaining the output of the sub-band model, let the output be Concatenate it into a full-band frequency spectrum O 2×513 , and send it into two complex fully-connected layers with a length of 513.

[0035] In step 2, after passing through two complex fully-connected layers, the composite sub-band model outputs the complex ideal ratio mask cIRM of the speech. For the training objective, the complex ideal ratio mask is multiplied by the complex spectrum Y of the original speech to be enhanced to reconstruct the enhanced signal The complex ideal ratio mask is as follows:

[0036]

[0037] where r and i represent the real part and the imaginary part respectively;

[0038] The cIRM is converted into polar coordinate form as follows:

[0039]

[0040]

[0041] where and represent the real part and the imaginary part of the estimated value cIRM respectively., and represent the amplitude and the phase of the estimated value cIRM respectively;

[0042] Finally, the output cIRM and the spectrum of the original speech to be enhanced are used together to reconstruct the clean speech signal;

[0043]

[0044]

[0045] where and represent the amplitude and the phase of the enhanced speech respectively, and represent the amplitude and the phase of the original speech to be enhanced respectively.

[0046] The advantages of the present invention compared with the prior art are as follows: on the one hand, compared with most neural network structures that only enhance the magnitude spectrum, the present invention directly adopts a complex network structure, improving the speech quality and intelligibility; on the other hand, the present invention adopts a network structure combining sub-bands and full-bands, which can capture both the local spectral patterns within the sub-bands (the full-band represents the whole and the sub-bands represent the parts) and the spectral patterns of the full-band and the cross-dependencies between sub-bands, making full use of the law of the auditory masking effect of the human ear, and having obvious advantages over general neural networks in terms of enhancement effect and calculation speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings and tables required in the description of the embodiments. Obviously, the drawings and tables in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 Simultaneous masking curve of sound at 1 KHz;

[0049] Figure 2 Specific methods and steps of the present invention;

[0050] Figure 3 Specific structure of the sub-band model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0052] As Figure 1 shown, the masking is a pure tone of 1 kHz, as shown by the "thick black line" in Figure 1. There are two voices with weaker sound intensities nearby, as shown by the "thick gray lines". These two weak sounds are below the masking threshold of the pure tone, so they are masked by it. Figure 1 The "thin black line" at the bottom in

[0053] As Figure 2 shown, a single-channel speech enhancement method based on the masking effect of a parallel sub-band complex convolutional recurrent neural network of the present invention

[0054] The pure human voices and noises are respectively from the public dataset WSJ0 and the musan noise set, including 160,000 mixed voices, from 39 males and 38 females respectively. There are 800 types of noises, with a total duration of 134 hours. The signal-to-noise ratio of the mixed voices is randomly set between -5dB and 20dB. The test set contains 27,000 mixed voices, from 6 untrained speakers and 12 types of noises, with a total duration of 38 hours.

[0055] The present invention specifically includes the following steps:

[0056] Step 1: Perform a discrete Fourier transform STFT with 1024 points on the single-channel original voice to be enhanced after framing and windowing, to obtain an initial vector with a length of 513; then divide the initial vector with 513 points into 22 sub-vectors with overlapping between adjacent sub-vectors according to the frequency division method of the Bark band, and define the overlapping points in the 22 overlapping sub-vectors as fault-tolerant points. The real part and the imaginary part of each sub-vector are used as two independent channels and fed into the corresponding 22 parallel complex structure sub-band models; the input vector dimension of the complex structure sub-band model corresponds one by one to the length of the 22 sub-vectors;

[0057] Frame and window the original voice with noise to be enhanced, where the frame length is 400 and the frame shift is 100 (the sampling rate of all audio is 16khz). After the STFT transform with 1024 points, two inputs of the real part and the imaginary part of the noisy signal are obtained, as follows:

[0058] Y(t,f) = S(t,f) + N(t,f)

[0059] Y = Y r +jY i

[0060] S = S r +jS i

[0061] Among them, Y(t,f) represents the spectrum of the single-channel original voice to be enhanced after the STFT transform, t represents the time dimension, f represents the frequency dimension. Similarly, S(t,f) and N(t,f) represent the clean voice and the background noise, Y and S represent the spectra of Y(t,f) and S(t,f), and the subscripts r and i respectively represent the real part and the imaginary part of the spectrum, Y r and Y iVectors of 513 dimensions each, the frequency range from 0 to 8000 Hz is equally spaced and cut, and the corresponding frequency resolution is 8000 (Hz) / 512 = 15.625 (Hz). It is known that within the range of 0 to 8000 Hz, the critical band ranges of 22 standard Bark bands are: (20 - 100), (100 - 200), (200 - 300), (300 - 400), (400 - 510), (510 - 630), (630 - 770), (770 - 920), (920 - 1080), (1080 - 1270), (1270 - 1480), (1480 - 1720), (1720 - 2000), (2000 - 2320), (2320 - 2700), (2700 - 3150), (3150 - 3700), (3700 - 4400), (4400 - 5300), (5300 - 6400), (6400 - 7700), (7700 - 8000).

[0062] According to the above standards, 22 sub - frequency bands are divided. Under the condition of a frequency resolution of 15.625 (Hz), in order to improve the model accuracy, the present invention defines the concept of a fault - tolerant point: in the above - mentioned standard Bark bands, there is a critical value between every two critical bands. For example, 100 Hz is the critical value between Bark1 and Bark2, and 200 Hz is the critical value between Bark1 and Bark2. When the difference between the corresponding frequency of the point after STFT and the critical value is less than 8 Hz, this point is called a fault - tolerant point. After calculation, there are 24 fault - tolerant points in the entire frequency band, which are 93.75, 203.125, 296.875, 312.5, 406.25, 515.625, 625, 765.625, 921.875, 1078.125, 1265.625, 1484.375, 1718.75, 2000, 2312.5, 2328.125, 2703.125, 3156.25, 3703.125, 4406.25, 5296.875, 6406.25, 7703.125, with the unit of Hz. Different from other points that are only sent into one sub - band model, the fault - tolerant points will be sent into both the left and right sub - band models for training at the same time. After calculation, there are 23 fault - tolerant points in the entire frequency band. For these points, 23 groups of optimizable weights between 0 and 1 are introduced to calculate their actual outputs. Each group has two weight coefficients and their sum is 1, as shown in the following formula:

[0063]

[0064] Among them, and respectively represent the outputs of the k - th fault - tolerant point in the i - th and j - th sub - band models, yk represents the actual output value of this point, and respectively represent the weights of the actual output value of this fault-tolerant point in the i-th and j-th sub-bands, and these weights are regarded as model parameters and automatically obtain the optimal values through gradient descent during training to ensure the reasonable output of the fault-tolerant point. Further, the ranges of the 22 sub-bands closest to the critical band are: (0 - 93.75), (93.75 - 203.125), (203.125 - 312.5), (296.875 - 406.25), (406.25 - 515.625), (515.625 - 625), (625 - 765.625), (765.625 - 921.875), (921.875 - 1078.125), (1078.125 - 1265.625), (1265.625 - 1484.375), (1484.375 - 1718.75), (1718.75 - 2000), (2000 - 2328.125), (2312.5 - 2703.125), (2703.125 - 3156.25), (3156.25 - 3703.125), (3703.125 - 4406.25), (4406.25 - 5296.875), (5296.875 - 6406.25), (6406.25 - 7703.125), (7703.125 - 8000), in the unit of Hz. Abbreviate Y(t, f) as Y 2×513 and divide it according to the above division method to obtain the inputs corresponding to 22 complex structure sub-band models where 2 represents the number of two channels, real part and imaginary part, 513 represents the frequency dimension, and the superscript represents the number of the complex structure sub-band model to be input.

[0065] Concatenate the outputs of the complex structure sub-band models and input them into two full-band complex fully connected layers for full-band feature modeling to obtain the speech to be enhanced as a complex ideal ratio mask (cIRM).

[0066] Such as Figure 3As shown, 22 complex structure sub-band models adopt the same encoding-decoding structure. There is a Long Short-Term Memory (LSTM) layer in the middle of the encoder and the decoder. An anti-convolution layer is used as the encoder, and a convolution layer is used as the decoder. All structures in the model are complex network structures. These complex network structures simulate complex number multiplication using real values to model the correlation between amplitude and phase. The specific structure of the sub-band model is: one layer of complex anti-convolution + one layer of complex LSTM + one layer of complex convolution. The output of the complex convolution layer is obtained by combining four traditional convolutions according to complex number multiplication. The complex filter matrix W = W r + jW i is convolved with the complex vector X = X r + jX i , where W r and W i are real matrices, and X r and X j are real vectors. The real part is used to simulate complex number calculations. By convolving the vector X with the filter W, we get:

[0067] F out =(X r * W r - X i * W i )+ j(X r * W i + X i * W r )

[0068] where F out is the output of the complex convolution layer. Similarly, there are also complex LSTM and complex fully connected layers, and the outputs F out are respectively defined as:

[0069] F rr = LSTM r (X r ) ; F ir = LSTM r (X i )

[0070] F ri = LSTM i (X r ) ; F ii = LSTM i (X i )

[0071] F out =(F rr - F ii )+ j(F ri + F ir )

[0072] F rr = Linear r (X r );F ir = Linear r (X i )

[0073] F ri = Linear i (X r );F ii = Linear i (X i )

[0074] F out = (F rr - F ii ) + j(F ri + F ir )

[0075] Among them, LSTM and Linear respectively represent the traditional LSTM and Linear neural network methods, and the subscripts r and i respectively represent the real part and the imaginary part of the corresponding network;

[0076] The real part and the imaginary part are two input channels, and the number of input and output channels of the transposed convolution and the convolution layer is the same. Each sub-band model consists of three parts: the first part is the complex transposed convolution layer + complex normalization + complex Relu activation function, the second part is the complex LSTM layer + complex fully connected layer + complex normalization + complex Relu activation function, and the third part is the complex convolution layer.

[0077] After obtaining the output of the sub-band model, let the output be Concatenate it into a full-band frequency spectrum O 2×513 , and send it into two complex fully connected layers with a length of 513.

[0078] Step 2: Use the complex ideal ratio mask (cIRM) as the training objective, use the output obtained in Step 2 as the cIRM to reconstruct the frequency spectrum of the clean speech together with the original speech to be enhanced, and then perform the inverse Fourier transform (ISTFT) to finally obtain the enhanced speech.

[0079] After passing through two complex fully connected layers, the model outputs the complex ideal ratio mask cIRM of the speech. For the training objective, use the complex ideal ratio mask to multiply the complex spectrum Y of the speech to be enhanced to reconstruct the enhanced signal The complex ideal ratio mask is an ideal mask defined in the complex domain:

[0080]

[0081] Among them, r and i respectively represent the real part and the imaginary part. Correspondingly, cIRM can be converted into the polar coordinate form as follows:

[0082]

[0083]

[0084] Among them, and respectively represent the real part and the imaginary part of the estimated value cIRM. and respectively represent the amplitude and the phase of the estimated value cIRM. Finally, the output cIRM and the spectrum of the original speech to be enhanced are used to reconstruct the clean speech signal;

[0085]

[0086]

[0087] Among them and respectively represent the amplitude and the phase of the enhanced speech, and respectively represent the amplitude and the phase of the original speech to be enhanced.

[0088] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A single-channel speech enhancement method based on a complex convolutional recurrent neural network with masking effect, characterized in that, Including: Step 1: Perform a discrete Fourier transform STFT with 1024 points on the single-channel original speech to be enhanced after framing and windowing to obtain an initial vector of length 513; then divide the initial vector of 513 points into 22 sub-vectors with overlapping between adjacent sub-vectors according to the frequency division method of the Bark band. The 22 sub-vectors correspond to 22 sub-bands, and the overlapping points in the 22 overlapping sub-vectors are defined as fault-tolerant points. The real and imaginary parts of each sub-vector are used as two independent channels and fed into the corresponding 22 parallel complex convolutional recurrent network sub-band models; the input vector dimension of the complex convolutional recurrent network sub-band model corresponds one-to-one with the length of the 22 sub-vectors; splice the outputs of the complex convolutional recurrent network sub-band models and input them into two full-band complex fully connected layers for full-band feature modeling; after the above operations, train the complex convolutional recurrent network sub-band model to obtain the complex ideal ratio mask cIRM of the clean speech. Step 2: Use the complex ideal ratio mask cIRM in Step 1 as the training target, reconstruct the frequency spectrum of the clean speech together with the cIRM and the original speech to be enhanced, and then perform the inverse Fourier transform ISTFT to finally obtain the enhanced speech.

2. The single-channel speech enhancement method based on the masking effect using a complex convolutional recurrent neural network according to claim 1, wherein In the said Step 1: The frame length in the original speech to be enhanced after framing and windowing is 400, the frame shift is 100, the sampling rate of all audio is 16 kHz. After the STFT transform with 1024 points, two inputs, namely the real and imaginary parts of the noisy signal, are obtained as follows: Y(t,f) = S(t,f) + N(t,f) Y = Y r + jY i S = S r + jS i Among them, Y(t,f) represents the single-channel original speech spectrum to be enhanced after STFT transformation, t represents the time dimension, f represents the frequency dimension, S(t,f) and N(t,f) represent the clean speech and background noise, Y and S represent the spectra of Y(t,f) and S(t,f), and the subscripts r and i represent the real and imaginary parts of the spectrum, respectively, Y r and Y i are both 513-dimensional vectors. The frequency range of 0 - 8000 Hz is equally spaced and cut, and the corresponding frequency resolution is 8000(Hz) / 512 = 15.625(Hz); it is known that within the range of 0 - 8000 Hz, 22 sub-bands are divided, and the specific division is as follows: the ranges of the 22 sub-bands closest to the critical band are: (0 - 93.75), (93.75 - 203.125), (203.125 - 312.5), (296.875 - 406.25), (406.25 - 515.625), (515.625 - 625), (625 - 765.625), (765.625 - 921.875), (921.875 - 1078.125), (1078.125 - 1265.625), (1265.625 - 1484.375), (1484.375 - 1718.75), (1718.75 - 2000), (2000 - 2328.125), (2312.5 - 2703.125), (2703.125 - 3156.25), (3156.25 - 3703.125), (3703.125 - 4406.25), (4406.25 - 5296.875), (5296.875 - 6406.25), (6406.25 - 7703.125), (7703.125 - 8000), with the unit of Hz; Under the condition of a frequency resolution of 15.625 (Hz), in order to improve the accuracy of the complex convolutional recurrent network sub-band model, the concept of a fault-tolerant point is defined: in the above standard Bark band, there is a critical value between every two critical bands. When the frequency corresponding to the point after STFT differs from the critical value by less than 8 Hz, this point is called a fault-tolerant point; after calculation, there are 23 fault-tolerant points in the entire frequency band, which are 93.75, 203.125, 296.875, 312.5, 406.25, 515.625, 625, 765.625, 921.875, 1078.125, 1265.625, 1484.375, 1718.75, 2000, 2312.5, 2328.125, 2703.125, 3156.25, 3703.125, 4406.25, 5296.875, 6406.25, 7703.125, with the unit of Hz; the fault-tolerant points will be fed into the left and right complex convolutional recurrent network sub-band models for training at the same time; for these points, introduce 23 sets of optimizable weights between 0 and 1 to calculate the outputs of the 23 points. Each set in the 23 sets has two weight coefficients and their sum is 1, as shown in the following formula: Among them, and represent the outputs of the k-th fault-tolerant point in the i-th and j-th complex convolutional recurrent network sub-band models respectively, and y k represents the actual output value of this point. and represent the weights of the actual output value of this fault-tolerant point in the i-th and j-th sub-bands respectively, and these weights are regarded as the parameters of the complex convolutional recurrent network sub-band model, and the optimal values are automatically obtained through gradient descent during training to ensure the reasonable output of the fault-tolerant point; Abbreviate Y(t, f) as Y 2×513 , and divide it according to the above division method to obtain the input corresponding to the 22 sub-complex-convolution cyclic network sub-band models where 2 represents the number of two channels of the real part and the imaginary part, 513 represents the frequency dimension; the superscript represents the number of the complex structure sub-band model to be input.

3. The single-channel speech enhancement method based on the masking effect of the complex convolutional recurrent neural network according to claim 1, characterized in that In the said step 1, the 22 complex convolutional recurrent network sub-band models adopt the same encoding-decoding structure. There is an LSTM layer in the middle of the encoder and the decoder. A deconvolution layer is used as the encoder and a convolution layer is used as the decoder. Moreover, all the structures in the model are complex network structures, and these complex network structures simulate complex multiplication using real values to model the correlation between amplitude and phase.

4. The method for single-channel speech enhancement based on a masked effect complex convolutional recurrent neural network according to claim 3, characterized in that: The specific structure of the complex convolutional recurrent network subband model is as follows: one layer of complex transposed convolution + one layer of complex LSTM + one layer of complex convolution. The output of the complex convolution layer is obtained by combining four convolutions according to complex multiplication. The complex filter matrix W = W r + jW i is convolved with the complex vector X = X r + jX i , where W r and W i are real matrices, and X r and X j are real vectors; the real part is used to simulate complex calculations. By convolving the vector X with the filter W, we get: F out = (X r * W r - X i * W i ) + j(X r * W i + X i * W r ) Among them, F out is the output of the complex convolutional layer. Similarly, there are also complex LSTM and complex fully connected layers, and the output F out is respectively defined as: F rr = LSTM r (X r ); F ir = LSTM r (X i ) F ri = LSTM i (X r ); F ii = LSTM i (X i ) F out = (F rr - F ii ) + j(F ri + F ir ) F rr = Linear r (X r );F ir = Linear r (X i ) F ri = Linear i (X r ) ; F ii = Linear i (X i ) F out = (F rr - F ii ) + j(F ri + F ir ) Among them, LSTM and Linear respectively represent the traditional LSTM and Linear neural network methods, and the subscripts r and i respectively represent the real part and the imaginary part of the corresponding network. The real part and the imaginary part are two input channels. The number of input and output channels of the deconvolution and convolution layers is the same. Each complex convolutional recurrent network sub-band model consists of three parts: the first part is a complex deconvolution layer + complex normalization + complex Relu activation function, the second part is a complex LSTM layer + complex fully connected layer + complex normalization + complex Relu activation function, and the third part is a complex convolution layer. After obtaining the output of the complex structure sub-band model, let the output be Concatenate them into a full-band frequency spectrum O 2×513 , and send it into two complex fully-connected layers with a length of 513.

5. The single-channel speech enhancement method based on the masking effect using a complex convolutional recurrent neural network according to claim 1, wherein In the said step 2, The complex ideal ratio mask cIRM is as follows: Among them, r and i respectively represent the real part and the imaginary part. The cIRM is converted into the polar coordinate form as follows: Among them, and represent the real part and the imaginary part of the estimated value cIRM, respectively., and represent the magnitude and the phase of the estimated value cIRM, respectively; Finally, the output cIRM and the spectrum of the original speech to be enhanced are used together to reconstruct the clean speech signal. where and represent the amplitude and phase of the enhanced speech respectively, and represent the amplitude and phase of the single-channel original speech to be enhanced respectively.

Citation Information

Patent Citations

  • Binaural speech enhancement method based on deep learning

    CN109448751A

  • Speaker-independent single-channel voice separation method

    CN111583954A