Single-channel speech noise reduction method and device based on deep learning
By using NetLSTM subnetwork and multi-layer stacked CNN in speech noise reduction, the problems of large parameters and poor ConvLSTM performance are solved, and efficient speech noise reduction effect is achieved.
Patent Information
- Application Number
- CN202210568276.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-24
AI Technical Summary
Traditional network structures have large parameters in speech timing information processing, while ConvLSTM has poor performance in the voice field, resulting in poor voice noise reduction effect.
Using the NetLSTM subnetwork structure, the matrix product in LSTM and the single-layer convolution in ConvLSTM are replaced with the network, and speech timing information is extracted through multi-layer stacking CNN to perform feature mapping and dimensionality reduction.
The number of network parameters is reduced, the consistency of noise reduction performance is maintained, and a more efficient voice noise reduction effect is achieved.
Smart Images

Figure CN114974282B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech processing, and in particular relates to a single-channel speech denoising method and device based on deep learning. Background Art
[0002] At present, single-channel deep learning denoising methods have made significant progress, and their performance is significantly better than traditional signal processing methods. In particular, the CRN model based on U-Net is one of its mainstream technical solutions. The CRN model is a network that combines a codec structure with time series modeling. Codecs usually use CNN networks to extract local features, while time series modeling is usually completed by networks such as LSTM or GRU. CNN networks are generally known for their high efficiency and have better performance than fully connected networks with comparable parameters.
[0003] Taking the network actually used by the applicant as an example, most of the model parameters are concentrated in the LSTM layer, and the original LSTM can be regarded as multiple groups of fully connected networks. Simply reducing the dimension of the LSTM model usually causes a rapid decline in performance. In the field of image processing technology, some people have proposed the ConvLSTM network, but this technology has never been successfully used in the field of speech. The applicant also tried the algorithm, but after verification, it was found that the performance declined too much, so he gave up this idea. In response to the problem of large number of LSTM network parameters and poor performance of ConvLSTM, the applicant conducted further research and development, and thus produced the technical solution of this application. Summary of the invention
[0004] To this end, the present invention provides a single-channel speech denoising method and device based on deep learning, which solves the problem that the traditional network structure has a large number of LSTM network parameters for obtaining speech timing information, while the ConvLSTM performance is poor.
[0005] In order to achieve the above object, the present invention provides the following technical solutions: a single-channel speech denoising method based on deep learning, in which the time series modeling adopts a NetLSTM subnetwork, and the NetLSTM subnetwork structure replaces the matrix product in LSTM and the single-layer convolution in ConvLSTM with a network;
[0006] The formula of the NetLSTM sub-network structure is:
[0007]
[0008]
[0009]
[0010]
[0011] H t =ot ·tanh(c t )
[0012] Among them, i t is the output of the input gate at time t, f t is the output of the forget gate at time t, o t is the output of the output gate at time t, c t is the state gate output at time t, H t is the network output at time t, X t refers to the input features, H t-1 is the network output at time t-1;
[0013] is the internal subnetwork of the input gate;
[0014] It is the internal sub-network of the forget gate;
[0015] is the internal subnetwork of the output gate;
[0016] It is the internal subnetwork of the state gate.
[0017] As the preferred solution for the single-channel speech denoising method based on deep learning, the Net network structure in this speech denoising method uses a multi-layer stacked CNN, which extracts speech timing information and completes feature mapping.
[0018] As a preferred solution for a single-channel speech denoising method based on deep learning, each one-dimensional convolutional network layer of the Net network structure in this speech denoising method performs feature dimension reduction and channel expansion, and maps the feature dimension and the channel dimension to each other for feature transformation and expansion.
[0019] As the preferred solution for the single-channel speech denoising method based on deep learning, the parameters of the one-dimensional convolutional network include filter length, stride and number of filters; the number of filters determines the number of output channels of the one-dimensional convolutional network, and the stride controls the dimension of the output of the one-dimensional convolutional network.
[0020] As a preferred solution of the single-channel speech noise reduction method based on deep learning, a training phase is also included, and the training phase includes:
[0021] Data generation: Mix the original clean speech data and the preset type of noise at a given signal-to-noise ratio, use the mixed speech as training input data, and the original clean speech as reference input data, and calculate the loss of the deep learning denoising model;
[0022] Feature extraction: The training input data and the original clean speech are framed and windowed, and each frame is transformed into a frequency domain feature by using a frequency domain transform;
[0023] Network training: The extracted frequency domain features are input into the deep learning denoising model for training. An implicit mask is estimated using the signal approximation method. The implicit mask is multiplied by the first frequency domain features of the noisy speech to estimate the second frequency domain features of the clean signal. The second frequency domain features are inversely transformed in the frequency domain, and then the enhanced speech in the time domain is obtained by overlapping and adding.
[0024] As a preferred solution for the single-channel speech denoising method based on deep learning, the error between the enhanced speech and the target speech is calculated using a loss function;
[0025] The loss function uses SNR, MSE or MAE;
[0026] The frequency domain transform uses DCT\FFT, and the inverse frequency domain transform uses IDCT\IFFT.
[0027] As a preferred solution of the single-channel speech denoising method based on deep learning, it also includes an application stage, in which the trained deep learning denoising model is used for reasoning.
[0028] The present invention also provides a single-channel speech noise reduction device based on deep learning, including a time series modeling unit, wherein the time series modeling unit adopts a NetLSTM sub-network, and the NetLSTM sub-network structure replaces the matrix product in the LSTM and the single-layer convolution in the ConvLSTM with a network;
[0029] The formula of the NetLSTM sub-network structure is:
[0030]
[0031]
[0032]
[0033]
[0034] H t =o t ·tanh(c t )
[0035] Among them, i t is the output of the input gate at time t, f t is the output of the forget gate at time t, o t is the output of the output gate at time t, c t is the state gate output at time t, H t is the network output at time t, X t refers to the input features, H t-1 is the network output at time t-1;
[0036] is the internal subnetwork of the input gate;
[0037] It is the internal sub-network of the forget gate;
[0038] is the internal subnetwork of the output gate;
[0039] It is the internal subnetwork of the state gate.
[0040] As a preferred solution for a single-channel speech denoising device based on deep learning, in this speech denoising device, the Net network structure uses a multi-layer stacked CNN, which extracts speech timing information and completes feature mapping through the multi-layer stacked CNN;
[0041] In the speech noise reduction device, each one-dimensional convolutional network layer of the Net network structure performs feature dimension reduction and channel expansion, and maps the feature dimension and the channel dimension to each other for feature transformation and expansion;
[0042] The parameters of a one-dimensional convolutional network include filter length, stride, and number of filters; the number of filters determines the number of output channels of the one-dimensional convolutional network, and the stride controls the dimension of the output of the one-dimensional convolutional network.
[0043] As a preferred solution of the single-channel speech noise reduction device based on deep learning, a training unit is also included, and the training unit includes:
[0044] Data generation subunit: used to mix the original clean speech data and the preset type of noise at a given signal-to-noise ratio, use the mixed speech as training input data and the original clean speech as reference input data, and calculate the loss of the deep learning denoising model;
[0045] Feature extraction subunit: used to frame and window the training input data and the original clean speech. Each frame uses frequency domain transformation to convert time domain features into frequency domain features.
[0046] Network training subunit: used to input the extracted frequency domain features into the deep learning denoising model for training, use the signal approximation method to estimate an implicit mask, multiply the implicit mask by the first frequency domain feature of the noisy speech, estimate the second frequency domain feature of the clean signal, perform inverse frequency domain transform on the second frequency domain feature, and then obtain the enhanced speech in the time domain by overlapping and adding.
[0047] As a preferred solution of the single-channel speech noise reduction device based on deep learning, it also includes an error analysis subunit for calculating the error between the enhanced speech and the target speech using a loss function;
[0048] The loss function uses SNR, MSE or MAE;
[0049] The frequency domain transform uses DCT\FFT, and the inverse frequency domain transform uses IDCT\IFFT.
[0050] As a preferred solution of a single-channel speech noise reduction device based on deep learning, it also includes an application unit, which uses a trained deep learning noise reduction model for reasoning.
[0051] The present invention has the following advantages: the timing modeling in the speech denoising method adopts the NetLSTM subnetwork, and the NetLSTM subnetwork structure replaces the matrix product in LSTM and the single-layer convolution in ConvLSTM with a network; the Net network structure uses a multi-layer stacked CNN, and extracts the speech timing information and completes the feature mapping through the multi-layer stacked CNN. Each layer of the one-dimensional convolutional network in the Net network structure performs feature dimension reduction and channel expansion, and maps the feature dimension to the channel dimension to perform feature transformation and expansion. The present invention draws on the ideas of U-Net and ConvLSTM, and uses multi-layer stacked convolution to replace a layer of convolutional network on the basis of ConvLSTM, and remaps the features at the same time to realize the full-link network function. The number of network parameters is greatly reduced compared with LSTM, while the performance remains consistent. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the implementation of the present invention or the technical solution in the prior art, the following briefly introduces the drawings required for the implementation or the prior art description. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0053] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0054] Figure 1 A schematic diagram of a single-channel speech noise reduction method based on deep learning provided in Example 1 of the present invention;
[0055] Figure 2 A schematic diagram showing the comparison between LSTM and ConvLSTM in the prior art;
[0056] Figure 3 A schematic diagram of the architecture of a single-channel speech noise reduction system based on deep learning provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0057] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] Since most of the actual model parameters are concentrated in the LSTM layer, and the original LSTM can be seen as multiple groups of fully connected networks, simply reducing the LSTM model dimension usually causes a rapid decline in performance. In order to solve the problem of large number of LSTM network parameters and poor performance of ConvLSTM, this embodiment deeply considers the essence of LSTM, and replaces the internal structure of LSTM with a network instead of a matrix or a single-layer convolution. In the denoising task, a multi-layer stacked convolution network is designed to replace the internal fully connected layer to achieve feature mapping. After replacing LSTM, the denoising performance of this network does not decrease, and the number of parameters is greatly reduced.
[0059] Example 1
[0060] See also Figure 1 , Embodiment 1 of the present invention provides a single-channel speech denoising method based on deep learning, in which the time series modeling adopts a NetLSTM subnetwork, and the NetLSTM subnetwork structure replaces the matrix product in LSTM and the single-layer convolution in ConvLSTM with a network;
[0061] The formula of the NetLSTM sub-network structure is:
[0062]
[0063]
[0064]
[0065]
[0066] H t =o t ·tanh(c t )
[0067] Among them, i t is the output of the input gate at time t, f t is the output of the forget gate at time t, o t is the output of the output gate at time t, c t is the state gate output at time t, H t is the network output at time t, Xt refers to the input features, H t-1 is the network output at time t-1;
[0068] is the internal subnetwork of the input gate;
[0069] It is the internal sub-network of the forget gate;
[0070] is the internal subnetwork of the output gate;
[0071] It is the internal subnetwork of the state gate.
[0072] See also Figure 2 In this embodiment, the matrix product in LSTM and the single-layer convolution in ConvLSTM are replaced by Net network, which essentially expands the internal structure of LSTM, that is, the network is more generalized. Both LSTM and ConvLSTM can be considered as a special case of NetLSTM. In practice, combined with the design of speech denoising network, the Net network structure uses multi-layer stacked CNN to extract speech timing information and complete feature mapping. Each layer of one-dimensional convolution network in the Net network structure performs feature dimension reduction and channel expansion, and maps feature dimension with channel dimension to perform feature transformation expansion.
[0073] For example, assuming that the input feature is a one-dimensional vector of dimension 64, the output feature is set to be a one-dimensional vector of dimension 128, the convolution kernel is set to 4, and the step size is 2, then except for the last layer, the feature dimension of each layer is reduced by 2, and 5 layers of Conv1d are used, and the number of channels in the last layer is set to 128. The number of channels in the middle layer can be set through experiments, and the final output feature dimension is 128, which meets the set conditions. This is just a special case. In essence, it is to map the feature dimension to the channel dimension to play a role in feature transformation and expansion. Based on the above ideas, Conv1d can also be replaced by any other suitable network structure.
[0074] In this embodiment, the parameters of the one-dimensional convolutional network include filter length, stride and number of filters; the number of filters determines the number of output channels of the one-dimensional convolutional network, and the stride controls the dimension of the output of the one-dimensional convolutional network.
[0075] Specifically, the details of the Net network are further explained: for example, if the 64-dimensional input vector x is to obtain an output vector y with a dimension of 128, the traditional method is to use a matrix W with a dimension of 64×128, then y=W×x. The disadvantage of this method is that there are many parameters. In this embodiment, W is replaced by multi-layer convolution. One-dimensional convolution has three main parameters, one is the filter length, one is the stride, and the other is the number of filters. The number of filters determines the number of output channels, and the stride controls the output dimension. Assuming the stride is 1, the output dimension remains unchanged. Assuming the stride is 2, the output dimension is halved. Taking a 1-channel input vector of dimension 64 as an example, the filter length is 4, the stride is 2, and the number of output channels is 8, then the output becomes 8-channel data of dimension 32. Therefore, the dimension can be compressed to 1 by stacking one-dimensional convolutions, and the channel is expanded to 128. Then the channel and dimension coordinates are interchanged, and then it becomes a vector with dimension 128 and channel 1, achieving the same effect as y=W×x.
[0076] In this embodiment, it also includes: a training phase, which includes:
[0077] Data generation: Mix the original clean speech data and the preset type of noise at a given signal-to-noise ratio, use the mixed speech as training input data, and the original clean speech as reference input data, and calculate the loss of the deep learning denoising model.
[0078] Assuming the original clean data s, the noise n, then the mixed signal x = s + alpha * n, where alpha controls the noise level. Because there are a lot of original clean data s, a lot of noise n, and by adjusting different alphas, a lot of mixed signals x can be obtained.
[0079] Feature extraction: The training input data and the original clean speech are framed and windowed, and each frame is transformed into a frequency domain feature using a frequency domain transform to convert the time domain features into frequency domain features.
[0080] Each speech segment of the training data is framed and windowed, and each frame is transformed using DCT\FFT to convert the time domain features into frequency domain features Fx(t,f); the clean speech is also subjected to this operation to obtain Fs(t,f) for use in calculating the loss of the deep learning denoising model.
[0081] Specifically, speech framing and windowing itself belongs to the prior art. The frame represents a short period of speech data, and the frame consists of N sampling points. Assuming that the frequency information in a very short period of time t' remains unchanged, Fourier transform of the frame with a length of t' can obtain an appropriate expression of the frequency domain and time domain information of the speech data. The frame length ranges from 20ms to 40ms, and adjacent frames overlap by 50%. Commonly used parameter settings are: frame length 25ms, step length 10ms (15ms overlap). The relationship between frame length (T), speech data sampling frequency (F) and frame sampling points (N): T = N / F.
[0082] Specifically, windowing is to divide the signal into frames, substitute each frame into the window function, and set the value outside the window to 0. The purpose is to eliminate the signal discontinuity that may be caused at both ends of each frame. Commonly used window functions include square window, Hamming window and Hanning window. According to the frequency domain characteristics of the window function, the Hamming window is often used.
[0083] In this embodiment, network training is also included: the extracted frequency domain features are input into the deep learning noise reduction model for training, an implicit mask is estimated using a signal approximation method, the implicit mask is multiplied by the first frequency domain features of the noisy speech, the second frequency domain features of the clean signal are estimated, the second frequency domain features are inversely transformed in the frequency domain, and then the enhanced speech in the time domain is obtained by overlapping and adding.
[0084] Specifically, an implicit mask Mask(t,f) is estimated, multiplied by the feature Fx(t,f) of the noisy speech, the feature Fs2(t,f) of the clean signal is estimated, iDCT (inverse DCT transform) or IFFT transform is performed on Fs2(t,f), and then the enhanced speech in the time domain is obtained by overlap-addition. Where t represents frame and f represents frequency.
[0085] In this embodiment, the voice is enhanced The error is calculated using the loss function with the target speech s. The loss function uses Scale-invariant SNR (SI-SNR), which is defined as follows:
[0086]
[0087] Among them, s and Represent clean speech and enhanced speech respectively, <,> represents the dot product of vectors, is the Euclidean norm; SNR is the ratio of the sound intensity of the clean signal to the noise, and SI is the effect of reducing signal changes through regularization. The loss function can also use MSE or MAE.
[0088] In this embodiment, an application phase is also included, in which the trained deep learning noise reduction model is used for reasoning. In actual use, the trained deep learning noise reduction model is used for reasoning. The unknown noisy speech data is framed, windowed, and feature extracted. The mask is obtained by the trained deep learning noise reduction model, and then multiplied with the feature extracted. After feature inverse transformation and overlap addition, the predicted clean speech is obtained. The specific principle is the same as network training, except that the error does not need to be calculated, and the enhanced signal is obtained. That's it.
[0089] In summary, the timing modeling in the speech denoising method disclosed in the present invention adopts the NetLSTM subnetwork. The NetLSTM subnetwork structure replaces the matrix product in LSTM and the single-layer convolution in ConvLSTM with a network; the Net network structure uses a multi-layer stacked CNN to extract the speech timing information and complete the feature mapping through the multi-layer stacked CNN. Each layer of the one-dimensional convolutional network in the Net network structure performs feature dimension reduction and channel expansion, and maps the feature dimension to the channel dimension to perform feature transformation and expansion. The present invention draws on the ideas of U-Net and ConvLSTM, and uses multi-layer stacked convolution to replace a layer of convolutional network on the basis of ConvLSTM, and remaps the features at the same time to realize the full-link network function. The number of network parameters is greatly reduced compared to LSTM, while the performance remains consistent.
[0090] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
[0091] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] Example 2
[0093] See also Figure 3, Embodiment 2 of the present invention also provides a single-channel speech noise reduction device based on deep learning, including a timing modeling unit 1, wherein the timing modeling unit 1 adopts a NetLSTM sub-network, and the NetLSTM sub-network structure replaces the matrix product in the LSTM and the single-layer convolution in the ConvLSTM with a network;
[0094] The formula of the NetLSTM sub-network structure is:
[0095]
[0096]
[0097]
[0098]
[0099] H t =o t ·tanh(c t )
[0100] Among them, i t is the output of the input gate at time t, f t is the output of the forget gate at time t, o t is the output of the output gate at time t, c t is the state gate output at time t, H t is the network output at time t, X t refers to the input features, H t-1 is the network output at time t-1;
[0101] is the internal subnetwork of the input gate;
[0102] It is the internal sub-network of the forget gate;
[0103] is the internal subnetwork of the output gate;
[0104] It is the internal subnetwork of the state gate.
[0105] In this embodiment, in the speech noise reduction device, the Net network structure uses a multi-layer stacked CNN, and the multi-layer stacked CNN is used to extract the speech timing information and complete the feature mapping;
[0106] In the speech noise reduction device, each one-dimensional convolutional network layer of the Net network structure performs feature dimension reduction and channel expansion, and maps the feature dimension and the channel dimension to each other for feature transformation and expansion;
[0107] The parameters of a one-dimensional convolutional network include filter length, stride, and number of filters; the number of filters determines the number of output channels of the one-dimensional convolutional network, and the stride controls the dimension of the output of the one-dimensional convolutional network.
[0108] In this embodiment, a training unit 2 is further included, and the training unit 2 includes:
[0109] Data generation subunit 21: used to mix the original clean speech data and the preset type of noise at a given signal-to-noise ratio, use the mixed speech as training input data and the original clean speech as reference input data, and calculate the loss of the deep learning denoising model;
[0110] Feature extraction subunit 22: used to frame and window the training input data and the original clean speech, and use frequency domain transformation for each frame to convert the time domain features into frequency domain features;
[0111] The network training subunit 23 is used to input the extracted frequency domain features into the deep learning noise reduction model for training, use the signal approximation method to estimate an implicit mask, multiply the implicit mask by the first frequency domain features of the noisy speech, estimate the second frequency domain features of the clean signal, perform an inverse frequency domain transform on the second frequency domain features, and then obtain the enhanced speech in the time domain by overlapping and adding.
[0112] In this embodiment, an error analysis subunit 24 is also included, which is used to calculate the error between the enhanced speech and the target speech using a loss function;
[0113] The loss function uses SNR, MSE or MAE;
[0114] The frequency domain transform uses DCT\FFT, and the inverse frequency domain transform uses IDCT\IFFT.
[0115] In this embodiment, an application unit 3 is also included, and the application unit 3 uses a trained deep learning denoising model for reasoning.
[0116] For the convenience of description, the above device is described in terms of functions and is described separately in various units. Of course, when implementing the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0117] It should be noted that the information interaction, execution process, etc. between the units / sub-units of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.
[0118] Example 3
[0119] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code of a single-channel speech noise reduction method based on deep learning is stored, and the program code includes instructions for executing the single-channel speech noise reduction method based on deep learning of embodiment 1 or any possible implementation thereof.
[0120] The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0121] Example 4
[0122] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0123] The processor and the memory communicate with each other through a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the single-channel speech noise reduction method based on deep learning of Example 1 or any possible implementation method thereof.
[0124] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor implemented by reading software codes stored in a memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0125] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.
[0126] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0127] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. Single-channel speech noise reduction method based on deep learning, It is characterized in that In the speech denoising method, the NetLSTM subnetwork is used for time series modeling. The NetLSTM subnetwork structure replaces the matrix product in LSTM and the single-layer convolution in ConvLSTM with the Net network. The Net network structure uses multi-layer stacked CNN to extract speech time series information and complete feature mapping. Each layer of the one-dimensional convolutional network in the Net network structure performs feature dimension reduction and channel expansion, and maps the feature dimension to the channel dimension for feature transformation and expansion. The formula of the NetLSTM sub-network structure is: H t = o t ·tanh(c t ) Among them, i t is the output of the input gate at time t, f t is the output of the forget gate at time t, o t is the output of the output gate at time t, c t is the state gate output at time t, H t is the network output at time t, X t refers to the input features, H t-1 is the network output at time t-1; is the internal subnetwork of the input gate; It is the internal sub-network of the forget gate; is the internal subnetwork of the output gate; It is the internal subnetwork of the state gate.
2. The single-channel speech noise reduction method based on deep learning according to claim 1, It is characterized in that The parameters of a one-dimensional convolutional network include filter length, stride, and number of filters; the number of filters determines the number of output channels of the one-dimensional convolutional network, and the stride controls the dimension of the output of the one-dimensional convolutional network.
3. The single-channel speech denoising method based on deep learning according to claim 1, It is characterized in that A training phase is also included, and the training phase includes: Data generation: Mix the original clean speech data and the preset type of noise at a given signal-to-noise ratio, use the mixed speech as training input data, and the original clean speech as reference input data, and calculate the loss of the deep learning denoising model; Feature extraction: The training input data and the original clean speech are divided into frames and windows, and each frame is transformed into a frequency domain feature by using a frequency domain transform; Network training: The extracted frequency domain features are input into the deep learning denoising model for training. An implicit mask is estimated using the signal approximation method. The implicit mask is multiplied by the first frequency domain features of the noisy speech to estimate the second frequency domain features of the clean signal. The second frequency domain features are inversely transformed in the frequency domain, and then the enhanced speech in the time domain is obtained by overlapping and adding.
4. The single-channel speech denoising method based on deep learning according to claim 3, It is characterized in that The enhanced speech and the target speech are compared using a loss function to calculate the error; The loss function uses SNR, MSE or MAE; The frequency domain transform uses DCT\FFT, and the inverse frequency domain transform uses IDCT\IFFT.
5. The single-channel speech noise reduction method based on deep learning according to claim 4, It is characterized in that It also includes an application phase, in which the trained deep learning denoising model is used for reasoning.
6. Single-channel speech noise reduction device based on deep learning, It is characterized in that It includes a time series modeling unit, wherein the time series modeling unit adopts a NetLSTM sub-network, and the NetLSTM sub-network structure replaces the matrix product in LSTM and the single-layer convolution in ConvLSTM with a Net network; the Net network structure uses a multi-layer stacked CNN, and extracts speech time series information and completes feature mapping through the multi-layer stacked CNN; each layer of the one-dimensional convolutional network in the Net network structure performs feature dimension reduction and channel expansion, and maps the feature dimension and the channel dimension to each other for feature transformation and expansion; The formula of the NetLSTM sub-network structure is: H t =o t ·tanh(c t ) Among them, i t is the output of the input gate at time t, f t is the output of the forget gate at time t, o t is the output of the output gate at time t, c t is the state gate output at time t, H t is the network output at time t, X t refers to the input features, H t-1 is the network output at time t-1; is the internal subnetwork of the input gate; It is the internal sub-network of the forget gate; is the internal subnetwork of the output gate; It is the internal subnetwork of the state gate.
7. The single-channel speech noise reduction device based on deep learning according to claim 6, It is characterized in that The parameters of a one-dimensional convolutional network include filter length, stride, and number of filters; the number of filters determines the number of output channels of the one-dimensional convolutional network, and the stride controls the dimension of the output of the one-dimensional convolutional network.
8. The single-channel speech noise reduction device based on deep learning according to claim 7, It is characterized in that Also included is a training unit, the training unit comprising: Data generation subunit: used to mix the original clean speech data and the preset type of noise at a given signal-to-noise ratio, use the mixed speech as training input data and the original clean speech as reference input data, and calculate the loss of the deep learning denoising model; Feature extraction subunit: used to frame and window the training input data and the original clean speech. Each frame uses frequency domain transformation to convert time domain features into frequency domain features. Network training subunit: used to input the extracted frequency domain features into the deep learning denoising model for training, use the signal approximation method to estimate an implicit mask, multiply the implicit mask by the first frequency domain feature of the noisy speech, estimate the second frequency domain feature of the clean signal, perform frequency domain inverse transform on the second frequency domain feature, and then obtain the enhanced speech in the time domain by overlap-addition; It also includes an error analysis subunit for calculating the error between the enhanced speech and the target speech using a loss function; The loss function uses SNR, MSE or MAE; Frequency domain transform uses DCT\FFT, and inverse frequency domain transform uses IDCT\IFFT; It also includes an application unit, which uses the trained deep learning denoising model for reasoning.
Citation Information
Patent Citations
Voice recognition method, apparatus, and equipment and storage medium
CN108550364A
Single-channel processing method and device for speech enhancement and readable storage medium
CN113192528A