A speech enhancement method of deep learning assisted spectral subtraction
By using deep learning-assisted spectral subtraction, combined with local feature extraction and noise estimation networks, the problem of insufficient noise estimation in classical spectral subtraction is solved, achieving better speech enhancement and network interpretability while reducing resource consumption.
Patent Information
- Application Number
- CN202310628286.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-05-30
AI Technical Summary
Classical spectral subtraction has shortcomings in noise estimation and parameter tuning, leading to a decline in speech enhancement performance. Furthermore, parameter settings rely on extensive experimental adjustments, which limits the actual noise reduction effect. The network interpretability of deep learning methods needs to be improved.
Deep learning-assisted spectral subtraction is employed. By combining local feature extraction, noise estimation, and parameter estimation networks with short-time Fourier transform and spectral subtraction operations, an enhanced amplitude spectrum is obtained, improving the interpretability and noise reduction effect of the network.
It achieves better speech denoising effect, while improving the interpretability of deep learning networks. The training method is simple and reliable, and the resource consumption is cost-effective.
Smart Images

Figure CN116504262B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental noise suppression technology, and more specifically, to a speech enhancement method using deep learning-assisted spectral subtraction. Background Technology
[0002] The difficulty of classical spectral subtraction lies in noise estimation and parameter tuning. Specifically, classical spectral subtraction noise estimation assumes that the first few frames of the noisy speech signal are environmental noise.
[0003] However, the performance of spectral subtraction speech enhancement deteriorates significantly when actual conditions deviate from the assumptions. Furthermore, the settings of smoothing and oversubtraction factors largely rely on manual adjustments through extensive experiments, which greatly limits the actual noise reduction effect of spectral subtraction. In recent years, thanks to the powerful nonlinear processing capabilities of deep learning, data-driven deep speech enhancement methods have demonstrated excellent noise reduction performance. However, the network interpretability of data-driven deep speech enhancement methods needs further improvement. Therefore, it is necessary to leverage the powerful adaptive learning capabilities of deep learning and employ data-driven deep learning to assist classic pattern-driven speech enhancement methods.
[0004] Therefore, it is of great significance to use data-driven deep learning to assist the classic pattern-driven speech enhancement method in order to improve the interpretability of deep learning networks. Summary of the Invention
[0005] The purpose of this invention is to provide a speech enhancement method based on deep learning-assisted spectral subtraction, which aims to improve the interpretability of networks in deep learning and achieve excellent noise reduction results.
[0006] The embodiments of the present invention are achieved through the following technical solutions:
[0007] A speech enhancement method using deep learning-assisted spectral subtraction includes the following steps:
[0008] Initial feature extraction is performed to obtain noisy speech signals in the discrete time domain. Logarithmic power spectrum characteristics and phase characteristics N, T, and F represent the number of sample points of the noisy speech signal in the discrete time domain, and the number of frames and frequency points of the noisy speech signal after transformation to the time-frequency domain, respectively.
[0009] Based on the logarithmic power spectrum characteristics of noisy speech signals And estimation and denoising network to obtain the enhanced amplitude spectrum ;
[0010] According to the enhanced amplitude spectrum and initial phase obtaining an enhanced time-domain speech signal by inverse short-time Fourier transform ;
[0011] The estimation and noise reduction network comprises a local feature extraction network, a noise estimation network and a parameter estimation network;
[0012] The method for obtaining the enhanced amplitude spectrum comprises the following steps:
[0013] According to the log power spectrum feature , a local refined feature is obtained by using the local feature extraction network :
[0014] ;
[0015] wherein, denotes a mapping function for executing the local feature extraction subnetwork, is a parameter for executing the local feature extraction subnetwork;
[0016] According to the local refined feature , a noise power spectrum is estimated in parallel by the noise estimation network and the parameter estimation network and an over-reduction factor , and a smoothing factor :
[0017] ;
[0018] wherein denotes a mapping function for executing the noise estimation subnetwork, is a parameter for executing the noise estimation subnetwork, denotes a mapping function for executing the parameter estimation subnetwork, is a parameter for executing the parameter estimation subnetwork;
[0019] According to the noise power spectrum , the over-reduction factor and the smoothing factor , the enhanced amplitude spectrum is obtained by a subnetwork for executing a spectral subtraction operation :
[0020] ;
[0021] wherein, denotes an amplitude spectrum value of the enhanced amplitude spectrum in the i-th row and the j-th column; denotes a noise power spectrum value of the noise power spectrum in the i-th row and the j-th column; The log power spectrum feature of the noisy speech signal The log power spectrum value of the ith row and jth column.
[0022] Preferably, the method for performing initial feature extraction is:
[0023] The discrete-time noisy speech signal is transformed into a time-frequency domain by a short-time Fourier transform;
[0024] The log power spectrum feature of the noisy speech signal is calculated And the phase feature .
[0025] Preferably, the local feature extraction network includes two convolutional layers and one linear layer.
[0026] The noise estimation network includes two linear layers, and the parameter estimation network includes one linear layer.
[0027] Preferably, each linear layer includes a fully connected layer, a batch normalization layer, and a ReLU activation function.
[0028] Each convolutional layer includes a two-dimensional convolutional layer, a batch normalization layer, and a PReLU activation function.
[0029] Preferably, the training method of the estimation and noise reduction network includes the following steps:
[0030] A set of training noisy speech signals and training clean speech signals in a discrete time domain are collected , wherein , ;
[0031] The set of training noisy speech signals and training clean speech signals are transformed into a time-frequency domain to obtain a set of log power spectrum features , wherein, is the input when training the estimation and noise reduction network, is the label for training the estimation and noise reduction network;
[0032] According to the set of log power spectrum features and the constructed estimation and noise reduction network, the estimation and noise reduction network is trained using a least mean square error loss function, and the network model and its parameters are saved after error convergence to obtain a trained estimation and noise reduction network.
[0033] The technical scheme of the embodiment of the present application has at least the following advantages and beneficial effects:
[0034] The present application can obtain better speech noise reduction effect compared with the classic spectral subtraction method;
[0035] The application adopts a deep learning assisted pattern driven spectral subtraction based on data driving, and can further improve the speech enhancement quality and the network interpretability in deep learning compared with a mapping based deep speech enhancement method.
[0036] The application trains a network based on deep learning, and the training method is not complicated, and the obtained estimation and denoising network are reliable;
[0037] The application has reasonable design, high cost performance of resource consumption for obtained denoising effect and algorithm, and is convenient for promotion and application. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some embodiments of the application, and should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the premise of the drawings.
[0039] Figure 1 A flowchart of a deep learning assisted spectral subtraction speech enhancement method provided by the application;
[0040] Figure 2 A local feature extraction network structure diagram of the application;
[0041] Figure 3 A noise estimation network structure diagram of the application;
[0042] Figure 4 A parameter estimation network structure diagram of the application. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical scheme and advantages of the embodiments of the application more clear, the following will combine the drawings in the embodiments of the application to clearly and completely describe the technical scheme in the embodiments of the application, and obviously, the described embodiments are a part of the embodiments of the application, not all the embodiments. The components of the embodiments of the application described and shown in the drawings here can be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the application provided in the drawings is not intended to limit the scope of the claimed application, but only represents selected embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0045] It should be noted that like reference numerals and characters refer to like elements throughout the following figures and the detailed description, and thus, definitions and explanations thereof need not be repeated for each figure. It should be noted that like reference numerals and characters refer to like elements throughout the following figures and the detailed description, and thus, definitions and explanations thereof need not be repeated for each figure.
[0046] Embodiment 1
[0047] With reference to Figure 1 The embodiment provides a speech enhancement method based on deep learning assisted spectral subtraction, and comprises the following steps:
[0048] Step S1: initial feature extraction is performed to obtain a log power spectrum feature and a phase feature of a noisy speech signal in a discrete time domain.
[0049] N, T and F respectively represent a sample point number of the noisy speech signal in the discrete time domain, and a frame number and a frequency point number of the noisy speech signal in the discrete time domain after the noisy speech signal is transformed into a time-frequency domain.
[0050] Step S2: an enhanced amplitude spectrum is obtained according to the log power spectrum feature and the phase feature of the noisy speech signal and an estimation and denoising network.
[0051] Step S3: an enhanced time-domain speech signal is obtained through short-time Fourier inverse transformation according to the enhanced amplitude spectrum and an initial phase.
[0052] The embodiment aims to provide a speech enhancement method based on deep learning assisted spectral subtraction, and compared with a classic speech enhancement method such as spectral subtraction and MMSE estimation, the embodiment method can obtain better speech denoising effect; compared with a deep speech enhancement method based on amplitude spectrum mapping, the embodiment method can obtain better denoising effect and stronger network interpretability.
[0053] In order to reflect the effect of the embodiment, the original noisy speech (indicated by original in Table 1), the speech obtained by the method of the embodiment (indicated by the embodiment in Table 1), the speech obtained by spectral subtraction (indicated by spectral subtraction in Table 1) and the speech obtained by the speech enhancement method based on amplitude spectrum mapping (indicated by amplitude spectrum mapping in Table 1) are compared. Table 1 below is the signal-to-noise ratio result of the above-mentioned different methods respectively applied to three kinds of noise (Babble, Factory1 and Destoryerengine) under PESQ and STOI index tests, and the intensities of the three kinds of noise are-5dB, 0dB and 5dB respectively.
[0054] Table 1
[0055]
[0056] As can be seen from Table 1, the method adopted in the embodiment can obtain the signal with the highest signal-to-noise ratio under the test of three noises and two indexes, which can intuitively reflect that the method adopted in the embodiment can produce better speech enhancement effect, that is, the method adopted in the embodiment can obtain better noise reduction effect.
[0057] Embodiment 2
[0058] In this embodiment, the method for performing initial feature extraction in step S1 is further described based on the technical solution of embodiment 1.
[0059] In this embodiment, the method for performing initial feature extraction is as follows:
[0060] Step S11: transforming the discrete time-domain noisy speech signal into time-frequency domain by short-time Fourier transform; Step S12: calculating the log power spectrum feature and phase feature of the noisy speech signal
[0061] Step S12: calculating the log power spectrum feature and phase feature of the noisy speech signal .
[0062] Short-time Fourier transform is a time-frequency analysis method, and the signal can be analyzed in time domain and frequency domain at the same time by using short-time Fourier transform.
[0063] When the short-time Fourier transform is performed, the signal is first segmented in time domain, and then Fourier transform is performed on each time segment to obtain frequency components. Finally, the frequency components of all time segments are combined together to obtain the joint representation of time domain and frequency domain.
[0064] It is particularly pointed out that when performing, the signal is divided into multiple equal-length time segments, and each time segment is regarded as a window. In short-time Fourier transform, the size of the window has an influence on the accuracy and time resolution of the short-time Fourier transform result. The smaller the window size, the higher the accuracy and the lower the time resolution.
[0065] Embodiment 3
[0066] In this embodiment, the related content of obtaining the enhanced amplitude spectrum in step S2 is further described based on the technical solution of embodiment 1.
[0067] In this embodiment, the estimation and noise reduction network includes a local feature extraction network, a noise estimation network and a parameter estimation network.
[0068] As a preferred solution, the method for obtaining the enhanced amplitude spectrum includes the following steps:
[0069] Step S21: obtaining local refined features according to the log power spectrum feature , by using the local feature extraction network :
[0070] ;
[0071] wherein, denotes a mapping function for executing the local feature extraction subnetwork, is a parameter for executing the local feature extraction subnetwork;
[0072] Step S22: estimating noise power spectrum , by the noise estimation network and the parameter estimation network in parallel, according to the local refined features and an over-reduction factor , and a smoothing factor :
[0073] ;
[0074] wherein denotes a mapping function for executing the noise estimation subnetwork, is a parameter for executing the noise estimation subnetwork, denotes a mapping function for executing the parameter estimation subnetwork, is a parameter for executing the parameter estimation subnetwork;
[0075] Step S23: obtaining the enhanced amplitude spectrum , by executing a subnetwork for performing a spectral subtraction operation, according to the noise power spectrum , the over-reduction factor , and the smoothing factor :
[0076] ;
[0077] wherein, denotes an amplitude spectrum value of the enhanced amplitude spectrum at the i-th row and the j-th column; denotes a noise power spectrum value of the noise power spectrum at the i-th row and the j-th column; denotes a log power spectrum value of the log power spectrum feature of the noisy speech signal at the i-th row and the j-th column.
[0078] Further, refer to Figure 2 , Figure 3 and Figure 4 :
[0079] The local feature extraction network can include two layers of convolutional layers (Convs, Convolutional Layers) and one layer of linear layers (LL, Linear Layer), in which Figure 2 Convs_1, Convs_2 and LL are used to represent, respectively;
[0080] The noise estimation network includes two layers of linear layers, in which Figure 3 LL_1 and LL_2 are used to represent, respectively;
[0081] The parameter estimation network includes one layer of linear layers, in which LL is used to represent. Figure 4
[0082] Further, each layer of the linear layer preferably includes a fully connected layer (FC, Fully Connected), a batch normalization layer (BN, Batch Normalization Layer) and a ReLU activation function;
[0083] On the other hand, each layer of the convolutional layer preferably includes a two-dimensional convolutional layer (2D-Conv, Two-Dimensional Convolutional Layer), a batch normalization layer (BN, Batch Normalization Layer) and a PReLU activation function.
[0084] The above fully connected layer, ReLU activation function, two-dimensional convolutional layer, batch normalization layer and PReLU activation function are represented by FC, ReLU, 2D-Conv, BN and PReLU, respectively. Figures 2-4
[0085] In the specific setting, the following methods are executed:
[0086] The kernel sizes of the two two-dimensional convolutional layers in the local feature extraction network are and ;
[0087] The input channel numbers of the two two-dimensional convolutional layers in the local feature extraction are 1 and c, and the output channel numbers are c and d, respectively;
[0088] The steps of the two two-dimensional convolutional layers in the local feature extraction are both (e, f); the input nodes and output nodes of the fully connected layer in the local feature extraction are both F; wherein the a, b, c, d, e and f are set according to engineering experience.
[0089] The input nodes and output nodes of the fully connected layers in the two layers of linear layers of the noise estimation network are both F;
[0090] The input nodes of the full connection layer of the parameter estimation network are F and g respectively, and the output nodes are g and 2 respectively, and g is set according to engineering experience.
[0091] Embodiment 4
[0092] This embodiment is based on the technical solution of embodiment 1, and further describes the related content of obtaining the enhanced amplitude spectrum in step S2.
[0093] As a preferred scheme of this embodiment, the training method of the estimation and noise reduction network comprises the following steps:
[0094] Collecting a set of training noisy speech signals and training clean speech signals in discrete time domain , wherein , ;
[0095] Transforming the set of training noisy speech signals and training clean speech signals into a time-frequency domain to obtain a set of log power spectrum features , wherein is the input when training the estimation and noise reduction network, is the label for training the estimation and noise reduction network;
[0096] According to the set of log power spectrum features and the constructed estimation and noise reduction network, the estimation and noise reduction network is trained using a least mean square error loss function, and the network model and its parameters are saved after error convergence to obtain the trained estimation and noise reduction network.
[0097] In this embodiment, the set of noisy speech and clean speech signals in discrete time domain is obtained according to actual application equipment recording or artificial synthesis.
[0098] It is particularly pointed out that the estimation and noise reduction network of this embodiment still comprises a local feature extraction network, a noise estimation network and a parameter estimation network, and the three sub-networks of the local feature extraction network, the noise estimation network and the parameter estimation network can be trained in a joint training manner.
[0099] The above is only a preferred embodiment of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A speech enhancement method using deep learning-assisted spectral subtraction, characterized in that, Includes the following steps: Initial feature extraction is performed to obtain noisy speech signals in the discrete time domain. Logarithmic power spectrum characteristics and phase characteristics N, T, and F represent the number of sample points of the noisy speech signal in the discrete time domain, and the number of frames and frequency points of the noisy speech signal after transformation to the time-frequency domain, respectively. Based on the logarithmic power spectrum characteristics of noisy speech signals And estimation and denoising network to obtain the enhanced amplitude spectrum ; Based on the enhanced amplitude spectrum and initial phase The enhanced time-domain speech signal is obtained through short-time inverse Fourier transform. ; The estimation and denoising network includes a local feature extraction network, a noise estimation network, and a parameter estimation network; The method for obtaining the enhanced amplitude spectrum includes the following steps: Based on the logarithmic power spectrum characteristics The local feature extraction network is used to obtain local refinement features. : ; in, This represents the mapping function that performs the local feature extraction subnetwork. The parameters for the subnetwork that performs local feature extraction; Based on the local refinement features The noise power spectrum is estimated in parallel using the noise estimation network and the parameter estimation network. With over-reduction factor and smoothing factor : ; in This represents the mapping function for the noise estimation subnetwork. To perform noise estimation of the subnetwork's parameters, This represents the mapping function of the subnetwork for parameter estimation. To perform parameter estimation of the subnetwork parameters; According to the noise power spectrum The over-subtraction factor With the smoothing factor The enhanced amplitude spectrum is obtained through a subnetwork that performs spectral subtraction. : ; in, Indicates the enhanced amplitude spectrum The amplitude spectrum value in the i-th row and j-th column; Represents the noise power spectrum The noise power spectrum value in the i-th row and j-th column; The logarithmic power spectrum characteristics representing noisy speech signals The logarithmic power spectrum value in the i-th row and j-th column.
2. The speech enhancement method using deep learning-assisted spectral subtraction according to claim 1, characterized in that, The method for performing initial feature extraction is as follows: The discrete-time domain noisy speech signal is transformed using short-time Fourier transform. Transform to the time-frequency domain; Calculate the logarithmic power spectral characteristics of a noisy speech signal and phase characteristics .
3. The speech enhancement method using deep learning-assisted spectral subtraction according to claim 1, characterized in that: The local feature extraction network comprises two convolutional layers and one linear layer; The noise estimation network comprises two linear layers; the parameter estimation network comprises one linear layer.
4. The speech enhancement method using deep learning-assisted spectral subtraction according to claim 3, characterized in that: Each linear layer includes a fully connected layer, a batch normalized layer, and a ReLU activation function; Each of the convolutional layers comprises a two-dimensional convolutional layer, a batch normalization layer, and a PReLU activation function.
5. The speech enhancement method using deep learning-assisted spectral subtraction according to claim 1, characterized in that, The training method for the estimation and denoising network includes the following steps: Collect a set of noisy and clean speech signals for training in the discrete time domain. ,in , ; The set of noisy training speech signals and clean training speech signals is transformed to the time-frequency domain to obtain their logarithmic power spectral feature set. ,in, The input is used to train the estimation and denoising network. To train the estimated and denoising network for labels; Based on the logarithmic power spectrum feature set The estimation and denoising network is constructed and trained using the minimum mean square error loss function. After the error converges, the network model and its parameters are saved to obtain the trained estimation and denoising network.
Citation Information
Patent Citations
Speech enhancement method based on Gaussian mixture model (GMM) noise estimation
CN104464728A
Method for spectral subtraction in speech enhancement
US20050071156A1