Speech noise reduction method based on deep learning
By extracting the characteristics of noise data and combining the LSTM and VAE methods, the problem of insufficient analysis of noisy voice data in the prior art is solved, and a more effective speech noise reduction effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/104005
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-07-05
- Publication Date
- 2025-06-19
AI Technical Summary
The existing deep learning-based speech noise reduction method fails to fully analyze and mine the characteristic information in noisy voice data, resulting in the inability to effectively reduce noise.
By extracting the characteristic information of the noise data and adding these characteristic information to the noise reduction process as a reference, data recoding and timing analysis are performed in combination with the LSTM and Variable Autoencoder (VAE) method, and finally the clear voice data after noise reduction is obtained.
It realizes more effective noise reduction for noisy voice data, can more fully analyze and characterize data characteristics, and improves the effect of voice noise reduction.
Smart Images

Figure CN2024104005_19062025_PF_FP_ABST
Abstract
Description
A speech noise reduction method based on deep learning
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311709812.9, filed on December 12, 2023, entitled “A Speech Noise Reduction Method Based on Deep Learning,” which is incorporated herein by reference in its entirety. Technical Field
[0003] The present application relates to the technical field of speech noise reduction, and specifically to a speech noise reduction method based on deep learning. Background Art
[0004] Existing speech noise reduction solutions, particularly those based on deep learning, typically focus on restoring clear speech data—that is, noise-free original data—from noisy speech data. These studies have overlooked the impact of noise in deep learning methods, and noisy speech data has not been fully analyzed and mined.
[0005] At the same time, when analyzing time-series data such as speech data, most methods only use recurrent neural networks such as RNN and LSTM. The network models are relatively simple and cannot fully analyze and characterize the data features.
[0006] Summary of the Invention
[0007] In order to overcome the shortcomings of the existing technology, this application provides a speech noise reduction method based on deep learning to solve the problems existing in the existing technology, such as the inability to fully analyze and mine noisy speech data and the inability to fully analyze and characterize data features.
[0008] The technical solution adopted by this application to solve the above problems is:
[0009] A speech denoising method based on deep learning extracts feature information from noise data during training and adds this feature information as a reference to the denoising process, ultimately obtaining clear original speech data after denoising.
[0010] As an optimal technical solution, a variational autoencoder calculation method is added to the LSTM algorithm to re-encode the data, mine the implicit feature information in the training data, and then put the implicit feature information into the LSTM network for training, and finally decode it to obtain the noise-reduced speech data.
[0011] As a preferred technical solution, the following steps are included:
[0012] S1, define the speech data set D, which includes N segments of speech data x with equal length and containing noise i , and N segments with x i Corresponding clear original voice data y i , x i with y i Equal length of time;
[0013] S2, for input data x i,n , respectively put in Perform attention mechanism calculation and re-represent the input data to obtain and
[0014] Among them, x i,n Represents x i Several data segments intercepted by sliding window, x i Represents a piece of speech data containing noise, x i ∈D1,1≤i≤N, sliding window size L window and sliding distance L length Set to L, that is, L window =L length =L, L is a hyperparameter; n represents the number of data segments intercepted, Represents the attention mechanism algorithm function used to extract clear speech data; represents the attention mechanism algorithm function used to extract noisy data, express The output feature information vector, express Output feature information vector;
[0015] S3, Input to the encoder structure Calculate and get Will Input to the encoder structure Calculate and get At this time Perform an attention mechanism calculation and add the output feature result to the feature middle;
[0016] in, represents the encoder calculation function used to extract feature information of clear speech data, represents the encoder calculation function used to extract the feature information of noise, express The output feature information vector, express Output feature information vector;
[0017] S4, will Input to LSTM network Train to get Will Input to LSTM network Train to get
[0018] in, Represents the LSTM calculation function used to extract feature information of clear speech data, express The output feature information vector, Represents the LSTM calculation function used to extract feature information of noise, express Output feature information vector.
[0019] S5, will Input to decoder structure Calculate and get Will Input to decoder structure Calculate and get At this time Perform an attention mechanism calculation and add the output feature result to the feature middle;
[0020] in, represents the decoder calculation function for extracting feature information of clear speech data, express The output feature information vector, represents a decoder calculation function for calculating feature information for extracting noise, express Output feature information vector;
[0021] S6, group the voice segments Splice into y in order i,1 , group the speech segments Splice into y in order i,2 , two output results y i,1 、y i,2 With input data x i The data dimensions are consistent;
[0022] Among them, y i,1 For the clear original voice data after noise reduction, y i,2 For input data xi The noise data contained in .
[0023] As a preferred technical solution, in step S2, the attention mechanism is calculated on the input data, and the calculation formula is as follows: a n1 (n2)=exp(score(x i,n1 ,x i,n2 ))=exp(x i,n1 ·x i,n2 ),n1≠n2
[0024] Among them, x i,n1 is the current input data x i,n A piece of data in x i,n2 For input data x i,n Divide by x i,n1 The data segment outside, a n1 (n2) is x i,n1 with x i,n2 The attention calculation value, a′ n1 (n2) is the weighted x i,n1 with x i,n2 The attention calculation value of
[0025] Get a′ n1 (n2) can be used to calculate the new x′ i,n1 , calculate x by the following formula i,n1 Re-characterization of x′ i,n1 =x i,n1 +W Att ∑a′ n1 (n2)x i,n2
[0026] Where x′ i,n1 Represents x i,n1 The feature vector calculated by the attention mechanism, W Att is a trainable parameter;
[0027] From the above calculations, we can get:
[0028] in, represents the attention mechanism function used to extract feature information of clear speech data, Represents the attention mechanism function used to extract feature information of noise.
[0029] As a preferred technical solution, in step S3, the obtained and Input to the encoder structure for calculation, the calculation formula is as follows: x=ReLU(Liner(x)) μ=Liner(x) σ=exp(Liner(x)) z=Sample(μ,σ)
[0030] Where x is the input vector, ReLU(·) is the activation function, Liner(·) is the linear function, μ is the mean, σ is the variance, Sample(·) is the sampling function, and z is the implicit feature vector with randomness;
[0031] From the above calculations, we can get:
[0032] As a preferred technical solution, in step S3, Perform an attention mechanism calculation and add the output feature result to the feature middle;
[0033] The calculation process is as follows:
[0034] Among them, β is a hyperparameter, The attention mechanism function for extracting the reference feature information of the noise for the first time.
[0035] As a preferred technical solution, in step S4, and Input into the LSTM network for training. The LSTM unit structure includes forget gate, input gate, output gate, and cell state.
[0036] As a preferred technical solution, in step S4,
[0037] Forget Gate f t Determines whether the information in the cell state is lost: the gate input data includes the input data x at time t t And the hidden vector h at the previous time t-1 t-1 , feedback to cell state c t-1 ; Forget gate f t The calculation formula is as follows: t =σ(W f [x t ,h t-1 ]+b f )
[0038] Input gate i t Determine the data that needs to be updated for the cell state; input gate i t The calculation formula is as follows: t =σ(W i [x t ,ht-1 ]+b i )
[0039] Update cell state c t , the cell state c at the previous moment t-1 t-1 With f t Multiply, discard the information that needs to be forgotten, and then add the new cell state; cell state c t The calculation formula is as follows: t =f t c t-1 +i t ×(tanh(W c [x t ,h t-1 ]+b c ))
[0040] Output gate o t Determine the output of some cell states, output gate o t The calculation formula is as follows: t =σ(W o [x t ,h t-1 ]+b o )
[0041] Finally, calculate the hidden vector h at the current time t t , h t The calculation formula is as follows: t =o t ×tanh(c t )
[0042] From the above calculation:
[0043] Among them, t represents the time in the time series data, t-1 is the time before t, the value range of t is [1,n], n represents the number of data segments in the input data, f t Represents the output information of the forget gate at time t, i t Represents the input gate output information at time t, o t Indicates the output information of the output gate at time t in LSTM, c t Represents the cell state output information at time t;
[0044] x t Represents the data at time t in the time series data, that is, the tth segment of the input data, with a dimension of n x ×1,n x is a hyperparameter;
[0045] h t-1Represents the hidden vector data at the previous moment t-1, c t-1 Represents the cell state vector data at the previous moment t-1; for t=1, h0 is defined as the dimension n h ×1 random initial vector, c0 is defined as the dimension is n h ×1 random initial vector, n h is a hyperparameter;
[0046] W f Represents the trainable parameter matrix for forget gate calculation, dimension is n h ×(n x +n h ), b f Represents the trainable parameter vector for forget gate calculation, with dimension n h ×1;
[0047] W i Represents the trainable parameter matrix of the input gate calculation, with dimension n h ×(n x +n h ), b i Represents the trainable parameter vector of the input gate calculation, with dimension n h ×1;
[0048] W c Represents the trainable parameter matrix for cell state calculation, dimension n h ×(n x +n h ), b c Represents a trainable parameter vector for cell state calculation, with dimension n h ×1;
[0049] W o Represents the trainable parameter matrix for output gate calculation, dimension n h ×(n x +n h ), b o Represents the trainable parameter vector of the output gate calculation, with dimension n h ×1.
[0050] As a preferred technical solution, in step S5,
[0051] Will and Input to the decoder structure for calculation, the calculation formula is as follows: x = ReLU (Liner (x)) x = Sigmoid (Liner (x))
[0052] Among them, ReLU(·) is the ReLU activation function, Sigmoid(·) is the Sigmoid activation function;
[0053] From the above calculations, we can get:
[0054] right Perform an attention mechanism calculation and add the output feature result to the feature middle;
[0055] The calculation process is as follows:
[0056] in, is the attention mechanism function for extracting the reference feature information of the noise for the second time, and γ is a hyperparameter.
[0057] As a preferred technical solution, in step S4, when training the LSTM network, the loss function formula is:
[0058] Among them, F loss represents the loss function, N represents the number of data instances in the dataset, and x i Represents a piece of speech data containing noise, y i Represents x i The corresponding clear original speech data, ‖·‖2 represents the Euclidean distance calculation, Represents the encoder calculation function used to extract feature information of clear speech data The resulting KL divergence, Represents the encoder calculation function used to extract feature information of noise The resulting KL divergence, F KL Represents the KL divergence calculation formula.
[0059] Compared with the prior art, this application has the following beneficial effects:
[0060] This application not only focuses on how to reduce the noise of speech data containing noise to clear original speech data, but also focuses on separating the noise data from the noise-containing speech data. At the same time, during the training process of this method, the feature information of the noise data is extracted and added as a reference to the noise reduction process, and finally the clear original speech data after noise reduction is obtained.
[0061] This application adds a VAE (Variational Autoencoder) calculation method to the LSTM algorithm, re-encodes the data, mines the implicit feature information in the training data, then puts this implicit feature information into the LSTM network for time series analysis, and finally decodes it to obtain the noise-reduced speech data;
[0062] This application adds an attention mechanism to perform feature screening on the data, extract important features, suppress or eliminate unimportant features or even interfering features, and improve the prediction effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is an overall flow chart of this application; and
[0064] Figure 2 is a diagram of the LSTM unit structure. DETAILED DESCRIPTION
[0065] The present application will be further described in detail below in conjunction with the embodiments and drawings, but the implementation methods of the present application are not limited thereto.
[0066] Example 1
[0067] As shown in Figures 1 and 2, the main purpose of this application is to provide a speech noise reduction method based on deep learning to solve the problems existing in the existing technology in speech noise reduction tasks.
[0068] In order to achieve the above-mentioned purpose, according to one aspect of the specific implementation of the present application, a speech noise reduction method based on deep learning is provided, which includes the following steps.
[0069] The training dataset is constructed by using noisy speech data and its corresponding clear original speech data.
[0070] For input data x i,n , respectively put in Perform attention mechanism calculation and re-represent the input data to obtain and
[0071] Will and Input to the encoder structure respectively and Calculate and get and At this time Perform an attention mechanism calculation and add the output feature result to the feature middle.
[0072] Will and Input into LSTM network respectively and Calculate and get and
[0073] Will and Input to the decoder structure respectively and Calculate and get and At this time Perform an attention mechanism calculation and add the output feature result to the feature middle.
[0074] Group voice clips Splice into y in order i,1 , group the speech segments Splice into y in order i,2 , two output results y i,1 、y i,2 With input data x i The data dimensions are consistent. Among them, y i,1 For the clear original voice data after noise reduction, y i,2 For input data x i The noise data contained in .
[0075] During the training process, the loss function F loss Calculate the loss value, train and update the model parameters; in the output results, y i,1 As the clear original voice data after noise reduction.
[0076] Example 2
[0077] As shown in FIG. 1 and FIG. 2 , as a further optimization of Example 1, based on Example 1, this embodiment also includes the following technical features.
[0078] In this application, the overall process is shown in Figure 1. The following will introduce them in detail.
[0079] Define a speech dataset D, which includes N segments of equal length speech data x containing noise. i , and N segments with x i Corresponding clear original voice data y i , x i with y i The time lengths are equal.
[0080] For a piece of noisy speech data x i ,x i ∈D,1≤i≤N, intercept several data segments x by sliding window i,n , n is the number of data segments intercepted. Sliding window size L window and sliding distance L length Set to L, that is, L window =L length=L, where L is a hyperparameter. The window size and sliding distance can be customized according to the specific task. From this, we can see that the data segment dimension is 1×L, which means that the adjacent data segments extracted are adjacent, continuous and non-overlapping. i The corresponding normal speech data without noise is defined as y i .
[0081] From the above, we can get a piece of voice data x i , a set of speech data segments of length L can be extracted, the number of which is n, that is, x i,n Similarly, the speech data set D can extract several groups of speech data segments of length L to form the training data set D Train ={(x i,n ,y i ),1≤i≤N}.
[0082] From the above data, the input data of the deep neural network model can be defined. Each input data x of the deep neural network model i,n Contains n segments of voice data.
[0083] For input data x i,n ∈D1 performs attention mechanism calculation, and The calculation formula is the same, as follows: a n1 (n2)=exp(score(x i,n1 ,x i,n2 ))=exp(x i,n1 ·x i,n2 ),n1≠n2
[0084] Among them, x i,n1 is the current input data x i,n A piece of data in x i,n2 For input data x i,n Divide by x i,n1 The data segment outside, a n1 (n2) is x i,n1 with x i,n2 The attention calculation value, a′ n1 (n2) is the weighted x i,n1 with x i,n2 The attention calculation value of . Get a′ n1 (n2) can be used to calculate the new x′ i,n1 , calculate x by the following formula i,n1 Re-characterization of x′ i,n1 =x i,n1 +W Att ∑a′n1 (n2)x i,n2
[0085] Among them, W Att is a trainable parameter.
[0086] From the above calculations, we can get:
[0087] get and Input to the encoder structure for calculation, and The calculation formula is the same as follows: x=ReLU(Liner(x)) μ=Liner(x) σ=exp(Liner(x)) z=Sample(μ,σ)
[0088] Among them, ReLU(·) is the activation function, Liner(·) is the linear function, mean μ and variance σ are intermediate parameters, Sample(·) is the sampling function, and z is the implicit feature vector with randomness.
[0089] From the above calculations, we can get:
[0090] At this time Perform an attention mechanism calculation and add the output feature result to the feature The calculation process is as follows:
[0091] Among them, β is a hyperparameter.
[0092] Will and Input to the LSTM network for calculation, and The calculation formula is the same. The LSTM unit structure is shown in Figure 2. For the input data x t The calculation of is as follows:
[0093] where f t 、i t 、o t 、c t They represent the forget gate, input gate, output gate and cell state in LSTM respectively.
[0094] By adding these gate structures, LSTM can control the flow of information, that is, selectively store, lose, and update information at a certain moment, thereby solving the long-distance information transmission problem of RNN. However, it also leads to more parameters and longer training time.
[0095] Forget Gate f t Determines whether the information in the cell state is lost. This gate will input the input layer x t And the data t of the previous hidden layer t-1 , the output is a value in the interval [0,1] and is fed back to the cell state c t-1 1 means "completely retain" and 0 means "completely discard". t =σ(W f x t +U f h t-1 +b f )
[0096] Input gate i t Determine the data that needs to be updated in the cell state. Update the cell state and compare the previous state with f t Multiply and discard the information that needs to be discarded, and then add the newly updated cell state: i t =σ(W i x t +U i h t-1 +b i )
[0097] Output gate o t Determine the output of some cell states, the calculation formula is as follows: c t =f i c t-1 +i t ×(tanh(W c x t +U c h t-1 +b c )) o t =σ(W o x t +U o h t-1 +b o ) h t =o t ×tanh(c t )
[0098] From the above calculations, we can get:
[0099] Will and Input to the decoder structure for calculation, and The calculation formula is the same, as follows: x = ReLU (Liner (x)) x = Sigmoid (Liner (x))
[0100] From the above calculations, we can get:
[0101] At this time Perform an attention mechanism calculation and add the output feature result to the feature The calculation process is as follows:
[0102] Among them, γ is a hyperparameter.
[0103] At this time, group the voice segments Splice into y in order i,1 , group the speech segments Splice into y in order i,2 , two output results y i,1 、y i,2 With input data x i The data dimensions are consistent.
[0104] When training the entire deep neural network model, the loss function F loss The formula is
[0105] During the training process, the loss function F loss Calculate the loss value, train and update the model parameters; in the output results, y i,1 As the clear original voice data after noise reduction.
[0106] As described above, the present application can be implemented well.
[0107] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.
[0108] The above is only a preferred embodiment of the present application and does not constitute any form of limitation to the present application. Based on the technical essence of the present application and within the principles of the present application, any simple modification, equivalent replacement and improvement of the above embodiment shall still fall within the scope of protection of the technical solution of the present application.
Claims
1. A speech noise reduction method based on deep learning, comprising: Extract feature information of noisy data during training; Add feature information as reference to the denoising process; Get clear original voice data after noise reduction.
2. The method according to claim 1, wherein: Based on the LSTM algorithm, the variational autoencoder calculation method is added to re-encode the data, mine the implicit feature information in the training data, and then put the implicit feature information into the LSTM network for training, and finally decode it to obtain the noise-reduced speech data.
3. The method according to claim 2, further comprising the following steps: S1, define a speech data set D, which includes N segments of equal length speech data x containing noise i , and N segments with x i Corresponding clear original voice data y i , x i With y i The duration is equal; S2, for input data x i,n , respectively put into Perform attention mechanism calculation and re-characterize the input data to obtain and Among them, x i,n Represents x i Several data segments intercepted by sliding window, x i Represents a piece of speech data containing noise, x i ∈D1,1≤i≤N, sliding window size L window and sliding distance L length Set to L, that is, L window =L length =L, L is a hyperparameter; n represents the number of data segments extracted, Represents the attention mechanism algorithm function used to extract clear speech data; represents the attention mechanism algorithm function used to extract noisy data, express The output feature information vector, express Output feature information vector; S3, Input to the encoder structure Calculate and get Will Input to the encoder structure Calculate and get At this time Perform an attention mechanism calculation and add the output feature result to the feature middle, in, represents the encoder calculation function for extracting feature information of clear speech data, represents the encoder calculation function used to extract the feature information of noise, express The output feature information vector, express Output feature information vector; S4, Input to LSTM network Train to get Will Input to LSTM network Train to get in, represents the LSTM calculation function used to extract feature information of clear speech data, express The output feature information vector, represents the LSTM calculation function used to extract the feature information of noise, express Output feature information vector; S5, Input to the decoder structure Calculate and get Will Input to the decoder structure Calculate and get At this time Perform an attention mechanism calculation and add the output feature result to the feature middle, in, represents a decoder calculation function for extracting feature information of clear speech data, express The output feature information vector, represents a decoder calculation function for calculating feature information for extracting noise, express Output feature information vector; S6, group the voice segments Concatenate into y in order i,1 , group the speech segments Concatenate into y in order i,2 , two output results y i,1 ,y i,2 With input data x i The data dimensions are consistent. Among them, y i,1 is the clear original voice data after noise reduction, y i,2 For the input data x i The noise data contained in .
4. The method according to claim 3, wherein: The calculation formula for the attention mechanism calculation of the input data in step S2 is as follows: a n1 (n2) = exp(score(x i,n1 ,x i,n2 ))=exp(x i,n1 ·x i,n2 ),n1≠n2, Among them, x i,n1 is the current input data x i,n A piece of data in x i,n2 For the input data x i,n Divide by x i,n1 The data segment outside, a n1 (n2) is x i,n1 With x i,n2 The attention calculation value, a′ n1 (n2) is the weighted x i,n1 With x i,n2 The attention calculation value of Get a′ n1 (n2) and the new x′ can be calculated i,n1 , the calculation of x is completed by the following formula i,n1 Re-listing Sign: x′ i,n1 =x i,n1 +W Att ∑a′ n1 (n2)x i,n2 , Among them, x′ i,n1 Represents x i,n1 The feature vector calculated by the attention mechanism, W Att is a trainable parameter; From the above calculation, we can get: in, represents the attention mechanism function used to extract feature information of clear speech data, Represents the attention mechanism function used to extract feature information of noise.
5. The method according to claim 4, wherein: In step S3, the obtained and Input to the encoder structure for calculation, the calculation formula is as follows: x = ReLU (Liner (x)); μ = Linear (x); σ = exp (Liner (x)); z = Sample (μ, σ), Where x is the input vector, ReLU(·) is the activation function, Liner(·) is the linear function, μ is the mean, σ is the variance, Sample(·) is the sampling function, and z is the implicit feature vector with randomness; From the above calculation, we can get:
6. The method according to claim 5, wherein: In step S3, Perform an attention mechanism calculation and add the output feature result to the feature middle, The calculation process is as follows: Among them, β is a hyperparameter, The attention mechanism function for extracting the reference feature information of the noise for the first time.
7. The method according to claim 6, wherein: In step S4, and Input into the LSTM network for training. The LSTM unit structure includes forget gate, input gate, output gate, and cell state.
8. The method according to claim 7, wherein: In step S4, Forget Gate t Determines whether the information in the cell state is lost: the gate input data includes the input data x at time t t And the hidden vector h at the previous time t-1 t-1 , feedback to the cell state c t-1 , Among them, the forget gate f t The calculation formula is as follows: f t =σ(W f [x t ,h t-1 ]+b f ); Input Gate i t Determine the data that needs to be updated in the cell state, Among them, the input gate i t The calculation formula is as follows: i t =σ(W i [x t ,h t-1 ]+b i ); Update cell state c t , the cell state c at the previous time t-1 t-1 With f t Multiply, discard the information that needs to be forgotten, and then add the new cell state. Among them, the cell state c t The calculation formula is as follows: c t =f t c t-1 +i t ×(tanh(W c [x t ,h t-1 ]+b c )); Output gate o t Determine the output of some cell states, Among them, the output gate o t The calculation formula is as follows: the t =σ(W o [x] t ,h t-1 ]+b o ); Finally, calculate the hidden vector h at the current time t t , Among them, h t The calculation formula is as follows: h t =o t ×tanh(c t ); From the above calculation: Where t represents the time in the time series data, t-1 is the time before t, the value range of t is [1,n], n represents the number of data segments in the input data, and f t Represents the output information of the forget gate at time t, i t Represents the input gate output information at time t, o t represents the output information of the output gate at time t in LSTM, c t Represents the cell state output information at time t; x t Represents the data at time t in the time series data, that is, the tth segment of the input data, with a dimension of n x ×1,n x is a hyperparameter; h t-1 represents the hidden vector data at the previous time t-1, c t-1 Represents the cell state vector data at the previous time t-1; for t = 1, h0 is defined as the dimension is n h ×1 random initial vector, c0 is defined as the dimension is n h ×1 random initial vector, n h is a hyperparameter; W f Represents the trainable parameter matrix for forget gate calculation, dimension is n h ×(n x +n h ), b f Represents the trainable parameter vector of the forget gate calculation, with dimension n h ×1; W i Represents the trainable parameter matrix of the input gate calculation, dimension n h ×(n x +n h ), b i Represents the trainable parameter vector of the input gate calculation, with dimension n h ×1; W c Represents the trainable parameter matrix for cell state calculation, dimension n h ×(n x +n h ), b c Represents a trainable parameter vector for cell state calculation, dimension n h ×1; W o Represents the trainable parameter matrix of the output gate calculation, dimension n h ×(n x +n h ), b o Represents the trainable parameter vector of the output gate calculation, with dimension n h ×1.
9. The method according to claim 8, wherein: In step S5, Will and Input to the decoder structure for calculation, The calculation formula is as follows: x = ReLU(Liner(x)); x = Sigmoid(Liner(x)), Among them, ReLU(·) is the ReLU activation function, Sigmoid(·) is the Sigmoid activation function; From the above calculation, we can get: right Perform an attention mechanism calculation and add the output feature result to the feature middle, The calculation process is as follows: in, is the attention mechanism function for extracting the reference feature information of the noise for the second time, and γ is a hyperparameter.
10. The method according to any one of claims 3 to 9, wherein: The loss function formula in step S4 is: Among them, F loss represents the loss function, N represents the number of data instances in the dataset, and x i Represents a piece of speech data containing noise, y i Represents x i The corresponding clear original speech data, ‖·‖2 represents the Euclidean distance calculation, Represents the encoder calculation function used to extract feature information of clear speech data The resulting KL divergence is, Represents the encoder calculation function used to extract feature information of noise The resulting KL divergence, F KL Represents the KL divergence calculation formula.
Citation Information
Patent Citations
Speech recognition method based on model pre-training and bidirectional LSTM
CN108682418A
Time series data clustering method based on noise reduction encoder and attention mechanism
CN112348068A
Audio noise reduction model training method, and audio noise reduction method and device
CN114974280A
Single-channel voice noise reduction method and device based on deep learning
CN114974282A
Intensive LSTM residual network denoising method based on dynamic attention
CN116563144A