Speech enhancement method based on PESQ-driven reinforcement learning to estimate prior signal-to-noise ratio
By introducing reinforcement learning and PESQ optimization in the Deep Xi-TCN network, the problem of poor speech enhancement effect in the existing technology in low signal-to-noise ratio and complex noise environments is solved, and higher speech quality perception evaluation scores and robustness are achieved.
Patent Information
- Application Number
- CN202111516319.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing voice enhancement technology has poor effect when dealing with low signal-to-noise ratio, non-steady state noise and strong reverberation environments, which can easily lead to speech distortion and is difficult to apply in short delay real-time processing.
Reinforcement learning is used to optimize the Deep Xi-TCN network, estimate the prior signal-to-noise ratio, and optimize it through PESQ indicators to improve the speech quality perception evaluation score.
Improve speech enhancement effect in complex scenarios such as low signal-to-noise ratio and non-steady state noise, improve robustness, and significantly improve perception scores.
Smart Images

Figure CN114141266B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech enhancement, and specifically relates to a method for optimizing a Deep Xi-TCN network estimation of a priori signal-to-noise ratio using reinforcement learning, which is used to improve the speech quality perception evaluation score. Background Art
[0002] In practical applications, ubiquitous noise and reverberation greatly impair the experience of voice interaction and the performance of automatic speech recognition (ASR). The purpose of speech enhancement is to extract clear speech from background interference to obtain higher speech intelligibility and perceptual quality. Spectral subtraction can be used to achieve noise suppression. This method estimates the noise power spectrum based on the minimum mean square error (MMSE) (GERKMANN T, HENDRIKS R C. Unbiased MMSE-Based Noise Power Estimation With Low Complexityand Low Tracking Delay [J]. IEEE Transactions on Audio Speech & Language Processing, 2012, 20 (4): 1383–1393), then subtracts the noise power spectrum from the noisy speech power spectrum to obtain the power spectrum of the enhanced speech, and then combines the phase information of the short-time Fourier spectrum of the noisy speech to obtain the short-time Fourier spectrum of the enhanced speech, and then obtains the enhanced speech signal through inverse Fourier transform. Spectral subtraction has achieved good noise suppression effects in many scenarios, but due to the limitations of its assumed noise and speech models, the algorithm has poor effects in processing certain speech in low signal-to-noise ratio (SNR) and non-steady-state noise scenarios, which can easily lead to speech distortion. The WPE algorithm is used for speech dereverberation (NAKATANI T, YOSHIOKA T, KINOSHITA K, et al. Speech Dereverberation Based on Variance-Normalized Delayed Linear Prediction [J]. IEEE Transactions on Audio Speech & Language Processing, 2010, 18 (7): 1717–1731). It establishes a time frame autoregressive model for the speech short-time Fourier spectrum, estimates the inverse filter coefficients and the power spectrum of the early reverberation in an iterative manner, and then obtains the short-time Fourier spectrum of the clear speech. The WPE algorithm has achieved excellent results in speech dereverberation, but the iterative nature of the algorithm makes it difficult to use in short-delay real-time processing.
[0003] In recent years, deep neural networks (DNNs) have achieved remarkable results in the field of speech enhancement due to their powerful nonlinear modeling capabilities (WANG, DL, CHEN J. Supervised speech separation based on deep learning: An overview [J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2018, 26 (10): 1702-1726.). For single-channel speech enhancement, end-to-end processing is the most direct approach, but it faces the challenge of generalization, that is, the output of DNNs may be severely deteriorated under noisy conditions not included in the training set. The recently proposed Deep Xi framework (ZHANG Q, NICOLSON A, WANG M, et al. DeepMMSE: A deep learning approach to MMSE-based noise power spectral density estimation [J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2020, 28: 1404-1415.) can be viewed as an effective hybrid approach that combines a rule-based MMSE speech enhancement strategy with a data-driven deep learning approach to estimate the prior signal-to-noise ratio. Unlike other noise power spectral density estimators, it does not make any assumptions about the characteristics of speech or noise, does not exhibit any tracking delay, and does not rely on bias compensation. In addition, DNN is only used to track the noise power spectral density (PSD) and signal-to-noise ratio, and its output is calculated according to a rule-based method, so the risk of end-to-end methods can be reduced.
[0004] For deep learning-based speech enhancement, some studies have pointed out that processing speech with a general standard such as the mean square error between the estimated signal and the clear speech does not guarantee high speech quality and intelligibility. Among the objective indicators related to human perception, Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI) are two popular indicators for evaluating speech quality and intelligibility. Therefore, it is a meaningful work to directly use these two functions to optimize the model. Some studies focus on the optimization of the STOI score to improve speech intelligibility, but the PESQ score cannot be improved by maximizing the STOI score. Other studies have simplified the calculation of the symmetric interference vector in PESQ and applied a center clipping operator on the absolute difference of the loudness spectrum, so that it can be included in the training objective. However, PESQ itself is not differentiable, and the derivative of backpropagation cannot be calculated, so it is difficult to get a general training scheme.
[0005] As a self-optimization method, reinforcement learning (RL) can be understood as taking actions in a feedback environment to let the machine learn the optimal strategy and maximize the cumulative reward. It has received widespread attention in the fields of robot behavior control, intelligent dialogue management, letting robots play games and speech recognition. The use of RL has been explored in end-to-end speech enhancement solutions, and it has been verified that enhanced speech can indeed lead to better PESQ scores, and RL enhancement solutions have the advantage of less training data. Summary of the invention
[0006] Traditional rule-based methods often have difficulty removing noise components when enhancing speech in low signal-to-noise ratio, non-steady-state noise, and strong reverberation environments, and may even cause serious speech distortion. The effect of pure end-to-end methods will be greatly deteriorated when faced with unfamiliar noise and reverberation environments. Deep Xi is a hybrid speech enhancement solution that combines rule-based methods and deep learning speech enhancement methods. However, since the training process relies on the mean square error standard between the estimated signal and the clear speech to achieve convergence, it cannot achieve the optimization of the perceptual evaluation of speech quality. To this end, the present invention proposes to further use reinforcement learning to introduce the PESQ indicator to optimize the estimation of the signal-to-noise ratio based on the prior signal-to-noise ratio estimated by the Deep Xi-TCN network, thereby obtaining clear speech with better perceptual scores, especially when the signal-to-noise ratio is low. The improvement effect is relatively obvious.
[0007] To achieve the above object, the technical solution adopted by the present invention is:
[0008] A speech enhancement method based on PESQ-driven reinforcement learning to estimate a priori signal-to-noise ratio, the method comprising the following steps:
[0009] Step 1, using the clear speech and noise in the training set to synthesize simulated noisy speech with a random signal-to-noise ratio, and performing short-time Fourier transform on the three to obtain a short-time Fourier spectrum of the clear speech, a short-time Fourier spectrum of the noise, and a short-time Fourier spectrum of the simulated noisy speech;
[0010] Step 2, using the short-time Fourier spectrum of the clear speech and the short-time Fourier spectrum of the simulated noisy speech to train the DeepXi-TCN network;
[0011] Step 3, dividing the short-time Fourier spectrum amplitude of the clear speech and the short-time Fourier spectrum amplitude of the noise, and mapping their range to [0,1], generating a mapping signal-to-noise ratio of the training set, and then generating a finite number of cluster centers through K-means clustering as a priori signal-to-noise ratio template;
[0012] Step 4, labeling each frame of the simulated noisy speech using the prior signal-to-noise ratio template to train the DQN network initialization parameters;
[0013] Step 5, the formal training phase, the DQN network selects a signal-to-noise ratio template at the frame level, the signal-to-noise ratio template is the signal-to-noise ratio inferred by the Deep Xi-TCN network trained in step 2 or the prior signal-to-noise ratio template generated in step 3; then the reward related to the PESQ value is calculated, reinforcement learning iterations are performed, and the DQN network parameters are updated;
[0014] Step 6: Input the short-time Fourier spectrum of the clear speech and the noisy speech synthesized from the test set into the DQN network trained in step 5, and perform an inverse short-time Fourier transform on the obtained short-time Fourier spectrum of the enhanced speech to obtain the time domain signal of the enhanced speech.
[0015] Furthermore, in step 2, the input data of the Deep Xi-TCN network first passes through a fully connected input layer, then passes through several residual blocks, and then outputs the estimated mapping signal-to-noise ratio through a fully connected output layer; each residual block includes three layers of one-dimensional convolutional networks with ReLU activation function and layer regularization, which can realize two-dimensional feature extraction of time-frequency domain blocks.
[0016] Furthermore, in step 4, the specific steps of labeling each frame of the simulated noisy speech using the priori signal-to-noise ratio template are as follows: using the mean square error rule to determine the distance between the ideal signal-to-noise ratio of all frequency points in each frame and the template signal-to-noise ratio, and taking the template with the smallest distance. The number m is the label of the corresponding frame; the index numbers from 1 to M are used as the labels of the corresponding frames in the training set.
[0017] The method of the present invention can enhance speech in a variety of complex noise scenarios such as low signal-to-noise ratio and non-steady-state noise, with high robustness, and the perceptual score is also significantly improved. As an effective hybrid method, the Deep Xi method combines the rule-based MMSE speech enhancement strategy and the data-driven deep learning method to estimate the prior signal-to-noise ratio. Unlike other noise power spectral density estimators, it does not make any assumptions about the characteristics of speech or noise, does not show any tracking delay, and does not rely on bias compensation. In addition, DNN is only used to track the noise power spectral density (Power Spectral Density, PSD) and signal-to-noise ratio, and its output is calculated according to a rule-based method, so the risk of the end-to-end method can be reduced. On this basis, by adding a double-layer fully connected network trained with a reinforcement learning strategy to select the signal-to-noise ratio template at the frame level, the optimization of introducing the PESQ indicator into the model is realized, and a better perceptual score of the estimated speech is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a processing flow chart of the method of the present invention.
[0019] Figure 2 It is a flow chart of the training phase in the method of the present invention.
[0020] Figure 3 Flowchart for reconstructing the time domain signal for the training phase.
[0021] Figure 4 Schematic diagram of the Deep Xi-TCN network structure.
[0022] Figure 5 Schematic diagram of the DQN network structure.
[0023] Figure 6 This is a curve chart showing the changes in PESQ scores during the training phase.
[0024] Figure 7 This is a comparison chart of speech enhancement results processed by the method of the present invention and Deep Xi-TCN, (a) clear speech signal, (b) noisy reverberation signal, (c) Deep Xi-TCN method processing result, (d) the processing result of the method of the present invention. DETAILED DESCRIPTION
[0025] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0026] This embodiment provides a speech enhancement method based on PESQ-driven reinforcement learning to estimate the prior signal-to-noise ratio. The PESQ scoring indicator is introduced into Deep Xi-TCN, the prior signal-to-noise ratio is regarded as the behavior in RL, and rewards related to PESQ are designed. Discrete actions are composed of a pre-trained frame-level prior signal-to-noise ratio template and the prior signal-to-noise ratio obtained by Deep Xi-TCN. Then the Double Q Learning strategy is used to select the best prior signal-to-noise ratio and PESQ reward function. The overall process is as follows: Figure 1 As shown, the following steps are included:
[0027] Step 1: Use the clear speech data set and the noise data set of the training set to synthesize simulated noisy speech with a random signal-to-noise ratio, and perform short-time Fourier transform on the three to obtain a short-time Fourier spectrum;
[0028] Step 2: Use the clear speech data set in step 1 and the short-time Fourier spectrum of simulated noisy speech to train the DeepXi-TCN network;
[0029] Step 3, using the short-time Fourier spectrum of the clear speech and noise corresponding to the simulated noisy speech synthesized in step 1 to generate an ideal mapping signal-to-noise ratio, and generating a finite number of cluster centers through K-means clustering as a priori signal-to-noise ratio template;
[0030] Step 4: Use the prior signal-to-noise ratio template to label each frame of the simulated noisy speech synthesized in step 1 to train the initialization parameters of the DQN network;
[0031] Step 5, the formal training phase is as follows Figure 2 As shown, the DQN network selects the signal-to-noise ratio or prior signal-to-noise ratio template inferred by the trained DeepXi-TCN network at the frame level, calculates the reward related to the PESQ value, and performs reinforcement learning feedback iteration to update the network parameters;
[0032] Step 6: Input the short-time Fourier spectrum of the simulated noisy speech obtained in step 1 into the trained model, and perform an inverse short-time Fourier transform on the obtained short-time Fourier spectrum of the enhanced speech to obtain a time domain signal of the enhanced speech.
[0033] 1. Deep Xi Hybrid Method
[0034] The signal model in the time-frequency domain can be obtained by short-time Fourier transform (STFT):
[0035] Y l [k] = S l [k]+D l [k] (1)
[0036] where Y l [k],S l [k] and D l [k] are the complex coefficients of the short-time Fourier transform of noisy speech, clear speech and noise, respectively. l is the time frame index and k is the discrete frequency index. Applying the standard assumptions of the Deep Xi framework, S l [k] and D l [k] is statistically independent in time and frequency frames and follows a conditional zero-mean Gaussian distribution with spectral variance λ s [l, k] and λ d [l, k]. Let R = |Y l [k]|, the prior signal-to-noise ratio ξ and the posterior signal-to-noise ratio γ are defined as:
[0037]
[0038] The Deep Xi framework is briefly described as follows. Theoretically, the range of the prior signal-to-noise ratio is [0, +∞], while DNN requires the training target to be in a limited interval. Therefore, an appropriate mapping is required. 10 (ξ l [k]) obeys the following Gaussian distribution:
[0039]
[0040] The mean and variance are μ k and σ k 2 The mapped signal-to-noise ratio is given by:
[0041]
[0042] erf(·) represents the Gaussian error function. The estimated prior signal-to-noise ratio Can be restored using:
[0043]
[0044] in is an estimate of the mapping signal-to-noise ratio.
[0045] The Deep Xi-TCN network replaces the ResLSTM network in the traditional Deep Xi framework with a temporal convolution network (TCN). Its structure is as follows: Figure 4 As shown in Figure 1, it consists of a fully connected layer FC connecting the input spectrum and several residual blocks, and then uses a fully connected layer of Sigmoidal units to connect the residual blocks and the output layer O. The input of the TCN network is the noisy speech spectrum R of the lth framel , connected to 40 residual blocks through a 256-node fully connected layer with ReLU activation function. Each residual block contains three one-dimensional causal dilated convolution units with dimensions (1, d f , 1), (k, d f , d), (1, d mod el , 1). The output dimension of the first and second units is d f =64, the output dimension of the third unit is d mod el = 256, the kernel size of the second unit is k = 3, and the expansion rate Where mod() is a modulus operation. The maximum dilation rate is set to 16, which means that the dimension of d will cycle through 1, 2, 4, 8, and 16 as the residual block labels increase. The existence of the causal dilated convolution unit allows the network to use contextual information (only the above if it is a causal network) and use temporal correlation to get better results. The last residual block is connected to an output layer with 256 nodes and a sigmoid activation function, which outputs the prior signal-to-noise ratio of the mapping of the lth frame.
[0046] After estimating the a priori SNR estimate, a corresponding gain function is needed to recover the estimated signal. The minimum mean square error log spectral amplitude (MMSE-LSA) estimator minimizes the MSE between the log spectra of the clear speech and the enhanced speech, which is one of the best performing gain functions. The instantaneous a posteriori SNR is estimated by the instantaneous a priori SNR as γ = ξ + 1, and the gain function is given by
[0047]
[0048] 2. XiDQN Model Framework
[0049] The reinforcement learning method proposed in this paper aims to improve the PESQ score. The Deep Q Network (DQN) is used to identify clear speech from the normalized power spectrum of noisy speech and select the highest reward prior signal-to-noise ratio, so it is called the XiDQN model, and the reward target uses the PESQ score.
[0050] Combination Figure 1 , 2 and 3. The processes of the initialization phase and the training phase are described in detail below.
[0051] In the initialization phase, the Deep Xi-TCN network obtains a frame-level mapped prior SNR and is considered as a candidate action, represented as In order to form a complete action template, the K-means clustering algorithm is used to obtain the ideal prior signal-to-noise ratio. M candidate actions are formed on the basis of the prior signal-to-noise ratio, which is generated by the ratio between the power spectra of the clear speech and the noise in the training set. In this way, a finite action template with M+1 candidate actions is generated. The DQN network can be viewed as an action-value function Q(R l , a l ), where R l =[R l [0],R l [1], ..., R l [K]] T is the amplitude spectrum of the noisy speech, a l =[a l [0],a l [1], ..., a l [K]] T is the prior signal-to-noise ratio, and K is the number of frequency points. In order to have reasonable initialization parameters for the DQN network before training, this embodiment pre-trains the network in the initialization stage.
[0052] Initialization parameter Θ of DQN q is trained in the following way. First, the prior signal-to-noise ratio of the training set is calculated and mapped to The distance between the ideal signal-to-noise ratio and the template signal-to-noise ratio of all frequencies in each frame is determined by the mean square error rule, and the number of the template with the smallest distance is taken as the label of the corresponding frame, as shown in the following formula (7). The index numbers from 1 to M are used as the labels of the corresponding frames in the training set.
[0053]
[0054] where ⊙ is the Hadamard product. This process can be viewed as a classification task. The network parameters are updated by back-propagation. The weights and biases of each fully connected layer are initialized with a normal distribution.
[0055] During the training phase, the parameters θ of DQN q The goal of training is to maximize the reward associated with PESQ. During training, a dual-Q learning strategy is used, which decouples selection from evaluation to prevent overestimation. This method does not require additional networks or parameters. This embodiment has two DQN networks with different update rates: the network updated in each iteration is called the evaluation DQN (Eval.DQN), and the network that periodically copies the parameters of Eval.DQN is called the target DQN (TargetDQN). The amplitude spectrum of the noisy speech is input into the two networks at the same time, and Q′ (R l , a l ) and Q(R l , a l). In addition to the update rate, another difference between the two DQNs is that the Target DQN directly follows the standard process of DQN to select an action, while the Evaluation DQN randomly picks an action with probability ∈ . After making an action selection, both produce their respective estimated speech and The reward is then calculated based on the difference between their PESQs and the DQN parameters are updated in a self-optimizing manner. Note that Figure 1 It focuses on action selection for a specific frame, while ignoring the context window size and block processing. The training details and reward settings are described below.
[0056] Use Q Learning strategy to select appropriate frame-level behavior a l As follows
[0057]
[0058]
[0059] where ⊙ is the Hadamard product. MMSE-LSA (.) Input vector or matrix to return the corresponding vector and matrix of each frequency point MMSE-LSA gain, as shown in formula (6). l-P , .., Y l , ..., Y l+P ] is the noisy speech spectrum, and 2P+1 is the length of the context window. is the prior signal-to-noise ratio matrix inferred by DQN, so is the inferred clear speech spectrum. It is the time domain waveform of the clear speech restored by inverse short time Fourier transform (iSTFT), which is needed for the calculation of the next reward.
[0060] The reward setting is very important. In order to give appropriate rewards for different signal-to-noise ratios and different noise types, the range of rewards needs to be constrained. The relative PESQ value between the evaluation network and the target network is used as the reward:
[0061]
[0062] Where α>0 is the scaling parameter. and It is the PESQ value calculated based on the estimated speech of the target DQN and the evaluation DQN. The PESQ value calculated by DQN respectively. Considering that the prior signal-to-noise ratio varies with time and the PESQ value cannot be calculated in one frame, it is necessary to calculate the reward that varies with time for multiple frames. The time weight E is used in the reward calculation l∈[0, 1], that is
[0063]
[0064]
[0065]
[0066] Once the ∈-greedy strategy changes from In the example, we randomly select an action a different from the one used to evaluate the DQN. ε , the expected Q value of the action-value function that evaluates the DQN iteration is updated according to the following rule
[0067]
[0068] Where Q(R l , a l ) is the Q value estimated by the target DQN, is the expected Q value for evaluating DQN. (At this time r l <0), the maximum Q value of the target DQN is subtracted from r l , the reward target DQN selects a better signal-to-noise ratio behavior than the evaluation DQN. In addition, in order to set an upper limit on the Q value of the evaluation DQN, the activation function of its output layer is softmax. Accordingly, It will also be normalized to satisfy
[0069] Finally, the parameter Θ is updated by minimizing the following equation q , so that the value Q′(R l , a l ) is close to the expected value
[0070]
[0071] In order to minimize formula (15), the present invention uses the RMSProp algorithm and the standard small batch stochastic gradient descent (SGD).
[0072] In the inference phase, only the trained Deep Xi-TCN and Target DQN are used. The trained target DQN also needs to determine which of the M+1 candidate SNR templates is best for a given frame.
[0073] 3. Datasets and Experimental Parameters
[0074] The method proposed in this invention is named XiDQN, and its performance is compared with the Deep Xi-TCN method. In the experiment, the clear speech corpus includes the TIMIT speech dataset (6289 corpora) and the train-clean-100 set of the Librispeech dataset (28539 corpora). The noisy audio includes the Nonspeech dataset, the environmental background noise dataset, and the noise part of the MUSAN corpus. The clear speech and noise are divided into training set, validation set, and test set, with ratios of 0.7, 0.1, and 0.2, respectively. In addition, white noise is added to the noise part of the training set. All speech and noise are unified to a sampling rate of 16kHz (recordings with a sampling frequency higher than 16kHz are downsampled to 16kHz). The generation rule of the noisy speech signal is as follows: each clear speech is mixed with a randomly selected noise signal, and the mixed signal-to-noise ratio is randomly sampled from -10dB to 15dB with an increment of 1dB.
[0075] The number of a priori SNR candidates in the template is 32. Figure 5 As shown, the DQN used in the framework consists of two fully connected hidden layers with 66 units and sigmoid activation function. The activation function of the output layer is softmax. The adjustable scale parameter in formula (9) is set to 20. The context half-window size P is set to 15. The dropout technique is used in training to avoid overfitting. The frame size of the STFT is 512 with a displacement of 256 samples. The greed parameter ∈ varies linearly from 0.20 to 0.01. The learning rate is set using the 1cycle learning rate method for training acceleration, increasing between 0.00001 and 0.0005 and then decreasing.
[0076] IV. Experimental Results
[0077] Figure 6 The evolution of the PESQ scores computed from the estimated utterances of the target DQN during training is shown. For comparison, a fixed average PESQ score computed from the trained Deep Xi-TCN is also depicted. A mini-batch of 8 training audios is used to iteratively update the evaluation DQN, whose parameters are periodically copied to the target DQN every 20 updates. Figure 6 It can be seen that the PESQ score increases with the number of iterations and exceeds the score of Deep Xi-TCN after about 160 iterations. XiDQN has an overall PESQ improvement of about 0.11 over Deep Xi-TCN after convergence. It should be noted that the convergence behavior of the PESQ score is not as smooth as the learning curve of Deep Xi, because PESQ is calculated over randomly selected samples from the training dataset.
[0078] On the test set, STOI is used as an evaluation metric in addition to PESQ. Table 1 lists the PESQ and STOI (%) scores for enhanced speech under -6dB, 0dB, 6dB, and 12dB SNR conditions. It can be seen that XiDQN has an advantage in STOI, although it pales in comparison to the advantage of PESQ. Note that at low SNR, the XiDQN method has a more obvious improvement over Deep Xi-TCN, indicating that the action selection made by the XiDQN network brings obvious gains when the noise energy is relatively high.
[0079] Table 1 PESQ and STOI (%) scores of the test set
[0080]
[0081] Figure 7 An example of a spectrogram of processed speech at 0dB SNR is shown. By comparing (c) and (d) of the graphs, the improvement of the proposed XiDQN can be seen. The two dashed boxes on the left of these two figures show the more effective noise suppression of XiDQN, while the dashed box on the right shows that XiDQN more clearly preserves the consonant syllables.
Claims
1. A speech enhancement method based on PESQ-driven reinforcement learning to estimate prior signal-to-noise ratio, characterized in that: The method comprises the following steps: Step 1, using the clear speech and noise in the training set to synthesize simulated noisy speech with a random signal-to-noise ratio, and performing short-time Fourier transform on the three to obtain a short-time Fourier spectrum of the clear speech, a short-time Fourier spectrum of the noise, and a short-time Fourier spectrum of the simulated noisy speech; Step 2, using the short-time Fourier spectrum of the clear speech and the short-time Fourier spectrum of the simulated noisy speech to train the Deep Xi-TCN network; Step 3, dividing the short-time Fourier spectrum amplitude of the clear speech and the short-time Fourier spectrum amplitude of the noise, and mapping their range to [0,1], generating a mapping signal-to-noise ratio of the training set, and then generating a finite number of cluster centers through K-means clustering as a priori signal-to-noise ratio template; Step 4, labeling each frame of the simulated noisy speech using the prior signal-to-noise ratio template to train the DQN network initialization parameters; Step 5, the formal training phase, the DQN network selects a signal-to-noise ratio template at the frame level, the signal-to-noise ratio template is the signal-to-noise ratio inferred by the Deep Xi-TCN network trained in step 2 or the prior signal-to-noise ratio template generated in step 3; then the reward related to the PESQ value is calculated, reinforcement learning iterations are performed, and the DQN network parameters are updated; Step 6: Input the short-time Fourier spectrum of the clear speech and the noisy speech synthesized from the test set into the DQN network trained in step 5, and perform an inverse short-time Fourier transform on the obtained short-time Fourier spectrum of the enhanced speech to obtain the time domain signal of the enhanced speech.
2. The method for speech enhancement based on PESQ-driven reinforcement learning estimation of prior signal-to-noise ratio according to claim 1, characterized in that: In step 2, the input data of the Deep Xi-TCN network first passes through a fully connected input layer, then passes through several residual blocks, and then outputs the estimated mapping signal-to-noise ratio through a fully connected output layer; each residual block includes three layers of one-dimensional convolutional networks with ReLU activation function and layer regularization, which can realize two-dimensional feature extraction of time-frequency domain blocks.
3. The speech enhancement method based on PESQ-driven reinforcement learning estimation of prior signal-to-noise ratio according to claim 1, characterized in that: In step 4, the specific steps of labeling each frame of the simulated noisy speech using the prior signal-to-noise ratio template are as follows: using the mean square error rule to determine the distance between the ideal signal-to-noise ratio of all frequency points in each frame and the template signal-to-noise ratio, and taking the template with the smallest distance. The number m is the label of the corresponding frame; the index numbers from 1 to M are used as the labels of the corresponding frames in the training set.
4. The method for speech enhancement based on PESQ-driven reinforcement learning estimation of prior signal-to-noise ratio according to claim 1, characterized in that: There are two DQN networks with different update rates in step 5: the network that is updated in each iteration is called the evaluation DQN network, while the network whose parameters are copied periodically is called the target DQN network; The double Q strategy is used to calculate the rewards associated with the PESQ value. The rewards are set as follows: Evaluate the relative PESQ value between the DQN network and the target DQN network Where a>0 is the scaling parameter, and It is the PESQ value calculated based on the estimated speech of the target DQN network and the evaluation DQN network; considering that the prior signal-to-noise ratio varies with time and the PESQ value cannot be calculated within one frame, it is necessary to calculate the time-varying rewards for multiple frames, and use the time weight E in the reward calculation l ∈[0,1], that is Where k is the discrete frequency domain number, l' is the frame number of the P frames before and after the lth frame, and 2P+1 is the size of the context window length; S l' [k] is the spectrum of clear speech, Y l' [k] is the spectrum of noisy speech, It is the a priori signal-to-noise ratio option inferred by the DQN network; By comparing the inference results of the current iterative evaluation DQN network with the inference results of the target DQN network with a delayed update, the corresponding node of the network will receive a corresponding reward if the result is better, otherwise it will be punished.
5. The method for speech enhancement based on PESQ-driven reinforcement learning estimation of prior signal-to-noise ratio according to claim 4, characterized in that: The behavior-value function of the DQN network iteration, i.e., the expected Q value of the Q function, is updated according to the following rules: Among them, Q(R l ,a l ) is the Q value estimated by the target DQN network, Q'(R l ,a l ) is the Q value estimated by the DQN network, is the expected Q value for evaluating the DQN network.
Citation Information
Patent Citations
Voice enhancing method based on generative adversarial network
CN110428849A
Speech enhancement method for generating adversarial network based on two-dimensional spectrogram and conditions
CN110718232A