In the low
signal-to-
noise ratio condition, aiming at the problems that the traditional neural network speech
feature extraction is insufficient, and the
speech enhancement effect needs to be improved, based on empirical mode
decomposition (EMD), temporal
convolution network (TCN) and gated
convolution recurrent neural network (GCRN), and combining with
feature fusion module (FFM), the application proposes a
speech enhancement model of adaptive mean median empirical mode
decomposition-multilayer gated
feature fusion module convolutional
recurrent neural network (ME-MGFCRN). The
network model adopts the frequency learning strategy to learn the
low frequency feature and the
high frequency feature, that is, the TCN and the MGFCRN network are used to obtain the
low frequency and the
high frequency feature, and the two groups of features are processed through the FMM, so as to realize the
speech enhancement in the
feature mapping mode. The model proposed in the application carries out the
ablation experiment and the comparison experiment on the
data set, and uses the
PESQ, fwSegSNR and STOI indexes to evaluate the speech enhancement effect. Research shows that under different
noise environments and different
signal-to-
noise ratios, the model proposed in the application is improved compared with other baseline models, especially under the low
signal-to-noise ratio condition of SNR of-5dB, the fwSegSNR and
PESQ are improved by more than 0.86dB and 0.02 respectively compared with other baseline models.