A speech enhancement method combining hearing loss compensation and speech noise reduction
By using a frequency-time convolution recursive metric generative adversarial network model, combined with hearing loss compensation and speech noise reduction, the problems of computational redundancy and noise amplification of hearing aids in noisy environments are solved, efficient hearing loss compensation and noise suppression are achieved, and the speech comprehension of hearing-impaired patients is improved.
Patent Information
- Application Number
- CN202310414553.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-04-17
AI Technical Summary
Existing hearing aids find it difficult to effectively combine hearing loss compensation and speech noise reduction in noisy environments, resulting in redundant calculations and amplification of ambient noise, affecting the compensation effect.
A metric generative adversarial network model with frequency-time convolutional recursion is adopted to achieve the combination of hearing loss compensation and speech noise reduction through hearing loss map embedding and metric discriminator adversarial training, and the compensation generator is used to generate compensated speech with better quality perceived by the human ear.
It stably and effectively improves the hearing loss compensation effect in noisy environments, and simultaneously completes noise suppression and spectral-level compensation specialized for audiograms through a single network, improving the speech comprehension of hearing-impaired patients.
Smart Images

Figure CN116434766B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech enhancement, and in particular relates to a speech enhancement method combining hearing loss compensation and speech noise reduction. Background Art
[0002] Hearing rehabilitation has become a significant challenge for people worldwide. It is estimated that approximately 500 million people worldwide suffer from hearing loss. Hearing aids are currently an effective means of hearing rehabilitation. However, existing hearing aid technologies still face significant challenges in compensating for hearing loss in noisy environments. Currently, noise reduction methods for hearing aids are mainly divided into spatial selection methods and non-spatial noise reduction methods. Spatial selection methods (such as beamforming) are designed based on the spatial differences between the target speech and interference sources, maintaining greater sensitivity to the direction of the target speech. However, in reality, noise can originate from multiple directions, and the orientation of the target speech and interference sources can also change over time. Single-microphone noise reduction technology primarily utilizes the time-frequency differences between the target speech and interference sources to suppress noise, but this provides limited improvement in speech clarity. Therefore, how to effectively improve speech intelligibility for hearing-impaired individuals in noisy environments is a key area of current hearing aid design.
[0003] In recent years, speech noise reduction technology based on deep neural networks has made breakthrough progress. Compared with traditional algorithms, deep learning-based methods have more obvious advantages in non-stationary and low signal-to-noise ratio environments. Deep learning-based noise reduction technology has also begun to be applied to the field of hearing aids. For example, auditory excitation features are extracted to feed the neural network and estimate the floating value masking to perform noise suppression for cochlear implants; neural architecture search is guided by speech intelligibility to optimize the neural network for noise reduction for hearing aids; deep filtering of broadband spectrograms is used to perform noise reduction for hearing aids. However, these methods do not reasonably combine noise suppression with hearing loss compensation. In fact, different hearing loss patients have different gain requirements for specific frequencies. While suppressing noise, the hearing loss compensation effect for specific patients should also be considered.
[0004] According to current research, speech processing algorithms based on hearing aids usually treat speech enhancement and hearing loss compensation as two separate algorithms. However, processing hearing loss compensation and speech enhancement separately will bring about computational redundancy and increase system latency. In addition, discussing hearing loss compensation without considering the acoustic environment may lead to the amplification of environmental noise, thereby affecting the compensation effect. Summary of the Invention
[0005] Purpose of the invention: In response to the problems existing in the prior art, the present invention discloses a speech enhancement method that combines hearing loss compensation and speech noise reduction, which can stably and effectively improve the effect of hearing loss compensation in noisy environments. The method is ingenious and novel and has good application prospects.
[0006] Technical solution: To achieve the above-mentioned purpose, the present invention adopts the following technical solution:
[0007] A speech enhancement method combining hearing loss compensation and speech noise reduction comprises the following steps:
[0008] S1: Extract complex spectrogram features from noisy training speech using short-time Fourier transform;
[0009] S2: Extending and embedding the hearing loss map along the frequency axis to obtain a hearing loss spectrum, and superimposing the hearing loss spectrum with the complex spectrum features of the noisy training speech obtained in step S1;
[0010] S3: Construct a metric generative adversarial network model based on frequency-time convolutional recursion. The main structure of the model includes a compensation generator and a metric discriminator.
[0011] S4: Alternately train the compensation generator and the metric discriminator, and optimize the metric generative adversarial network model through adversarial loss and compensation loss related to perceptual quality;
[0012] S5: The complex spectrogram features of the speech to be tested are superimposed with the hearing loss spectrum and then input into the trained compensation generator, and the enhanced speech waveform of the speech to be tested is reconstructed according to the output of the compensation generator.
[0013] Preferably, the step S2 includes:
[0014] First, the hearing loss map is Fourier transformed and then sampled at frequencies of [250, 500, 1000, 2000, 4000, 8000] Hz to obtain a 6-dimensional hearing loss map gain. For each dimension of the hearing loss map gain, it is normalized and used as the gain value for each frequency interval to form a hearing loss spectrum.
[0015] The hearing loss spectrum is then superimposed with the complex spectrogram features of the noisy training speech to achieve alignment in frequency bins.
[0016] Preferably, the compensation generator includes a convolution encoder, a frequency-time recursive processing module and a convolution decoder;
[0017] The step S3 comprises:
[0018] S31: Construct a convolutional encoder for the compensation generator. The convolutional encoder uses five convolutional blocks to extract local patterns from its input features along the frequency axis and reduce the feature resolution. Each convolutional block includes a two-dimensional convolutional layer followed by batch normalization and a PReLU activation function. The convolutional layers of the first two convolutional blocks have a convolution stride of 2.
[0019] The input of the convolutional encoder is the input of the compensation generator;
[0020] S32: constructing a frequency-time recursive processing module of the compensation generator, the frequency-time recursive processing module including a first recursive processing module and a second recursive processing module having the same structure;
[0021] The input of the first recursive processing module is the output of the convolutional encoder. The first recursive processing module first slices its input features along the time axis and independently models the frequency axis correlation through the BLSTM network for each frame. After the BLSTM network, a fully connected layer and a normalization layer are sequentially set for post-processing. The features obtained after post-processing are superimposed with the input features of the first recursive processing module to form the first feature;
[0022] The first feature is then sliced along the frequency axis, and the time correlation of each frequency stream is modeled using an LSTM network. A fully connected layer and a normalization layer are sequentially set after the LSTM network for post-processing. The features obtained after post-processing are superimposed with the first feature to form the output of the first recursive processing module.
[0023] The input of the second recursive processing module is the output of the first recursive processing module. The second recursive processing module first slices its input features along the time axis and independently models the frequency axis correlation through the BLSTM network for each frame. After the BLSTM network, a fully connected layer and a normalization layer are sequentially set for post-processing. The features obtained after post-processing are superimposed with the input features of the second recursive processing module to form the second feature.
[0024] The second feature is then sliced along the frequency axis, and the time correlation of each frequency stream is modeled using an LSTM network. A fully connected layer and a normalization layer are sequentially set after the LSTM network for post-processing. The features obtained after post-processing are superimposed with the second feature to form the output of the second recursive processing module.
[0025] The output of the second recursive processing module is the output of the frequency-time recursive processing module;
[0026] S33: Construct a convolutional decoder for the compensation generator. The convolutional decoder is a mirror image of the convolutional encoder. It uses five convolutional blocks to restore its input features to the input feature size of the convolutional encoder. Each convolutional block includes a two-dimensional transposed convolutional layer and a batch normalization process and a PReLU activation function set in sequence after the transposed convolutional layer. The transposed convolutional layers of the last two convolutional blocks are set with a convolution stride of 2.
[0027] The output of the convolution encoder and the output of the frequency-time recursive processing module are spliced in the channel dimension and input into the convolution decoder; the output of the fourth convolution block of the convolution encoder and the output of the first convolution block of the convolution decoder are spliced in the channel dimension and input into the second convolution block of the convolution decoder; the output of the third convolution block of the convolution encoder and the output of the second convolution block of the convolution decoder are spliced in the channel dimension and input into the third convolution block of the convolution decoder; the output of the second convolution block of the convolution encoder and the output of the third convolution block of the convolution decoder are spliced in the channel dimension and input into the fourth convolution block of the convolution decoder; the output of the first convolution block of the convolution encoder and the output of the fourth convolution block of the convolution decoder are spliced in the channel dimension and input into the fifth convolution block of the convolution decoder.
[0028] The output of the convolutional decoder is the output of the compensation generator;
[0029] S34: Construct a metric discriminator. First, extract the input features through four two-dimensional convolution blocks. Each two-dimensional convolution block consists of a two-dimensional convolution layer, an instance normalization layer, and a PReLU activation function connected in sequence. After the fourth two-dimensional convolution block, a global average pooling operation is performed. Finally, the output is activated by two fully connected layers and a sigmoid function.
[0030] Preferably, the convolution kernel size and step size of the transposed convolution layer of the first convolution block, the second convolution block, the third convolution block, the fourth convolution block, and the fifth convolution block of the convolution decoder are respectively the same as the convolution kernel size and step size of the convolution layer of the fifth convolution block, the fourth convolution block, the third convolution block, the second convolution block, and the first convolution block of the convolution encoder.
[0031] Preferably, step S4 includes:
[0032] S41: Inputting the features obtained by superimposing the hearing loss spectrum in step S2 with the complex spectrogram features of the noisy training speech into a compensation generator, and multiplying the output of the compensation generator by the complex spectrogram features of the noisy training speech to obtain the complex spectrogram features of the enhanced training speech;
[0033] S42: Train the metric discriminator:
[0034] First, the two inputs of the metric discriminator are set to the clean compensated speech and the enhanced training speech after the noisy training speech is compensated by the fitting formula. The loss is calculated by comparing the prediction score output by the metric discriminator with the HASQI label obtained by the HASQI calculation between the enhanced training speech and the clean compensated speech. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech and the enhanced training speech are superimposed with the hearing loss spectrum, and the superimposed features are input into the metric discriminator.
[0035] In addition, both inputs of the metric discriminator are set to clean compensated speech, and the prediction score output by the metric discriminator is compared with the all-one vector to calculate the loss. Before the clean compensated speech is input into the metric discriminator, its complex spectrogram features are superimposed with the hearing loss spectrum, and the superimposed features are input into the metric discriminator.
[0036] The total loss of the discriminator in this step is measured The description is as follows:
[0037]
[0038] in, and HL are clean compensated speech, enhanced training speech and corresponding hearing loss map, respectively. D represents the metric discriminator network. HASQI Indicates the HASQI label calculated by HASQI;
[0039] The total loss Backpropagation to the metric discriminator to update the parameters of the metric discriminator;
[0040] S43: Training the compensation generator:
[0041] In this step, the metric discriminator is frozen and its parameters are not updated. The two inputs of the metric discriminator are set to the clean compensated speech and the enhanced training speech, respectively. The prediction score output by the metric discriminator is compared with the all-one vector to calculate the loss. Perceptual quality-related losses are also set to enhance the model's noise suppression, including PMSQE loss and PASE loss. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech and the enhanced training speech are superimposed with the hearing loss spectrum, and the superimposed features are input to the metric discriminator.
[0042] The total loss of compensating the generator in this step is described as follows:
[0043]
[0044] Among them, PMSQE is calculated based on the power spectrum of the clean compensated speech and the enhanced training speech, while PASE uses a pre-trained speech feature encoder to measure the feature distance between the enhanced training speech and the clean compensated speech. α and β are the weight coefficients of the perceptual quality related loss;
[0045] The total loss Back propagation to the compensation generator to update the parameters of the compensation generator;
[0046] S44: Repeat steps S41 to S43 to alternately train the compensation generator and the metric discriminator until a preset number of training rounds is reached or the total loss of the compensation generator is less than a preset value, and the training ends.
[0047] Preferably, α is set to 1 and β is set to 0.25.
[0048] Preferably, step S5 includes:
[0049] First, the complex spectrogram features of the speech under test are multiplied by the output of the compensation generator to obtain the complex spectrogram features of the enhanced speech under test. Then, the complex spectrogram features of the enhanced speech under test are restored to the time domain waveform of the enhanced speech under test through inverse Fourier transform. Finally, the enhanced speech waveform of the speech under test is synthesized through an overlap-add algorithm. The overlap-add algorithm refers to the superposition of overlapping parts between frames of the time domain waveform of the enhanced speech under test.
[0050] Beneficial effects: Compared with the prior art, the present invention has the following significant beneficial effects:
[0051] The speech enhancement method combining hearing loss compensation and speech noise reduction described in the present invention simultaneously achieves noise suppression and audiogram-specific spectral compensation through a single network:
[0052] First, an audiogram embedding method is introduced, enabling the network to learn the correspondence between audiograms and frequency bins. Second, adversarial training of a metric discriminator and a compensation generator is employed. The metric discriminator guides the compensation generator to produce compensated speech with improved perceptual quality. Finally, an additional compensation loss is set to further enhance the model's noise suppression. This invention utilizes a metric generative adversarial network to simultaneously achieve noise reduction and hearing loss compensation for specific audiograms. This method can stably and effectively improve the effectiveness of hearing loss compensation in noisy environments. The method is ingenious and novel, with promising application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flow chart of a speech enhancement method combining hearing loss compensation and speech noise reduction according to the present invention;
[0054] Figure 2 It is a structural diagram of the compensation generator in the present invention;
[0055] Figure 3 Schematic diagram of alternating training of the compensation generator and metric discriminator of the present invention. DETAILED DESCRIPTION
[0056] The present invention will be further described below with reference to the accompanying drawings.
[0057] The speech enhancement method combining hearing loss compensation and speech noise reduction described in the present invention achieves both noise suppression and spectral compensation specialized for the audiogram through a single network. First, a hearing loss map embedding method is introduced to enable the network to learn the correspondence between the hearing loss map and the frequency interval. Second, adversarial training of the metric discriminator and the compensation generator is used to guide the compensation generator to generate compensated speech with better quality perceived by the human ear. Finally, additional compensation loss is set to further improve the model's noise suppression. Figure 1 As shown, the present invention specifically includes the following steps:
[0058] Step S1: extracting complex spectrogram features from the noisy training speech through short-time Fourier transform.
[0059] Step S2: Extend and embed the hearing loss map along the frequency axis to obtain a hearing loss spectrum, and superimpose the hearing loss spectrum with the complex spectrum features of the noisy training speech obtained in step S1.
[0060] A hearing loss chart is a hearing curve drawn based on the subject's hearing threshold during pure tone audiometry. The horizontal axis usually represents frequency and the vertical axis represents hearing level (dB).
[0061] Specifically, the hearing loss map is first subjected to a 512-point Fourier transform at 16kHz. The map is then upsampled at frequencies of [250, 500, 1000, 2000, 4000, 8000] Hz to obtain a 6-dimensional hearing loss map gain. Each dimension of the hearing loss map gain is normalized and used as the gain value for each frequency interval. Finally, a multi-dimensional (257-dimensional) gain-embedded hearing loss spectrum is formed. The corresponding relationship is shown in Table 1 below:
[0062] Table 1
[0063]
[0064] Among them, the normalized gain value of the hearing loss graph sampled at 250Hz is HL(1), and HL(1) is used as the gain in the frequency range of 0 to 250Hz; the normalized gain value of the hearing loss graph sampled at 500Hz is HL(2), and HL(2) is used as the gain in the frequency range of 250Hz to 500Hz (excluding 250Hz); the normalized gain value of the hearing loss graph sampled at 1000Hz is HL(3), and HL(3) is used as the gain in the frequency range of 500Hz to 1000Hz (excluding 500Hz); the normalized gain value of the hearing loss graph sampled at 2000Hz is HL(4), and HL(4) is used as the gain in the frequency range of 500Hz to 1000Hz (excluding 500Hz); The normalized gain value of the hearing loss graph obtained by sampling is HL(4), and HL(4) is used as the gain in the frequency range of 1000Hz to 2000Hz (excluding 1000Hz); the normalized gain value of the hearing loss graph obtained by sampling at a frequency of 4000Hz is HL(5), and HL(5) is used as the gain in the frequency range of 2000Hz to 4000Hz (excluding 2000Hz); the normalized gain value of the hearing loss graph obtained by sampling at a frequency of 8000Hz is HL(6), and HL(6) is used as the gain in the frequency range of 4000Hz to 8000Hz (excluding 4000Hz). Convert the frequency range of 0 to 8000 Hz to the frequency points of 1 to 257. Then: the frequency range corresponding to the frequency points of 1 to 8 is 0 to 250 Hz, and the gain is HL (1); the frequency range corresponding to the frequency points of 10 to 17 is 250 Hz to 500 Hz (excluding 250 Hz), and the gain is HL (2); the frequency range corresponding to the frequency points of 18 to 33 is 500 Hz to 1000 Hz (excluding 500 Hz), and the gain is HL (3); the frequency range corresponding to the frequency points of 34 to 65 is 1000 Hz to 2000 Hz (excluding 1000 Hz), and the gain is HL (4); the frequency range corresponding to the frequency points of 66 to 129 is 2000 Hz to 4000 Hz (excluding 2000 Hz), and the gain is HL (5); the frequency range corresponding to the frequency points of 130 to 257 is 4000 Hz to 8000 Hz (excluding 4000 Hz), and the gain is HL (6).
[0065] The hearing loss spectrum is then superimposed with the complex spectrogram features of the noisy training speech to achieve alignment in frequency bins.
[0066] Step S3: Construct a metric generative adversarial network model based on frequency-time convolutional recursion. The main structure of the model includes a compensation generator and a metric discriminator. The compensation generator includes a convolutional encoder, a frequency-time recursive processing module, and a convolutional decoder. The specific construction steps are as follows:
[0067] Step S31: Construct a convolutional encoder for the compensation generator. The convolutional encoder uses five convolutional blocks to extract local patterns from its input features along the frequency axis and reduce the feature resolution. Each convolutional block includes a two-dimensional convolutional layer and batch normalization processing and PReLU activation function set in sequence after the convolutional layer. The convolutional layers of the first two convolutional blocks are set with a convolution step size of 2.
[0068] The input of the convolutional encoder is the input of the compensation generator.
[0069] Step S32: Construct a frequency-time recursive processing module of the compensation generator. The frequency-time recursive processing module includes a first recursive processing module and a second recursive processing module with the same structure. The first recursive processing module and the second recursive processing module are both processed along dual paths. In addition, a skip connection is set to transmit information on the two paths.
[0070] The input of the first recursive processing module is the output of the convolutional encoder. The first recursive processing module first slices its input features along the time axis and independently models the frequency axis correlation through a BLSTM (Bidirectional Long Short Term Memory) network for each frame. After the BLSTM network, a fully connected layer and a normalization layer are sequentially set for post-processing. The features obtained after post-processing are superimposed with the input features of the first recursive processing module to form the first feature.
[0071] The first feature is then sliced along the frequency axis, and the time correlation of each frequency stream is modeled using an LSTM (Long Short Term Memory) network. A fully connected layer and a normalization layer are sequentially set after the LSTM network for post-processing. The features obtained after post-processing are superimposed on the first feature to form the output of the first recursive processing module.
[0072] The input of the second recursive processing module is the output of the first recursive processing module. The second recursive processing module first slices its input features along the time axis and independently models the frequency axis correlation through the BLSTM network for each frame. After the BLSTM network, a fully connected layer and a normalization layer are sequentially set for post-processing. The features obtained after post-processing are superimposed with the input features of the second recursive processing module to form the second feature.
[0073] The second feature is then sliced along the frequency axis, and the time correlation of each frequency stream is modeled using the LSTM network. A fully connected layer and a normalization layer are sequentially set after the LSTM network for post-processing. The features obtained after post-processing are superimposed on the second feature to form the output of the second recursive processing module.
[0074] The output of the second recursive processing module is the output of the frequency-time recursive processing module.
[0075] Step S33: Construct a convolutional decoder for the compensation generator. The convolutional decoder is a mirror image of the convolutional encoder and uses five convolutional blocks to restore its input features to the original input feature size of the convolutional encoder. Each convolutional block includes a two-dimensional transposed convolutional layer and a batch normalization process and a PReLU activation function set in sequence after the transposed convolutional layer. The transposed convolutional layers of the last two convolutional blocks are set with a convolution step of 2. Preferably, the convolution kernel size and step size of the transposed convolution layers of the first convolutional block, the second convolutional block, the third convolutional block, the fourth convolutional block, and the fifth convolutional block of the convolutional decoder are the same as the convolution kernel size and step size of the convolution layers of the fifth convolutional block, the fourth convolutional block, the third convolutional block, the second convolutional block, and the first convolutional block of the convolutional encoder, respectively.
[0076] The input of the convolutional decoder includes the output of the convolutional encoder and the output of the frequency-time recursive processing module. To improve the information flow, a skip connection strategy is used between the convolutional encoder and the convolutional decoder to overlap the output of each convolutional block of the convolutional encoder with the input of each convolutional block of the convolutional decoder. Specifically:
[0077] The output of the convolutional encoder (i.e., the output of the fifth convolutional block of the convolutional encoder) and the output of the frequency-time recursive processing module are concatenated in the channel dimension and input into the convolutional decoder (i.e., the first convolutional block of the convolutional decoder);
[0078] The output of the fourth convolution block of the convolution encoder and the output of the first convolution block of the convolution decoder are concatenated in the channel dimension and input into the second convolution block of the convolution decoder;
[0079] The output of the third convolution block of the convolution encoder and the output of the second convolution block of the convolution decoder are concatenated in the channel dimension and input into the third convolution block of the convolution decoder;
[0080] The output of the second convolution block of the convolution encoder and the output of the third convolution block of the convolution decoder are concatenated in the channel dimension and input into the fourth convolution block of the convolution decoder;
[0081] The output of the first convolution block of the convolution encoder and the output of the fourth convolution block of the convolution decoder are concatenated in the channel dimension and input into the fifth convolution block of the convolution decoder.
[0082] The output of the convolutional decoder (i.e., the output of the fifth convolutional block of the convolutional decoder) is the output of the compensation generator.
[0083] The specific structural diagram of the compensation generator is as follows: Figure 2 shown.
[0084] Step S34: Construct a metric discriminator. The metric discriminator first extracts features from its input through four 2D convolutional blocks. Each 2D convolutional block consists of a 2D convolutional layer, an instance normalization layer, and a Pre-ReLU activation function. After the fourth 2D convolutional block, global average pooling is performed, and finally the output is activated through two fully connected layers and a sigmoid function.
[0085] Step S4: Alternately train the compensation generator and the metric discriminator, and optimize them through adversarial loss and compensation loss related to perceptual quality to achieve noise suppression and spectral level compensation specialized for the audiogram. For the entire noise reduction and hearing loss compensation framework, the training of the compensation generator and the metric discriminator will be performed alternately, such as Figure 3 The specific steps are described as follows:
[0086] Step S41: Input the features obtained by superimposing the hearing loss spectrum in step S2 with the complex spectrogram features of the noisy training speech into a compensation generator, and multiply the output of the compensation generator and the complex spectrogram features of the noisy training speech to obtain the inferred complex spectrogram features of the enhanced training speech.
[0087] Step S42: Train the metric discriminator:
[0088] First, the two inputs of the metric discriminator are set to the clean compensated speech and the enhanced training speech after the noisy training speech has been compensated using the fitting formula. The loss is calculated by comparing the predicted score output by the metric discriminator with the HASQI (Hearing-Aid Speech Quality Index) label obtained by calculating the HASQI between the enhanced training speech and the clean compensated speech. The loss can be the mean square error (MSE) loss. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech and the enhanced training speech are superimposed with the hearing loss spectrum, and the superimposed features are input to the metric discriminator.
[0089] In addition, both inputs of the metric discriminator are set to the clean compensated speech after the noisy training speech has been compensated by the fitting formula. The target at this point is set to an all-ones vector as the ideal upper limit of the estimate. The loss is calculated by combining the predicted score output by the metric discriminator with the all-ones vector. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech are superimposed with the hearing loss spectrum, and the superimposed features are then input into the metric discriminator.
[0090] The total loss of the discriminator in this step is measured The description is as follows:
[0091]
[0092] in, and HL are the clean compensated speech, enhanced training speech and corresponding hearing loss map after the noisy training speech sampled in each batch is compensated by the fitting formula, D represents the metric discriminator network, M HASQI Indicates the HASQI label calculated by HASQI.
[0093] The total loss Backpropagate to the metric discriminator and update the parameters of the metric discriminator.
[0094] Step S43: Training the compensation generator:
[0095] In this step, the metric discriminator is frozen, and no parameters are updated. Instead, the two inputs to the metric discriminator are set to the clean compensated speech and the enhanced training speech, respectively. The predicted scores output by the metric discriminator are optimized toward an all-ones vector. This involves calculating a loss between the predicted scores and the all-ones vector. Additionally, perceptual quality-related losses, including PMSQE and PASE, are applied to enhance the model's noise suppression. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech and the enhanced training speech are superimposed with the hearing loss spectrum, and the superimposed features are then fed into the metric discriminator.
[0096] The total loss of compensating the generator in this step is described as follows:
[0097]
[0098] Among them, PMSQE (PESQ-inspired loss function) is designed based on the perceptual evaluation of the speech quality (PESQ), taking into account human auditory masking and threshold effects, and is calculated based on the power spectrum of the clean compensated speech and the enhanced training speech. PASE (problem-agnostics speech encoder) uses a pre-trained speech feature encoder to measure the feature distance between the enhanced training speech and the clean compensated speech. α and β are the weight coefficients of the two perceptual quality-related losses, set to 1 and 0.25 respectively.
[0099] The total loss Backpropagate to the compensation generator and update the parameters of the compensation generator.
[0100] Step S44: Repeat steps S41 to S43 to alternately train the compensation generator and the metric discriminator until a preset number of training rounds is reached or the total loss of the compensation generator is less than a preset value, and the training ends.
[0101] Step S5: Superimpose the complex spectrogram features of the speech to be tested with the hearing loss spectrum and input them into the trained compensation generator to reconstruct the enhanced speech waveform. Specifically:
[0102] First, the ideal complex-valued mask output by the compensation generator is multiplied by the complex spectrogram features of the test speech to obtain the complex spectrogram features of the enhanced test speech. Then, the complex spectrogram features of the enhanced test speech are restored to the time-domain waveform of the enhanced test speech through an inverse Fourier transform. Finally, the enhanced speech waveform of the test speech is synthesized using an overlap-and-add algorithm. The overlap-and-add algorithm superimposes the frames of the time-domain waveform of the enhanced test speech. Because the short-time Fourier transform calculation requires the speech to be framed and windowed, there is overlap between the frames. Therefore, after the inverse Fourier transform, the frames must be overlapped and added to restore the enhanced speech waveform of the test speech to remove the effects of the window.
[0103] In summary, the speech enhancement method of the present invention, which combines hearing loss compensation and speech noise reduction, simultaneously completes noise reduction and hearing loss compensation for a specific audiogram by measuring the generative adversarial network. It can stably and effectively improve the effect of hearing loss compensation in noisy environments. The method is ingenious and novel and has good application prospects.
[0104] In order to fully compare the compensation effect of the method described in the present invention, a comparative experiment was conducted on the public dataset DNS. The DNS dataset contains 500 hours of clean corpus from 2,150 speakers and 65,000 noise clips totaling about 180 hours. More than 60,000 clean corpus were randomly divided into training set and validation set with a ratio of 60:1. Noisy speech was generated by randomly selecting segments from clean corpus and noise clips and mixing them at a random SNR (Signal-Noise Ratio) between -5dB and 15dB. During training, 3-second pronunciation segments were randomly clipped as input signals. The performance indicators corresponding to the test set are shown in Table 2, where the comparison algorithms include the DCCRN (Deep Complex Convolution Recurrent Network) model that won the championship on the DNS dataset, the FTCRN model that does not use a metric discriminator, and the FTCRN-MGAN model described in the present invention that introduces a metric discriminator and a compensation generator. From the perspective of performance indicators, compared with using the fitting formula to compensate for noisy speech, the model proposed in this invention has great advantages in all indicators. The HASQI index, WB-PESQ index and STOI index are improved by 0.2089, 1.094 and 0.0756 respectively.
[0105] Table 2
[0106]
[0107] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A speech enhancement method combining hearing loss compensation and speech noise reduction, characterized in that: The steps include: S1: Extract complex spectrogram features from noisy training speech using short-time Fourier transform; S2: Extending and embedding the hearing loss map along the frequency axis to obtain a hearing loss spectrum, and superimposing the hearing loss spectrum with the complex spectrum features of the noisy training speech obtained in step S1; S3: Construct a metric generative adversarial network model based on frequency-time convolutional recursion. The main structure of the model includes a compensation generator and a metric discriminator. S4: Alternately train the compensation generator and the metric discriminator, and optimize the metric generative adversarial network model through adversarial loss and compensation loss related to perceptual quality; S5: The complex spectrogram features of the speech to be tested are superimposed with the hearing loss spectrum and then input into the trained compensation generator, and the enhanced speech waveform of the speech to be tested is reconstructed according to the output of the compensation generator.
2. The method for speech enhancement combining hearing loss compensation and speech noise reduction according to claim 1, characterized in that: The step S2 comprises: First, the hearing loss map is Fourier transformed and then sampled at the frequencies of [250, 500, 1000, 2000, 4000, 8000] Hz to obtain a 6-dimensional hearing loss map gain. For each dimension of the hearing loss map gain, it is normalized and used as the gain value for each frequency interval, finally forming a hearing loss spectrum. The hearing loss spectrum is then superimposed with the complex spectrogram features of the noisy training speech to achieve alignment in frequency bins.
3. The method for speech enhancement combining hearing loss compensation and speech noise reduction according to claim 1, characterized in that: The compensation generator includes a convolution encoder, a frequency-time recursive processing module and a convolution decoder; The step S3 comprises: S31: Construct a convolutional encoder for the compensation generator. The convolutional encoder uses five convolutional blocks to extract local patterns and reduce the feature resolution of its input features along the frequency axis. Each convolutional block consists of a 2D convolutional layer followed by batch normalization and a PReLU activation function. The input of the convolutional encoder is the input of the compensation generator; S32: constructing a frequency-time recursive processing module of the compensation generator, the frequency-time recursive processing module including a first recursive processing module and a second recursive processing module having the same structure; The input of the first recursive processing module is the output of the convolutional encoder. The first recursive processing module first slices its input features along the time axis and independently models the frequency axis correlation through the BLSTM network for each frame. After the BLSTM network, a fully connected layer and a normalization layer are sequentially set for post-processing. The features obtained after post-processing are superimposed with the input features of the first recursive processing module to form the first feature; The first feature is then sliced along the frequency axis, and the time correlation of each frequency stream is modeled using an LSTM network. A fully connected layer and a normalization layer are sequentially set after the LSTM network for post-processing. The features obtained after post-processing are superimposed with the first feature to form the output of the first recursive processing module. The input of the second recursive processing module is the output of the first recursive processing module. The second recursive processing module first slices its input features along the time axis and independently models the frequency axis correlation through the BLSTM network for each frame. After the BLSTM network, a fully connected layer and a normalization layer are sequentially set for post-processing. The features obtained after post-processing are superimposed with the input features of the second recursive processing module to form the second feature. The second feature is then sliced along the frequency axis, and the time correlation of each frequency stream is modeled using an LSTM network. A fully connected layer and a normalization layer are sequentially set after the LSTM network for post-processing. The features obtained after post-processing are superimposed with the second feature to form the output of the second recursive processing module. The output of the second recursive processing module is the output of the frequency-time recursive processing module; S33: Construct a convolutional decoder for the compensation generator. The convolutional decoder is a mirror image of the convolutional encoder and uses five convolutional blocks to restore its input features to the input feature size of the convolutional encoder. Each convolutional block includes a 2D transposed convolution layer followed by batch normalization and a PReLU activation function. The output of the convolution encoder and the output of the frequency-time recursive processing module are spliced in the channel dimension and input into the convolution decoder; the output of the fourth convolution block of the convolution encoder and the output of the first convolution block of the convolution decoder are spliced in the channel dimension and input into the second convolution block of the convolution decoder; the output of the third convolution block of the convolution encoder and the output of the second convolution block of the convolution decoder are spliced in the channel dimension and input into the third convolution block of the convolution decoder; the output of the second convolution block of the convolution encoder and the output of the third convolution block of the convolution decoder are spliced in the channel dimension and input into the fourth convolution block of the convolution decoder; the output of the first convolution block of the convolution encoder and the output of the fourth convolution block of the convolution decoder are spliced in the channel dimension and input into the fifth convolution block of the convolution decoder; The output of the convolutional decoder is the output of the compensation generator; S34: Construct a metric discriminator. First, extract the input features through four two-dimensional convolution blocks. Each two-dimensional convolution block consists of a two-dimensional convolution layer, an instance normalization layer, and a PReLU activation function connected in sequence. After the fourth two-dimensional convolution block, a global average pooling operation is performed. Finally, the output is activated by two fully connected layers and a sigmoid function.
4. The method for speech enhancement combining hearing loss compensation and speech noise reduction according to claim 3, characterized in that: The convolution kernel size and stride of the transposed convolution layer of the first convolution block, the second convolution block, the third convolution block, the fourth convolution block, and the fifth convolution block of the convolution decoder are the same as the convolution kernel size and stride of the convolution layer of the fifth convolution block, the fourth convolution block, the third convolution block, the second convolution block, and the first convolution block of the convolution encoder, respectively.
5. The method for speech enhancement combining hearing loss compensation and speech noise reduction according to claim 1, characterized in that: The step S4 comprises: S41: Inputting the features obtained by superimposing the hearing loss spectrum in step S2 with the complex spectrogram features of the noisy training speech into a compensation generator, and multiplying the output of the compensation generator by the complex spectrogram features of the noisy training speech to obtain the complex spectrogram features of the enhanced training speech; S42: Train the metric discriminator: First, the two inputs of the metric discriminator are set to the clean compensated speech and the enhanced training speech after the noisy training speech is compensated by the fitting formula. The loss is calculated by comparing the prediction score output by the metric discriminator with the HASQI label obtained by the HASQI calculation between the enhanced training speech and the clean compensated speech. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech and the enhanced training speech are superimposed with the hearing loss spectrum, and the superimposed features are input into the metric discriminator. In addition, both inputs of the metric discriminator are set to clean compensated speech, and the prediction score output by the metric discriminator is compared with the all-one vector to calculate the loss. Before the clean compensated speech is input into the metric discriminator, its complex spectrogram features are superimposed with the hearing loss spectrum, and the superimposed features are input into the metric discriminator. The total loss of the discriminator in this step is measured The description is as follows: , in, 、 and are clean compensated speech, enhanced training speech and corresponding hearing loss map, D represents the metric discriminator network, Indicates the HASQI label calculated by HASQI; The total loss Backpropagation to the metric discriminator to update the parameters of the metric discriminator; S43: Training the compensation generator: In this step, the metric discriminator is frozen and its parameters are not updated. The two inputs of the metric discriminator are set to the clean compensated speech and the enhanced training speech, respectively. The prediction score output by the metric discriminator is compared with the all-one vector to calculate the loss. Perceptual quality-related losses are also set to enhance the model's noise suppression, including PMSQE loss and PASE loss. Before entering the metric discriminator, the complex spectrogram features of the clean compensated speech and the enhanced training speech are superimposed with the hearing loss spectrum, and the superimposed features are input to the metric discriminator. The total loss of compensating the generator in this step is described as follows: , Among them, PMSQE is calculated based on the power spectrum of the clean compensated speech and the enhanced training speech, while PASE uses a pre-trained speech feature encoder to measure the feature distance between the enhanced training speech and the clean compensated speech. and is the weight coefficient of the perceived quality-related loss; The total loss Back propagation to the compensation generator to update the parameters of the compensation generator; S44: Repeat steps S41 to S43 to alternately train the compensation generator and the metric discriminator until a preset number of training rounds is reached or the total loss of the compensation generator is less than a preset value, and the training ends.
6. The method for speech enhancement combining hearing loss compensation and speech noise reduction according to claim 5, characterized in that: Set to 1, Set to 0.
25.
7. The method for speech enhancement combining hearing loss compensation and speech noise reduction according to claim 1, characterized in that: The step S5 comprises: First, the complex spectrogram features of the speech under test are multiplied by the output of the compensation generator to obtain the complex spectrogram features of the enhanced speech under test. Then, the complex spectrogram features of the enhanced speech under test are restored to the time domain waveform of the enhanced speech under test through inverse Fourier transform. Finally, the enhanced speech waveform of the speech under test is synthesized through an overlap-add algorithm. The overlap-add algorithm refers to the superposition of overlapping parts between frames of the time domain waveform of the enhanced speech under test.
Citation Information
Patent Citations
Method for reducing noise by using hearing threshold of impaired hearing
CN101901602A
Speech enhancement hearing aid method
CN109147808A