A speech enhancement method and system based on neural homomorphic synthesis and phase estimation
By combining neural homomorphic synthesis and phase estimation with multi-decoder neural networks and phase discriminators, the problem of insufficient or excessive noise suppression in existing speech enhancement methods under low signal-to-noise ratio environments is solved, achieving higher quality speech enhancement results.
Patent Information
- Application Number
- CN202410425822.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-04-10
AI Technical Summary
Existing speech enhancement methods are prone to insufficient or excessive noise suppression in low signal-to-noise ratio environments, ignoring phase information and resulting in decreased perception quality and intelligibility.
A method based on neural homomorphic synthesis and phase estimation is adopted. The decomposed speech components and phase spectrum are predicted by a multi-decoder neural network structure. The excitation and vocal tract are accurately separated by a neural homomorphic filter. The phase estimation is optimized by a phase discriminator. The speech enhancement effect is optimized by adversarial neural network and multi-resolution loss function.
It achieves higher quality speech enhancement in low signal-to-noise ratio environments, improves perception quality and intelligibility, accurately separates excitation and vocal tract, and reduces problems of insufficient or excessive noise suppression.
Smart Images

Figure CN118366469B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech enhancement technology, and in particular to a speech enhancement method and system based on neural homomorphic synthesis and phase estimation. Background Technology
[0002] In daily life, voice is one of the most frequently encountered and used types of information. Due to noisy environments or various factors affecting network transmission, voice signals inevitably contain noise, such as car horns on the road or static caused by network signal problems during transmission. With the rapid development of intelligent networks and information transmission, and the increasing penetration of artificial intelligence technology into people's lives, the accuracy of voice information transmission is becoming increasingly important for daily life, making the optimization and improvement of voice enhancement models increasingly crucial.
[0003] Speech enhancement refers to reducing background noise in noisy speech to improve speech quality and intelligibility. Due to the diversity of environments, it is one of the most challenging tasks in the field of signal processing. This task has profound value for applications such as speech recognition, communication, and hearing aids. Classical statistical methods have been extensively studied for decades. In recent years, with the significant improvement in computing power, deep learning has become one of the most important application trends in various information processing fields. Speech enhancement has gradually been widely incorporated into supervised learning problems in the field of machine learning, and combining deep neural networks (DNNs) for speech enhancement has become the mainstream method in this research field.
[0004] Thanks to significant advancements in deep neural networks, data-driven speech enhancement methods have garnered considerable attention and achieved rapid progress. Generally, existing data-driven DNN-based speech enhancement methods can be categorized into two types: 1. Time-domain methods: using the original speech waveform as both input and output of the neural network. 2. Frequency-domain (TF) methods: using the original speech waveform after undergoing a short-time Fourier transform (STFT) as both input and output of the neural network.
[0005] Researchers have developed a speech enhancement algorithm using neural homomorphic synthesis (NHS-SE). This algorithm uses a neural homomorphic signal processing module to process noisy, framed speech to obtain its excitation and vocal tracts. Then, two complex convolutional recurrent networks are used to estimate the clean spectra of the excitation and vocal tracts, respectively. Finally, the enhanced speech is synthesized using the denoised components. However, the existing NHS-SE technique employs traditional homomorphic filtering signal processing methods, utilizing simple homomorphic filters to obtain the excitation and vocal tracts. Simple homomorphic filters are not suitable for practical applications.
[0006] In summary, existing speech enhancement methods often rely on amplitude-based processing techniques, neglecting the importance of phase information. This leads to a decrease in perceptual quality and intelligibility, even though methods with complex spectrograms indirectly enhance the phase. Furthermore, mainstream deep learning-based speech enhancement methods are prone to undersuppression or oversuppression of noise, especially under conditions of very low signal-to-noise ratio.
[0007] Therefore, in view of the shortcomings of the existing technology, we propose a technical solution to address the deficiencies and inadequacies of the existing technology. Summary of the Invention
[0008] In view of this, it is indeed necessary to provide a speech enhancement method and system based on neural homomorphic synthesis and phase estimation. This method uses neural homomorphic synthesis and phase estimation for speech enhancement and employs a multi-decoder neural network structure to simultaneously predict the decomposed speech components and phase spectrum. It not only achieves amplitude and phase prediction but also focuses on phase optimization within the complex spectrogram, solving the problems of excessive speech suppression and insufficient noise suppression in low signal-to-noise ratio environments. Furthermore, it employs a neural homomorphic filter to achieve more accurate separation of the excitation and vocal tract, thereby addressing the technical problem of homomorphic filters not matching reality.
[0009] In order to solve the technical problems existing in the prior art, the technical solution of the present invention is as follows:
[0010] A speech enhancement method based on neural homomorphic synthesis and phase estimation includes the following steps:
[0011] Step S1: Construct a homomorphic filtering module to receive noisy speech signals, process the signals, and output noisy speech features. The noisy speech features include at least phase information, excitation information, and vocal tract information.
[0012] Step S2: Construct an enhancement module to receive noisy speech features, perform signal processing, and output enhanced phase information, excitation information, and vocal tract information. The enhancement module includes a phase estimation module, a first neural network cepstral inverse system module, and a second neural network cepstral inverse system module. The phase estimation module processes the received phase information to output enhanced phase information; the first neural network cepstral inverse system module processes the received excitation information to output enhanced excitation information; and the second neural network cepstral inverse system module processes the received vocal tract information to output enhanced vocal tract information.
[0013] Step S3: Construct a post-processing module to synthesize the enhanced phase information, excitation information, and vocal tract information, and output the enhanced speech signal.
[0014] As a further improvement, in step S1, the homomorphic filtering module consists of a cepstral processing module and a neural network homomorphic filter. The cepstral processing module is used to perform signal processing on the input noisy speech to extract phase information and cepstral information.
[0015] As a further improvement, the cepstral processing module includes at least discrete-time Fourier transform processing, logarithmic amplitude processing, and inverse discrete-time Fourier transform processing. That is, after the speech signal x(n) is processed by the cepstral processing module, the cepstral information of the speech is represented as follows: Where FFT represents Discrete-Time Fourier Transform, ln represents Logarithmic Amplitude Transform, and IFFT represents Inverse Discrete-Time Fourier Transform.
[0016] As a further improvement, a neural network homomorphic filter is used to filter the speech cepstral information processed by the cepstral processing module and output excitation information and vocal tract information. It is implemented using a deep neural network and specifically includes the following steps:
[0017] S11: Construct and train a deep neural network to implement a neural network homomorphic filter l(n);
[0018] S12: The trained neural network homomorphic filter is used to process the speech processed by the cepstral processing module to achieve the separation of excitation and vocal tract.
[0019] As a further improvement, in step S2, the enhancement module is based on a neural network implementation, specifically including the following steps:
[0020] S21: Construct and train an enhanced neural network model;
[0021] S22: Use the trained enhanced neural network model to process the input noisy speech features and output the enhanced phase information, excitation information and vocal tract information.
[0022] As a further improvement, the augmented neural network model employs an adversarial neural network, with the following loss function:
[0023]
[0024] Where, α PD α Metric To adjust the parameters, For multi-resolution loss, Phase generator training loss, The metric generator measures the loss.
[0025] As a further improvement plan, The multi-resolution loss is as follows:
[0026]
[0027] Where y and Let _____ represent the true value and the estimated value of the time-domain speech, respectively; let ‖·‖1 represent the L1 norm; and let ∑ represent the summation. Due to resolution amplitude loss Complex spectral loss Anti-winding loss The following is a representation composed of different parameter weights. Set the values to 0.5, 0.1, and 0.3 respectively.
[0028]
[0029] This indicates a loss in resolution amplitude. Complex spectral loss Anti-winding loss as follows:
[0030]
[0031]
[0032]
[0033] Among them, Y i and These are time-domain speech y and The i-th resolution STFT, ‖·‖ F And ||·||1 are the Frobenius norm and L1 norm Y. pi and Represents the phase spectrum. To express the expectation, Δ DF and Δ DT Represents the derivative along the frequency axis and the time axis, and the dewinding function f. AW as follows:
[0034]
[0035] The training loss of the phase generator is explained as follows:
[0036]
[0037] in, To combat the losses, f D f represents the phase discriminator. θ For phase generator, Y in This refers to the compressed speech feature information;
[0038] The metric generator's metric loss is explained as follows:
[0039]
[0040] Where ||·||² represents the L2 norm, f MD Represents the metric discriminator, These are time-domain speech y and The amplitude information.
[0041] This invention also discloses a speech enhancement system based on neural homomorphic synthesis and phase estimation, comprising at least a homomorphic filtering module, an enhancement construction module, and a post-processing module, wherein...
[0042] The homomorphic filtering module is used to receive noisy speech signals, process the signals, and output noisy speech features, wherein the noisy speech features include at least phase information, excitation information, and vocal tract information.
[0043] An enhancement module is used to receive noisy speech features, perform signal processing, and output enhanced phase information, excitation information, and vocal tract information. The enhancement module includes a phase estimation module, a first neural network cepstral inverse system module, and a second neural network cepstral inverse system module. The phase estimation module processes the received phase information to output enhanced phase information; the first neural network cepstral inverse system module processes the received excitation information to output enhanced excitation information; and the second neural network cepstral inverse system module processes the received vocal tract information to output enhanced vocal tract information.
[0044] The post-processing module synthesizes the enhanced phase information, excitation information, and vocal tract information to output the enhanced speech signal.
[0045] As a further improvement, the homomorphic filtering module consists of a cepstral processing module and a neural network homomorphic filter. The cepstral processing module is used to perform signal processing on the input noisy speech to extract phase information and cepstral information. The signal processing in the cepstral processing module includes at least discrete-time Fourier transform processing, logarithmic amplitude processing, and inverse discrete-time Fourier transform processing.
[0046] As a further improvement, the enhancement module is implemented using a neural network model.
[0047] Compared with existing technologies, this invention proposes a speech enhancement algorithm that combines neural homomorphic synthesis and phase estimation, which has competitive performance and enhancement compared with previous methods using neural homomorphic synthesis and other methods, and has at least the following technical effects:
[0048] 1. This invention implements a neural network homomorphic filter to achieve more accurate separation of the excitation and the audio channel; while in the prior art, most of the existing technologies use traditional homomorphic filters for filtering in the separation of the audio channel and the excitation.
[0049] 2. Existing methods for speech enhancement using neural networks (DNNs) typically rely on amplitude-based processing techniques, often neglecting the importance of phase information, which leads to a decrease in perceptual quality and intelligibility. This invention specifically sets up a phase estimation module for phase information, utilizing complex spectral loss and anti-winding loss to enhance phase recovery capability.
[0050] 3. This invention uses a phase discriminator to minimize the Wasserstein distance between the generated speech phase direction vector and the clean speech phase direction vector, thereby achieving higher quality phase estimation and obtaining better noise reduction effect. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating the steps of a speech enhancement method based on neural homomorphic synthesis and phase estimation proposed in this invention.
[0052] Figure 2 This is a system architecture diagram of a speech enhancement system based on neural homomorphic synthesis and phase estimation proposed in this invention.
[0053] Figure 3 This is a block diagram of the neural network model structure used to implement the enhancement module in this invention.
[0054] The following specific embodiments will further illustrate the present invention in conjunction with the above-described accompanying drawings. Detailed Implementation
[0055] The technical solution provided by the present invention will be further described below with reference to the accompanying drawings.
[0056] Most existing speech enhancement methods based on deep neural networks typically operate in the frequency domain, using neural networks to directly estimate a clean speech spectrogram or predict a real or complex mask representing the target spectrogram. However, these methods often rely on amplitude-based processing techniques, neglecting the importance of phase information. This can easily lead to under-suppression or over-suppression of noise, resulting in a decrease in perceptual quality and intelligibility. Even methods with complex spectrograms inevitably enhance the phase indirectly.
[0057] To address the technical shortcomings of existing technologies, this invention provides a speech enhancement method based on neural homomorphic synthesis and phase estimation. (See [link to relevant documentation]). Figure 1 and Figure 2 The diagram shows the flowchart and system structure of the method of the present invention, which includes the following steps:
[0058] Step S1: Construct a homomorphic filtering module to receive noisy speech signals, process the signals, and output noisy speech features. The noisy speech features include at least phase information, excitation information, and vocal tract information.
[0059] Step S2: Construct an enhancement module to receive noisy speech features, perform signal processing, and output enhanced phase information, excitation information, and vocal tract information. The enhancement module includes a phase estimation module, a first neural network cepstral inverse system module, and a second neural network cepstral inverse system module. The phase estimation module processes the received phase information to output enhanced phase information; the first neural network cepstral inverse system module processes the received excitation information to output enhanced excitation information; and the second neural network cepstral inverse system module processes the received vocal tract information to output enhanced vocal tract information.
[0060] Step S3: Construct a post-processing module to synthesize the enhanced phase information, excitation information, and vocal tract information, and output the enhanced speech signal.
[0061] In step S1, the homomorphic filtering module consists of a cepstral processing module and a neural network homomorphic filter. The cepstral processing module is used to process the input noisy speech signal to extract phase and cepstral information. The signal processing in the cepstral processing module includes at least discrete-time Fourier transform processing, logarithmic amplitude processing, and inverse discrete-time Fourier transform processing. That is, after the speech signal x(n) undergoes the above series of operations, the cepstral information of the processed speech is represented as follows: Where FFT represents Discrete-Time Fourier Transform, ln represents Logarithmic Amplitude Transform, and IFFT represents Inverse Discrete-Time Fourier Transform.
[0062] The neural network homomorphic filter is used to filter the speech cepstral information processed by the cepstral processing module and output excitation and vocal tract information. It is implemented using a deep neural network and includes the following steps:
[0063] S11: Construct and train a deep neural network (DNN) to implement the neural network homomorphic filter l(n);
[0064] S12: The trained neural network homomorphic filter is used to process the speech after it has been processed by the cepstral processing module to achieve the separation of excitation and vocal tract;
[0065] Step S11 can be summarized as follows:
[0066] S111: Obtain a large training set of paired noisy and clean speech samples, and preprocess the data. Specifically, a noisy speech dataset is simulated using a hybrid dataset of the multilingual VoiceBank dataset and the noisy dataset DEMAND. The clean speech dataset from VoiceBank is used, with noisy speech samples paired with corresponding clean speech samples to form the training set and the set of speech samples to be denoised. All speech samples in the training set are downsampled to 16kHz.
[0067] S112: Use the preprocessed noisy speech dataset as the feature input for the cepstral processing module to extract speech signal;
[0068] S113: Train a neural network homomorphic filter using a deep neural network (DNN) to process the features.
[0069] Step S113 further includes:
[0070] S1131: Use the speech features in S112 as input to the input layer of a deep neural network (DNN);
[0071] S1132: The input layer data is estimated through two long short-term memory (LSTM) layers in two frequency dimensions and compressed into the output layer through a linear layer in one recovery compression dimension;
[0072] S1133: Pass the output layer data through a sigmoid function y to obtain an output in the range of 0 to 1.
[0073]
[0074] A well-trained neural network homomorphic filter can be obtained through multiple iterations of training. in f represents the speech signal after the cepstral processing module. l This indicates that it has been processed through a deep neural network.
[0075] Step S122, following training in step S11, involves separating the input speech into the excitation and vocal tract components:
[0076]
[0077]
[0078] Where l(n) represents the filter, This represents the speech signal after passing through the cepstral processing module.
[0079] In step S2, the enhancement module is implemented based on a neural network; see [link / reference]. Figure 3 The diagram shows the neural network model structure used to implement the enhancement module, which includes the following steps:
[0080] S21: Construct and train an enhanced neural network model;
[0081] S22: Use the trained augmented neural network model to process the input noisy speech features and output the enhanced phase information, excitation information and vocal tract information;
[0082] In step S21, the steps for training the augmented neural network model are as follows:
[0083] S211: Extract features from the training set to obtain a training set consisting of noisy speech and corresponding clean speech. The processing method for this step is the same as that for step S111.
[0084] S212: The training set speech is separated and processed by the homomorphic filtering module, and then feature compression is performed, including the vocal tract spectrogram. Excitation spectrum Phase direction vector These feature information are stacked to obtain in, This indicates information containing time, frequency, and i dimensions.
[0085] S213: Stacked Information Y in The information is then compressed into a condensed time-frequency (TF) domain latent space representation using a shared encoder.
[0086] S214: The latent space representation is processed by two levels of Conformer in the successive stages to model the time and frequency dependencies;
[0087] S215: The original temporal space is reconstructed from the feature information through three parallel decoders. The phase decoder is used to output the enhanced phase information, the vocal tract decoder is used to output the enhanced vocal tract information, and the excitation decoder is used to output the enhanced excitation information. At the same time, the output mask is made using the noisy speech features.
[0088] S216: Train the metric discriminator and phase discriminator to optimize phase estimation and speech estimation.
[0089] S217: Optimize the loss of multi-resolution short-time Fourier transform on the enhanced speech;
[0090] S218: The amplitude spectrum mask for the predicted vocal tract and excitation is obtained, where, and and phase shift
[0091] In step S213 above, the shared encoder consists of three 2D convolutional (Conv2D) blocks, responsible for encoding the input features into a latent space representation. Each Conv2D block contains a Conv2D layer, a normalization layer, and a PReLu activation function f(y). i Each convolutional block is activated by the PReLu activation function, and the input noisy channel spectrum, activation spectrum, and phase spectrum are used as features respectively. The feature dimension is increased by increasing the number of channels in the convolutional layer. Each 2D convolutional block in the encoder contains 3 convolutional layers with the number of channels set to 32, 64, and 128 respectively.
[0092]
[0093] In step S215 above, each decoder consists of three 2D deconvolution (ConTrans2D) blocks, responsible for reconstructing the original space from the latent space representation. Each deconvolution block contains a ConTrans2D layer, a normalization layer, and a PReLu activation function. The decoder deconvolution block contains three deconvolution layers; the number of channels in the channel decoder and excitation decoder deconvolution layers are set to 64, 32, and 1, respectively, while the number of channels in the phase decoder is 64, 32, and 2.
[0094] Step S216 follows Wasserstein GAN and Metric GAN respectively. The flowchart is shown in Table 1. The specific steps are as follows:
[0095] S2161: Training the metric discriminator. The metric discriminator determines the source of the speech by comparing the PESQ scores of the generated speech with those of clean speech.
[0096] S2162: Training the phase discriminator. The discriminator determines the source of the phase direction vector by comparing the Wasserstein distance between the generated phase spectrum direction vector and the clean phase spectrum direction vector, and optimizes the discriminator parameters accordingly.
[0097] In S2161, the system measures the discriminator training loss. and the corresponding generator loss measurement They are respectively:
[0098]
[0099]
[0100] Where ||·||2 represents the L2 norm, Q PESQ This is a perceptual evaluation indicator for speech quality. It expresses expectation.
[0101] In S2162, the losses in the system include the phase discriminator loss. Corresponding adversarial loss of the generator as follows:
[0102]
[0103]
[0104] in, To combat the losses, For gradient penalty, f represents expectation D Represents the discriminator and penalty weights. Set to 10.
[0105] Table 1 Training process for phase discriminator and metric discriminator
[0106]
[0107]
[0108] In step S217 above, the multi-resolution loss function includes: amplitude resolution loss, complex spectrum loss, and anti-winding loss, specifically as follows:
[0109] The amplitude component of the resolution STFT loss is defined as:
[0110]
[0111] Among them, Y i and These are time-domain speech y and The i-th resolution STFT, ‖·‖ F And ||·||1 are the Frobenius norm and L1 norm, respectively;
[0112] Complex spectral loss and anti-winding loss are used for the wrapping property of the phase and are defined as follows:
[0113]
[0114]
[0115] ‖·‖1 is the L1 norm, Y i and These are time-domain speech y and The i-th resolution STFT, Y pi and Represents the phase spectrum. To express the expectation, Δ DF and Δ DT Represents the derivatives along the frequency and time axes and the dewinding function.
[0116] Therefore, the resolution loss is obtained by assigning different parameters to the amplitude resolution loss, complex spectrum loss, and anti-winding loss:
[0117]
[0118] in We set the parameters to 0.5, 0.1, and 0.3 respectively, and finally added a temporal L1 speech loss to further improve speech quality. Based on the above, we obtain the final multi-resolution loss as follows:
[0119]
[0120] I represents the total number of resolutions, represents the summation, and ||·||1 represents the L1 norm.
[0121] Combining the loss functions in steps S216 and S217, we can summarize as follows:
[0122]
[0123] Where α PD α Metric Set them to 0.005 and 0.05 respectively. For multi-resolution loss, Phase generator training loss, The metric generator measures the loss.
[0124] Step S23 is described in detail below:
[0125] The noisy speech is enhanced by training modules, which includes the enhanced vocal tract. Enhanced incentives Enhanced phase
[0126]
[0127]
[0128] Step S3 will synthesize the enhanced phase information, excitation information, and vocal tract information to output the enhanced speech signal. The specific steps are as follows:
[0129] S31: Vocal Tract and Activation Convolution Synthesis Feature Part 1
[0130] S32: Feature part 1 and the fully enhanced speech after phase synthesis.
[0131] This invention also provides a speech enhancement system based on neural homomorphic synthesis and phase estimation, see [link to relevant documentation]. Figure 2The diagram shown is a system structure block diagram of the present invention, which includes at least a homomorphic filtering module, a construction and enhancement module, and a post-processing module.
[0132] The homomorphic filtering module is used to receive noisy speech signals, process the signals, and output noisy speech features, wherein the noisy speech features include at least phase information, excitation information, and vocal tract information.
[0133] An enhancement module is used to receive noisy speech features, perform signal processing, and output enhanced phase information, excitation information, and vocal tract information. The enhancement module includes a phase estimation module, a first neural network cepstral inverse system module, and a second neural network cepstral inverse system module. The phase estimation module processes the received phase information to output enhanced phase information; the first neural network cepstral inverse system module processes the received excitation information to output enhanced excitation information; and the second neural network cepstral inverse system module processes the received vocal tract information to output enhanced vocal tract information.
[0134] The post-processing module synthesizes the enhanced phase information, excitation information, and vocal tract information to output the enhanced speech signal.
[0135] In the above technical solution, the homomorphic filtering module consists of a cepstral processing module and a neural network homomorphic filter. The cepstral processing module is used to perform signal processing on the input noisy speech to extract phase information and cepstral information. The signal processing in the cepstral processing module includes at least discrete-time Fourier transform processing, logarithmic amplitude processing, and inverse discrete-time Fourier transform processing.
[0136] The enhancement module is implemented using a neural network model.
[0137] In the above technical solution of this invention, we use the VoiceBank+DEMAND dataset and the Interspeech2021DNS Challenge dataset. The VoiceBank+DEMAND dataset includes a test set and a training set. The test set contains 824 utterances from two unknown speakers, with signal-to-noise ratios (SNRs) of 2.5, 7.5, 12.5, and 17.5 dB, respectively. In the training set, noisy speech is mixed at 5 dB intervals within a SNR range of 0–15 dB. Therefore, we randomly selected 1000 utterances from the training subset as the validation dataset, and the remaining 10572 utterances as the training dataset, downsampling them to 16 kHz.
[0138] The DNS challenge dataset contains 760.53 hours of clean speech, 181 hours of balanced noise, and provides 3076 real and 115,000 synthetic room impulse responses (RIRs). We generated 300 hours of unereverberant noise-clean pairs and 200 hours of reverberant noise-clean pairs using synthesis tools, with signal-to-noise ratios (SNRs) randomly distributed between -5 and 20 dB. A total of 60,000 noise-clean pairs were synthesized, from which 1000 pairs were randomly selected for validation. The evaluation set includes 150 synthesized noise-clean pairs, divided into "reverberant" and "unereverberant" groups, with SNRs ranging from 0 to 25 dB.
[0139] These datasets were then used for training and evaluation, and the speech enhancement quality was assessed using common objective speech metrics. Evaluations on both datasets demonstrate that our proposed neural homomorphic phase speech enhancement can more accurately estimate clean speech. See Tables 3, 4, and 5. Table 1 compares the speech metrics of our invention with other speech enhancement methods on the VoiceBank+DEMAND dataset; Table 2 compares the speech metrics of our invention with other speech enhancement methods on the DNS challenge dataset; Table 3 also evaluates a neural homomorphic amplitude domain speech enhancement method (NHS-MagSE), demonstrating that phase optimization plays a crucial role in homomorphic synthesis speech enhancement.
[0140] Table 3 compares the present invention with some representative speech enhancement methods on the VoiceBank+DEMAND dataset.
[0141]
[0142] In Table 3, higher values indicate better performance, and "-" indicates that the original paper did not provide results. PESQ, STOI, CSIG, CBAK, and COVL are all common speech evaluation metrics. As shown in the table, the proposed Neural Homomorphic Phase Speech Enhancement (NHS-PE) method achieved the highest scores on the CSIG and STOI metrics on the test set. Furthermore, compared to methods using amplitude or phase as input (MetricGAN, MetricGAN+, PHASEN), the NHS-PE method achieved the highest scores.
[0143] Table 4 compares the present invention with some representative speech enhancement methods on the DNS Challenge dataset.
[0144]
[0145] Table 4 presents the results comparing our invention with other methods on the DNS Challenge. Our method demonstrates superior performance in reverberation-free denoising evaluations, showing significant advantages in both PESQ and STOI, and maintaining its advantage even with reverberation. The gap between the complete neural homomorphic synthesis speech enhancement method system and the phase detector removal system is due to the instability of generative adversarial networks (GANs).
[0146] In addition, to demonstrate the significant impact of phase optimization in neural homomorphic synthesis speech enhancement, we evaluate a neural homomorphic amplitude domain speech enhancement method (NHS-MagSE) on the VoiceBank+DEMAND dataset and assess various variations in the score of this method with a parallel phase enhancement network (PM) and a filter model (LM). The results are shown in Table 5.
[0147] Table 5. Scores of the amplitude-domain speech enhancement method based on neural homomorphism (NHS-MagSE) on the VoiceBank+DEMAND dataset.
[0148]
[0149] Table 5 shows that, under the same dataset conditions, our innovations on amplitude and phase inputs outperform the neural homomorphic amplitude domain speech enhancement method (NHSMagSE). Furthermore, when the phase discriminator is removed, only eSTOI and SI–SNR metrics increase, while the others decrease. When all discriminators are removed, the metrics decrease further. This demonstrates the crucial role of phase optimization in homomorphic synthesized speech enhancement.
[0150] In summary, this invention proposes a speech enhancement algorithm combining neural homomorphic synthesis and phase estimation. A multi-decoder network with amplitude and phase inputs is designed, utilizing phase-aware anti-winding loss and adversarial strategies, along with a phase discriminator, to achieve better phase estimation capabilities. Furthermore, the use of a neural homomorphic filter for speech separation yields improved performance. The experimental results validate the effectiveness of the phase spectrum enhancement data path and phase discriminator in the proposed method and system. The technical solution of this invention achieves comparable performance on two popular benchmarks, particularly demonstrating significant improvements to the neural homomorphic synthesis method.
[0151] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0152] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A speech enhancement method based on neural homomorphic synthesis and phase estimation, characterized in that, Includes the following steps: Step S1: Construct a homomorphic filtering module to receive noisy speech signals, process the signals, and output noisy speech features. The noisy speech features include at least phase information, excitation information, and vocal tract information. Step S2: Construct an enhancement module to receive noisy speech features, perform signal processing, and output enhanced phase information, excitation information, and vocal tract information. The enhancement module includes a phase estimation module, a first neural network cepstral inverse system module, and a second neural network cepstral inverse system module. The phase estimation module processes the received phase information to output enhanced phase information; the first neural network cepstral inverse system module processes the received excitation information to output enhanced excitation information; and the second neural network cepstral inverse system module processes the received vocal tract information to output enhanced vocal tract information. Step S3: Construct a post-processing module to synthesize the enhanced phase information, excitation information, and vocal tract information, and output the enhanced speech signal; In step S2, the enhancement module is implemented based on a neural network, specifically including the following steps: S21: Construct and train an augmented neural network model; the augmented neural network model reconstructs the original temporal space from feature information through three parallel decoders, where the phase decoder is used to output the augmented phase information, the vocal tract decoder is used to output the augmented vocal tract information, and the excitation decoder is used to output the augmented excitation information. At the same time, a mask is output using noisy speech features; simultaneously, a metric discriminator and a phase discriminator are trained to optimize speech estimation and phase estimation. S22: Use the trained augmented neural network model to process the input noisy speech features and output the enhanced phase information, excitation information and vocal tract information; The augmented neural network model employs an adversarial neural network, and its loss function is: ; in, , To adjust the parameters, For multi-resolution loss, Phase generator training loss, The metric generator measures the loss.
2. The speech enhancement method based on neural homomorphic synthesis and phase estimation according to claim 1, characterized in that, In step S1, the homomorphic filtering module consists of a cepstral processing module and a neural network homomorphic filter. The cepstral processing module is used to perform signal processing on the input noisy speech to extract phase information and cepstral information.
3. The speech enhancement method based on neural homomorphic synthesis and phase estimation according to claim 1, characterized in that, The cepstral processing module includes at least discrete-time Fourier transform processing, logarithmic amplitude processing, and inverse discrete-time Fourier transform processing; that is, processing of speech signals. After processing by the cepstral processing module, the cepstral information of the speech is represented as follows: Where FFT represents discrete-time Fourier transform processing, ln represents logarithmic magnitude processing, and IFFT represents inverse discrete-time Fourier transform processing.
4. The speech enhancement method based on neural homomorphic synthesis and phase estimation according to claim 1, characterized in that, The neural network homomorphic filter is used to filter the speech cepstral information after the cepstral processing module and output excitation and vocal tract information. It is implemented using a deep neural network and includes the following steps: S11: Construct and train a deep neural network to implement a neural network homomorphic filter ; S12: The trained neural network homomorphic filter is used to process the speech processed by the cepstral processing module to achieve the separation of excitation and vocal tract.
5. The speech enhancement method based on neural homomorphic synthesis and phase estimation according to claim 1, characterized in that, The multi-resolution loss is as follows: ; in, and These represent the true and estimated values of the time-domain speech, respectively. Describing the L1 norm, To express summation, Due to resolution amplitude loss Complex spectral loss Anti-winding loss The following is a representation composed of different parameter weights. Set the values to 0.5, 0.1, and 0.3 respectively. ; This indicates a loss in resolution amplitude. Complex spectral loss Anti-winding loss as follows: ; ; ; in, and They are time-domain speech and The i-th resolution STFT It is the Frobenius norm and the L1 norm. Represents the phase spectrum. Expressing expectations, Represents the derivatives along the frequency and time axes and the dewinding function. as follows: ; The training loss of the phase generator is explained as follows: ; in, To combat the losses, Indicates phase discriminator, For phase generator, This refers to the compressed speech feature information; The metric generator's metric loss is explained as follows: ; in Describing the L2 norm, Represents the metric discriminator, They are time-domain speech and The amplitude information.
6. A speech enhancement system based on neural homomorphic synthesis and phase estimation using any one of the methods described in claims 1-5, characterized in that, It includes at least a homomorphic filtering module, a construction and enhancement module, and a post-processing module, among which, The homomorphic filtering module is used to receive noisy speech signals, process the signals, and output noisy speech features, wherein the noisy speech features include at least phase information, excitation information, and vocal tract information. An enhancement module is used to receive noisy speech features, perform signal processing, and output enhanced phase information, excitation information, and vocal tract information. The enhancement module includes a phase estimation module, a first neural network cepstral inverse system module, and a second neural network cepstral inverse system module. The phase estimation module processes the received phase information to output enhanced phase information; the first neural network cepstral inverse system module processes the received excitation information to output enhanced excitation information; and the second neural network cepstral inverse system module processes the received vocal tract information to output enhanced vocal tract information. The post-processing module synthesizes the enhanced phase information, excitation information, and vocal tract information to output the enhanced speech signal.
7. The speech enhancement system based on neural homomorphic synthesis and phase estimation according to claim 6, characterized in that, The homomorphic filtering module consists of a cepstral processing module and a neural network homomorphic filter. The cepstral processing module is used to process the input noisy speech signal to extract phase information and cepstral information. The cepstral processing module includes at least discrete-time Fourier transform processing, logarithmic amplitude processing, and inverse discrete-time Fourier transform processing.
8. The speech enhancement system based on neural homomorphic synthesis and phase estimation according to claim 6, characterized in that, The enhancement module is implemented using a neural network model.
Citation Information
Patent Citations
Speech enhancement method, electronic equipment and storage medium
CN114678036A