Light-weight binaural speech enhancement method based on Fourier network

By combining a global adaptive Fourier modulation mechanism and a dynamic optimization gate, the problem of balancing performance and complexity in existing binaural speech enhancement models on resource-constrained devices is solved, achieving efficient speech enhancement and spatial cue preservation.

CN121148409APending Publication Date: 2025-12-16EAST CHINA NORMAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511662703.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing binaural speech enhancement models struggle to balance high performance and low complexity on resource-constrained devices, and existing methods are inadequate in modeling long-term temporal dependencies, resulting in unsatisfactory enhancement effects.

Method used

A lightweight binaural speech enhancement method is proposed, which combines a global adaptive Fourier modulation mechanism with dynamic optimization gates. By combining dual-path feature encoding and fusion, a global adaptive Fourier modulator, a binaural decoder, and dynamic optimization gates, a balance between performance and efficiency is achieved.

Benefits of technology

While maintaining high enhancement performance and spatial cues integrity, it achieves lightweight design and high computational efficiency, improves the model's robustness to complex acoustic environments, and effectively suppresses processing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148409A_ABST
    Figure CN121148409A_ABST
Patent Text Reader

Abstract

The invention discloses a light-weight binaural speech enhancement method based on a Fourier network, which recovers pure speech from a noisy binaural signal through a light-weight deep complex network architecture. The method comprises the following steps: firstly, generating robust acoustic representation by utilizing a double-path feature coding and fusion module and combining short-time Fourier transform features and psychoacoustic-based Gammatone features; then, through a global adaptive Fourier modulator backbone network, a long-time-sequence context is efficiently modeled with extremely low calculation cost, and meanwhile, phase information is kept through a real number gating mechanism; and finally, performing intelligent correction on an enhancement result through a dynamic optimization gate. According to the method, a long time sequence dependence modeling task is converted into a Fourier domain adaptive filtering process, so that the speech enhancement performance and the spatial clue fidelity are maintained, the model parameter quantity and the calculation complexity are reduced, and the problem that the performance and the efficiency of resource-constrained equipment are difficult to consider at the same time in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of audio signal processing and artificial intelligence technology, and is applicable to binaural hearing devices with strict limitations on computing resources, such as hearing aids, smart headphones, virtual reality (VR) and augmented reality (AR) devices. Specifically, it refers to a lightweight binaural speech enhancement method based on Fourier networks that aims to balance speech enhancement performance with computational efficiency while effectively preserving spatial auditory cues. Background Technology

[0002] Speech enhancement technology aims to improve speech quality and intelligibility by suppressing background noise. For binaural hearing devices such as hearing aids, speech enhancement not only needs to reduce noise, but more importantly, it must preserve spatial perception information encoded by binaural cues, such as binaural sound level differences and binaural time differences, thereby ensuring the user's accurate ability to locate sound sources. Traditional single-channel speech enhancement methods are unsuitable for binaural scenarios because their design principles disrupt the correlation between vocal channels, thus severely impairing spatial cues.

[0003] To address this issue, early research primarily employed statistical signal processing methods, such as minimum variance distortionless response beamformers or multi-channel Wiener filters. These methods performed well and were computationally efficient in stationary noise environments, but their performance deteriorated sharply in the non-stationary, dynamically changing noise scenarios common in the real world. Therefore, the focus of technological development shifted to end-to-end modeling methods based on deep learning, particularly deep, complex networks capable of jointly optimizing amplitude and phase in the complex domain, thereby generating higher-fidelity speech.

[0004] However, existing deep learning methods face a severe performance-efficiency trade-off in the field of binaural speech enhancement. On the one hand, in pursuit of the ultimate enhancement effect, complex architectures, including parallel networks and complex Transformers, have emerged. While these models have achieved industry-leading performance, their large model size and high computational cost make them difficult to deploy in real time on resource-constrained portable devices. On the other hand, lightweight models designed to meet deployment requirements often sacrifice significant performance, especially in their ability to model long-term temporal dependencies, resulting in unsatisfactory enhancement effects. Furthermore, most existing methods rely on a single Short-Time Fourier Transform (STFT) feature, whose inherent time-frequency resolution trade-off limits the robustness of the model's acoustic representation. Therefore, developing a novel technical solution that can maintain high enhancement performance and spatial cue integrity while being lightweight and computationally efficient has become a critical technical bottleneck that urgently needs to be overcome in this field. Summary of the Invention

[0005] The purpose of this invention is to provide a lightweight binaural speech enhancement method based on Fourier networks, aiming to solve the technical challenge of achieving both high performance and low complexity in existing binaural speech enhancement models. By employing a globally adaptive Fourier modulation mechanism and combining it with dynamic optimization gates for post-processing, a balance is achieved between performance, efficiency, and spatial cue preservation.

[0006] The technical solution for achieving the objective of this invention is as follows:

[0007] A lightweight binaural speech enhancement method based on Fourier networks includes the following steps:

[0008] Step 1: Perform dual-channel feature encoding and fusion on the noisy binaural signal to generate fused complex spectral features;

[0009] Step 2: Input the complex spectral features into the network backbone composed of a global adaptive Fourier modulator to perform long-term temporal context modeling and obtain deep features containing global dynamic information.

[0010] Step 3: Input the deep features into the binaural decoder to estimate the relative acoustic transfer function of the target speech and noise;

[0011] Step 4: Calculate the preliminary enhanced pure speech complex spectrum based on the relative acoustic transfer function;

[0012] Step 5: Using a dynamic optimization gate module, generate a frequency-related confidence-gated signal based on the deep features;

[0013] Step 6: Use the gated signal to perform weighted fusion of the preliminary enhanced spectrogram obtained in Step 4 and the original noisy spectrogram to obtain the final optimized binaural complex spectrogram;

[0014] Step 7: Perform an inverse short-time Fourier transform on the final optimized binaural complex spectrogram to obtain the enhanced time-domain binaural speech signal.

[0015] Furthermore, step one specifically includes:

[0016] 1a. The noisy binaural signal is input into a dual-channel encoder to extract features. The dual-channel encoder includes a main path and an auxiliary path. The main path uses a short-time Fourier transform (STFT) to generate a complex spectrum. The auxiliary path utilizes a Gammatone filter bank to generate a perceptual feature inspired by the human auditory system. The two feature paths are each passed through an independent encoder composed of lightweight one-dimensional convolutions to extract high-level abstract representations.

[0017] 1b. A cross-channel attention mechanism is employed for feature fusion. An attention map is generated using the amplitude information of the Gammatone features to dynamically modulate the STFT feature path, thereby enhancing key frequency components. The calculation formula is as follows:

[0018]

[0019] in, Represents the convolution operation. This represents the modulo operation. It is the Sigmoid activation function. Element-wise multiplication; characteristics after fusion After undergoing a complex squeezing and excitation (SE) module for channel-level recalibration, an information-rich and robust fused complex spectral feature is finally formed for use by the subsequent network backbone.

[0020] Furthermore, step two specifically includes:

[0021] The fused complex spectral features generated in step one are input into the network backbone composed of a Global Adaptive Fourier Modulator (GAFM). For the input complex spectral features... GAFM first performs average pooling on the feature amplitudes of each frequency band along the time dimension, thereby extracting a compact global context vector. :

[0022]

[0023] in Indicates along the time dimension The average operation of vectors It is then fed into a multilayer perceptron (MLP) to generate a set of mixing coefficients. The mixing coefficients are used to adjust a predefined and fixed Fourier basis matrix. A linear combination is performed to synthesize a content-adaptive real-valued gated signal. :

[0024]

[0025] in As a learnable scaling factor, since the gated signal is a real value, when it is element-wise multiplied with the original complex input features, only the amplitude of the features is modulated, while its phase remains strictly unchanged. This property is crucial for preserving the interaural time difference (ITD), a key spatial localization cue. Finally, the modulated features are integrated through a complex residual concatenation block to output deep features containing global dynamic information. .

[0026] Furthermore, steps three and four specifically include:

[0027] Two parallel decoder heads, each consisting of multiple lightweight 2D convolutional modules, are configured to estimate the relative acoustic transfer function of the target speech. Relative acoustic transfer function of noise Using the closed-form solution, based on the estimated two transfer functions and the original noisy signal spectrum... and The spectrogram of the preliminarily enhanced clean speech was obtained by calculation. and The calculation formula is as follows:

[0028]

[0029] Furthermore, step five specifically includes:

[0030] The deep features output in step two By aggregating information over time, a frequency-dependent confidence-gated signal is generated. ,right Average pooling is performed along the time dimension, followed by a confidence-gated signal generated through a 1x1 convolutional layer and a sigmoid activation function. The calculation formula is as follows:

[0031]

[0032] This gating signal The value of is between [0,1], and can be understood as the model's confidence in its enhancement results in each frequency band.

[0033] Step six specifically includes:

[0034] The gating signal generated in step five Used for the preliminary enhanced spectrum obtained in step four. Compared with the original noisy spectrum Weighted fusion is performed to obtain the final optimized binaural complex spectrogram. The calculation formula is as follows:

[0035]

[0036] in, , representing the left and right ear channels. This mechanism allows the model to trust its augmentation results in the high-confidence frequency region, while falling back to the original signal in the low-confidence region, thus achieving a dynamic balance between denoising and fidelity.

[0037] Furthermore, step seven specifically includes:

[0038] The final optimized binaural complex spectrogram generated in step six Perform an inverse short-time Fourier transform (iSTFT) to convert the signal from the time-frequency domain back to the time domain, thereby obtaining the final enhanced time-domain binaural speech signal.

[0039] The beneficial effects of this invention lie in achieving a balance between performance and efficiency through a novel architectural design. First, the dual-path feature encoding integrates standard acoustic features and psychoacoustic features, enhancing the model's robustness in representing complex acoustic environments. Second, the core global adaptive Fourier modulator can efficiently model global temporal dynamics with extremely low computational cost, while its real number gating mechanism strictly preserves phase information, thus accurately retaining crucial binaural spatial cues. Finally, the dynamic optimization gate intelligently balances noise reduction intensity and signal fidelity, effectively suppressing processing artifacts. Experiments demonstrate that this invention, using only a minimal number of parameters, outperforms existing large and complex models in key metrics such as binaural intelligibility and spatial cue preservation. Attached Figure Description

[0040] Figure 1 This is a flowchart of the method of the present invention;

[0041] Figure 2 This is a schematic diagram of a global adaptive Fourier network according to an embodiment of the present invention;

[0042] Figure 3 This refers to the lightweight 1D / 2D module in the global adaptive Fourier network of this invention.

[0043] Figure 4 This refers to the feature fusion module in the global adaptive Fourier network of this invention.

[0044] Figure 5 This refers to the globally adaptive Fourier modulation module in the globally adaptive Fourier network of this invention. Detailed Implementation

[0045] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0046] An embodiment of the present invention provides a lightweight binaural speech method based on Fourier networks, such as... Figure 1 , Figure 2 As shown, it adopts an encoder-decoder architecture, including the following specific steps:

[0047] Step 1: Dual-path feature encoding and fusion

[0048] To construct a comprehensive acoustic characterization, this invention employs a dual-path encoder architecture. The dual-path encoder includes a main path and an auxiliary path. The main path uses a standard short-time Fourier transform (STFT) to generate a complex spectrum. The auxiliary path then utilizes a Gammatone filter bank to generate a perceptual feature inspired by the human auditory system. The two feature paths are each passed through an independent encoder consisting of lightweight one-dimensional convolutions to extract high-level abstract representations.

[0049] Subsequently, as Figure 4 As shown, this invention employs a cross-channel attention mechanism for feature fusion. Specifically, an attention map is generated using the amplitude information of the Gammatone feature. This map is used to dynamically modulate the STFT feature path to enhance key frequency components. The calculation formula is as follows:

[0050]

[0051] in, Represents the convolution operation. This represents the modulo operation. It is the Sigmoid activation function. This is element-wise multiplication. The merged features... After undergoing a complex squeeze-and-excitation (SE) module for channel-dimensional recalibration, an informative and robust fused complex spectral feature is finally formed for use by the subsequent network backbone.

[0052] Step 2: Global Adaptive Fourier Modulation

[0053] The fused features generated in step one are input into the network backbone composed of a Global Adaptive Fourier Modulator (GAFM). For the input complex features... GAFM first performs average pooling on the feature amplitudes of each frequency band along the time dimension, thereby extracting a compact global context vector. :

[0054]

[0055] in Indicates along the time dimension The average operation is then performed. This vector is subsequently fed into a small multilayer perceptron (MLP) to generate a set of mixing coefficients. These coefficients are used to iterate over a predefined and fixed Fourier basis matrix. A linear combination is performed to synthesize a content-adaptive real-valued gated signal. :

[0056]

[0057] in This is a learnable scaling factor. Since the gated signal is real-valued, when it is element-wise multiplied with the original complex input features, only the amplitude of the features is modulated, while its phase remains strictly unchanged. This property is crucial for preserving the key spatial localization cue of interaural time difference. Finally, the modulated features are integrated through a standard complex residual concatenation block to output deep features containing global dynamic information. .

[0058] Step 3: Relative Acoustic Transfer Function Estimation

[0059] The goal of the decoding phase is to extract deep features from the network backbone. The enhanced binaural signal is reconstructed. This invention employs two parallel decoder heads, each consisting of multiple lightweight two-dimensional convolutional modules, used to estimate the relative acoustic transfer function of the target speech. Relative acoustic transfer function of noise These two transfer functions describe the relationship between the target speech and noise signals in both ears.

[0060] Step 4: Preliminary Speech Spectrum Recovery

[0061] Based on the target speech relative acoustic transfer function estimated in step three Acoustic transfer function relative to noise And the spectrum of the original noisy input signal. and (Representing the left and right ears respectively), the spectrogram of the preliminarily enhanced pure speech can be recovered through closed-form solution. and The calculation formula is as follows:

[0062]

[0063] in, This is a minimal constant used to ensure numerical stability. Guided by the physical model, this step effectively separates speech and noise.

[0064] Step 5: Dynamically optimize gating signal generation

[0065] To further suppress potential processing artifacts and improve the model's robustness, this invention introduces a dynamic optimization gate. This module optimizes the deep features output from step two. A frequency-dependent confidence-gated signal is generated by aggregating information over time. Specifically, for Average pooling is performed along the time dimension, followed by generation through a 1x1 convolutional layer and a sigmoid activation function. The calculation formula is as follows:

[0066]

[0067] This gating signal The value of is between [0,1], and can be understood as the model's confidence in its enhancement results in each frequency band.

[0068] Step Six: Gated Weighted Fusion

[0069] The gating signal generated in step five Used to enhance the preliminary spectrum obtained in step four. Compared with the original noisy spectrum Weighted fusion is performed to obtain the final optimized binaural complex spectrogram. The calculation formula is as follows:

[0070]

[0071] in, , representing the left and right ear channels. This mechanism allows the model to trust its augmentation results in the high-confidence frequency region, while falling back to the original signal in the low-confidence region, thus achieving a dynamic balance between denoising and fidelity.

[0072] Step 7: Time-domain signal reconstruction

[0073] The final optimized binaural complex spectrogram generated in step six Perform an inverse short-time Fourier transform (iSTFT) to convert it from the time-frequency domain back to the time domain, thereby obtaining the final enhanced time-domain binaural speech signal.

[0074] As a preferred embodiment, the present invention further includes an experimental verification step to demonstrate the beneficial effects of the method of the present invention, as follows:

[0075] 1) Dataset and Implementation Details: Training and evaluation were performed on a strictly partitioned dataset with no overlap between speaker and head-related impulse response (HRIR) data. The audio sampling rate was 16 kHz, and the STFT used a 256-point FFT and a 128-point frame shift. The Gammatone filter bank was configured with 64 channels. The model was trained using the AdamW optimizer.

[0076] 2) Loss Function: To collaboratively optimize noise reduction performance, speech intelligibility, and binaural cue preservation, a composite loss function is used during training. :

[0077]

[0078] Task loss function It is a weighted sum of four objectives:

[0079]

[0080] in, It is a scale-invariant signal-to-noise ratio loss. It is a short-term loss of objective intelligibility. and These represent the losses for interaural level difference and interaural phase difference, respectively. Regularization term. This is applied to the gate control signal of the dynamically optimized gate. This is to encourage them to make clear decisions.

[0081] 3) Evaluation Indicators: Four standard objective indicators are used for comprehensive evaluation: ∆PESQ for quantifying the improvement in perceived quality, and ILD error for measuring the accuracy of spatial cue reconstruction. ) and IPD error ( ), and MBSTOI, which comprehensively assesses speech intelligibility and binaural cue integrity.

Claims

1. A lightweight binaural speech enhancement method based on Fourier networks, characterized in that, The method includes the following specific steps: Step 1: Perform dual-channel feature encoding and fusion on the noisy binaural signal to generate fused complex spectral features; Step 2: Input the complex spectral features into the network backbone composed of a global adaptive Fourier modulator to perform long-term temporal context modeling and obtain deep features containing global dynamic information. Step 3: Input the deep features into the binaural decoder to estimate the relative acoustic transfer function of the target speech and noise; Step 4: Calculate the preliminary enhanced pure speech complex spectrum based on the relative acoustic transfer function; Step 5: Using a dynamic optimization gate module, generate a frequency-related confidence-gated signal based on the deep features; Step 6: Use the gated signal to perform weighted fusion of the preliminary enhanced spectrogram obtained in Step 4 and the original noisy spectrogram to obtain the final optimized binaural complex spectrogram; Step 7: Perform an inverse short-time Fourier transform on the final optimized binaural complex spectrogram to obtain the enhanced time-domain binaural speech signal.

2. The lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step one specifically includes: 1a. The noisy binaural signal is input into a dual-channel encoder to extract features. The dual-channel encoder includes a main path and an auxiliary path. The main path uses a short-time Fourier transform (STFT) to generate a complex spectrum. The auxiliary path utilizes a Gammatone filter bank to generate a perceptual feature inspired by the human auditory system. The two feature paths are each passed through an independent encoder composed of lightweight one-dimensional convolutions to extract high-level abstract representations. 1b. A cross-channel attention mechanism is employed for feature fusion. An attention map is generated using the amplitude information of the Gammatone feature, which is used to dynamically modulate the STFT feature path. The calculation formula is as follows: ; in, Represents the convolution operation. This represents the modulo operation. It is the Sigmoid activation function. Element-wise multiplication; characteristics after fusion After channel-dimensional recalibration via a complex squeeze and excitation (SE) module, fused complex spectral features are formed.

3. The lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step two specifically includes: The fused complex spectral features generated in step one are input into the network backbone composed of a Global Adaptive Fourier Modulator (GAFM). For the input complex spectral features... GAFM first performs average pooling on the feature amplitudes of each frequency band along the time dimension to extract a global context vector. : ; in Indicates along the time dimension The average operation of vectors It is then fed into a multilayer perceptron (MLP) to generate a set of mixing coefficients. The mixing coefficients are relative to a predefined and fixed Fourier basis matrix. By performing linear combination, a content-adaptive real-valued gated signal is synthesized. : ; in The real-valued gated signal is a learnable scaling factor. With the input complex value features Element-wise multiplication is performed to obtain the modulated features; finally, the modulated features are integrated through a complex residual connect block to output deep features containing global dynamic information. .

4. The lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step three specifically includes: Two parallel decoder heads are used to extract the deep features from the output of step two. The reconstructed and enhanced binaural signals are used to estimate the relative acoustic transfer function of the target speech. Each head segment consists of multiple lightweight two-dimensional convolutional modules. Relative acoustic transfer function of noise .

5. The lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step four specifically includes: Based on the target speech relative acoustic transfer function estimated in step three Acoustic transfer function relative to noise And the spectrum of the original noisy input signal. and , representing the left and right ears respectively, can be used to recover the preliminary enhanced pure speech spectrogram through closed-form decoding. and The calculation formula is as follows: ; in, It is a very small constant used to ensure numerical stability.

6. A lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step five specifically includes: The deep features output in step two Average pooling is performed along the time dimension, followed by a confidence-gated signal generated through a 1x1 convolutional layer and a sigmoid activation function. The calculation formula is as follows: ; This gating signal The value of is between [0, 1].

7. A lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step six specifically includes: The gating signal generated in step five Used for the preliminary enhanced spectrum obtained in step four. Compared with the original noisy spectrum Weighted fusion is performed to obtain the final optimized binaural complex spectrogram. The calculation formula is as follows: ; in, , representing the left and right ear canals; This is element-wise multiplication.

8. A lightweight binaural speech enhancement method based on Fourier networks according to claim 1, characterized in that, Step seven specifically includes: The final optimized binaural complex spectrogram generated in step six Perform an inverse short-time Fourier transform (iSTFT) to convert it from the time-frequency domain back to the time domain, and obtain the final enhanced time-domain binaural speech signal.