GRU-Based RNN Noise Reduction for Real-Time Conference Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current RNN-based noise reduction methods for real-time conferences face challenges with poor real-time performance and high computation burdens, making them unsuitable for real-time applications.
Innovation Solution
The method employs a GRU-based RNN model for noise reduction, which involves training with a frame-and-window approach, using logarithmic spectra in the frequency domain, and applying a noise reduction suppression coefficient to process speech signals efficiently, reducing computation burden and enhancing real-time performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional RNN-based noise reduction methods are used, then noise reduction effectiveness is improved, but real-time performance deteriorates and computation burden increases
Solution Approach 1:
The patent segments the noise reduction process into distinct stages: frame-and-window operation to divide continuous speech into frames, Fourier transform to convert to frequency domain, and RNN processing applied selectively to specific frequency bins. This segmentation allows parallel processing of multiple frames and frequency components, significantly improving real-time performance while maintaining noise reduction effectiveness.
Solution Approach 2:
The patent extracts only the necessary components for noise reduction by applying frame-and-window operations to obtain logarithmic spectra, then selectively processing specific frequency bins through the RNN model. By extracting and processing only relevant features rather than entire speech signals, the computation burden is reduced while preserving noise reduction effectiveness.
2Reliability
If RNN-based noise reduction method is implemented, then noise reduction capability is improved, but computation burden increases making it unsuitable for real-time systems
Solution Approach 1:
The patent applies local quality by processing only specific frequency bins through the computationally intensive RNN model, while other frequency components are handled with simpler operations. The frame-and-window operation and Fourier transform are applied locally to individual frames, and the RNN is applied selectively to frequency bins where noise reduction is most needed, reducing overall computation burden while maintaining noise reduction capability.
Solution Approach 2:
The patent implements partial action by applying the RNN model to only a subset of frequency bins rather than processing the entire spectrum. By performing partial processing on selected frequency components and using simpler methods for others, the computation burden is significantly reduced while the noise reduction capability is preserved in the critical frequency regions.
Data Source
AI summary
Disclosed herein is a method for RNN-based noise reduction in a real-time conference, comprising: performing frame-and-window for a speech signal to obtain a logarithmic spectrum of the speech signal, and placing the logarithmic spectrum into the RNN model to determine a noise reduction suppression coefficient, and then obtaining the denoised speech signal by applying the noise reduction suppression coefficient to the logarithmic spectrum of the original signal, thereby achieving utilization of the RNN noise reduction method in real-time conferences. In the present disclosure, when inputting the RNN model for estimation, only the logarithmic spectrum of the current frame needs to be inputted. The RNN model of the present disclosure has few requirements on inputted information, without performing huge preprocessing on the received speech signal, which in turn reduces computation burden, increases response speed, and enhances real-time performance.


