GRU-Based RNN Noise Reduction for Real-Time Conference Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current RNN-based noise reduction methods for real-time conferences face challenges with poor real-time performance and high computation burdens, making them unsuitable for real-time applications.

Innovation Solution

The method employs a GRU-based RNN model for noise reduction, which involves training with a frame-and-window approach, using logarithmic spectra in the frequency domain, and applying a noise reduction suppression coefficient to process speech signals efficiently, reducing computation burden and enhancing real-time performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional RNN-based noise reduction methods are used, then noise reduction effectiveness is improved, but real-time performance deteriorates and computation burden increases

Engineering Contradiction:
Improvenoise reduction effectivenessVSAvoidreal-time performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the noise reduction process into distinct stages: frame-and-window operation to divide continuous speech into frames, Fourier transform to convert to frequency domain, and RNN processing applied selectively to specific frequency bins. This segmentation allows parallel processing of multiple frames and frequency components, significantly improving real-time performance while maintaining noise reduction effectiveness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the necessary components for noise reduction by applying frame-and-window operations to obtain logarithmic spectra, then selectively processing specific frequency bins through the RNN model. By extracting and processing only relevant features rather than entire speech signals, the computation burden is reduced while preserving noise reduction effectiveness.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If RNN-based noise reduction method is implemented, then noise reduction capability is improved, but computation burden increases making it unsuitable for real-time systems

Engineering Contradiction:
Improvenoise reduction capabilityVSAvoidcomputation burden
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by processing only specific frequency bins through the computationally intensive RNN model, while other frequency components are handled with simpler operations. The frame-and-window operation and Fourier transform are applied locally to individual frames, and the RNN is applied selectively to frequency bins where noise reduction is most needed, reducing overall computation burden while maintaining noise reduction capability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements partial action by applying the RNN model to only a subset of frequency bins rather than processing the entire spectrum. By performing partial processing on selected frequency components and using simpler methods for others, the computation burden is significantly reduced while the noise reduction capability is preserved in the critical frequency regions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11024324B2Methods and devices for RNN-based noise reduction in real-time conferences
Publication Date: 2021.06.01 YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
  • US11024324B2 patent drawing
  • US11024324B2 patent drawing
  • US11024324B2 patent drawing

AI summary

Disclosed herein is a method for RNN-based noise reduction in a real-time conference, comprising: performing frame-and-window for a speech signal to obtain a logarithmic spectrum of the speech signal, and placing the logarithmic spectrum into the RNN model to determine a noise reduction suppression coefficient, and then obtaining the denoised speech signal by applying the noise reduction suppression coefficient to the logarithmic spectrum of the original signal, thereby achieving utilization of the RNN noise reduction method in real-time conferences. In the present disclosure, when inputting the RNN model for estimation, only the logarithmic spectrum of the current frame needs to be inputted. The RNN model of the present disclosure has few requirements on inputted information, without performing huge preprocessing on the received speech signal, which in turn reduces computation burden, increases response speed, and enhances real-time performance.