Kalman Filter Noise Reduction for Nonstationary Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing noise reduction algorithms in audio systems, particularly in mobile voice applications and single microphone devices, struggle to effectively handle nonstationary noise environments, leading to speech distortion and residual noise, as they rely on inaccurate noise power spectral density estimation.

Innovation Solution

The use of two Kalman filters to adaptively control the bandwidth for noise reduction, with one filter smoothing the noisy speech power spectral density and the other estimating the noise power spectral density, allowing for improved speech detection and reduced noise in dynamic environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional noise reduction algorithms are used, then stationary noise can be reduced, but nonstationary noise cannot be effectively handled leading to speech distortion and residual noise

Engineering Contradiction:
Improvenoise reduction effectivenessVSAvoidadaptability to nonstationary noise
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic adaptation by continuously updating noise power spectral density estimates in real-time using speech activity detection. The system transitions from static noise estimation to dynamic tracking, adjusting noise parameters adaptively as noise conditions change, enabling effective handling of nonstationary noise environments while maintaining speech quality

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses the input signal itself to estimate noise characteristics by identifying speech activity periods and using those periods to update noise models. The algorithm serves itself by extracting noise information from the mixed signal without requiring separate noise reference inputs, enabling autonomous adaptation to changing noise conditions

Inventive Principle:
Principle #25Self-service

2Measurement precision

If noise estimation is performed during speech pause periods, then noise can be estimated, but speech detection accuracy decreases in continuous speech scenarios

Engineering Contradiction:
Improvenoise power spectral density estimation accuracyVSAvoidspeech detection accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where speech activity detection results are used to control noise estimation updates. The system continuously monitors speech activity and uses this information to determine when and how to update noise parameters, creating a closed-loop system that adapts to speech presence and maintains both accurate noise estimation and reliable speech detection

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically changes estimation parameters based on speech activity detection. When speech is detected, the system adjusts its noise estimation approach to avoid corrupting speech signals. This parameter adaptation allows accurate noise modeling during pauses while maintaining speech detection accuracy during continuous speech by modifying estimation behavior according to speech presence

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8244523B1Systems and methods for noise reduction
Publication Date: 2012.08.14 ROCKWELL COLLINS INC
  • US8244523B1 patent drawing
  • US8244523B1 patent drawing
  • US8244523B1 patent drawing

AI summary

An apparatus is shown for detecting speech in an audio signal obtained from an input device, the audio including speech and noise. The apparatus includes a processing circuit which includes a filter configured to smooth the audio signal. The processing circuit is configured to control the bandwidth of the filter based on characteristics of the audio signal and to provide a smoothed signal obtained from the filter to a voice activity detector configured to determine whether the smoothed signal represents speech.