Adaptive Noise Pre-processor for Speech Codec Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing noise pre-processors in voice codecs for wireless telephones fail to effectively improve speech quality under noisy conditions, particularly in suppressing background noise during speech transitions and in regions with fricative and nasal sounds.

Innovation Solution

The noise pre-processor employs adaptive smoothing constants, conditional boosting of signal-to-noise ratio estimates, and adaptive gain calculation to enhance noise suppression, using templates for voice metric computation and long-term prediction coefficients to refine SNR estimates, thereby improving noise reduction without degrading speech quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional noise pre-processor is used in speech codec, then device complexity is kept low, but speech quality deteriorates under noisy conditions particularly during speech transitions and fricative/nasal sounds

Engineering Contradiction:
Improvespeech qualityVSAvoidnoise pre-processor complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic adaptation by updating smoothing constants based on speech activity detection and signal-to-noise ratio estimates. The noise pre-processor transitions from static parameters to dynamic parameters that adapt to changing speech conditions, particularly during speech transitions and fricative/nasal sounds, thereby improving speech quality without requiring a complete redesign of the system architecture.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes key parameters including smoothing constants (alpha and beta), noise update rates, and gain factors based on detected speech conditions. By modifying these parameters dynamically according to speech activity and SNR estimates, the system achieves better noise suppression performance under varying conditions while maintaining the same fundamental processing structure, thus avoiding excessive complexity increase.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If adaptive smoothing constants are used to improve noise suppression, then speech quality improves, but computational complexity increases

Engineering Contradiction:
Improvenoise suppression effectivenessVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies different smoothing constants to different frequency channels based on local signal characteristics. Instead of using a single global smoothing constant, the system computes channel-specific smoothing parameters adapted to the local SNR conditions in each frequency band, improving noise suppression effectiveness while keeping computations localized and efficient.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs preliminary speech activity detection and SNR estimation before applying adaptive smoothing. By pre-processing the signal to identify speech regions and estimate noise levels, the system can selectively apply adaptive smoothing only where needed, avoiding unnecessary computations in silent or purely noisy segments, thus reducing overall computational complexity.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If signal-to-noise ratio estimation is enhanced with conditional boosting, then noise reduction performance improves, but processing time increases

Engineering Contradiction:
Improvenoise reduction performanceVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements conditional boosting that applies gain factors selectively based on detected speech conditions. Instead of always applying maximum boosting, the system applies partial boosting only when speech activity is detected and SNR conditions warrant it, avoiding excessive processing in regions where full boosting is not needed, thus reducing average processing time while maintaining performance where required.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent updates SNR estimates and applies conditional boosting at specific intervals (every speech frame) rather than continuously. By periodic updating of noise parameters and selective application of boosting gains, the system achieves good noise reduction performance while avoiding continuous heavy computation, thereby controlling processing time.

Inventive Principle:
Principle #19Periodic action

4Reliability

If adaptive gain calculation with variable minimum gain is used, then speech quality in noisy conditions improves, but algorithm complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidalgorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces fixed minimum gain with adaptive minimum gain that varies based on overall SNR conditions and speech activity. The minimum gain parameter becomes dynamic, adjusting automatically to preserve speech quality in different noise environments without requiring manual configuration or complex lookup tables, thus improving performance with minimal additional complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses feedback from SNR estimates and speech activity detection to adjust gain parameters adaptively. By continuously monitoring signal conditions and adjusting minimum gain accordingly, the system optimizes speech quality in noisy conditions while using simple feedback loops rather than complex optimization algorithms, keeping algorithm complexity low.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS7366658B2Noise pre-processor for enhanced variable rate speech codec
Publication Date: 2008.04.29 TEXAS INSTRUMENTS INC
  • US7366658B2 patent drawing
  • US7366658B2 patent drawing
  • US7366658B2 patent drawing

AI summary

An enhanced noise pre-processor in a speech codec smoothes channel energy estimate moving toward a first smoothing constant if a prior signal to noise ratio estimate for more than five channels are above a threshold and toward a second smaller smoothing constant otherwise. Forming a signal to noise ratio estimate for each channel includes conditionally boosting if a signal energy estimate is more than a predetermined factor of a noise energy estimate and signal to noise ratio estimates are above a threshold for more than five channels. The estimated signal to noise ratio is conditionally modified if two long term prediction coefficients are above a predetermined factor. The estimated signal to noise ratio is not modified and a voice metric is set greater than a voice metric threshold upon matching templates corresponding to the fricative and nasal speech sounds. An adaptive minimum channel gain is chosen based on a current signal to noise ratio estimate.