Voice Quality Enhancement via Speech-Noise Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech encoders have rudimentary noise suppressors that misclassify audio signals, leading to poor noise suppression and quality issues due to their monaural nature and conservative classification approaches, which can result in choppy speech and inefficient data rate allocation.

Innovation Solution

A high-quality noise suppressor is integrated to classify audio signals and provide classification data to the speech encoder, allowing for improved noise suppression and data rate management through the use of scaling factors and acoustic cues, enabling more accurate classification and smoother transitions between encoding modes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a monaural noise suppressor with conservative classification is used, then device complexity is reduced, but noise suppression quality and speech classification accuracy deteriorate

Engineering Contradiction:
Improvenoise suppressor structureVSAvoidspeech classification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transitions from monaural to binaural noise suppression by utilizing audio signals from both left and right microphones. This dimensional change enables the system to perform spatial analysis and classify speech and noise more accurately across different frequency ranges, resolving the contradiction between device simplicity and classification accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent divides the audio frequency spectrum into multiple frequency ranges (e.g., low, mid, high frequencies) and applies different noise suppression and classification strategies to each range. This segmentation allows the system to achieve high classification accuracy across diverse speech and noise characteristics without requiring a single overly complex algorithm.

Inventive Principle:
Principle #1Segmentation

2Object-affected harmful factors

If aggressive noise suppression is applied, then noise reduction improves, but speech quality deteriorates due to misclassification

Engineering Contradiction:
Improvebackground noise levelVSAvoidspeech quality
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent applies different noise suppression strengths to different frequency ranges based on local characteristics. Speech frequencies receive minimal suppression to preserve quality, while noise-dominated frequencies receive stronger suppression. This localized approach resolves the contradiction by adapting the suppression intensity to the specific spectral content of each frequency band.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system continuously monitors classification confidence and adjusts noise suppression intensity accordingly. When speech is detected with high confidence, suppression is reduced to preserve quality; when noise dominates, suppression is increased. This feedback mechanism dynamically balances noise reduction and speech quality preservation.

Inventive Principle:
Principle #23Feedback

3Reliability

If data rate is increased to improve speech encoding quality, then voice quality improves, but resource wastage increases during noise periods

Engineering Contradiction:
Improvevoice qualityVSAvoiddata rate resource consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent implements dynamic data rate adjustment based on real-time speech-noise classification. During speech periods, the system uses higher data rates to maintain quality; during noise periods, it switches to lower data rates or silence suppression modes. This dynamic adaptation resolves the contradiction by matching resource consumption to actual speech presence.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes encoding parameters (data rate, bit allocation) based on classification results. When noise is detected, parameters are adjusted to reduce transmission resources; when speech is detected, parameters are optimized for quality. This parameter adaptation allows the system to maintain quality during speech while conserving resources during noise periods.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8311817B2Systems and methods for enhancing voice quality in mobile device
Publication Date: 2012.11.13 SAMSUNG ELECTRONICS CO LTD
  • US8311817B2 patent drawing
  • US8311817B2 patent drawing
  • US8311817B2 patent drawing

AI summary

Provided are methods and systems for enhancing the quality of voice communications. The method and corresponding system may involve classifying an audio signal into speech, and speech and noise and creating speech-noise classification data. The method may further involve sharing the speech-noise classification data with a speech encoder via a shared memory or by a Least Significant Bit (LSB) of a Pulse Code Modulation (PCM) stream. The method and corresponding system may also involve sharing acoustic cues with the speech encoder to improve the speech noise classification and, in certain embodiments, sharing scaling transition factors with the speech encoder to enable the speech encoder to gradually change data rate in the transitions between the encoding modes.