Speech Signal Processing for Intelligibility via Harmonic Gain

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound quality enhancing algorithms face a tradeoff between minimizing residual background noise and speech distortion, which can deteriorate speech intelligibility.

Innovation Solution

A speech signal processing apparatus and method that determines a gain for input signals based on harmonic characteristics of voiced speech using a comb filter, preserves voiced speech by applying this gain, and uses linear predictive coefficients to preserve unvoiced speech, effectively reducing background noise while minimizing distortion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If background noise is removed from input signal, then residual background noise is reduced, but speech distortion is intensified and speech intelligibility deteriorates

Engineering Contradiction:
Improveresidual background noiseVSAvoidspeech distortion
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent segments the speech signal into voiced speech and unvoiced speech components, applying different processing strategies to each. Voiced speech undergoes harmonic characteristic-based gain determination to preserve musicality, while unvoiced speech is processed separately using linear predictive coefficients. This segmentation allows selective noise reduction without uniformly distorting all speech components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by determining gain values specifically for harmonic components of voiced speech based on their spectral characteristics. The comb filter is designed to target specific frequency regions where harmonic structures exist, allowing noise reduction in those regions while preserving the local quality and intelligibility of the voiced speech portions.

Inventive Principle:
Principle #3Local quality

2Loss of information

If gain is applied to preserve harmonic components, then speech intelligibility is enhanced, but speech distortion may occur

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidspeech distortion
Core Design Contradiction:
Loss of informationVSManufacturing precision

Solution Approach 1:

The patent employs feedback mechanisms where the processed speech signal is continuously monitored and fed back into the system. The gain values are dynamically adjusted based on the feedback from harmonic characteristic analysis, allowing the system to learn from previous processing results and refine its performance to minimize distortion while maintaining intelligibility.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes parameters dynamically by adjusting gain values based on the detected harmonic characteristics of the input signal. The comb filter parameters are modified according to the spectral content, and linear predictive coefficients are updated in real-time. This parameter adaptation allows the system to optimize speech intelligibility while minimizing distortion for each specific input condition.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9767829B2Speech signal processing apparatus and method for enhancing speech intelligibility
Publication Date: 2017.09.19 SAMSUNG ELECTRONICS CO LTD
  • US9767829B2 patent drawing
  • US9767829B2 patent drawing
  • US9767829B2 patent drawing

AI summary

A speech signal processing apparatus and a speech signal processing method for enhancing speech intelligibility are provided. The speech signal processing apparatus includes an input signal gain determiner to determine a gain of an input signal based on a harmonic characteristic of a voiced speech, a voiced speech output unit to output a voiced speech in which a harmonic component is preserved by applying the gain to the input signal, a linear predictive coefficient determiner to determine a linear predictive coefficient based on the voiced speech, and an unvoiced speech preserver to preserve an unvoiced speech of the input signal based on the linear predictive coefficient.