Formant Boost Estimator for Speech Intelligibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing noise reduction technologies, such as Active Noise Cancellation (ANC), are limited in improving speech intelligibility in noisy environments without headsets and face challenges in manipulating resonances due to computational complexity, leading to artificially sounding speech.

Innovation Solution

A device with a processor and memory that calculates noise and speech spectral estimates, formant signal-to-noise ratio (SNR) estimates, and applies gain factors to enhance speech intelligibility by targeting a pre-selected SNR within each formant, using a low-order linear prediction filter and Levinson-Durbin algorithm, while maintaining naturalness and spectral contrast.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If resonances are manipulated through line spectral pair representation to reduce computational complexity, then computational cost decreases, but resonances close to each other become difficult to manipulate separately due to interaction

Engineering Contradiction:
Improvecomputational complexityVSAvoidease of manipulating resonances
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent segments the speech spectrum into distinct formant regions and processes each formant independently using separate gain factors. This segmentation allows resonances to be manipulated separately without interaction problems, while still achieving the desired formant boosting effect through the segmented processing approach.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If resonances are strengthened by moving poles closer to the unit circle, then speech intelligibility improves, but bandwidth narrows resulting in artificially-sounding speech

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidbandwidth
Core Design Contradiction:
Measurement precisionVSShape

Solution Approach 1:

The patent applies different gain factors to different formant regions, allowing selective strengthening of resonances in specific frequency bands while preserving the natural bandwidth characteristics in other regions. This local quality approach ensures that each formant is enhanced independently without unnecessarily narrowing the overall bandwidth, maintaining natural speech quality.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If polynomial root-finding algorithms are used to obtain resonances from LPC coefficients, then resonance manipulation is enabled, but computational expense increases

Engineering Contradiction:
Improveresonance manipulation capabilityVSAvoidcomputational expense
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts formant center frequencies and bandwidths directly from the speech spectrum using peak detection methods, bypassing the need for polynomial root-finding algorithms. This extraction approach enables resonance manipulation while significantly reducing computational expense by working directly with spectral peaks rather than solving complex polynomial equations.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10043533B2Method and device for boosting formants from speech and noise spectral estimation
Publication Date: 2018.08.07 GOODIX TECH HK CO LTD
  • US10043533B2 patent drawing
  • US10043533B2 patent drawing
  • US10043533B2 patent drawing

AI summary

A device including a processor and a memory is disclosed. The memory includes a noise spectral estimator to calculate noise spectral estimates from a sampled environmental noise, a speech spectral estimator to calculate speech spectral estimates from the input speech, a formant signal to noise ratio (SNR) estimator to calculate SNR estimates using the noise spectral estimates and speech spectral estimates within each formant detected in a speech spectrum. The memory also includes a formant boost estimator to calculate and apply a set of gain factors to each frequency component of the input speech such that the resulting SNR within each formant reaches a pre-selected target value.