Formant Boost Estimator for Speech Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise reduction technologies, such as Active Noise Cancellation (ANC), are limited in improving speech intelligibility in noisy environments without headsets and face challenges in manipulating resonances due to computational complexity, leading to artificially sounding speech.
Innovation Solution
A device with a processor and memory that calculates noise and speech spectral estimates, formant signal-to-noise ratio (SNR) estimates, and applies gain factors to enhance speech intelligibility by targeting a pre-selected SNR within each formant, using a low-order linear prediction filter and Levinson-Durbin algorithm, while maintaining naturalness and spectral contrast.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If resonances are manipulated through line spectral pair representation to reduce computational complexity, then computational cost decreases, but resonances close to each other become difficult to manipulate separately due to interaction
Solution Approach 1:
The patent segments the speech spectrum into distinct formant regions and processes each formant independently using separate gain factors. This segmentation allows resonances to be manipulated separately without interaction problems, while still achieving the desired formant boosting effect through the segmented processing approach.
2Measurement precision
If resonances are strengthened by moving poles closer to the unit circle, then speech intelligibility improves, but bandwidth narrows resulting in artificially-sounding speech
Solution Approach 1:
The patent applies different gain factors to different formant regions, allowing selective strengthening of resonances in specific frequency bands while preserving the natural bandwidth characteristics in other regions. This local quality approach ensures that each formant is enhanced independently without unnecessarily narrowing the overall bandwidth, maintaining natural speech quality.
3Adaptability or versatility
If polynomial root-finding algorithms are used to obtain resonances from LPC coefficients, then resonance manipulation is enabled, but computational expense increases
Solution Approach 1:
The patent extracts formant center frequencies and bandwidths directly from the speech spectrum using peak detection methods, bypassing the need for polynomial root-finding algorithms. This extraction approach enables resonance manipulation while significantly reducing computational expense by working directly with spectral peaks rather than solving complex polynomial equations.
Data Source
AI summary
A device including a processor and a memory is disclosed. The memory includes a noise spectral estimator to calculate noise spectral estimates from a sampled environmental noise, a speech spectral estimator to calculate speech spectral estimates from the input speech, a formant signal to noise ratio (SNR) estimator to calculate SNR estimates using the noise spectral estimates and speech spectral estimates within each formant detected in a speech spectrum. The memory also includes a formant boost estimator to calculate and apply a set of gain factors to each frequency component of the input speech such that the resulting SNR within each formant reaches a pre-selected target value.


