Speech Intelligibility Optimization via Psycho-Acoustic Gain Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech signal processing technologies fail to effectively enhance speech intelligibility in noisy or distorted environments, particularly for individuals with hearing impairments, as they are insensitive to smaller distortion types and do not account for hearing loss.
Innovation Solution
The method employs psycho-acoustic variables from a model of speech perception, such as Fletcher's Articulation Index, to determine optimal frequency-band specific gain adjustments, which are iteratively calculated and applied to maximize speech intelligibility by minimizing noise interference and distortion, while ensuring the loudness limit is not exceeded.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If simplified AI metrics are used to evaluate communication systems, then the calculation is easy to use, but the metrics are insensitive to smaller but significant differences and fail in comparisons of different distortion types
Solution Approach 1:
The patent modifies the AI calculation by incorporating hearing loss parameters and distortion-type-specific weighting factors. This changes the parameters of the AI metric to make it sensitive to different distortion types while maintaining the basic calculation framework, thus resolving the contradiction between ease of use and measurement precision.
2Measurement precision
If Fletcher's 1950 finely tuned AI metric is used, then the prediction power is superior, but the concepts are difficult and at odds with current research trends
Solution Approach 1:
The patent segments the AI calculation into distinct components: hearing loss compensation, distortion-type identification, and weighted scoring. This segmentation maintains the prediction power of Fletcher's metric while organizing the complexity into manageable, modular sections that align with current research trends.
3Reliability
If background noise and distortions are present in speech transmission, then speech transmission occurs, but the speech becomes partially or completely unintelligible
Solution Approach 1:
The patent implements a feedback mechanism where the AI metric is calculated for different frequency bands and gain adjustments, and this information is used to iteratively optimize the frequency response. This feedback loop compensates for noise and distortion effects, maintaining speech intelligibility while preserving transmission capability.
4Reliability
If cochlear damage occurs, then sound detection is impeded, but the spectral and temporal smearing increases masking effectiveness of background noise
Solution Approach 1:
The patent applies local quality by implementing frequency-specific gain adjustments tailored to the listener's hearing loss profile. Different frequency bands receive different compensation levels based on the localized damage in the cochlea, which reduces noise masking in affected frequencies while preserving sound detection capability in less damaged regions.
Data Source
AI summary
Methods and apparatus for maximizing speech intelligibility use psycho-acoustic variables of a model of speech perception to control the determination of optimal frequency-band specific gain adjustments. Speech signals (or other audio input) whose intelligibility is to be improved are characterized by parameters which are applied to the model. These include measurements or estimates of speech intensity level, average noise spectrum of the incoming audio signal, and/or the current frequency-gain characteristic of the hearing compensation device. Characterizations of listeners based on hearing test results, for example, may also be applied to the model. Frequency-band specific gain adjustments generated by use of the model can be used for hearing aids, assistive listening devices, telephones, cellular telephones, or other speech delivery systems, personal music delivery systems, public-address systems, sound systems, speech generating systems, or other devices or mediums which project, transfer or assist in the detection or recognition of speech.


