Monaural Speech Intelligibility Predictor Using Clean Reference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in predicting the intelligibility of noisy or processed speech signals without a clean reference, which is essential for improving speech understanding in hearing aids and similar devices.
Innovation Solution
A monaural speech intelligibility predictor unit that processes both clean and noisy speech signals using time-frequency representations, envelope extraction, and normalization to estimate intelligibility, incorporating a hearing loss model for personalized adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If monaural speech intelligibility prediction is performed without a clean reference signal, then the system can operate in purely acoustic environments, but the prediction accuracy deteriorates significantly
Solution Approach 1:
The patent introduces a clean reference signal as an intermediary element that mediates between the noisy speech signal and the intelligibility prediction process. This reference signal, obtained through wireless transmission or separate recording, serves as a mediator that enables accurate prediction by providing a noise-free baseline for comparison with the degraded speech signal
Solution Approach 2:
The system performs preliminary acquisition and processing of the clean reference signal before the actual intelligibility prediction takes place. By obtaining the reference signal in advance through wireless transmission or separate recording, the system prepares the necessary clean baseline data beforehand, enabling more accurate prediction when combined with the noisy speech signal
2Measurement precision
If complex signal processing operations are applied to improve prediction accuracy, then the intelligibility estimation becomes more precise, but the computational complexity and processing time increase
Solution Approach 1:
The patent segments the speech signal into time-frequency units and processes different aspects of the signal separately. By dividing the complex prediction task into manageable segments (time-frequency analysis, envelope extraction, spectral comparison), the system achieves accurate prediction while keeping each processing stage computationally tractable
Solution Approach 2:
The system extracts specific critical features from the speech signal, such as temporal envelopes and spectral characteristics, rather than processing the entire signal in full detail. By taking out and focusing on the most informative features for intelligibility prediction, the system maintains high accuracy while reducing overall computational complexity
3Measurement precision
If the system processes both clean and noisy speech signals simultaneously, then accurate intelligibility prediction is achieved, but the system requires additional input channels and processing resources
Solution Approach 1:
The patent makes the reference signal acquisition mechanism multi-functional by enabling it to serve dual purposes: providing clean speech signals for intelligibility prediction and potentially serving as a wireless communication channel. This universal approach allows the same input mechanism to fulfill multiple functions, reducing the need for separate dedicated channels
Data Source
Figure 1A~1B
Figure 2A~2C
Figure 3A~3C
AI summary
A monaural intrusive speech intelligibility predictor unit comprises • First and second input units for providing time-frequency representations s(k,m) and x(k,m) of noise-free and noisy and/or processed versions of a target signal, respectively, k being a frequency bin index, k=1, 2, ..., K, and m being a time index; • First and second envelope extraction units for providing time-frequency sub-band representations of the signals sj(m) and xj(m), j being a frequency sub-band index, j=1, 2, ..., J; • First and second time-frequency segment division units for dividing the time-frequency sub-band representations sj(m) and xj(m) into time-frequency segments Sm and Xm corresponding to a number N of successive samples of the sub-band signals; • An intermediate speech intelligibility calculation unit adapted for providing intermediate speech intelligibility coefficients dm estimating an intelligibility of said time-frequency segment Xm, based on said time-frequency segments Sm and Xm or normalized and/or transformed versions S̃m, and X̃m thereof; and • A final monaural speech intelligibility calculation unit for calculating a final monaural speech intelligibility predictor d estimating an intelligibility of said noisy and/or processed version x of the target signal by combining said intermediate speech intelligibility coefficients dm, or a transformed version thereof, over time. A hearing aid comprises a monaural, intrusive intelligibility predictor unit, and a configurable signal processor adapted to control or influence the processing of one or more electric input signals representing environment sound to maximize the final speech intelligibility predictor d. A binaural hearing aid system comprises first and second hearing aids.