Speech Intelligibility Predictor Using Time-Frequency Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech intelligibility measures are not suitable for time-frequency weighted noisy speech and are often complex, making them less transparent and ineffective for evaluating localized signal degradations in noisy environments.

Innovation Solution

A method providing a time-frequency representation of both target and noisy speech signals to calculate intermediate speech intelligibility coefficients, which are then averaged to produce a final predictor value, allowing for objective intelligibility assessment in noisy conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing objective intelligibility measures (AI, SII, STI) are used, then they can assess speech intelligibility for certain types of degradation, but they become less transparent and less appropriate when time-frequency weighting is applied to noisy speech

Engineering Contradiction:
Improveintelligibility measurement accuracyVSAvoidmeasure structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech signal into time-frequency units (TFUs) and calculates intelligibility coefficients for each unit independently. This segmentation allows the measure to handle time-frequency weighted noisy speech by analyzing local characteristics rather than relying on global statistics, thereby maintaining transparency and accuracy for localized signal degradations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces local speech intelligibility coefficients that assess the quality of each time-frequency unit individually. This local quality approach enables the measure to evaluate specific regions of the spectrogram where degradation occurs, making it more transparent and appropriate for time-frequency weighted processing compared to global measures.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If existing OIMs based on long-term statistics of entire speech signals are used, then they can provide overall intelligibility assessment, but they cannot effectively evaluate localized short-time TF-region degradations

Engineering Contradiction:
Improvelocalized degradation detection accuracyVSAvoidtime resolution
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the speech signal into short time frames and frequency bands to create time-frequency units. This segmentation provides both temporal and spectral resolution, enabling the measure to detect localized degradations in specific TF-regions while maintaining overall intelligibility assessment capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from analyzing only temporal statistics or only spectral statistics to analyzing the joint time-frequency domain. By introducing the time-frequency dimension, the measure can simultaneously assess local degradations in both time and frequency, providing comprehensive intelligibility evaluation that previous measures could not achieve.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If complex OIMs with extensively trained parameters are used, then they can handle certain degradation types, but they become less transparent and harder to interpret for evaluating signal processing effects

Engineering Contradiction:
Improvedegradation type coverageVSAvoidparameter training complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses simple, physically meaningful parameters (signal power, noise power, signal-to-noise ratio) that can be directly calculated from the time-frequency representation without extensive training. These parameters are changed and combined in a straightforward manner to produce the intelligibility coefficient, maintaining transparency while covering multiple degradation types through the time-frequency analysis framework.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9064502B2Speech intelligibility predictor and applications thereof
Publication Date: 2015.06.23 OTICON
  • US9064502B2 patent drawing
  • US9064502B2 patent drawing
  • US9064502B2 patent drawing

AI summary

The application relates to a method of providing a speech intelligibility predictor value for estimating an average listener's ability to understand of a target speech signal when said target speech signal is subject to a processing algorithm and/or is received in a noisy environment. The application further relates to a method of improving a listener's understanding of a target speech signal in a noisy environment and to corresponding device units. The object of the present application is to provide an alternative objective intelligibility measure, e.g. a measure that is suitable for use in a time-frequency environment. The invention may e.g. be used in audio processing systems, e.g. listening systems, e.g. hearing aid systems.