Speech Intelligibility Evaluation via Disturbance Density Functions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech quality assessment algorithms, such as POLQA, fail to accurately evaluate the intelligibility of degraded speech signals, which is a critical factor in information transfer and differs from sound quality perception.
Innovation Solution
A method that samples both reference and degraded speech signals into frames, forms frame pairs, and calculates a difference function to derive an overall quality parameter, compensating for disturbances based on human auditory perception models, particularly focusing on consonant-vowel-consonant signal-to-noise ratios to improve intelligibility assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current speech quality assessment algorithms (POLQA, PESQ) are used to evaluate degraded speech signals, then sound quality can be assessed, but intelligibility evaluation is inaccurate and mismatches human assessment
Solution Approach 1:
The patent changes the evaluation parameters by introducing disturbance density functions that specifically model consonant-vowel-consonant signal-to-noise ratios, rather than using general quality metrics. This parameter change enables the algorithm to capture intelligibility-specific characteristics that traditional algorithms miss, resolving the contradiction between measurement precision and alignment with human assessment
Solution Approach 2:
The patent segments the speech signal into distinct phonetic components (consonants and vowels) and evaluates disturbances in each segment separately using CVC-specific SNR calculations. This segmentation allows the algorithm to focus on intelligibility-critical segments, improving both measurement precision and reliability compared to holistic quality assessment
2Loss of information
If general speech quality algorithms focus on sound quality parameters, then audio bandwidth quality can be measured, but intelligibility information transfer quality is lost
Solution Approach 1:
The patent extracts intelligibility-specific information from the degraded speech signal by calculating disturbance density functions that isolate consonant-vowel-consonant patterns. This extraction process separates intelligibility metrics from general sound quality parameters, preventing loss of intelligibility information while maintaining measurement precision for the extracted features
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to a method of evaluating intelligibility of a degraded speech signal received from an audio transmission system conveying a reference speech signal. The method comprises sampling said signals into reference and degraded signal frames, and forming frame pairs by associating reference and degraded signal frames with each other. For each frame pair a difference function representing disturbance is provided, which is then compensated for specific disturbance types for providing a disturbance density function. Based on the density function of a plurality of frame pairs, an overall quality parameter is determined. The method provides for compensating the overall quality parameter for the effect that the assessment of intelligibility of CVC words is dominated by the intelligibility of consonants.