Speech Intelligibility Evaluation via Adaptive Disturbance Functions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech quality assessment algorithms, such as POLQA, PESQ, and PSQM, fail to accurately evaluate the intelligibility of degraded speech signals, which is crucial for information transfer quality, as they primarily focus on sound quality rather than information clarity.
Innovation Solution
A method that samples reference and degraded speech signals into frames, forms frame pairs, pre-processes them to compare differences, and uses disturbance density functions adapted to human auditory perception, considering audio power levels to assess intelligibility by switching between different difference functions based on a threshold disturbance level optimized for the audio power conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current speech quality assessment algorithms (POLQA, PESQ, PSQM) are used to evaluate degraded speech signals, then sound quality can be assessed, but intelligibility evaluation accuracy deteriorates
Solution Approach 1:
The patent segments the speech signal into frames and further divides the frequency spectrum into critical bands. It separates the evaluation into multiple components: distortion type identification, disturbance level measurement, and intelligibility scoring. This segmentation allows the complex evaluation task to be broken down into manageable steps, improving accuracy while maintaining computational feasibility through structured processing stages.
Solution Approach 2:
The patent applies different evaluation criteria and weighting factors to different frequency bands and distortion types. Instead of uniform evaluation, it assigns local quality weights based on the specific characteristics of each frequency region and disturbance type, recognizing that human perception varies across different spectral regions. This localized approach enhances intelligibility assessment accuracy by matching the specific perceptual characteristics of each signal segment.
2Measurement precision
If disturbance density functions are adapted to human auditory perception model, then evaluation accuracy improves, but processing time increases
Solution Approach 1:
The patent performs preliminary actions by pre-calculating and storing disturbance density functions that are adapted to human auditory perception. These functions are prepared in advance and can be directly applied to new signals without re-computing the entire perception model. This preliminary preparation significantly reduces processing time during actual evaluation while maintaining high accuracy in quality assessment.
Solution Approach 2:
The patent changes parameters dynamically based on the detected disturbance type and audio power level. Instead of using a fixed set of evaluation parameters, it adjusts the weighting factors, threshold values, and processing depth according to the specific characteristics of the degraded signal. This adaptive parameter adjustment optimizes the balance between evaluation accuracy and processing efficiency for each specific case.
3Measurement precision
If multiple difference functions are used to compensate for different disturbance types, then evaluation accuracy improves, but algorithm complexity increases
Solution Approach 1:
The patent implements a dynamic selection mechanism that chooses the appropriate difference function based on the detected disturbance type and audio power level. Rather than processing all possible disturbance types simultaneously, the system dynamically adapts its behavior by selecting the most relevant evaluation function for each specific case. This dynamic approach maintains high accuracy across diverse disturbance conditions while reducing overall algorithmic complexity through conditional processing paths.
Solution Approach 2:
The patent changes algorithmic parameters including the selection of difference functions, weighting factors, and processing thresholds based on the detected audio power level and disturbance characteristics. By adjusting these parameters according to the specific signal conditions, the system can use multiple sophisticated evaluation functions when needed while falling back to simpler processing when appropriate, thus managing complexity adaptively rather than uniformly.
Data Source
AI summary
The present invention relates to a method of evaluating intelligibility of a degraded speech signal received from an audio transmission system conveying a reference signal. The method comprises sampling said reference and degraded signal into frames, and forming frame pairs. For each pair one or more difference functions representing a difference between the degraded and reference signal are provided. A difference function is selected and compensated for different disturbance types, such as to provide a disturbance density function adapted to human auditory perception. An overall quality parameter is determined indicative of the intelligibility of the degraded signal. The method comprises determining a switching parameter indicative of audio power level of said degraded signal, for performing said selecting.


