Gammachirp Envelope Distortion Index for Speech Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech intelligibility calculating methods, such as sEPSM and dcGC-sEPSM, are limited in their ability to accurately estimate speech intelligibility due to dependency on speech enhancement techniques and inability to simulate non-linearity of the peripheral auditory system, particularly for hearing-impaired individuals.
Innovation Solution
A speech intelligibility calculating method using a gammachirp envelope distortion index (GEDI) that finds the feature of a distortion component between clean and enhanced speech signals, allowing for accurate estimation without relying on specific speech enhancement methods and reflecting features of both hearing and hearing-impaired auditory systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional sEPSM or dcGC-sEPSM methods are used to calculate speech intelligibility, then the calculation can be performed using established models, but the calculation is limited to speech enhancement techniques that can estimate residual noise and cannot accurately reflect non-linear features of the auditory system
Solution Approach 1:
The invention extracts and analyzes only the necessary components (enhanced speech and clean speech) while eliminating the need for residual noise estimation. By focusing on the distortion component between clean and enhanced speech envelopes, the method removes the limitation of requiring residual noise from enhancement techniques.
Solution Approach 2:
Instead of trying to estimate residual noise as done in conventional methods, the invention inverts the approach by directly analyzing the distortion component between clean and enhanced speech. This reversal of the problem-solving approach eliminates the dependency on residual noise estimation capabilities.
2Device complexity
If gammatone auditory filter bank is used in sEPSM, then the model is computationally simpler, but it cannot simulate the non-linearity of the peripheral auditory system
Solution Approach 1:
The invention changes the filter bank parameters from linear gammatone filters to dynamic compressive gammachirp filters that can adapt their characteristics. This parameter change enables the model to simulate non-linear auditory system behavior while maintaining computational feasibility through the dynamic adjustment of filter properties.
3Ease of manufacture
If residual noise estimation is required for speech intelligibility calculation, then conventional models can be applied, but the method becomes dependent on specific speech enhancement techniques and their ability to estimate residual noise
Solution Approach 1:
The invention makes the speech intelligibility calculation self-sufficient by directly computing the distortion component from clean and enhanced speech without requiring residual noise estimation from the enhancement technique. This self-service approach eliminates dependency on enhancement method capabilities.
Data Source
AI summary
A speech intelligibility calculating method is a method executed by a speech intelligibility calculating apparatus, the speech intelligibility calculating method including: a speech intelligibility calculating step of calculating a speech intelligibility that is an objective assessment index of a speech quality, based on a difference component between features found through an analysis of an input clean speech and an input enhanced speech, using one or more filter banks; and a step of outputting the speech intelligibility calculated at the speech intelligibility calculating step. This speech intelligibility calculating method is capable of calculating a speech intelligibility without any dependency on a speech enhancement method.


