Voice Quality Evaluation Using Human Auditory Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing objective voice quality evaluation tools, particularly non-intrusive signal domain models, face challenges in accuracy due to their reliance on oral phonation mechanisms that differ from auditory perception, leading to inaccurate voice quality assessments.
Innovation Solution
A method and apparatus that perform human auditory modeling processing, variable resolution time-frequency analysis, and feature extraction on voice signals to obtain a voice quality evaluation result, using a band-pass filter bank and discrete wavelet transform to mimic human auditory processing and enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If non-intrusive signal domain model based on oral phonation mechanism is used, then device complexity is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent replaces the oral phonation mechanism model with a human auditory system model. Instead of modeling how sound is produced (mechanical/physical process), the system models how sound is perceived (biological/psychological process). This substitution fundamentally changes the evaluation approach from source-oriented to receiver-oriented, improving accuracy while maintaining non-intrusive operation.
Solution Approach 2:
The patent transforms the evaluation parameters from phonation-based parameters (glottal pulse, vocal tract characteristics) to auditory perception parameters (temporal envelope, spectral characteristics, psychoacoustic features). This parameter transformation aligns the evaluation metrics with human perception, resolving the accuracy issue while keeping the model complexity manageable through selective feature extraction.
2Measurement precision
If subjective testing is used, then measurement precision is improved, but loss of time increases and productivity decreases
Solution Approach 1:
The patent creates an objective copy of the subjective evaluation process by modeling the human auditory system's perception mechanisms. Instead of requiring actual human listeners (subjective testing), the system uses computational models that replicate human auditory processing, including temporal envelope extraction, spectral analysis, and psychoacoustic feature computation. This copying approach maintains high measurement precision while eliminating time loss and productivity issues associated with organizing human testers.
Solution Approach 2:
The patent substitutes the mechanical process of organizing human testers, scheduling sessions, and collecting manual feedback with an automated computational system. The auditory modeling processing, time-frequency analysis, and feature extraction are performed automatically by algorithms, replacing the entire subjective testing workflow with an objective computational equivalent that delivers comparable accuracy without the organizational overhead.
Data Source
AI summary
A method for evaluating voice quality includes performing human auditory modeling processing on a voice signal to obtain a first signal; performing variable resolution time-frequency analysis on the first signal to obtain a second signal; and performing, based on the second signal, feature extraction and analysis to obtain a voice quality evaluation result of the voice signal. According to the foregoing technical solutions, a problem that accuracy of a voice quality evaluation is not high can be solved. A voice quality evaluation result with relatively high accuracy is finally obtained by performing human auditory modeling processing, then converting a to-be-detected signal into a multi-resolution signal, further analyzing the time-frequency signal of variable resolution, extracting a feature corresponding to the signal, and performing further analysis.


