Voice Quality Evaluation Using Human Auditory Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing objective voice quality evaluation tools, particularly non-intrusive signal domain models, face challenges in accuracy due to their reliance on oral phonation mechanisms that differ from auditory perception, leading to inaccurate voice quality assessments.

Innovation Solution

A method and apparatus that perform human auditory modeling processing, variable resolution time-frequency analysis, and feature extraction on voice signals to obtain a voice quality evaluation result, using a band-pass filter bank and discrete wavelet transform to mimic human auditory processing and enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If non-intrusive signal domain model based on oral phonation mechanism is used, then device complexity is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improvemodel complexityVSAvoidvoice quality evaluation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces the oral phonation mechanism model with a human auditory system model. Instead of modeling how sound is produced (mechanical/physical process), the system models how sound is perceived (biological/psychological process). This substitution fundamentally changes the evaluation approach from source-oriented to receiver-oriented, improving accuracy while maintaining non-intrusive operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the evaluation parameters from phonation-based parameters (glottal pulse, vocal tract characteristics) to auditory perception parameters (temporal envelope, spectral characteristics, psychoacoustic features). This parameter transformation aligns the evaluation metrics with human perception, resolving the accuracy issue while keeping the model complexity manageable through selective feature extraction.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If subjective testing is used, then measurement precision is improved, but loss of time increases and productivity decreases

Engineering Contradiction:
Improvevoice quality evaluation accuracyVSAvoidtesting efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates an objective copy of the subjective evaluation process by modeling the human auditory system's perception mechanisms. Instead of requiring actual human listeners (subjective testing), the system uses computational models that replicate human auditory processing, including temporal envelope extraction, spectral analysis, and psychoacoustic feature computation. This copying approach maintains high measurement precision while eliminating time loss and productivity issues associated with organizing human testers.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent substitutes the mechanical process of organizing human testers, scheduling sessions, and collecting manual feedback with an automated computational system. The auditory modeling processing, time-frequency analysis, and feature extraction are performed automatically by algorithms, replacing the entire subjective testing workflow with an objective computational equivalent that delivers comparable accuracy without the organizational overhead.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10049674B2Method and apparatus for evaluating voice quality
Publication Date: 2018.08.14 HUAWEI TECH CO LTD
  • US10049674B2 patent drawing
  • US10049674B2 patent drawing
  • US10049674B2 patent drawing

AI summary

A method for evaluating voice quality includes performing human auditory modeling processing on a voice signal to obtain a first signal; performing variable resolution time-frequency analysis on the first signal to obtain a second signal; and performing, based on the second signal, feature extraction and analysis to obtain a voice quality evaluation result of the voice signal. According to the foregoing technical solutions, a problem that accuracy of a voice quality evaluation is not high can be solved. A voice quality evaluation result with relatively high accuracy is finally obtained by performing human auditory modeling processing, then converting a to-be-detected signal into a multi-resolution signal, further analyzing the time-frequency signal of variable resolution, extracting a feature corresponding to the signal, and performing further analysis.