Single-Sided Speech Quality Measurement Using PLP and GMM
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current non-intrusive speech quality measurement techniques are computationally intensive, making them impractical for widespread deployment outside of test facilities, despite their potential for real-time network-wide assessment without the need for clean speech signals.
Innovation Solution
A single-ended speech quality measurement method that extracts perceptual features from received speech signals using a feature extraction module, assesses these features with statistical models to form indicators of quality, and produces a speech quality score, reducing processing requirements while maintaining performance through the use of Perceptual Linear Prediction (PLP) coefficients and Gaussian Mixture Models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If non-intrusive speech quality measurement techniques are used, then real-time network-wide assessment without clean speech signals is achieved, but computational complexity increases making deployment impractical
Solution Approach 1:
The patent extracts only the essential perceptual features from speech signals that are necessary for quality assessment, rather than processing the entire signal. By identifying and extracting specific features (spectral characteristics, temporal envelope, zero-crossing rate) that correlate with perceived quality, the system achieves accurate measurement with reduced computational load, enabling practical deployment
Solution Approach 2:
The patent transforms the speech signal into a different parameter space using perceptual linear prediction (PLP) coefficients and other perceptual transformations. This parameter transformation simplifies the complexity of raw speech signals while preserving quality-relevant information, making the measurement computationally tractable for real-time network-wide deployment
2Measurement precision
If traditional vector quantization and hidden Markov models are used for non-intrusive measurement, then quality estimation accuracy is improved, but processing time increases by up to 40%
Solution Approach 1:
The patent applies partial action by using simplified statistical models (Gaussian distributions) instead of complex hidden Markov models for certain aspects of the analysis. By applying the appropriate level of model complexity only where necessary and using simpler models elsewhere, the system maintains adequate measurement precision while significantly reducing processing time
Solution Approach 2:
The patent segments the speech signal into short frames and processes each frame independently using parallel computations. This segmentation allows for efficient batch processing and reduces the computational burden compared to analyzing the entire signal sequence, thereby reducing processing time while maintaining accuracy through sufficient frame sampling
3Productivity
If intrusive double-ended measurement is used, then processing requirements are reduced, but requirement for clean transmitted speech signal becomes problematic in working networks
Solution Approach 1:
The patent enables the measurement system to serve itself by extracting all necessary quality assessment information from the received speech signal alone, without requiring the original transmitted signal. The system uses inherent properties of the received signal (spectral characteristics, temporal features, perceptual parameters) to assess quality, making it adaptable to working networks where only the received signal is available
Data Source
AI summary
A non-intrusive speech quality estimation technique is based on statistical or probability models such as Gaussian Mixture Models (“GMMs”). Perceptual features are extracted from the received speech signal and assessed by an artificial reference model formed using statistical models. The models characterize the statistical behavior of speech features. Consistency measures between the input speech features and the models are calculated to form indicators of speech quality. The consistency values are mapped to a speech quality score using a mapping optimized using machine learning algorithms, such as Multivariate Adaptive Regression Splines (“MARS”). The technique provides competitive or better quality estimates relative to known techniques while having lower computational complexity.


