Speech Quality Estimation via Signal Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech quality estimation methods face challenges in accurately compensating for time and intensity differences between reference and test speech signals, leading to incorrect matching and inadequate handling of speech interruptions, which affects the accuracy of perceived speech quality assessment.
Innovation Solution
A method that aligns signal parts of the reference and test speech signals by matching those with similar lengths and intensities, computes a performance measure for each pair, and re-matches poorly performing pairs, while identifying perceptually dominant frequency sub-bands and scaling the test speech spectrum to align intensity differences, and computes a combined measure of distortion for interrupted signal parts to improve speech quality estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional alignment procedures are used to compensate time and intensity differences, then the spectral representations can be compared, but incorrect matching of signal parts occurs leading to reduced measurement precision
Solution Approach 1:
The patent applies preliminary action by performing multiple alignment attempts before final comparison. The method tries different alignment hypotheses (different delay values and intensity compensations) in advance, evaluates their performance, and selects the best alignment before computing the final spectral difference. This ensures correct matching of signal parts before quality assessment.
Solution Approach 2:
The patent uses feedback by computing a performance measure (such as correlation or similarity metric) for each alignment attempt and using this feedback to select the optimal alignment. The alignment that maximizes the performance measure is chosen, ensuring that the subsequent spectral comparison is based on correctly matched signal parts.
2Measurement precision
If simple intensity alignment is applied, then processing is simplified, but speech interruptions are not properly handled reducing measurement precision
Solution Approach 1:
The patent applies segmentation by dividing the speech signal into multiple segments or frames and processing each segment separately. This allows the method to handle speech interruptions by identifying segments with low energy or high distortion and treating them differently from normal speech segments. The overall quality metric is computed by combining results from all segments.
Solution Approach 2:
The patent uses dynamics by making the alignment procedure adaptive rather than static. The method dynamically adjusts alignment parameters based on local signal characteristics, such as energy levels and spectral content. This allows the system to handle varying conditions including speech interruptions by adapting the alignment strategy to each local segment.
3Measurement precision
If spectral difference is computed directly without sophisticated alignment, then processing is faster, but time-varying intensity differences are not compensated reducing measurement precision
Solution Approach 1:
The patent applies parameter changes by adjusting intensity parameters (gain factors) for different segments of the signal to compensate for time-varying intensity differences. The method computes optimal intensity scaling factors that minimize the spectral difference between aligned signals, thereby accurately compensating for intensity variations without requiring complex temporal alignment.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
The invention relates to a method for estimating speech quality, wherein a reference speech signal (301) enters a telecommunication network resulting in a test speech signal (302) and wherein the method comprises the following steps of aligning the reference speech signal (301) and the test speech signal (302) by matching signal parts of the reference speech signal (301) with signal parts of the test speech signal (302), wherein matched signal parts (303, 304) are of similar length in the time domain and have similar intensity summed over their length, and computing and comparing the speech spectra of the reference speech signal (301) and the test speech signal (302) that are aligned, resulting in a difference measure, the difference measure being indicative of the speech quality of the test speech signal.