Voice Quality Scoring via Spectrogram Similarity Across Wideband Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice quality evaluation methods, such as PESQ, are slow due to complex algorithms and limited sampling frequency bands, making it difficult to evaluate signals with wide frequency bands efficiently.
Innovation Solution
A method involving recording a playback of a standard audio to obtain a to-be-evaluated signal, determining a first spectrogram, and a second spectrogram, and an image similarity between the first spectrogram and the second spectrogram, and an image similarity between the first spectrogram and the second spectrogram, to obtain a voice quality score of the to-be-evaluated signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the PESQ method is used for voice quality evaluation, then the evaluation can handle network transmission issues such as time misalignment and spectral distortion, but the evaluation process becomes slow due to multiple iterations of alignment processing and complex parameter filtering algorithms
Solution Approach 1:
The patent extracts and removes the complex alignment processing and parameter filtering steps from the traditional PESQ method. Instead of performing multiple iterations of time alignment and spectral parameter filtering, the invention directly compares spectrograms of the test signal and reference signal, significantly reducing computational complexity while maintaining evaluation reliability for network transmission issues.
Solution Approach 2:
The patent uses spectrogram representation as a visual copy of the signal's time-frequency characteristics. By converting audio signals into spectrogram images and comparing these visual representations directly, the method avoids complex mathematical operations while preserving the essential features needed for quality evaluation, including spectral distortion and temporal misalignment detection.
2Reliability
If the PESQ method is used for voice quality evaluation, then the evaluation can assess spectral distortion, but the sampling frequency band is limited to less than 16 kHz, making it difficult to evaluate signals with wide frequency bands
Solution Approach 1:
The patent creates a universal evaluation method that works across different sampling rates by using spectrogram representation. The spectrogram technique naturally adapts to different frequency bands regardless of the original sampling rate, allowing the same comparison-based approach to evaluate both narrowband (less than 16 kHz) and wideband (greater than 16 kHz) signals effectively, thus achieving multi-functionality in terms of sampling rate compatibility.
3Measurement precision
If multiple iterations of alignment processing and complex parameter filtering algorithms are performed, then the voice quality evaluation can be comprehensive, but the evaluation time becomes long
Solution Approach 1:
The patent replaces the mechanical computational process of iterative alignment and filtering with a more efficient image-based comparison approach. By substituting complex signal processing operations with spectrogram image comparison, the method maintains measurement precision for voice quality assessment while dramatically reducing the computational time required, as image comparison operations are inherently more efficient than iterative mathematical optimization.
Data Source
AI summary
A method and an apparatus for evaluating voice quality are provided. In the method, a playback of a standard audio is recorded to obtain a to-be-evaluated signal. Then a first power spectrum of the to-be-evaluated signal on a critical frequency band is determined to obtain a first spectrogram. Then a second power spectrum of a reference signal corresponding to the standard audio on a critical frequency band is determined to obtain a second spectrogram, where the reference signal is a sampled signal corresponding to the standard audio. Then an image similarity between the first spectrogram and the second spectrogram is determined to obtain a voice quality score of the to-be-evaluated signal.


