Harmonic-to-Non-Harmonic Ratio for Single-Channel Speech Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack a method to effectively measure reverberance in speech communication systems using only observed speech signals from a single audio channel, which is crucial for assessing speech quality in various environments with background noise.
Innovation Solution
The method employs the harmonic to non-harmonic ratio (HnHR) to estimate speech quality by decomposing the observed signal into harmonic and non-harmonic components, using the room impulse response and acoustic parameters like reverberation time and clarity, and provides feedback to users when the quality falls below a threshold.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional speech quality evaluation methods are used, then they require clean reference signals for comparison, but this makes them inapplicable to single-channel audio systems without reference signals
Solution Approach 1:
The patent applies self-service by enabling the speech quality estimator to evaluate quality using only the observed speech signal itself, without requiring external reference signals. The system processes the single-channel audio signal through harmonic-to-noise ratio analysis, allowing the signal to serve its own evaluation purpose independently.
Solution Approach 2:
The patent extracts the harmonic components from the observed speech signal and separates them from noise components. By computing the harmonic-to-noise ratio through spectral analysis and temporal envelope extraction, the system isolates quality-relevant features from the mixed signal to enable evaluation without reference signals.
2Measurement precision
If harmonicity-based methods are used to estimate speech quality, then they can work with observed signals alone, but they struggle to distinguish between harmonicity caused by speech and harmonicity caused by reverberation
Solution Approach 1:
The patent applies dynamics by analyzing the temporal evolution of harmonic components. The system computes the temporal envelope of harmonic amplitudes and compares it against expected speech patterns, allowing dynamic differentiation between harmonicity from voiced speech (which follows glottal pulse patterns) and harmonicity from reverberation (which decays exponentially).
Solution Approach 2:
The patent applies preliminary action by first estimating the fundamental frequency and identifying harmonic frequencies before analyzing the temporal envelope. This preliminary spectral analysis establishes a foundation for subsequent temporal characterization, enabling the system to distinguish speech-related harmonics from reverberation-related harmonics through their different temporal behaviors.
3Ease of operation
If the system provides continuous feedback on speech quality, then users can adjust their speech to improve quality, but this requires real-time processing of audio frames
Solution Approach 1:
The patent applies partial action by processing only the necessary features for quality estimation rather than analyzing the entire audio signal in detail. The system extracts fundamental frequency, identifies harmonic frequencies, and computes temporal envelopes only at these specific frequencies, reducing computational load while maintaining estimation accuracy for real-time feedback.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Speech quality estimation technique embodiments are described which generally involve estimating the human speech quality of an audio frame in a single-channel audio signal. A representation of a harmonic component of the frame is synthesized and used to compute a non-harmonic component of the frame. The synthesized harmonic component representation and the non-harmonic component are then used to compute a harmonic to non-harmonic ratio (HnHR). This HnHR is indicative of the quality of a user's speech and is designated as an estimate of the speech quality of the frame. In one implementation, the HnHR is used to establish a minimum speech quality threshold below which the quality of the user's speech is considered unacceptable. Feedback to the user is then provided based on whether the HnHR falls below the threshold.