Speech Signal Processing Circuit Spectral Balance Ratio Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current objective measures for assessing speech signal quality, particularly for artificial bandwidth extension (ABE)-processed signals, lack reliability and language dependency, as they fail to accurately predict subjective listening test scores and require time-consuming and costly subjective tests.
Innovation Solution
A speech-signal-processing circuit that calculates spectral balance ratio (SBR) features and other distortion metrics from time-frequency domain representations of reference and degraded signals, using a cognitive model to determine an output score indicative of perceived quality, eliminating the need for phonetic transcription and enhancing language independence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If subjective listening tests are used to assess speech signal quality, then measurement precision is improved, but loss of time and cost increase
Solution Approach 1:
The patent creates a computational model that copies and simulates human listening test evaluations. The system uses feature extraction and cognitive models to replicate subjective listening test results automatically, providing precise quality assessment without requiring actual human listeners. This allows the system to achieve measurement precision equivalent to subjective tests while eliminating the time and cost losses associated with human evaluation.
2Productivity
If existing objective measures are used for speech quality assessment, then productivity is improved, but reliability deteriorates due to language dependency and inaccuracy
Solution Approach 1:
The patent changes the parameters used in objective measures by incorporating spectral balance ratio (SBR) features and cognitive models that are specifically trained to handle multiple languages. The system adjusts feature extraction parameters and model weights to be language-independent, thereby improving reliability while maintaining productivity. The cognitive model processes various speech characteristics in a way that generalizes across different languages, eliminating the language dependency that plagues existing objective measures.
3Measurement precision
If phonetic transcription is required for quality assessment, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent extracts and processes only the essential spectral and temporal features from speech signals directly in the time-frequency domain, eliminating the need for phonetic transcription. The system extracts features such as spectral balance ratio, temporal characteristics, and spectral envelope information, which are sufficient for accurate quality assessment without requiring complex phonetic analysis. This extraction approach reduces device complexity while maintaining measurement precision.
4Adaptability or versatility
If artificial bandwidth extension algorithms are applied, then adaptability is improved for wideband speech, but object-generated harmful factors increase due to signal degradation
Solution Approach 1:
The patent implements a feedback mechanism that uses the extracted spectral and temporal features to evaluate and optimize the bandwidth extension process. The cognitive model provides feedback on the quality of extended bandwidth speech, allowing the system to identify and correct distortions introduced by the extension algorithm. This feedback loop enables the system to maintain adaptability for wideband speech while minimizing harmful signal degradation through continuous monitoring and adjustment.
Data Source
AI summary
A speech-signal-processing-circuit configured to receive a time-frequency-domain-reference-speech-signal and a time-frequency-domain-degraded-speech-signal. The time-frequency-domain-reference-speech-signal comprises: an upper-band-reference-component with frequencies that are greater than a frequency-threshold-value; and a lower-band-reference-component with frequencies that are less than the frequency-threshold-value. The time-frequency-domain-degraded-speech-signal comprises: an upper-band-degraded-component with frequencies that are greater than the frequency-threshold-value; and a lower-band-degraded-component with frequencies that are less than the frequency-threshold-value. The speech-signal-processing-circuit comprises: a disturbance calculator configured to determine one or more SBR-features based on the time-frequency-domain-reference-speech-signal and the time-frequency-domain-degraded-speech-signal by: for each of a plurality of frames: determining a reference-ratio based on the ratio of (i) the upper-band-reference-component to (ii) the lower-band-reference-component; determining a degraded-ratio based on the ratio of (i) the upper-band-degraded-component to (ii) the lower-band-degraded-component; and determining a spectral-balance-ratio based on the ratio of the reference-ratio to the degraded-ratio; and (ii) determining the one or more SBR-features based on the spectral-balance-ratio for the plurality of frames.


