Neural Network Speech Quality Assessment Without Reference Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for evaluating speech quality in audio and video signals, such as VoIP, are limited in their ability to accurately predict subjective quality across a wide range of conditions, including coding distortions, noise, and variable delays, especially for wide-band audio signals.
Innovation Solution
An apparatus and method utilizing a neural network trained with specific transmission standards and coding data to generate score signals representing speech quality, which can produce scores according to ITU-T methods like PESQ, PEAQ, or POLQA without requiring a reference signal, using supervised learning and objective analytic quality testing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If classical quality measurement techniques (signal-to-noise ratio) are used, then the measurement method is simple, but the accuracy of speech quality prediction is insufficient for new technologies like VoIP
Solution Approach 1:
The patent transforms the measurement approach by changing from classical physical parameters (signal-to-noise ratio) to perceptual parameters that model human speech quality perception. This includes using spectral distortion measures, temporal distortion measures, and other parameters that better reflect how humans perceive speech quality in modern communication systems with coding distortions, packet loss, and variable delays.
Solution Approach 2:
The patent replaces classical mechanical/electrical measurement systems with a perceptual model that simulates human auditory processing. This includes implementing models of human speech perception, auditory filtering, and quality assessment that better match subjective human evaluations rather than relying on traditional engineering metrics.
2Measurement precision
If double-ended objective models are used, then the speech quality assessment accuracy is improved, but the requirement for reference signals increases system complexity
Solution Approach 1:
The patent extracts and removes the requirement for reference signals from the measurement system. By developing single-ended objective models that assess speech quality using only the degraded signal, the system eliminates the need for complex reference signal management, synchronization, and comparison infrastructure while maintaining assessment accuracy through alternative perceptual modeling approaches.
Solution Approach 2:
The patent enables the degraded signal to assess its own quality without external reference. The single-ended objective models analyze the degraded signal's inherent characteristics, spectral properties, and temporal features to determine speech quality, allowing the system to be self-sufficient and avoid the complexity of requiring pristine reference signals from the original source.
3Measurement precision
If human/subjective tests are used to produce LQS scores, then the speech quality evaluation reflects actual human perception, but the testing process is time-consuming and requires human involvement
Solution Approach 1:
The patent creates objective models that copy and simulate human subjective quality assessment behavior. By training objective models on large datasets of paired subjective evaluations and objective measurements, the system learns to replicate human perception patterns, producing scores that closely match subjective LQS evaluations but can be generated automatically without human test participants.
Solution Approach 2:
The patent replaces the human testing system with an automated computational model. Instead of requiring human listeners to evaluate speech quality in controlled experiments, the system uses trained objective models that automatically predict subjective quality scores from signal characteristics, eliminating the need for human involvement while maintaining the perceptual accuracy of subjective tests.
4Adaptability or versatility
If existing objective models (PESQ, PEAQ) are used, then the assessment covers a wide range of conditions, but the models are limited to specific bandwidths and coding types
Solution Approach 1:
The patent develops universal objective models that can assess speech quality across multiple bandwidths (narrow-band, wide-band, ultra-wide-band) and various coding types (codec standards, packet loss conditions, noise environments). The models are designed to be bandwidth-agnostic and coding-agnostic, providing consistent accurate assessments regardless of the specific transmission conditions or signal characteristics.
Solution Approach 2:
The patent implements dynamic models that adapt to different transmission conditions and signal characteristics. The objective models can dynamically adjust their analysis parameters, frequency resolution, and processing strategies based on the input signal properties and transmission conditions, providing optimized assessment for each specific scenario rather than using fixed parameters.
Data Source
AI summary
An apparatus for generating a score signal representing the quality of an audio or video signal supplied to the apparatus is proposed. The apparatus comprises: an input for supplying an audio or video signal, a computing unit implementing a neural network, the computing unit being supplied with the audio or video signal, and producing a score signal representing the quality of an audio or video signal supplied representing at least one predefined quality parameter of the audio or video signal, the neural network being set up by being trained with training data of a specific transmission standard and/or codec used for generating the audio or video data.


