Neural Network Speech Quality Assessment Without Reference Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for evaluating speech quality in audio and video signals, such as VoIP, are limited in their ability to accurately predict subjective quality across a wide range of conditions, including coding distortions, noise, and variable delays, especially for wide-band audio signals.

Innovation Solution

An apparatus and method utilizing a neural network trained with specific transmission standards and coding data to generate score signals representing speech quality, which can produce scores according to ITU-T methods like PESQ, PEAQ, or POLQA without requiring a reference signal, using supervised learning and objective analytic quality testing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If classical quality measurement techniques (signal-to-noise ratio) are used, then the measurement method is simple, but the accuracy of speech quality prediction is insufficient for new technologies like VoIP

Engineering Contradiction:
Improvespeech quality prediction accuracyVSAvoidmeasurement system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the measurement approach by changing from classical physical parameters (signal-to-noise ratio) to perceptual parameters that model human speech quality perception. This includes using spectral distortion measures, temporal distortion measures, and other parameters that better reflect how humans perceive speech quality in modern communication systems with coding distortions, packet loss, and variable delays.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces classical mechanical/electrical measurement systems with a perceptual model that simulates human auditory processing. This includes implementing models of human speech perception, auditory filtering, and quality assessment that better match subjective human evaluations rather than relying on traditional engineering metrics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If double-ended objective models are used, then the speech quality assessment accuracy is improved, but the requirement for reference signals increases system complexity

Engineering Contradiction:
Improvespeech quality assessment accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the requirement for reference signals from the measurement system. By developing single-ended objective models that assess speech quality using only the degraded signal, the system eliminates the need for complex reference signal management, synchronization, and comparison infrastructure while maintaining assessment accuracy through alternative perceptual modeling approaches.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent enables the degraded signal to assess its own quality without external reference. The single-ended objective models analyze the degraded signal's inherent characteristics, spectral properties, and temporal features to determine speech quality, allowing the system to be self-sufficient and avoid the complexity of requiring pristine reference signals from the original source.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If human/subjective tests are used to produce LQS scores, then the speech quality evaluation reflects actual human perception, but the testing process is time-consuming and requires human involvement

Engineering Contradiction:
Improvesubjective quality reflectionVSAvoidtesting efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates objective models that copy and simulate human subjective quality assessment behavior. By training objective models on large datasets of paired subjective evaluations and objective measurements, the system learns to replicate human perception patterns, producing scores that closely match subjective LQS evaluations but can be generated automatically without human test participants.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the human testing system with an automated computational model. Instead of requiring human listeners to evaluate speech quality in controlled experiments, the system uses trained objective models that automatically predict subjective quality scores from signal characteristics, eliminating the need for human involvement while maintaining the perceptual accuracy of subjective tests.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If existing objective models (PESQ, PEAQ) are used, then the assessment covers a wide range of conditions, but the models are limited to specific bandwidths and coding types

Engineering Contradiction:
Improvecoverage of transmission conditionsVSAvoidprediction accuracy for wide-band signals
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent develops universal objective models that can assess speech quality across multiple bandwidths (narrow-band, wide-band, ultra-wide-band) and various coding types (codec standards, packet loss conditions, noise environments). The models are designed to be bandwidth-agnostic and coding-agnostic, providing consistent accurate assessments regardless of the specific transmission conditions or signal characteristics.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamic models that adapt to different transmission conditions and signal characteristics. The objective models can dynamically adjust their analysis parameters, frequency resolution, and processing strategies based on the input signal properties and transmission conditions, providing optimized assessment for each specific scenario rather than using fixed parameters.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11322173B2Evaluation of speech quality in audio or video signals
Publication Date: 2022.05.03 ROHDE & SCHWARZ GMBH & CO KG
  • US11322173B2 patent drawing
  • US11322173B2 patent drawing
  • US11322173B2 patent drawing

AI summary

An apparatus for generating a score signal representing the quality of an audio or video signal supplied to the apparatus is proposed. The apparatus comprises: an input for supplying an audio or video signal, a computing unit implementing a neural network, the computing unit being supplied with the audio or video signal, and producing a score signal representing the quality of an audio or video signal supplied representing at least one predefined quality parameter of the audio or video signal, the neural network being set up by being trained with training data of a specific transmission standard and/or codec used for generating the audio or video data.