Dialog Quality Estimation from Mixed Audio Without Clean Reference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dialog quality metering methods require a clean dialog reference signal, which is not always available, leading to reduced flexibility and efficiency in estimating dialog quality in noisy environments.

Innovation Solution

A method and system for training a dialog separator using a dialog separation model to estimate dialog quality metrics by minimizing a loss function based on the difference between estimated and reference values, allowing for dialog quality estimation in mixed audio signals without a separate clean dialog reference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional quality metering methods are used that compare clean dialog and noisy dialog, then measurement precision of dialog quality is improved, but device complexity increases and ease of operation deteriorates due to requiring separate clean reference signals

Engineering Contradiction:
Improvedialog quality measurement precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary functional components from the traditional two-signal comparison method. Instead of requiring both clean and noisy signals as separate inputs, the system processes only the noisy signal through a trained neural network that internally performs the separation and quality assessment functions, thereby reducing system complexity while maintaining measurement precision

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a learned copy of the clean dialog signal through the neural network's estimated dialog output. This estimated dialog serves as a virtual reference that replicates the quality assessment properties of a true clean signal without requiring physical access to the original clean recording, enabling quality measurement from noisy input alone

Inventive Principle:
Principle #26Copying

2Measurement precision

If traditional quality metering methods are used that require clean dialog reference, then measurement precision is improved, but adaptability deteriorates because clean reference signals are not always available

Engineering Contradiction:
Improvedialog quality measurement precisionVSAvoidflexibility in noisy environments
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs self-service by using the noisy signal itself as the input for quality assessment. The neural network is trained to extract quality metrics directly from the noisy signal without external clean reference assistance, enabling the system to serve itself in environments where clean references are unavailable, thus improving adaptability while maintaining precision

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the operational parameters of the quality assessment system from requiring two distinct signal inputs (clean and noisy) to processing a single noisy signal input. This parameter change in the input signal configuration, enabled by the trained neural network model, allows the system to adapt to real-world conditions where clean references are absent while preserving measurement precision through learned separation

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If dialog separation model is trained to minimize loss function based on quality metric difference, then manufacturing precision of dialog separation is improved, but use of energy increases due to iterative training process

Engineering Contradiction:
Improvedialog separation accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-training the dialog separation model offline using large datasets and iterative optimization. This preliminary training phase performs the computationally intensive energy-consuming operations in advance, allowing the deployed system to achieve high separation accuracy with minimal real-time energy consumption during actual quality assessment operations

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4275206B1Determining dialog quality metrics of a mixed audio signal
Publication Date: 2025.12.31 DOLBY LABORATORIES LICENSING CORP
  • EP4275206B1 patent drawingFigure 1
  • EP4275206B1 patent drawingFigure 2
  • EP4275206B1 patent drawingFigure 3

AI summary

Disclosed is a method for determining one or more dialog quality metrics of a mixed audio signal comprising a dialog component and a noise component, the method comprising separating an estimated dialog component from the mixed audio signal by means of a dialog separator using a dialog separating model determined by training the dialog separator based on the one or more quality metrics; providing the estimated dialog component from the dialog separator to a quality metrics estimator; and determining the one or more quality metrics by means of the quality metrics estimator based on the mixed signal and the estimated dialog component. Further disclosed is a method for training a dialog separator, a system comprising circuitry configured to perform the method, and a non-transitory computer-readable storage medium.