Adaptive Dialog Enhancement Smoothing for Music-Correlated Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional dialog enhancement algorithms face challenges with accuracy and latency issues, particularly in distinguishing speech from music and low-SNR content, leading to false positives and abrupt dialog boosts, which are exacerbated by fixed smoothing factors that fail to adapt to varying content and context.

Innovation Solution

An adaptive smoothing algorithm dynamically adjusts the smoothing factor based on music confidence scores, signal-to-noise ratio, and latency to reduce false positives and ensure a natural dialog boost, using a higher smoothing factor for music-correlated content and a lower factor for pure speech, thereby enhancing dialog intelligibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a fixed smoothing factor is used in dialog enhancement, then the processing is simple and fast, but it causes false positives in music content and abrupt dialog boosts

Engineering Contradiction:
Improveprocessing speedVSAvoidaccuracy of speech detection
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies dynamics by making the smoothing factor adaptive rather than fixed. The smoothing factor dynamically changes based on the detected audio content type (speech, music, or other) and confidence levels. This allows the system to optimize processing for each content type: using lower smoothing factors for speech to maintain responsiveness, and higher smoothing factors for music to reduce false positives and artifacts.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of smoothing factor based on content analysis. By analyzing audio content and confidence scores, the system adjusts the smoothing factor parameter to match the current content type. This parameter adaptation resolves the contradiction by allowing fast processing for clear speech while preventing artifacts in music through higher smoothing factors when music is detected.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a higher smoothing factor is used to reduce false positives in music, then artifacts are reduced, but dialog enhancement becomes less responsive to actual speech changes

Engineering Contradiction:
Improvereduction of false positivesVSAvoidresponsiveness to speech changes
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The system dynamically adjusts the smoothing factor based on real-time content detection. When speech is detected with high confidence, a lower smoothing factor is applied to maintain responsiveness. When music or uncertain content is detected, a higher smoothing factor is applied to reduce false positives. This dynamic adaptation allows the system to optimize both reliability and speed for different content types.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The smoothing factor parameter is changed based on content type and confidence analysis. The system selects from multiple smoothing factor values depending on whether the current content is speech, music, or other, and based on the confidence score. This parameter switching enables the system to achieve high responsiveness for speech while maintaining low false positive rates for music.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If conventional dialog enhancement is applied without content-based adaptation, then the processing is simple, but it produces abrupt dialog boosts and unwanted artifacts

Engineering Contradiction:
Improvealgorithm simplicityVSAvoidquality of dialog enhancement
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the audio processing based on content type. Instead of applying a single uniform enhancement algorithm, the system divides processing into different paths based on detected content (speech, music, other) and confidence levels. This segmentation allows simple processing for clear speech cases while applying more sophisticated adaptive processing only when needed, thus maintaining overall simplicity while improving quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adapts processing parameters based on content analysis. By changing the smoothing factor and gain application based on detected content type and confidence, the system improves dialog enhancement quality without requiring completely complex algorithms. The adaptation is achieved through parameter modification rather than fundamental algorithmic changes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3803861B1Dialog enhancement using adaptive smoothing
Publication Date: 2022.01.19 DOLBY LABORATORIES LICENSING CORP
  • EP3803861B1 patent drawingFigure 1
  • EP3803861B1 patent drawingFigure 2
  • EP3803861B1 patent drawingFigure 3

AI summary

A method of enhancing dialog intelligibility in an audio signal, comprising determining a speech confidence score that the audio content includes speech content, determining a music confidence score that the audio content includes music correlated content, in response to the speech confidence score, and applying a user selected gain of selected frequency bands of the audio signal to obtain a dialogue enhanced audio signal. The user selected gain is smoothed by an adaptive smoothing algorithm, an impact of past frames in said smoothing algorithm being determined by a smoothing factor, the smoothing factor being calculated in response to the music confidence score, and having a relatively higher value for content having a relatively higher music confidence score and a relatively lower value for speech content having a relatively lower music confidence score, so as to increase the impact of past frames on the dialogue enhancement of music correlated content.