Dialogue Enhancement Adaptive Smoothing for Music-Linked False Positives

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional dialogue enhancement algorithms suffer from inaccuracies in speech classification, leading to "false positives" and fluctuations in dialogue boost, especially in the presence of music or low signal-to-noise ratio (SNR) content. Additionally, latency issues in speech detectors can result in missed initial utterances and abrupt dialogue boosts.

Innovation Solution

The implementation of an adaptive smoothing algorithm that adjusts the smoothing factor based on music confidence scores, SNR, and latency. This algorithm increases the impact of past frames for music-correlated content, reducing sensitivity to false positives, and adjusts the smoothing factor dynamically to optimize dialogue enhancement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a fixed smoothing factor is used in dialogue enhancement, then the processing is simple and fast, but the system is sensitive to false positives and produces fluctuating dialogue boost in music or low SNR environments

Engineering Contradiction:
Improvedialogue enhancement stabilityVSAvoidsmoothing algorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by making the smoothing factor adaptive rather than fixed. The smoothing factor dynamically adjusts based on music confidence scores and speech activity detection, allowing the system to handle different audio environments optimally. When music is detected, a larger smoothing factor reduces sensitivity to false positives, while in speech-only regions, a smaller factor preserves responsiveness.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of smoothing factor from a constant value to a variable that depends on music confidence scores and speech activity. This parameter adaptation allows the system to automatically adjust its behavior based on the detected audio content, resolving the contradiction between stability and complexity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a large smoothing factor is used, then false positives are reduced and dialogue boost becomes more stable, but the response to speech changes becomes slower and latency increases

Engineering Contradiction:
Improvefalse positive reductionVSAvoiddialogue enhancement latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The smoothing factor dynamically adapts based on the detected audio environment. In music-heavy regions, a larger smoothing factor is applied to reduce false positives. In speech-dominated regions, the smoothing factor decreases to improve response time. This dynamic adjustment resolves the time-stability tradeoff by optimizing for the current context.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the smoothing factor parameter based on music confidence scores and speech activity detection. This parameter modulation allows the system to achieve both false positive reduction and low latency by selecting appropriate smoothing levels for different audio scenarios.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If speech classification is made more sensitive to detect all speech, then speech detection accuracy improves, but false positives increase in music or low SNR conditions

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidfalse positives
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent introduces music confidence score as an intermediary parameter that mediates between speech detection sensitivity and false positive reduction. The music confidence score provides additional context that helps distinguish true speech from false positives in music or low SNR environments, allowing the system to maintain high detection accuracy while reducing erroneous detections.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system adjusts the smoothing factor parameter based on music confidence scores to compensate for false positives. When music is detected, the increased smoothing reduces the impact of false positive speech detections, allowing the speech classifier to operate with high sensitivity without suffering from increased false positives.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12272376B2Dialog enhancement using adaptive smoothing which depends exponentially on a smoothing factor
Publication Date: 2025.04.08 DOLBY LABORATORIES LICENSING CORP
  • US12272376B2 patent drawing
  • US12272376B2 patent drawing
  • US12272376B2 patent drawing

AI summary

A method of enhancing dialog intelligibility in an audio signal, comprising determining a speech confidence score that the audio content includes speech content, determining a music confidence score that the audio content includes music correlated content, in response to the speech confidence score, and applying a user selected gain of selected frequency bands of the audio signal to obtain a dialogue enhanced audio signal. The user selected gain is smoothed by an adaptive smoothing algorithm, an impact of past frames in said smoothing algorithm being determined by a smoothing factor, the smoothing factor being calculated in response to the music confidence score, and having a relatively higher value for content having a relatively higher music confidence score and a relatively lower value for speech content having a relatively lower music confidence score, so as to increase the impact of past frames on the dialogue enhancement of music correlated content.