Dialogue Enhancement Adaptive Smoothing for Music-Linked False Positives
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional dialogue enhancement algorithms suffer from inaccuracies in speech classification, leading to "false positives" and fluctuations in dialogue boost, especially in the presence of music or low signal-to-noise ratio (SNR) content. Additionally, latency issues in speech detectors can result in missed initial utterances and abrupt dialogue boosts.
Innovation Solution
The implementation of an adaptive smoothing algorithm that adjusts the smoothing factor based on music confidence scores, SNR, and latency. This algorithm increases the impact of past frames for music-correlated content, reducing sensitivity to false positives, and adjusts the smoothing factor dynamically to optimize dialogue enhancement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a fixed smoothing factor is used in dialogue enhancement, then the processing is simple and fast, but the system is sensitive to false positives and produces fluctuating dialogue boost in music or low SNR environments
Solution Approach 1:
The patent applies dynamics by making the smoothing factor adaptive rather than fixed. The smoothing factor dynamically adjusts based on music confidence scores and speech activity detection, allowing the system to handle different audio environments optimally. When music is detected, a larger smoothing factor reduces sensitivity to false positives, while in speech-only regions, a smaller factor preserves responsiveness.
Solution Approach 2:
The patent changes the parameter of smoothing factor from a constant value to a variable that depends on music confidence scores and speech activity. This parameter adaptation allows the system to automatically adjust its behavior based on the detected audio content, resolving the contradiction between stability and complexity.
2Reliability
If a large smoothing factor is used, then false positives are reduced and dialogue boost becomes more stable, but the response to speech changes becomes slower and latency increases
Solution Approach 1:
The smoothing factor dynamically adapts based on the detected audio environment. In music-heavy regions, a larger smoothing factor is applied to reduce false positives. In speech-dominated regions, the smoothing factor decreases to improve response time. This dynamic adjustment resolves the time-stability tradeoff by optimizing for the current context.
Solution Approach 2:
The patent changes the smoothing factor parameter based on music confidence scores and speech activity detection. This parameter modulation allows the system to achieve both false positive reduction and low latency by selecting appropriate smoothing levels for different audio scenarios.
3Measurement precision
If speech classification is made more sensitive to detect all speech, then speech detection accuracy improves, but false positives increase in music or low SNR conditions
Solution Approach 1:
The patent introduces music confidence score as an intermediary parameter that mediates between speech detection sensitivity and false positive reduction. The music confidence score provides additional context that helps distinguish true speech from false positives in music or low SNR environments, allowing the system to maintain high detection accuracy while reducing erroneous detections.
Solution Approach 2:
The system adjusts the smoothing factor parameter based on music confidence scores to compensate for false positives. When music is detected, the increased smoothing reduces the impact of false positive speech detections, allowing the speech classifier to operate with high sensitivity without suffering from increased false positives.
Data Source
AI summary
A method of enhancing dialog intelligibility in an audio signal, comprising determining a speech confidence score that the audio content includes speech content, determining a music confidence score that the audio content includes music correlated content, in response to the speech confidence score, and applying a user selected gain of selected frequency bands of the audio signal to obtain a dialogue enhanced audio signal. The user selected gain is smoothed by an adaptive smoothing algorithm, an impact of past frames in said smoothing algorithm being determined by a smoothing factor, the smoothing factor being calculated in response to the music confidence score, and having a relatively higher value for content having a relatively higher music confidence score and a relatively lower value for speech content having a relatively lower music confidence score, so as to increase the impact of past frames on the dialogue enhancement of music correlated content.


