Adaptive Dialog Enhancement Smoothing for Music-Correlated Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional dialog enhancement algorithms face challenges with accuracy and latency issues, particularly in distinguishing speech from music and low-SNR content, leading to false positives and abrupt dialog boosts, which are exacerbated by fixed smoothing factors that fail to adapt to varying content and context.
Innovation Solution
An adaptive smoothing algorithm dynamically adjusts the smoothing factor based on music confidence scores, signal-to-noise ratio, and latency to reduce false positives and ensure a natural dialog boost, using a higher smoothing factor for music-correlated content and a lower factor for pure speech, thereby enhancing dialog intelligibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a fixed smoothing factor is used in dialog enhancement, then the processing is simple and fast, but it causes false positives in music content and abrupt dialog boosts
Solution Approach 1:
The patent applies dynamics by making the smoothing factor adaptive rather than fixed. The smoothing factor dynamically changes based on the detected audio content type (speech, music, or other) and confidence levels. This allows the system to optimize processing for each content type: using lower smoothing factors for speech to maintain responsiveness, and higher smoothing factors for music to reduce false positives and artifacts.
Solution Approach 2:
The patent changes the parameter of smoothing factor based on content analysis. By analyzing audio content and confidence scores, the system adjusts the smoothing factor parameter to match the current content type. This parameter adaptation resolves the contradiction by allowing fast processing for clear speech while preventing artifacts in music through higher smoothing factors when music is detected.
2Reliability
If a higher smoothing factor is used to reduce false positives in music, then artifacts are reduced, but dialog enhancement becomes less responsive to actual speech changes
Solution Approach 1:
The system dynamically adjusts the smoothing factor based on real-time content detection. When speech is detected with high confidence, a lower smoothing factor is applied to maintain responsiveness. When music or uncertain content is detected, a higher smoothing factor is applied to reduce false positives. This dynamic adaptation allows the system to optimize both reliability and speed for different content types.
Solution Approach 2:
The smoothing factor parameter is changed based on content type and confidence analysis. The system selects from multiple smoothing factor values depending on whether the current content is speech, music, or other, and based on the confidence score. This parameter switching enables the system to achieve high responsiveness for speech while maintaining low false positive rates for music.
3Device complexity
If conventional dialog enhancement is applied without content-based adaptation, then the processing is simple, but it produces abrupt dialog boosts and unwanted artifacts
Solution Approach 1:
The patent segments the audio processing based on content type. Instead of applying a single uniform enhancement algorithm, the system divides processing into different paths based on detected content (speech, music, other) and confidence levels. This segmentation allows simple processing for clear speech cases while applying more sophisticated adaptive processing only when needed, thus maintaining overall simplicity while improving quality.
Solution Approach 2:
The system adapts processing parameters based on content analysis. By changing the smoothing factor and gain application based on detected content type and confidence, the system improves dialog enhancement quality without requiring completely complex algorithms. The adaptation is achieved through parameter modification rather than fundamental algorithmic changes.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method of enhancing dialog intelligibility in an audio signal, comprising determining a speech confidence score that the audio content includes speech content, determining a music confidence score that the audio content includes music correlated content, in response to the speech confidence score, and applying a user selected gain of selected frequency bands of the audio signal to obtain a dialogue enhanced audio signal. The user selected gain is smoothed by an adaptive smoothing algorithm, an impact of past frames in said smoothing algorithm being determined by a smoothing factor, the smoothing factor being calculated in response to the music confidence score, and having a relatively higher value for content having a relatively higher music confidence score and a relatively lower value for speech content having a relatively lower music confidence score, so as to increase the impact of past frames on the dialogue enhancement of music correlated content.