Dialogue Component Compression for Clearer Audio Codec Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codec systems provide subtle dialogue enhancement, and existing solutions are not directly applicable or effective for improving dialogue clarity in audio signals.
Innovation Solution
The proposed solution involves additional processing of estimated dialogue components using compression and equalization, where dialogue enhancement parameters are used to estimate and enhance the dialogue component separately from the audio signal, increasing its signal-to-noise ratio and making it more audible by applying a compressor and user-determined gain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If dialogue enhancement parameters are applied as linear gains directly to the audio signal, then the dialogue enhancement is subtle, but the processing complexity is low
Solution Approach 1:
The audio signal is segmented into dialogue components and non-dialogue components using dialogue enhancement parameters. This segmentation allows separate processing of the dialogue portion, enabling application of compression and equalization specifically to enhance dialogue clarity without processing the entire audio signal, thus resolving the contradiction between dialogue clarity and processing complexity
Solution Approach 2:
The dialogue components are estimated and separated from the audio signal before applying enhancement processing. This preliminary action of component separation enables subsequent compression and equalization to be applied more effectively to the dialogue portion, improving dialogue clarity while maintaining reasonable processing complexity through targeted processing
2Power
If compression is applied to the estimated dialogue component, then the average power of dialogue increases, but the processing complexity increases
Solution Approach 1:
Compression is applied locally only to the estimated dialogue component rather than the entire audio signal. This local processing approach increases the average power of dialogue components specifically, while limiting the overall processing complexity by avoiding global processing of all audio content
Solution Approach 2:
Compression parameters are applied to the dialogue component to change its power characteristics. By modifying the compression ratio, threshold, and attack/release parameters specifically for the dialogue portion, the average power of dialogue is increased while the processing complexity remains manageable through parameter optimization
Data Source
Figure 1~3
Figure 4~6
Figure 7a~7b
AI summary
Dialogue enhancement of an audio signal, comprising obtaining a set of time-varying parameters configured to estimate a dialogue component present in said audio signal, estimating the dialogue component from the audio signal, applying a compressor only to the estimated dialogue component, to generate a processed dialogue component, applying a user-determined gain to the processed dialogue component, to provide an enhanced dialogue component. The processing of the estimated dialogue may be performed on the decoder side or encoder side. The invention enables an improved dialogue enhancement.