Audio Codec Dialogue Compression for Clearer Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codec systems provide subtle dialogue enhancement, and existing solutions are not directly applicable or effective for improving dialogue clarity in audio encoding and decoding processes.
Innovation Solution
The proposed solution involves additional processing of an estimated dialogue component using compression and optional equalization on the decoder or encoder side, allowing for enhanced dialogue enhancement by increasing the signal-to-noise ratio and making the dialogue component more audible, while maintaining the peak level of the audio signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If dialogue enhancement parameters are applied as linear gains directly to the audio signal, then the dialogue enhancement effect is achieved, but the enhancement effect is too subtle and not perceptually effective
Solution Approach 1:
The patent changes the processing parameters from simple linear gains to dynamic compression and equalization parameters. The compression ratio, threshold, attack time, and release time are adjusted to enhance dialogue components selectively, while equalization parameters modify frequency characteristics to improve perceptual clarity. This transforms the enhancement from subtle linear scaling to pronounced perceptual improvement.
Solution Approach 2:
The system transitions from static linear gain application to dynamic processing where compression and equalization parameters adapt to the input signal characteristics. The compressor dynamically adjusts gain based on signal level, and the equalizer adapts frequency response based on content analysis, creating a more responsive and perceptually effective enhancement system.
2Measurement precision
If the average power of the dialogue component is increased to improve clarity, then dialogue enhancement is improved, but signal clipping may occur
Solution Approach 1:
The patent applies compression parameters (ratio, threshold, attack, release) to control the dynamic range of the dialogue component. By adjusting these parameters, the system increases average power through make-up gain while the compressor prevents peak levels from exceeding the clipping threshold, thus improving clarity without causing distortion.
Solution Approach 2:
The compressor acts as a protective mechanism that preemptively reduces peak levels before they can cause clipping. By applying compression before the signal reaches its maximum amplitude, the system prepares the signal to accommodate increased average power without risking distortion, effectively cushioning against potential harmful clipping effects.
3Measurement precision
If compression and equalization processing is applied to the estimated dialogue component, then dialogue enhancement is significantly improved, but the processing complexity increases
Solution Approach 1:
The patent separates the audio signal into estimated dialogue components and other components using dialogue estimation parameters. This segmentation allows independent processing of the dialogue portion with compression and equalization, while the rest of the signal remains unaffected. The separation simplifies the overall system by applying complex processing only where needed.
Solution Approach 2:
The estimated dialogue component acts as an intermediary that bridges the original audio signal and the enhanced output. The compression and equalization processing are applied to this intermediate representation rather than the full signal, reducing computational complexity while maintaining enhancement quality. The intermediary allows selective processing without requiring complex full-signal analysis.
Data Source
AI summary
Dialogue enhancement of an audio signal, comprising obtaining a set of time-varying parameters configured to estimate a dialogue component present in said audio signal, estimating the dialogue component from the audio signal, applying a compressor only to the estimated dialogue component, to generate a processed dialogue component, applying a user-determined gain to the processed dialogue component, to provide an enhanced dialogue component. The processing of the estimated dialogue may be performed on the decoder side or encoder side. The invention enables an improved dialogue enhancement.


