Audio Encoding Downmix Gain Compensation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multi-channel audio signals experience sound level loss when downmixed to mono or stereo, leading to degradation in sound quality due to limited bit size and high sound levels, especially when a large number of channels are downmixed.
Innovation Solution
Applying an arbitrary downmix gain to the downmix signal to modify its sound level, using spatial information such as channel level difference, inter-channel coherence, and channel time difference to maintain audio quality during encoding and decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a multi-channel audio signal is downmixed to mono or stereo, then the signal can be encoded with limited bit size, but sound level loss occurs and sound quality degrades
Solution Approach 1:
The patent applies preliminary action by calculating and storing the downmix gain value before the downmixing process. This pre-calculated gain information is then used to compensate for sound level loss during or after downmixing, preventing quality degradation before it occurs rather than correcting it afterward.
Solution Approach 2:
The patent changes the parameter of gain compensation by introducing a downmix gain value that is calculated based on the relationship between original multi-channel signals and downmixed signals. This parameter is applied to adjust the amplitude of downmixed channels, thereby maintaining sound quality while enabling compact encoding.
2Productivity
If downmixing is applied to reduce channel count, then data compression is achieved, but spatial information and sound level are lost
Solution Approach 1:
The patent implements feedback by calculating the downmix gain based on the actual downmixed signal characteristics and using this calculated gain to adjust the downmixed signal. This closed-loop approach ensures that spatial information is preserved by continuously adapting the gain compensation to the specific signal characteristics.
Solution Approach 2:
The downmix gain is calculated in advance based on the relationship between original and downmixed signals, preserving spatial information characteristics before they are lost in the compression process. This preliminary calculation of gain parameters enables efficient encoding while maintaining quality.
3Use of energy by moving object
If high sound levels are present in multi-channel audio, then dynamic range is maintained, but downmixing causes clipping and distortion
Solution Approach 1:
The patent changes the amplitude parameter of downmixed signals by applying the calculated downmix gain, which prevents clipping and distortion while maintaining the dynamic range characteristics of the original multi-channel audio. This parameter adjustment occurs in a controlled manner based on signal relationships.
Solution Approach 2:
The patent converts the potential harm of sound level loss into a benefit by calculating the downmix gain from the relationship between original and downmixed signals. This approach uses the downmixing process itself to generate the compensation information needed to prevent distortion, turning a problematic effect into a useful feature.
Data Source
Figure 1
Figure 2(a)~2(c)
Figure 3
AI summary
A method and/or apparatus for encoding and/or decoding an audio signal is disclosed, in which a downmix gain is applied to a downmix signal in an encoding apparatus which, in turn, transmits, to a decoding apparatus, a bitstream containing information as to the applied downmix gain. The decoding apparatus recovers the downmix signal, using the downmix gain information. A method and/or apparatus for encoding and/or decoding an audio signal is also disclosed, in which the encoding apparatus can apply an arbitrary downmix gain (ADG) to the downmix signal, and can transmit a bitstream containing information as to the applied ADG to the decoding apparatus. The decoding apparatus recovers the downmix signal, using the ADG information. A method and/or apparatus for encoding and/or decoding an audio signal is also disclosed, in which the method and/or apparatus can also vary the energy level of a specific channel, and can recover the varied energy level .