Adaptive Multi-Channel Audio Downmixing for Correlated Signal Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding technologies struggle to efficiently encode multi-channel audio signals, particularly when channels are not dominated by a single dominant channel and are highly correlated, leading to inefficient data usage and decoding challenges.
Innovation Solution
Adaptive downmixing of audio signals involves forming a primary output channel from a sum of scaled non-primary channels and prediction channels, using mixing and prediction gains to create an output multi-channel audio signal that is efficiently encoded by allocating fewer bits to less dominant channels or discarding them entirely, while ensuring the primary channel contains most sonic elements and channels are largely uncorrelated.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multi-channel audio signals with highly correlated channels are encoded using existing technologies, then decoding challenges arise and data usage becomes inefficient, but maintaining channel correlation preserves audio quality
Solution Approach 1:
The audio signal is segmented into a dominant channel and multiple non-dominant channels. The dominant channel is encoded with full detail while non-dominant channels are derived through prediction based on the dominant channel, reducing the data required for encoding while maintaining audio quality through the prediction relationship.
Solution Approach 2:
The patent applies parameter changes by using mixing coefficients and prediction gains to transform the multi-channel signal into a form where one channel dominates. This parameter transformation allows efficient encoding by allocating more bits to the dominant channel and fewer bits to non-dominant channels that can be predicted from it.
2Quantity of substance
If bits are allocated equally to all channels, then decoding quality is maintained, but data usage becomes inefficient when channels are highly correlated
Solution Approach 1:
Different bit allocation strategies are applied to different channels based on their importance. The dominant channel receives higher bit allocation to preserve critical audio information, while non-dominant channels receive fewer bits since they can be reconstructed through prediction from the dominant channel, optimizing the overall data usage to quality ratio.
3Device complexity
If all non-primary channels are fully encoded, then audio quality is preserved, but encoding complexity and data requirements increase
Solution Approach 1:
The patent extracts the essential audio information into a single dominant channel that contains most of the sonic elements. The non-dominant channels are not fully encoded but are instead derived through prediction from the dominant channel using mixing coefficients and prediction gains, significantly reducing encoding complexity while preserving audio quality.
Data Source
AI summary
Systems, methods, and computer program products are disclosed for adaptive downmixing of audio signals with improved continuity. An audio encoding system receives an input multi-channel audio signal including a primary input audio channel and L non-primary input audio channels. The system determines a set of L input gains. For each of the channels and gains, the system forms a respective scaled non-primary input audio channel. The system forms a primary output audio channel from the sum of the primary input audio channel and the scaled non-primary input audio channels. The system determines a set of L prediction gains. The system forms a prediction channel from the primary output audio channel. The system forms L non-primary output audio channels. The system forms an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels.


