Multi-Channel Audio Encoding Using Complex Power Ratios
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio compression techniques struggle to maintain high-quality multi-channel audio at low bitrates, as they often compromise on quality due to limited computational resources in computers and networks.
Innovation Solution
The proposed solution involves encoding and decoding strategies that use a combined channel and power ratio parameters to reconstruct individual physical channels, allowing for efficient compression and decompression of multi-channel audio data, even at low bitrates, by maintaining second-order statistics and using complex transforms for channel extension processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If lossy compression techniques are used to reduce bitrate, then transmission cost decreases, but audio quality deteriorates
Solution Approach 1:
The patent applies parameter changes by transforming audio data from time domain to frequency domain using Fourier transforms, then applying perceptual weighting to modify the importance of different frequency components. This allows selective preservation of perceptually important information while compressing less important components, achieving bitrate reduction without significant quality loss.
Solution Approach 2:
The patent implements local quality through perceptual weighting that assigns different importance levels to different frequency bands based on human auditory characteristics. Critical bands with higher perceptual importance receive better preservation, while less critical bands are more aggressively compressed, creating non-uniform quality distribution optimized for human perception.
2Manufacturing precision
If multi-channel audio is transmitted with high quality, then audio fidelity is maintained, but bitrate consumption increases
Solution Approach 1:
The patent merges multiple audio channels into a combined channel representation using linear transformation. By encoding the combined channel along with correlation parameters rather than encoding each channel independently, the system achieves significant bitrate reduction while preserving inter-channel spatial relationships and audio fidelity.
Solution Approach 2:
The patent transitions from encoding channels in the time domain to encoding in the frequency domain using Fourier transforms. This dimensional change allows for perceptual weighting and more efficient compression by exploiting frequency-domain characteristics and human auditory perception properties.
3Measurement precision
If complex transforms are used for channel extension processing, then reconstruction accuracy improves, but computational complexity increases
Solution Approach 1:
The patent performs preliminary frequency transformation and perceptual weighting at the encoder before compression. By pre-processing the audio data in the frequency domain and applying perceptual models upfront, the system reduces the computational burden on the decoder while maintaining high reconstruction accuracy through complex transforms.
Data Source
AI summary
An audio encoder encodes a combined channel (e.g., a sum channel) for a group of plural physical audio channels. The encoder determines plural parameters for representing individual physical channels of the group as modified versions of the encoded combined channel. The plural parameters comprise ratios of power in each individual channel to power in the combined channel (e.g., a ratio of the power of a right channel to the power of the combined channel, and a ratio of the power of the left channel to the power of the combined channel). The plural parameters can include a complex parameter. The combined channel and the plural parameters facilitate reconstruction at the audio decoder of source channels. An audio decoder performs a forward complex transform on the multi-channel audio data and reconstructs plural channels from the multi-channel audio data. The decoder can maintain second-order statistics for the source channels.


