MDCT Complex Prediction Stereo Coding With Lower Decode Complexity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing stereo audio coding methods, particularly at higher bitrates, suffer from high computational complexity due to the use of QMF-based approaches, which are not optimal for efficient coding.
Innovation Solution
Implement a decoder and encoder system using complex prediction stereo coding based on critically sampled MDCT transforms, allowing for efficient coding by generating stereo signals from downmix and residual signals through modules that compute oversampled frequency-domain representations, enabling flexible switching between direct, joint, and complex prediction coding modes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If QMF-based approaches are used for stereo audio coding, then audio quality can be maintained, but computational complexity increases significantly at higher bitrates
Solution Approach 1:
The patent changes the fundamental parameter of the transform basis from QMF (quadrature mirror filter) to MDCT (modified discrete cosine transform). This parameter change allows the system to achieve similar audio quality through a different mathematical foundation that is computationally more efficient, particularly at higher bitrates. The MDCT-based complex prediction coding maintains the ability to represent stereo information accurately while reducing the computational burden of the coding process.
2Device complexity
If complex prediction stereo coding with MDCT is implemented, then computational complexity is reduced, but coding flexibility across different modes may be limited
Solution Approach 1:
The patent implements a universal MDCT-based complex prediction coding framework that can operate in multiple modes (direct stereo, joint stereo, and complex prediction modes). The same MDCT transform and complex prediction mechanism serve all coding modes, providing versatility without requiring separate specialized processors for each mode. This multi-functional approach allows the system to adapt to different coding scenarios while maintaining a unified, computationally efficient architecture.
Data Source
AI summary
The invention provides methods and devices for stereo encoding and decoding using complex prediction in the frequency domain. In one embodiment, a decoding method, for obtaining an output stereo signal from an input stereo signal encoded by complex prediction coding and comprising first frequency-domain representations of two input channels, comprises the upmixing steps of:(i) computing a second frequency-domain representation of a first input channel; and(ii) computing an output channel on the basis of the first and second frequency-domain representations of the first input channel, the first frequency-domain representation of the second input channel and a complex prediction coefficient. The upmixing can be suspended responsive to control data.


