Stereo Audio Decoding With Reversible Mid-Side Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-channel audio processing techniques face challenges in achieving high coding gain while maintaining audio quality and reducing computational complexity, particularly in situations where mid/side coding does not yield significant coding gains due to phase shifts or equal waveforms, and parametric coding introduces artifacts and information loss.
Innovation Solution
A method and apparatus for stereo audio decoding using a prediction of a second combination signal from a first combination signal, employing a frequency-domain approach with a critically sampled transform like MDCT, and adaptively reversing the prediction direction based on energy distribution between mid and side signals to enhance coding gain and reduce computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If mid/side coding is applied to stereo audio signals, then coding gain is improved when left and right signals are similar, but coding gain is lost or bit rate increases when signals have phase shifts or equal waveforms
Solution Approach 1:
The patent applies dynamics by making the prediction direction adaptive rather than fixed. The decoder determines whether to predict the side signal from the mid signal or vice versa based on signal characteristics, allowing the coding system to dynamically adjust to different signal conditions (phase shifts, equal waveforms, or similar signals) to maintain coding gain across various scenarios
Solution Approach 2:
The patent changes the parameter of prediction direction (from fixed to variable) to resolve the contradiction. By introducing a prediction direction indicator that can take different values (predicting side from mid, or mid from side), the system adapts its coding strategy based on the actual signal conditions, thereby maintaining coding gain whether signals are similar or have phase shifts
2Productivity
If parametric coding techniques are used to reduce bit rate, then productivity is improved, but audio quality deteriorates due to artifacts and information loss
Solution Approach 1:
The patent substitutes parametric coding (which relies on binaural cues and mental reconstruction) with waveform-based predictive coding. Instead of using complex parametric models that introduce artifacts, the patent uses a simpler prediction mechanism operating directly on the waveform data in the MDCT domain, replacing the mechanical parametric processing with a more direct signal processing approach that preserves audio fidelity
3Loss of information
If frequency-selective mid/side coding is applied to maintain coding gain, then coding gain is preserved in certain bands, but device complexity increases
Solution Approach 1:
The patent applies universality by creating a coding framework that handles multiple signal conditions (phase-shifted signals, equal waveforms, similar signals) through a single unified mechanism. The variable prediction direction approach serves as a universal solution that adapts to different frequency bands and signal characteristics without requiring separate frequency-selective processing paths, thereby reducing overall device complexity while maintaining coding gain
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
An audio or video encoder and an audio or video decoder are based on a combination of two audio or video channels (201, 202) to obtain a first combination signal (204) as a mid signal and a residual signal (205) which can be derived using a predicted side signal derived from the mid signal. The first combination signal and the prediction residual signal are encoded (209) and written (212) into a data stream (213) together with the prediction information (206) derived by an optimizer (207) based on an optimization target (208) and a prediction direction indicator indicating a prediction direction associated with the residual signal. A decoder uses the prediction residual signal, the first combination signal, the prediction direction indicator and the prediction information to derive a decoded first channel signal and a decoded second channel signal. In an encoder example or in a decoder example, a real-to-imaginary transform can be applied for estimating the imaginary part of the spectrum of the first combination signal. For calculating the prediction signal used in the derivation of the prediction residual signal, the real-valued first combination signal is multiplied by a real portion of the complex prediction information and the estimated imaginary part of the first combination signal is multiplied by an imaginary portion of the complex prediction information.