Complex Prediction Audio Coding for Phase-Shifted Stereo Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-channel audio processing techniques face challenges in achieving coding gain, particularly when channel signals are phase-shifted or have similar waveforms with different amplitudes, leading to increased bit rates and computational complexity in decoding and encoding processes.
Innovation Solution
The use of a prediction-based approach in the modified discrete cosine transform (MDCT) domain, where prediction information is calculated and used to estimate the imaginary part of the combination signal, allowing for efficient encoding and decoding of multi-channel audio signals by reducing bit rates without compromising audio quality, and simplifying computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If mid/side coding is applied when channel signals are phase-shifted by 90 degrees, then the mid and side signals become similar in range, but coding gain is lost and bit rate increases
Solution Approach 1:
The patent applies frequency-selective mid/side coding by changing the coding parameter (mid/side mode) based on the phase relationship between channels in different frequency bands. When phase shift is detected (e.g., 90 degrees), the system switches away from mid/side coding in that frequency band, thereby maintaining coding gain where applicable and avoiding bit rate increase where phase shifting occurs.
2Reliability
If mid/side coding is applied when left and right signals are identical, then maximum coding gain is achieved, but computational complexity increases in decoding
Solution Approach 1:
The patent applies mid/side coding selectively rather than uniformly across all frequency bands. By applying the coding technique only where it provides benefit (when channels are similar) and using alternative methods where phase shifts occur, the system achieves sufficient coding gain without the full computational overhead of universal mid/side decoding.
3Quantity of substance
If parametric processing based on binaural cues is used, then bit rate is reduced, but information loss increases due to lossy coding
Solution Approach 1:
The patent segments the audio signal into different frequency bands and applies different coding strategies to each band. In frequency bands where phase relationships are stable and coding gain is achievable, mid/side coding is applied. In bands with phase shifts, traditional channel coding is used. This segmentation allows the system to minimize overall bit rate while preserving audio information where critical.
Data Source
AI summary
An encoder, based on a combination of two audio channels, obtains a first combination signal as a mid-signal and a residual signal derivable using a predicted side signal derived from the mid signal. The first combination signal and the prediction residual signal are encoded and written into a data stream together with the prediction information. A decoder generates decoded first and second channel signals using the prediction residual signal, the first combination signal and the prediction information. A real-to-imaginary transform may be applied for estimating the imaginary part of the spectrum of the first combination signal. For calculating the prediction signal used in the derivation of the prediction residual signal, the real-valued first combination signal is multiplied by a real portion of the complex prediction information and the estimated imaginary part of the first combination signal is multiplied by an imaginary portion of the complex prediction information.


