Complex Prediction Audio Coding for Phase-Shifted Stereo Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multi-channel audio processing techniques face challenges in achieving coding gain, particularly when channel signals are phase-shifted or have similar waveforms with different amplitudes, leading to increased bit rates and computational complexity in decoding and encoding processes.

Innovation Solution

The use of a prediction-based approach in the modified discrete cosine transform (MDCT) domain, where prediction information is calculated and used to estimate the imaginary part of the combination signal, allowing for efficient encoding and decoding of multi-channel audio signals by reducing bit rates without compromising audio quality, and simplifying computational complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If mid/side coding is applied when channel signals are phase-shifted by 90 degrees, then the mid and side signals become similar in range, but coding gain is lost and bit rate increases

Engineering Contradiction:
Improvecoding gainVSAvoidbit rate
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies frequency-selective mid/side coding by changing the coding parameter (mid/side mode) based on the phase relationship between channels in different frequency bands. When phase shift is detected (e.g., 90 degrees), the system switches away from mid/side coding in that frequency band, thereby maintaining coding gain where applicable and avoiding bit rate increase where phase shifting occurs.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If mid/side coding is applied when left and right signals are identical, then maximum coding gain is achieved, but computational complexity increases in decoding

Engineering Contradiction:
Improvecoding gainVSAvoiddecoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies mid/side coding selectively rather than uniformly across all frequency bands. By applying the coding technique only where it provides benefit (when channels are similar) and using alternative methods where phase shifts occur, the system achieves sufficient coding gain without the full computational overhead of universal mid/side decoding.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If parametric processing based on binaural cues is used, then bit rate is reduced, but information loss increases due to lossy coding

Engineering Contradiction:
Improvebit rateVSAvoidaudio information loss
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the audio signal into different frequency bands and applies different coding strategies to each band. In frequency bands where phase relationships are stable and coding gain is achievable, mid/side coding is applied. In bands with phase shifts, traditional channel coding is used. This segmentation allows the system to minimize overall bit rate while preserving audio information where critical.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8655670B2Audio encoder, audio decoder and related methods for processing multi-channel audio signals using complex prediction
Publication Date: 2014.02.18 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US8655670B2 patent drawing
  • US8655670B2 patent drawing
  • US8655670B2 patent drawing

AI summary

An encoder, based on a combination of two audio channels, obtains a first combination signal as a mid-signal and a residual signal derivable using a predicted side signal derived from the mid signal. The first combination signal and the prediction residual signal are encoded and written into a data stream together with the prediction information. A decoder generates decoded first and second channel signals using the prediction residual signal, the first combination signal and the prediction information. A real-to-imaginary transform may be applied for estimating the imaginary part of the spectrum of the first combination signal. For calculating the prediction signal used in the derivation of the prediction residual signal, the real-valued first combination signal is multiplied by a real portion of the complex prediction information and the estimated imaginary part of the first combination signal is multiplied by an imaginary portion of the complex prediction information.