Stereo-Based Immersive Coding for Low-Bandwidth Spatial Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio coding techniques struggle to deliver immersive audio content efficiently over limited bandwidth while maintaining high audio quality, particularly when using multi-channel formats like Dolby Atmos or MPEG-H, due to the need for larger bandwidth and the challenges of spectral distortions and image shifts during spatial rendering.

Innovation Solution

A stereo-based immersive audio coding system that uses a two-channel stereo signal and directional parameters to recreate immersive audio experiences, applying weighting factors in the frequency domain to minimize spectral distortions and decorrelate channel pairs, while reducing the number of channels transmitted.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multi-channel audio formats (Dolby Atmos, MPEG-H) are used to deliver immersive audio content, then audio quality and immersive experience are improved, but bandwidth requirements increase

Engineering Contradiction:
Improveaudio qualityVSAvoidbandwidth
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential spatial information (directional parameters) from the full multi-channel audio signal. Instead of transmitting all channels, it separates the audio into a stereo downmix and extracts directional metadata that describes the spatial positioning of sound sources. This extracted metadata is then used at the decoder to reconstruct the immersive spatial experience, significantly reducing bandwidth while maintaining audio quality.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If bandwidth is reduced to lower data rates, then transmission efficiency is improved, but audio quality deteriorates

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the representation parameters of the audio signal by transforming it into a different domain. Instead of transmitting raw multi-channel audio data, it uses perceptual audio coding to represent the audio in terms of perceptually relevant parameters (spectral coefficients, directional parameters, spatial metadata). This parameter transformation allows for efficient compression while preserving the perceptual quality of the audio experience.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If spatial rendering is applied to recreate immersive audio from stereo signals, then bandwidth is reduced, but spectral distortions and image shifts occur

Engineering Contradiction:
ImprovebandwidthVSAvoidspatial accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the encoder analyzes the stereo downmix signal and adjusts the directional parameters based on the actual spatial characteristics present in the downmixed signal. The decoder uses this feedback information along with the transmitted directional metadata to accurately reconstruct the spatial positioning. This feedback loop ensures that spectral distortions and image shifts are minimized by adapting the spatial rendering to the actual content being transmitted.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12417773B2Stereo-based immersive coding
Publication Date: 2025.09.16 APPLE INC
  • US12417773B2 patent drawing
  • US12417773B2 patent drawing
  • US12417773B2 patent drawing

AI summary

Disclosed is an audio codec that represents an immersive signal by a two-channel stereo signal that is a stereo rendering of the immersive signal and directional parameters. The directional parameters may be based on a perceptual model describing the direction of virtual speaker pairs to recreate the perceived location of dominant sounds. Audio processing at the decoder may be performed on the stereo signal in the frequency domain for multiple channel pairs using time-frequency tiles. Spatial localization of the audio signals may use a panning approach by applying weightings to the time-frequency tiles of the stereo signal for each output channel pair. The weightings for the time-frequency tiles may be derived based on the directional parameters, an analysis of the stereo signal, and the output channel layout. The weightings may be used to adaptively process the time-frequency tiles using a de-correlator to reduce or minimize spectral distortions from spatial rendering.