Binaural Split Rendering for Low-Latency AR Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Immersive audio rendering in small form factor devices like AR glasses is computationally demanding, leading to inefficiencies and increased power consumption due to the need for significant processing capabilities, and latency issues in split-rendering topologies result in a perceivable loss of quality in the immersive experience.

Innovation Solution

A method for low complexity low bitrate prediction-based split rendering, where a main device with high processing resources performs downmixing and binaural representation generation, and a lightweight device reconstructs binaural audio using metadata and updated pose information, reducing the need for real-time binaural rendering on the lightweight device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If immersive audio rendering is performed on small form factor devices like AR glasses, then high-quality immersive audio experience is achieved, but computational complexity and power consumption increase significantly

Engineering Contradiction:
Improveaudio qualityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The audio rendering process is segmented into two parts: complex binaural rendering is performed on a remote server, while only lightweight downmixing and HRTF application are performed on the AR device. This segmentation allows the computationally intensive tasks to be offloaded, reducing power consumption while maintaining audio quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A remote server acts as an intermediary to perform the computationally intensive binaural rendering tasks. The server processes the immersive audio content and sends pre-processed audio signals to the AR device, which then only needs to apply simple HRTF filtering based on head tracking data, significantly reducing on-device computational requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If split-rendering topology is used to reduce computational complexity on end-device, then processing requirements are reduced, but transmission latency increases to 100 ms

Engineering Contradiction:
Improveprocessing complexityVSAvoidtransmission latency
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The remote server performs preliminary binaural rendering and prepares audio signals in advance based on received head tracking data. By pre-processing the audio content before transmission, the system reduces the amount of real-time processing needed on the AR device and minimizes the impact of transmission latency on audio quality.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If traditional split-rendering is used with 100 ms latency, then computational load on end-device is reduced, but motion-to-sound latency causes perceivable quality loss

Engineering Contradiction:
Improveprocessing loadVSAvoidimmersive experience quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system changes the parameter being transmitted from raw immersive audio content to compact HRTF filters and downmix signals. This parameter transformation allows the AR device to reconstruct high-quality binaural audio locally using minimal data and simple processing, maintaining immersive experience quality while reducing both computational load and the impact of transmission latency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12604152B2Binarual rendering
Publication Date: 2026.04.14 DOLBY LABORATORIES LICENSING CORP
  • US12604152B2 patent drawing
  • US12604152B2 patent drawing
  • US12604152B2 patent drawing

AI summary

An aspect of the present disclosure relates to processing audio comprising decoding a first bitstream (b1) to obtain decoded immersive audio content (A), decoding a second bitstream (bp) to obtain pose information (P, V, V′) associated with a user of a lightweight processing device, determining a first head-pose (P′) based on the pose information, providing a downmix representation (Dmx) of the immersive audio content (A) corresponding to the first head pose (P′), rendering a set of binaural representations (BINn) of the immersive audio content (A), wherein the binaural representations correspond to a second set of head poses (Pn), computing reconstruction metadata (M) to enable reconstruction of the set of binaural representations from the downmix representation (Dmx), the metadata (M) including the first head pose (P′), and encoding the downmix representation (Dmx) and the reconstruction metadata (M) in a third bitstream (b2).