Immersive Audio Asset Pre-rendering for Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional immersive audio systems experience high motion-to-sound latency due to the need for real-time head tracking and audio data transmission, which can diminish the immersive experience, especially in systems with limited processing resources.

Innovation Solution

A system that pre-renders immersive audio assets for known listener poses, shifting the resource burden to devices with greater availability, such as servers, and uses local storage and pre-fetching to reduce latency, allowing for efficient generation and playback of output audio signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-time head tracking and audio data transmission are used to update audio scenes, then immersive audio quality is maintained, but motion-to-sound latency increases unnaturally

Engineering Contradiction:
Improveimmersive audio qualityVSAvoidmotion-to-sound latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system pre-renders audio scenes for anticipated head positions before the user actually reaches those positions. By predicting future head poses and preparing audio content in advance, the system eliminates the latency inherent in real-time rendering, allowing seamless transitions without noticeable delay between head movement and audio update.

Inventive Principle:
Principle #10Preliminary action

2Use of energy by moving object

If audio scene updates and binauralization are performed at the remote server, then processing resources on the headphone device are conserved, but transmission latency increases

Engineering Contradiction:
Improveprocessing power consumptionVSAvoidaudio data transmission latency
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The system performs audio scene updates and binauralization processing in advance at the remote server before the user actually needs the audio content. By pre-processing audio scenes for anticipated head positions and storing them locally, the system transfers the computational burden to the server while avoiding real-time transmission delays during actual playback.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates local copies of pre-rendered audio scenes for different head positions. Instead of transmitting and processing audio data in real-time, the headphone device stores multiple pre-prepared audio scene copies locally, allowing instant retrieval and playback without transmission latency or heavy local processing requirements.

Inventive Principle:
Principle #26Copying

3Loss of time

If pre-rendered audio assets are used for known listener poses, then latency is reduced and resources are conserved, but adaptability to unexpected head movements decreases

Engineering Contradiction:
Improveaudio response latencyVSAvoidhandling of unexpected head poses
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The system divides the continuous audio scene into discrete segments corresponding to different head positions. By creating separate pre-rendered audio assets for specific head poses, the system enables efficient retrieval and playback for anticipated positions while maintaining the ability to handle unexpected movements through selective rendering or interpolation between segments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250024219A1Sound field adjustment
Publication Date: 2025.01.16 QUALCOMM INC
  • US20250024219A1 patent drawing
  • US20250024219A1 patent drawing
  • US20250024219A1 patent drawing

AI summary

A device includes a memory configured to store audio data associated with an immersive audio environment. The device also includes one or more processors configured to obtain a listener pose in the immersive audio environment associated with a first time and determine whether the listener pose is associated with a pre-rendered asset. The one or more processors are configured to obtain a rendered asset by selecting, based on the determination, between obtaining the pre-rendered asset and performing a rendering operation to generate the rendered asset. The one or more processors are also configured to generate an output audio signal based on the rendered asset.