3DOF/6DOF Audio Rendering With Chained Digested Parameters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing immersive audio rendering technologies face high computational complexity, particularly in power-limited devices like AR glasses, necessitating a more efficient distribution of computational burden and minimizing motion-to-sound latency.
Innovation Solution
A method and system for rendering audio using a chain of renderers, where canonical rendering parameters are digested into less computationally intensive digested parameters, allowing processing on servers and end devices, with metadata and audio data being processed and transmitted to minimize latency and complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If immersive audio rendering is performed on power-limited end devices, then audio quality and user experience are maintained, but computational complexity and power consumption increase
Solution Approach 1:
The rendering process is divided into two distinct stages: pre-rendering on powerful servers and final rendering on end devices. The server performs computationally intensive tasks including generating HOA coefficients, applying HRTF filters, and creating intermediate audio representations. The end device only performs lightweight final rendering operations, thus maintaining audio quality while reducing computational complexity and power consumption.
Solution Approach 2:
Computationally intensive rendering operations are performed in advance on servers before the audio is delivered to end devices. The server pre-processes audio content, pre-calculates spatial parameters, and prepares intermediate representations that can be efficiently finalized on the end device. This preliminary action transfers the computational burden from power-limited devices to powerful servers.
2Adaptability or versatility
If computational tasks are distributed to end devices, then rendering flexibility is improved, but motion-to-sound latency increases
Solution Approach 1:
The system performs preliminary rendering operations on servers, generating intermediate audio representations and spatial parameters in advance. This pre-processing reduces the computational workload on end devices, enabling faster final rendering with reduced motion-to-sound latency while maintaining rendering flexibility through client-side parameter adjustment.
Solution Approach 2:
The server acts as an intermediary that performs the bulk of computational work and delivers pre-processed audio data to end devices. This intermediary processing stage reduces the time required for final rendering on client devices, thereby reducing motion-to-sound latency while preserving the ability to adapt to user movements.
3Measurement precision
If full canonical rendering parameters are processed on end devices, then rendering accuracy is maintained, but power consumption increases
Solution Approach 1:
The rendering parameter processing is segmented between server and end device. The server processes complex canonical rendering parameters including HOA coefficient generation and HRTF filter application. The end device receives pre-processed data and only performs lightweight final rendering calculations, thus maintaining rendering accuracy while significantly reducing power consumption.
Solution Approach 2:
The server creates simplified copies or representations of the audio data in intermediate formats (such as HOA coefficients and spatial parameters) that can be efficiently processed on end devices. These copied representations preserve the essential acoustic information needed for accurate rendering while requiring minimal computational resources to process.
Data Source
AI summary
Described herein is a method of rendering audio, the method including: receiving, at a first renderer, first audio data and first metadata for the first audio data, the first metadata including one or more canonical rendering parameters; processing, at the first renderer, the first metadata and optionally the first audio data for generating second metadata and optionally second audio data, wherein the processing includes generating one or more first digested rendering parameters based on the one or more canonical rendering parameters; providing, by the first renderer, the second metadata and optionally the second audio data for further processing by a second renderer, the second metadata including the one or more first digested rendering parameters and optionally a first portion of the one or more canonical rendering parameters. Described is also a further method of rendering audio, respective systems and computer program products.


