6DoF MPEG-I Immersive Audio Edge Rendering for Low-Latency Delivery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently rendering complex 6-degree-of-freedom (6DoF) MPEG-I immersive audio due to computational demands, especially for resource-constrained devices, and delivering it with low latency and responsiveness to user position and orientation changes over bandwidth-constrained networks.
Innovation Solution
Utilizing edge computing to offload complex rendering tasks, transforming MPEG-I audio to an intermediate format suitable for low-latency delivery, and encoding spatial parameters with IVAS codec to maintain spatial audio cues and responsiveness, enabling efficient delivery to devices with limited resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If complex 6DoF MPEG-I immersive audio rendering is performed on resource-constrained devices, then rendering quality is improved, but computational complexity and processing time increase
Solution Approach 1:
The rendering process is divided into two parts: complex rendering operations are performed on the network side (server/edge computing), while simpler decoding and local rendering operations are performed on the user device. This segmentation allows high-quality rendering without requiring complex computational capabilities on resource-constrained devices.
Solution Approach 2:
A communication network acts as an intermediary between the rendering server and the user device. The network transmits rendered audio data from the server to the device, enabling the device to access high-quality rendered audio without performing the complex rendering computations itself.
2Ease of operation
If rendering operations are offloaded to network side, then device computational burden is reduced, but transmission latency increases
Solution Approach 1:
The rendering operations are performed in advance on the network side before the audio data needs to be consumed by the user device. This preliminary rendering allows the system to prepare audio data proactively, reducing the overall latency when the user actually needs to listen to the audio content.
Solution Approach 2:
The system dynamically adjusts the degree of rendering offloading based on network conditions and device capabilities. When network conditions are favorable, more rendering can be performed on the server side; when latency is critical, more rendering is performed locally on the device.
3Manufacturing precision
If full 6DoF rendering is performed, then spatial audio quality is improved, but bandwidth requirements increase
Solution Approach 1:
Instead of transmitting complete 6DoF rendered audio data, the system transmits only the essential spatial audio information and parameters needed for local rendering. This allows high spatial audio quality to be maintained while significantly reducing the bandwidth required for transmission.
Solution Approach 2:
The system changes the representation parameters of audio data from complete rendered waveforms to compressed spatial parameters and metadata. This parameter transformation enables the transmission of high-quality spatial audio information using much less bandwidth.
Data Source
AI summary
An apparatus configured to: obtain a user position value; obtain at least one input audio signal and associated metadata enabling a rendering of the at least one input audio signal; generate an intermediate format immersive audio signal based on the at least one input audio signal, the metadata, and the user position value, wherein the intermediate format immersive audio signal is configured to be rendered into a spatial audio output based, at least partially, on a user orientation value; and encode the intermediate format immersive audio signal.


