Audio Decoder Scene Packet Segmentation for VR Spatial Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current VR and AR technologies face challenges in providing an immersive audio experience with high definition and efficient data transmission and decoding, particularly in achieving six degrees of freedom audio while maintaining feasible bandwidths.

Innovation Solution

The use of three distinct packet types - scene configuration packets, scene update packets, and scene payload packets - within a bitstream, conforming to MPEG-H MHAS packet definitions, allows for efficient transmission and rendering of spatial audio by defining renderer configurations, updating metadata, and providing necessary scene information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high definition spatial audio rendering is implemented for immersive VR/AR experience, then audio quality and immersion are improved, but data transmission bandwidth and decoding complexity increase

Engineering Contradiction:
Improveaudio qualityVSAvoiddecoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio scene is segmented into multiple packets of different types (scene configuration packets, scene update packets, scene payload packets). Each packet type carries specific renderer configuration information for different aspects of the audio scene, allowing the decoder to process and render audio in a structured, modular manner that reduces overall decoding complexity while maintaining high definition quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The renderer configuration information dynamically adapts to temporal evolution of the audio scene through timestamp information. The system evaluates timestamps and adjusts rendering configurations in real-time, enabling high definition spatial audio rendering that responds to scene changes without requiring complete re-decoding of entire audio sequences

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If complete renderer configuration information is provided for temporal evolution of audio scenes, then audio rendering accuracy is improved, but data transmission volume increases

Engineering Contradiction:
Improverendering accuracyVSAvoiddata transmission volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts and separates renderer configuration information into distinct packet types that can be independently transmitted and processed. Scene configuration packets contain essential rendering parameters, while scene update packets carry only the changes needed, reducing redundant data transmission while maintaining complete rendering accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Scene configuration packets are transmitted in advance to provide renderer configuration information before the actual audio rendering is needed. This preliminary provision of configuration data allows the decoder to be pre-configured for accurate rendering without requiring all configuration details to be present simultaneously, optimizing data transmission efficiency

Inventive Principle:
Principle #10Preliminary action

3Productivity

If real-time audio scene rendering is implemented, then user experience responsiveness is improved, but processing latency increases

Engineering Contradiction:
Improverendering speedVSAvoidprocessing latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system uses periodic scene update packets with timestamp information to refresh renderer configurations at regular intervals. This periodic updating mechanism enables real-time audio scene rendering by synchronizing configuration updates with the audio playback timeline, maintaining responsiveness while managing processing latency through predictable, periodic processing cycles

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20240420706A1Audio decoder, audio encoder, method for decoding, method for encoding and bitstream, using a plurality of packets, the packets comprising one or more scene configuration packets defining a temporal evolution of a rendering scenario and comprising a timestamp information
Publication Date: 2024.12.19 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US20240420706A1 patent drawing
  • US20240420706A1 patent drawing
  • US20240420706A1 patent drawing

AI summary

Embodiments create an audio decoder which spatially renders one or more audio signals. The audio decoder receives a plurality of packets of different packet types, comprising one or more scene configuration packets providing a renderer configuration information defining a usage of scene objects and/or a usage of scene characteristics, and comprising one or more scene update packets defining a update of scene metadata for the rendering, and comprising one or more scene payload packets comprising definitions of one or more of the scene objects and/or definitions of one or more of the scene characteristics. The audio decoder selects definitions of one or more scene objects and/or definitions of one or more scene characteristics, which are in included in the scene payload packets, for the rendering in dependence on the renderer configuration information. The audio decoder updates one or more scene metadata in dependence on a content of the one or more scene update packets. Further embodiments are related to encoders, methods and bitstreams.Further embodiments create decoders, encoders, methods and bitstreams with scene update packets with update conditions, with scene configuration packets providing a renderer configuration information defining a temporal evolution of a rendering scenario and with a timestamp information and/or with subscene cell information, wherein the cell information defines an association between the one or more cells and respective one or more data structures.