Layered Audio Space Description for Immersive VR Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio processing technologies in virtual and augmented reality struggle to create immersive audio experiences by accurately rendering sounds based on a user's location and movement within a virtual scene, often resulting in unrealistic audio interactions.

Innovation Solution

The use of a layered description of a space of interest in an audio scene, where a first layer contains common attribute values shared by multiple subspaces and a second layer contains individual attribute values, allows for efficient audio coding, transmission, and rendering of audio inputs based on the location of a subject within the scene, using processing circuitry to determine the subspaces and render audio outputs accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a layered description with common nodes and individual nodes is used to represent audio spaces, then the coding efficiency and transmission bandwidth are improved, but the processing complexity to retrieve and compute attribute values increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The audio space description is segmented into two layers: a first layer containing common nodes with attribute values shared by multiple subspaces, and a second layer containing individual nodes for specific subspaces. This segmentation allows efficient representation by separating common information from specific information, reducing overall data redundancy while maintaining the ability to retrieve precise attributes for any subspace through a structured lookup process.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If individual attribute values are stored for each subspace, then the precision of audio rendering is improved, but the data transmission bandwidth and storage requirements increase

Engineering Contradiction:
Improveaudio rendering precisionVSAvoiddata quantity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

Attribute values that are common across multiple subspaces are merged into shared common nodes in the first layer. Instead of storing duplicate copies of the same attribute value in each individual subspace node, the system consolidates these repeated values into a single reference point that multiple subspaces can point to, thereby reducing total data quantity while preserving precise attribute information for accurate audio rendering.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If a comprehensive description of all subspaces is provided, then the adaptability of audio rendering to different user locations is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveaudio rendering adaptabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary organization of audio space data by pre-structuring the layered description with common nodes and individual nodes before runtime processing. Attribute values are pre-computed and stored in the appropriate layer, and the hierarchical structure is established in advance, enabling rapid retrieval and processing during actual audio rendering without requiring complex computations at runtime for each user location query.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11937070B2Layered description of space of interest
Publication Date: 2024.03.19 TENCENT AMERICA LLC
  • US11937070B2 patent drawing
  • US11937070B2 patent drawing
  • US11937070B2 patent drawing

AI summary

Aspects of the disclosure provide methods and apparatuses for audio processing. In some examples, an apparatus for media processing includes processing circuitry. The processing circuitry receive audio inputs associated with a layered description for a space of interest in an audio scene. The space of interest includes a plurality of subspaces. The layered description includes a first layer and a second layer. The first layer has a common node with a first value that is a common attribute value of two or more subspaces in the plurality of subspaces. The second layer has individual nodes respectively associated with each of the plurality of subspaces. The processing circuitry determines the plurality of subspaces of the space of interest based on the layered description, and renders an audio output based on the audio inputs in response to a location of a subject of the audio scene being in the space of interest.