Rendering Metadata for User Movement Audio Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer-mediated reality systems, such as VR, AR, and MR, face challenges in providing an immersive auditory experience due to the lack of controls that allow content creators to maintain a consistent and realistic audio rendering based on user movement, leading to reduced immersion and limited content availability.

Innovation Solution

The implementation of rendering metadata that enables or disables adaptations in audio rendering based on user movement, allowing for the selection and application of appropriate renderers to generate speaker feeds that align with the user's position and orientation, using techniques like six degrees of freedom rendering and binaural rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio rendering is dynamically adapted based on user movement, then auditory immersion is improved, but device complexity increases

Engineering Contradiction:
Improveauditory immersionVSAvoidrendering control complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio rendering system is segmented into distinct controllable components through rendering metadata that can independently enable or disable specific adaptations (e.g., distance rendering, Doppler effect, delay effects) based on user movement, allowing selective complexity management while maintaining immersion

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The rendering system dynamically adapts audio output based on real-time user movement detection, adjusting rendering parameters such as distance attenuation, Doppler shifts, and time delays to maintain realistic auditory immersion as the user moves through the virtual environment

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If rendering metadata is obtained from bitstream, then content availability increases, but manufacturing precision requirements increase

Engineering Contradiction:
Improvecontent availabilityVSAvoidrendering metadata specification
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The rendering metadata structure is designed as a universal control mechanism that can be embedded in audio bitstreams across different content formats and delivery platforms, enabling consistent rendering control while accommodating various content creation workflows and distribution channels

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If renderer adaptations are enabled based on user movement, then auditory experience realism is improved, but processing requirements increase

Engineering Contradiction:
Improveauditory experience realismVSAvoidprocessing power
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The system applies rendering adaptations selectively to specific audio elements based on their importance and the user's movement characteristics, rather than uniformly processing all audio data, thereby reducing overall processing requirements while maintaining realism where it matters most

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11184731B2Rendering metadata to control user movement based audio rendering
Publication Date: 2021.11.23 QUALCOMM INC
  • US11184731B2 patent drawing
  • US11184731B2 patent drawing
  • US11184731B2 patent drawing

AI summary

In general, techniques are described for rendering metadata to control user movement based audio rendering. A device comprising a memory and one or more processors may be configured to perform the techniques. The memory may be configured to store audio data representative of a soundfield. The one or more processors may be coupled to the memory, and configured to obtain rendering metadata indicative of controls for enabling or disabling adaptations, based on an indication of a movement of a user of the device, of a renderer used to render audio data representative of a soundfield, specify, in a bitstream representative of the audio data, the rendering metadata, and output the bitstream.