Pre-rendered Audio Elements for VR Acoustic Fidelity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 6DoF renderers in VR/AR/MR environments fail to accurately reproduce a content creator's desired soundfield due to insufficient metadata and limited capabilities, leading to bitrate limitations and artistic intent discrepancies, especially in low-power decoding scenarios.

Innovation Solution

A method and system that encode and decode audio scene content using effective audio elements, which encapsulate the impact of the acoustic environment, allowing for efficient compression and rendering with predetermined modes that enhance acoustics and preserve artistic intent, even on low-power decoders, by determining effective audio elements and rendering modes based on listener position and orientation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If 6DoF renderers use parametrized metadata to describe sound sources and VR/AR/MR environment, then the rendering flexibility and adaptability improve, but bitrate limitations and data availability constraints worsen

Engineering Contradiction:
Improverendering flexibilityVSAvoidbitrate
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent pre-calculates and stores impulse responses at multiple positions in the VR/AR/MR environment during encoding. These pre-computed acoustic data are then transmitted to the decoder, eliminating the need for real-time computation of complex acoustic environments. This preliminary action resolves the contradiction by providing rendering flexibility through pre-stored positional data while keeping bitrate manageable through selective transmission.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the acoustic environment into discrete impulse responses at specific positions. Instead of transmitting continuous or exhaustive environmental data, the system divides the acoustic space into manageable segments (individual impulse responses at key positions). This segmentation allows the decoder to reconstruct accurate acoustic environments using only the transmitted segments, resolving the bitrate limitation while maintaining adaptability.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If 6DoF renderers attempt to reproduce accurate sound fields in all positions, then audio quality and fidelity improve, but computational complexity and resource requirements worsen

Engineering Contradiction:
Improveaudio qualityVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs complex acoustic calculations and sound field reproductions in advance during the encoding phase. The encoder pre-computes impulse responses for multiple positions and stores them in a lookup table. During decoding, the system simply retrieves pre-computed data based on listener position, dramatically reducing computational complexity while maintaining high audio quality and fidelity across all positions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates copies of acoustic impulse responses at different positions and stores them for retrieval. Instead of computing sound fields in real-time for each listener position, the system creates and stores copies of acoustic data at key positions during encoding. The decoder then selects the appropriate copy based on current listener position, maintaining audio quality while minimizing computational complexity during playback.

Inventive Principle:
Principle #26Copying

3Productivity

If pre-rendered audio elements are used to encapsulate acoustic environment impact, then compression efficiency and decoding speed improve, but rendering accuracy and artistic intent preservation worsen

Engineering Contradiction:
Improvecompression efficiencyVSAvoidrendering accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the acoustic environment into discrete impulse responses at specific positions rather than using a single pre-rendered audio element. The encoder computes and transmits multiple impulse responses corresponding to different listener positions. The decoder selects the appropriate segment based on current position, maintaining rendering accuracy while achieving compression efficiency through selective transmission of only necessary acoustic data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects which pre-computed impulse response to use based on the listener's current position in the VR/AR/MR environment. Rather than using a static pre-rendered audio element, the decoder chooses from multiple pre-computed segments according to real-time positional information. This dynamic selection maintains rendering accuracy across different positions while preserving compression efficiency through the use of pre-computed data.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230262407A1Methods, apparatus and systems for a pre-rendered signal for audio rendering
Publication Date: 2023.08.17 DOLBY INTERNATIONAL AB
  • US20230262407A1 patent drawing
  • US20230262407A1 patent drawing
  • US20230262407A1 patent drawing

AI summary

The present disclosure relates to a method of decoding audio scene content from a bitstream by a decoder that includes an audio renderer with one or more rendering tools. The method comprises receiving the bitstream, decoding a description of an audio scene from the bitstream, determining one or more effective audio elements from the description of the audio scene, determining effective audio element information indicative of effective audio element positions of the one or more effective audio elements from the description of the audio scene, decoding a rendering mode indication from the bitstream, wherein the rendering mode indication is indicative of whether the one or more effective audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using a predetermined rendering mode, and in response to the rendering mode indication indicating that the one or more effective audio elements represent the sound field obtained from pre-rendered audio elements and should be rendered using the predetermined rendering mode, rendering the one or more effective audio elements using the predetermined rendering mode, wherein rendering the one or more effective audio elements using the predetermined rendering mode takes into account the effective audio element information, and wherein the predetermined rendering mode defines a predetermined configuration of the rendering tools for controlling an impact of an acoustic environment of the audio scene on the rendering output. The disclosure further relates to a method of generating audio scene content and a method of encoding audio scene content into a bitstream.