Virtual Scene Audio Level Control via Anchor Source Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Virtual reality systems face challenges in maintaining consistent audio levels across different virtual worlds, leading to user discomfort and potential hearing damage due to varying hardware capabilities of consumer devices.

Innovation Solution

The method involves determining an anchor audio source and its target perceived loudness for each virtual scene, which is encoded as metadata, allowing virtual scene rendering devices to automatically adjust audio levels and prioritize audio sources based on relevance and hardware constraints, ensuring seamless transitions and safe listening levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If different content providers create virtual worlds with different audio levels, then content diversity and creativity are improved, but audio level consistency and user experience deteriorate

Engineering Contradiction:
Improvecontent diversityVSAvoidaudio level consistency
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by embedding audio metadata (including anchor audio source identification and target perceived loudness values) into the virtual world content during the content creation phase. This allows the rendering device to automatically adjust audio levels based on pre-provided guidelines, ensuring consistency across different content providers without restricting their creative freedom.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If manual volume adjustment is enabled for each virtual world, then audio level precision is improved, but user convenience and immersion deteriorate

Engineering Contradiction:
Improveaudio level precisionVSAvoiduser convenience
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling the virtual scene rendering device to automatically control its own audio output levels based on the embedded metadata. The device independently identifies the anchor audio source, retrieves the target perceived loudness, and adjusts the volume without requiring user intervention, thus maintaining precision while improving convenience.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses feedback by continuously monitoring the current audio level and comparing it with the target perceived loudness from the metadata. The rendering device automatically adjusts the volume to minimize the difference between current and target levels, ensuring precise audio control while operating autonomously.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If complex audio processing is applied to all virtual worlds, then audio quality is improved, but device compatibility and accessibility deteriorate

Engineering Contradiction:
Improveaudio qualityVSAvoiddevice compatibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by providing differentiated audio metadata for different virtual worlds, specifying which audio sources are anchors and what their target loudness should be. This allows each virtual world to receive customized audio processing instructions appropriate to its specific content, while the rendering device can adapt the level of processing based on its hardware capabilities.

Inventive Principle:
Principle #3Local quality

4Ease of operation

If high audio levels are used in virtual worlds, then immersive experience is improved, but listener safety and hearing protection deteriorate

Engineering Contradiction:
Improveimmersive experienceVSAvoidhearing damage risk
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent implements preliminary anti-action by pre-defining target perceived loudness values in the audio metadata that are designed to balance immersion with safety. The system proactively prevents harmful audio levels by automatically adjusting volume based on these pre-calculated targets, countacting the potential harm before it can affect the user.

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentEP4492828A1Methods, apparatus, and systems for performing audio level control of virtual reality scenes
Publication Date: 2025.01.15 DOLBY LABORATORIES LICENSING CORP
  • EP4492828A1 patent drawingFigure 1
  • EP4492828A1 patent drawingFigure 2
  • EP4492828A1 patent drawingFigure 3

AI summary

Systems, methods, and computer program products for providing audio metadata for a virtual scene are provided. A representation of the virtual scene is obtained, wherein the representation of the virtual scene comprises at least one audio source. An anchor sound source is determined from the at least one audio source. A target perceived loudness of the anchor sound source is determined. The anchor sound source and the target perceived loudness of the anchor sound source are provided as the audio metadata for the virtual scene for encoding by an encoder.