Virtual Scene Audio Level Control via Anchor Source Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Virtual reality systems face challenges in maintaining consistent audio levels across different virtual worlds, leading to user discomfort and potential hearing damage due to varying hardware capabilities of consumer devices.
Innovation Solution
The method involves determining an anchor audio source and its target perceived loudness for each virtual scene, which is encoded as metadata, allowing virtual scene rendering devices to automatically adjust audio levels and prioritize audio sources based on relevance and hardware constraints, ensuring seamless transitions and safe listening levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If different content providers create virtual worlds with different audio levels, then content diversity and creativity are improved, but audio level consistency and user experience deteriorate
Solution Approach 1:
The patent applies preliminary action by embedding audio metadata (including anchor audio source identification and target perceived loudness values) into the virtual world content during the content creation phase. This allows the rendering device to automatically adjust audio levels based on pre-provided guidelines, ensuring consistency across different content providers without restricting their creative freedom.
2Manufacturing precision
If manual volume adjustment is enabled for each virtual world, then audio level precision is improved, but user convenience and immersion deteriorate
Solution Approach 1:
The patent implements self-service by enabling the virtual scene rendering device to automatically control its own audio output levels based on the embedded metadata. The device independently identifies the anchor audio source, retrieves the target perceived loudness, and adjusts the volume without requiring user intervention, thus maintaining precision while improving convenience.
Solution Approach 2:
The system uses feedback by continuously monitoring the current audio level and comparing it with the target perceived loudness from the metadata. The rendering device automatically adjusts the volume to minimize the difference between current and target levels, ensuring precise audio control while operating autonomously.
3Manufacturing precision
If complex audio processing is applied to all virtual worlds, then audio quality is improved, but device compatibility and accessibility deteriorate
Solution Approach 1:
The patent applies local quality by providing differentiated audio metadata for different virtual worlds, specifying which audio sources are anchors and what their target loudness should be. This allows each virtual world to receive customized audio processing instructions appropriate to its specific content, while the rendering device can adapt the level of processing based on its hardware capabilities.
4Ease of operation
If high audio levels are used in virtual worlds, then immersive experience is improved, but listener safety and hearing protection deteriorate
Solution Approach 1:
The patent implements preliminary anti-action by pre-defining target perceived loudness values in the audio metadata that are designed to balance immersion with safety. The system proactively prevents harmful audio levels by automatically adjusting volume based on these pre-calculated targets, countacting the potential harm before it can affect the user.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems, methods, and computer program products for providing audio metadata for a virtual scene are provided. A representation of the virtual scene is obtained, wherein the representation of the virtual scene comprises at least one audio source. An anchor sound source is determined from the at least one audio source. A target perceived loudness of the anchor sound source is determined. The anchor sound source and the target perceived loudness of the anchor sound source are provided as the audio metadata for the virtual scene for encoding by an encoder.