Audio Scene Processing for Immersive Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In virtual reality and augmented reality systems, the existing audio processing methods fail to provide a realistic audio experience as users move closer to or farther from sound objects, as the amplitude of background sounds often needs to be reduced to avoid overwhelming the user, leading to continuous changes in sound levels that detract from the immersive experience.
Innovation Solution
The system employs a spatial audio capture apparatus and additional audio capture devices to identify objects of interest, allowing their audio signals to increase in amplitude independently of combined background sounds, even when the overall amplitude is maximized, by using dynamic range compression and spatial repositioning techniques to enhance the perceived distance and movement effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the amplitude of audio signals from objects of interest is increased as the user approaches, then the realism and immersive experience is improved, but the background sound levels become disturbed and overwhelming
Solution Approach 1:
The audio processing system segments the audio scene into distinct audio objects, each with independent amplitude control. This allows the amplitude of objects of interest to be adjusted separately from background sounds, enabling realistic volume changes as users approach without disturbing background sound levels.
Solution Approach 2:
Different amplitude modulation characteristics are applied to different audio objects based on their importance and spatial relationship to the user. Objects of interest receive dynamic amplitude adjustment based on user proximity, while background objects maintain stable levels, creating localized quality differences in the audio experience.
2Stability of the object's composition
If dynamic range compression is applied to control overall amplitude, then the background sound levels are stabilized, but the audio signals from objects of interest cannot increase independently
Solution Approach 1:
The system divides the audio processing into separate channels for different audio objects. Dynamic range compression is applied selectively to background objects to stabilize their levels, while objects of interest are processed through separate amplitude control mechanisms that allow independent adjustment based on user proximity.
Solution Approach 2:
The system implements dynamic amplitude control where the processing characteristics change based on real-time conditions. As users approach objects of interest, the system dynamically adjusts the amplitude of these objects independently from the compressed background mix, allowing adaptability while maintaining overall stability.
3Measurement precision
If continuous changes in sound levels are applied to reflect user movement, then the spatial awareness is improved, but the immersive experience is detracted from
Solution Approach 1:
Amplitude changes are applied locally to specific audio objects rather than uniformly across the entire audio scene. This allows precise spatial awareness through directional volume changes of objects of interest while maintaining stable background levels, preventing the continuous fluctuations that detract from immersion.
Solution Approach 2:
The system implements conditional dynamic processing where amplitude changes are triggered by specific conditions (user proximity to objects of interest) rather than continuous movement. This creates realistic spatial cues only when relevant, maintaining immersion during stable periods.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
An apparatus is disclosed, comprising means for identifying, from a plurality of audio objects in an audio scene, one or more audio objects of interest, and means for processing first audio signals associated with the plurality of objects for provision to a user device. The processing may be based on the position of the user device in the audio scene. The processing may comprise combining the first audio signals associated with the audio objects to form combined first audio signals, modifying the amplitude of the combined first audio signals and limiting to a first level the maximum amplitude of the combined first audio signals , and modifying the amplitude of one or more individual first audio signals, associated with the one or more audio objects of interest, said modifying being independent of that for the combined first audio signals.