Audio Object Location Adjustment for VR Immersion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multimedia content production faces challenges in harmonizing video and audio elements, particularly in next-generation contents like VR, where matching visual and audio object locations is difficult, leading to a lack of immersion due to inconsistencies between visual and auditory stimuli.
Innovation Solution
An audio signal processing apparatus that includes a matching unit to select an audio object corresponding to a visual object, a location adjusting unit to adjust the sound image based on the selected audio and visual object locations, and an output unit to synchronize the audio signal, ensuring harmony between visual and audio elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional multi-channel stereo audio signal is used, then the audio content can be reproduced, but the audio location cannot be adjusted according to the direction of the user's head, resulting in deteriorated sense of immersion
Solution Approach 1:
The patent transforms the audio signal from traditional multi-channel stereo format to object-based audio format, changing the fundamental parameter structure. This allows audio objects to be independently positioned in three-dimensional space and adjusted according to user head direction, thereby achieving adaptive audio location adjustment while maintaining manageable processing complexity through standardized transformation processes.
Solution Approach 2:
The patent implements dynamic audio location adjustment by continuously tracking the user's head direction and repositioning audio objects in real-time according to the head's orientation. This dynamic adaptation enables the audio system to respond to user movement, enhancing immersion without requiring overly complex processing by using efficient spatial transformation algorithms.
2Reliability
If visual and audio objects are not matched in location, then content production is simpler, but the user experiences heterogeneity and loss of immersion due to inconsistency between visual and auditory stimuli
Solution Approach 1:
The patent introduces an object-based audio framework as an intermediary layer between traditional audio signals and the final rendered output. This intermediary system automatically matches audio objects with corresponding visual objects in three-dimensional space, ensuring synchronization and consistency between visual and auditory stimuli while simplifying the content production process through automated spatial alignment algorithms.
Solution Approach 2:
The patent performs preliminary spatial positioning of audio objects during the content production phase, establishing their three-dimensional locations in advance. This preliminary action enables automatic matching with visual objects during playback, ensuring synchronization without requiring complex real-time adjustments and reducing content production difficulty through pre-configured spatial relationships.
Data Source
AI summary
Provided are an audio signal processing method and apparatus for adjusting a location of an audio object in correspondence to a location of a visual object. The audio signal processing apparatus includes a matching unit configured to select an audio object corresponding to a visual object extracted from a video signal among at least one audio object extracted from an audio signal, a location adjusting unit configured to adjust a location of a sound image of the audio signal based on a location of the selected audio object and a location of a visual object corresponding to the selected audio, and an output unit configured to output an audio signal whose the location of the sound image is adjusted.


