Spatial Congruency Adjustment in Omnidirectional Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing systems face challenges in spatial congruency between auditory and visual scenes, leading to an unnatural conferencing experience due to discrepancies in sound field alignment with the visual scene, especially when using multi-channel audio formats like 5.1 or 7.1 surround sound.
Innovation Solution
A method and system that adjust spatial congruency by capturing and processing audio and visual scenes using omnidirectional devices, converting B-format signals, and applying head-related transfer functions to synchronize the auditory and visual scenes through transducers and displays, ensuring alignment within predefined thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-channel audio formats (5.1 or 7.1 surround) are used to enhance spatial audio experience, then audio quality and immersion are improved, but spatial congruency discrepancies between auditory and visual scenes worsen
Solution Approach 1:
The patent introduces an audio object as an intermediary element that bridges the auditory and visual scenes. The audio object carries spatial position information that can be independently controlled and adjusted, serving as a mediator to align the auditory scene with the visual scene in omnidirectional video conferencing systems
Solution Approach 2:
The patent changes the spatial position parameter of the audio object to adjust spatial congruency. By modifying the position parameter of the audio object within a predefined threshold, the system aligns the auditory scene with the visual scene without requiring complex reprocessing of the entire multi-channel audio signal
2Reliability
If audio object position is adjusted to align auditory scene with visual scene, then spatial congruency is improved, but system complexity increases
Solution Approach 1:
The patent adjusts the spatial position parameter of the audio object to achieve spatial congruency alignment between auditory and visual scenes. This parameter-based adjustment approach simplifies the processing complexity compared to comprehensive scene reprocessing
Solution Approach 2:
The system compares the spatial position of the audio object with the visual scene to detect spatial congruency discrepancies. This feedback mechanism enables automatic adjustment of the audio object position to maintain alignment, reducing the need for complex manual processing
Data Source
Figure 1~3
Figure 4
Figure 5~6
AI summary
Example embodiments disclosed herein relate to spatial congruency adjustment. A method for adjusting spatial congruency in a video conference is disclosed. The method includes unwarping a visual scene captured by a video endpoint device into at least one rectilinear scene, the video endpoint device being configured to capture the visual scene in an omnidirectional manner, detecting spatial congruency between the at least one rectilinear scene and an auditory scene captured by an audio endpoint device that is positioned in relation to the video endpoint device. The spatial congruency being a degree of alignment between the auditory scene and the at least one rectilinear scene and in response to the detected spatial congruency being below the threshold, adjusting the spatial congruency. Corresponding system and computer program products are also disclosed.