Multi-Camera Video-Audio Alignment via Spatial Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-camera devices, such as mobile communication devices, face challenges in synchronizing video and audio data captured from multiple cameras, leading to mismatches between visual and audio representations, which can confuse users and detract from immersive experiences in applications like virtual or augmented reality.
Innovation Solution
The method involves using multiple cameras to capture video data from different regions and spatial microphones to capture directional audio, with video and audio mappings that align the data to ensure synchronized and immersive outputs, allowing for side-by-side or stacked presentation of video and audio data, and optionally including or stretching audio from outside regions to provide a 360-degree experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If multiple cameras are used to capture video data from different regions, then the field of view and coverage are improved, but the synchronization and alignment between video and audio data become more difficult
Solution Approach 1:
The patent divides the audio field into multiple directional regions (first audio region, second audio region, third audio region) corresponding to different camera fields of view. Each region is processed independently with its own audio mapping, allowing precise alignment between video and audio data from different cameras while maintaining overall synchronization.
Solution Approach 2:
Different audio mappings are applied to different spatial regions. The first audio mapping aligns audio from the first region with the first camera video, the second audio mapping aligns audio from the second region with the second camera video, and the third audio mapping handles audio from outside both regions. This local processing ensures precise video-audio alignment for each camera while maintaining global coherence.
2Adaptability or versatility
If audio data from outside the camera regions is included to provide 360-degree experience, then the immersive quality is improved, but the complexity of audio mapping and synchronization increases
Solution Approach 1:
The audio field is segmented into three distinct regions: inside the first camera region, inside the second camera region, and outside both regions. Each segment is processed with a dedicated audio mapping (first, second, or third audio mapping), simplifying the overall complexity by breaking down the 360-degree audio problem into manageable regional components.
Solution Approach 2:
The system uses a universal audio mapping framework that can handle multiple types of audio data (directional audio from within camera regions and ambient audio from outside regions) through a unified process. The same multi-camera video system serves both standard viewing and immersive 360-degree experiences.
Data Source
AI summary
This specification describes: using a first camera of a multi-camera device to obtain first video data of a first region; using a second camera of the multi-camera device to obtain second video data of a second region; generating a multi-camera video output from the first and second video data using a first video mapping to map the first video data to a first portion of the multi-camera video output and using a second video mapping to map the second video data to a second portion of the multi-camera video output; and generating an audio output from obtained audio data, the audio output comprising an audio output having a directional component within the first portion of the video output and an audio output having a directional component within the second portion of the video output, wherein generating the audio output comprises using a first audio mapping to map audio data having a directional component within the first region to the audio output having a directional component within the first portion of the video output and using a second audio mapping to map audio data having a directional component within the second region to the audio output having a directional component within the second portion of the video output.


