Multi-Camera Video-Audio Alignment via Spatial Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-camera devices, such as mobile communication devices, face challenges in synchronizing video and audio data captured from multiple cameras, leading to mismatches between visual and audio representations, which can confuse users and detract from immersive experiences in applications like virtual or augmented reality.

Innovation Solution

The method involves using multiple cameras to capture video data from different regions and spatial microphones to capture directional audio, with video and audio mappings that align the data to ensure synchronized and immersive outputs, allowing for side-by-side or stacked presentation of video and audio data, and optionally including or stretching audio from outside regions to provide a 360-degree experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If multiple cameras are used to capture video data from different regions, then the field of view and coverage are improved, but the synchronization and alignment between video and audio data become more difficult

Engineering Contradiction:
Improvefield of viewVSAvoidvideo-audio alignment
Core Design Contradiction:
Area of stationary objectVSManufacturing precision

Solution Approach 1:

The patent divides the audio field into multiple directional regions (first audio region, second audio region, third audio region) corresponding to different camera fields of view. Each region is processed independently with its own audio mapping, allowing precise alignment between video and audio data from different cameras while maintaining overall synchronization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different audio mappings are applied to different spatial regions. The first audio mapping aligns audio from the first region with the first camera video, the second audio mapping aligns audio from the second region with the second camera video, and the third audio mapping handles audio from outside both regions. This local processing ensures precise video-audio alignment for each camera while maintaining global coherence.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If audio data from outside the camera regions is included to provide 360-degree experience, then the immersive quality is improved, but the complexity of audio mapping and synchronization increases

Engineering Contradiction:
Improve360-degree audio coverageVSAvoidaudio mapping complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio field is segmented into three distinct regions: inside the first camera region, inside the second camera region, and outside both regions. Each segment is processed with a dedicated audio mapping (first, second, or third audio mapping), simplifying the overall complexity by breaking down the 360-degree audio problem into manageable regional components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses a universal audio mapping framework that can handle multiple types of audio data (directional audio from within camera regions and ambient audio from outside regions) through a unified process. The same multi-camera video system serves both standard viewing and immersive 360-degree experiences.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11503226B2Multi-camera device
Publication Date: 2022.11.15 NOKIA TECHNOLOGIES OY
  • US11503226B2 patent drawing
  • US11503226B2 patent drawing
  • US11503226B2 patent drawing

AI summary

This specification describes: using a first camera of a multi-camera device to obtain first video data of a first region; using a second camera of the multi-camera device to obtain second video data of a second region; generating a multi-camera video output from the first and second video data using a first video mapping to map the first video data to a first portion of the multi-camera video output and using a second video mapping to map the second video data to a second portion of the multi-camera video output; and generating an audio output from obtained audio data, the audio output comprising an audio output having a directional component within the first portion of the video output and an audio output having a directional component within the second portion of the video output, wherein generating the audio output comprises using a first audio mapping to map audio data having a directional component within the first region to the audio output having a directional component within the first portion of the video output and using a second audio mapping to map audio data having a directional component within the second region to the audio output having a directional component within the second portion of the video output.