Compositing Angularly Separated Sub-Scenes in Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videoconferencing technologies face challenges in effectively capturing and displaying angularly separated sub-scenes within a wide scene, leading to difficulties in following audio and non-verbal cues, especially in multi-person settings, where audio feedback and skewed camera perspectives become significant issues.
Innovation Solution
A method involving the recording of a panoramic video signal with a wide camera, subsampling sub-scene video signals at specific bearings of interest, and compositing them side-by-side to form a stage scene video signal, with additional sub-scene signals being transitioned or removed to maintain an aspect ratio of 2:1 or less, ensuring seamless and automatic viewing for remote participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single camera with limited horizontal field of view (70 degrees) is used, then the camera setup is simple, but it is difficult for remote party to follow audio, body language, and non-verbal cues from participants at sharp angles to the camera
Solution Approach 1:
The patent divides the wide scene into multiple angularly separated sub-scenes, each captured by the single wide camera. These sub-scenes are then composited together to create a multi-person view that preserves body language and non-verbal cues from all participants, resolving the information loss while maintaining simple camera setup
Solution Approach 2:
The patent transitions from a single perspective view to a multi-perspective composited view by capturing sub-scenes at different bearings and combining them. This dimensional transformation allows remote participants to see participants at sharp angles clearly, as if multiple cameras were used, while actually using a single camera
2Loss of information
If cameras of two or more mobile devices are used to capture multi-person scene, then more participants can be captured, but audio feedback and crosstalk increase significantly
Solution Approach 1:
The patent merges multiple sub-scenes captured by a single camera into one composited video output. This consolidation achieves multi-person coverage equivalent to multiple cameras while using only one camera, thereby eliminating the audio feedback and crosstalk issues that arise from multiple devices
Solution Approach 2:
The single wide camera performs multiple functions that would traditionally require multiple cameras: capturing wide scenes, capturing individual participants at sharp angles, and providing seamless audio-visual synchronization. This multi-functionality eliminates the need for multiple devices and their associated audio problems
3Loss of information
If multiple sub-scene video signals are composited side-by-side, then visibility of non-verbal cues improves, but the aspect ratio increases beyond standard single camera format
Solution Approach 1:
The patent dynamically adjusts parameters including sub-scene width, bearing angles, and compositing layout to fit within standard aspect ratios. By changing these parameters, the system preserves body language visibility while conforming to conventional video formats for seamless remote viewing
4Loss of information
If additional sub-scene video signals are added to the composited output, then participant coverage increases, but the complexity of compositing and tracking increases
Solution Approach 1:
The system automatically performs compositing and tracking operations without manual intervention. The processor autonomously identifies participants, determines optimal bearings, composites sub-scenes, and adjusts layouts in real-time, reducing operational complexity despite handling multiple participants
Data Source
Figure 1A~1B
Figure 2A~2D
Figure 2L~2G
AI summary
A densely composited single camera signal may be formed from a panoramic video signal having an aspect ratio of substantially 2.4: 1 or greater, captured from a wide camera. Two or more sub-scene video signals are subsampled at respective bearings of interest, and may be composited side-by-side to form a stage scene video signal having an aspect ratio of substantially 2: 1 or less. 80% or more of the area of the stage scene video signal may be subsampled from the panoramic video signal.