Multi-Camera Video Conference System Frontal View Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional panoramic cameras in video conferencing setups are typically installed in the middle of a conference room, resulting in limited viewing angles that restrict the capture of facial expressions and emotions, leading to incomplete or side/rear views of participants, which can hinder remote attendees' understanding of the communication.
Innovation Solution
A multi-camera video conference image processing system comprising a first panoramic camera and at least one second camera positioned at the front of the conference room, with a system-on-chip that processes images to generate panoramic and photographic frames, selects frames based on physical features, and corresponds them to ensure all participants are displayed with a frontal view, overcoming the limitations of the panoramic camera's installation position.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If a panoramic camera is installed in the middle of the conference room to capture all participants, then the coverage area is improved, but the viewing angle and facial feature capture are worsened
Solution Approach 1:
The system divides the conference room monitoring into two segments: a panoramic camera for overall coverage and front cameras for detailed facial capture. This segmentation allows each camera to specialize in its optimal function, resolving the contradiction between coverage area and facial feature precision.
Solution Approach 2:
The system introduces front cameras as intermediary devices that capture high-quality facial images, which are then integrated with panoramic views through image processing. This intermediary approach enables both wide coverage and precise facial feature capture simultaneously.
2Quantity of substance
If a panoramic camera is used to capture all participants, then the quantity of participants covered is improved, but the quality of individual facial expressions is worsened
Solution Approach 1:
The system merges images from multiple cameras (panoramic and front cameras) into a single integrated output. This combining approach allows the system to maintain high participant coverage while enhancing individual facial expression quality through the contribution of front camera images.
Solution Approach 2:
The system transitions from a single-dimension panoramic view to a multi-dimensional imaging approach by incorporating front cameras at different positions. This dimensional expansion enables simultaneous capture of both group coverage and individual facial details.
3Measurement precision
If multiple cameras are deployed to improve facial view quality, then the image quality is improved, but the system complexity is worsened
Solution Approach 1:
The system design allows cameras to serve multiple functions: front cameras can capture both individual facial expressions and contribute to overall room coverage when combined with panoramic images. This multi-functionality reduces the need for additional specialized devices, thereby limiting complexity growth.
Solution Approach 2:
The image processing system automatically selects and integrates the most appropriate views from different cameras based on real-time conditions, without requiring manual intervention. This self-service approach simplifies operation despite the multi-camera configuration.
Data Source
AI summary
The present invention relates to a multi-camera video conference image processing system, including: a first panoramic camera for capturing physical features of the conferees in a panoramic manner to generate a first panoramic image; at least one second camera for capturing the physical features of the conferees to generate at least one second image; and a system-on-chip for: receiving the first panoramic image and the at least one second image; processing the first panoramic image to generate a panoramic frame for each conferee; processing the at least one second image to generate a photographic frame for each conferee; corresponding the panoramic frame to the photographic frame for each conferee; selecting the panoramic frame or the photographic frame of each conferee based on the physical features; and processing the selected frames of each conferee, so as to generate and output a video frame.


