Video Conference Camera System Using Digital Image Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems using multiple cameras and audio tracking limit the view to only one participant at a time, causing an unsatisfactory experience for remote participants and are expensive.
Innovation Solution
A video communication system employing a single wide-angle, high-resolution digital camera that automatically adjusts its pan, tilt, and zoom based on analytics and context, such as participant location and meeting content, to provide multiple views of the meeting room to remote participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple cameras and audio tracking are used to track active speakers, then speaker tracking capability is improved, but device complexity and cost increase, and the view is limited to only one participant at a time
Solution Approach 1:
The patent divides the high-resolution digital image into multiple regions of interest (ROIs) corresponding to different participants in the meeting room. Instead of using multiple cameras to track different speakers, a single camera captures the entire scene and digital segmentation identifies and extracts individual participant regions, allowing the system to switch between different participants' views without requiring multiple physical cameras.
Solution Approach 2:
The patent creates digital copies or extracts of different participant regions from a single captured image. By processing the image to identify and extract multiple ROIs, the system generates virtual views of different participants that can be displayed or transmitted, eliminating the need for multiple physical cameras while achieving multi-speaker tracking capability.
2Adaptability or versatility
If multiple cameras are used to provide different views, then view diversity is improved, but cost and device complexity increase
Solution Approach 1:
The patent transitions from a spatial solution (multiple cameras positioned at different locations) to a digital processing solution (single camera with image segmentation and ROI extraction). By moving the diversity generation from the physical dimension to the digital processing dimension, the system achieves multiple view perspectives from a single camera capture, reducing hardware complexity while maintaining view diversity.
3Manufacturing precision
If manual camera adjustment is performed for every meeting, then view optimization is improved, but time consumption increases
Solution Approach 1:
The patent implements automatic camera view optimization through digital image processing that autonomously identifies participants, determines optimal regions of interest, and adjusts the displayed view without manual intervention. The system self-adjusts by processing the captured image, identifying ROIs based on participant positions and meeting context, and automatically selecting or switching between different participant views, eliminating the need for manual camera adjustment while maintaining optimized viewing.
4Device complexity
If a single camera is used, then device complexity and cost are reduced, but the ability to provide multiple participant views simultaneously is worsened
Solution Approach 1:
The patent implements dynamic view switching and composition using a single camera by continuously processing images to identify active speakers and relevant participants, then dynamically adjusting which regions are displayed or emphasized. The system dynamically switches between different participant views based on meeting context, speaker activity, and participant positions, providing adaptive multi-participant viewing capability without requiring multiple simultaneous camera feeds.
Data Source
AI summary
A video communication system that includes a computer readable medium and a processor, coupled with a wide angle and high resolution digital camera and the computer readable medium. The processor causes the wide angle and high resolution digital camera to acquire a digital image of a local participant during a video communication session. The processor extracts a first image of a first set of objects and a second image of a second set of objects from the digital image and provides the extracted first and second images to a remote endpoint for display to another participant.


