Automatic Video Framing Using Single Camera and Digital Compositing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conferencing systems require multiple cameras and complex processing units to automatically frame participants, resulting in high costs and computational inefficiencies.
Innovation Solution
A single high-resolution camera with a wide-angle lens, such as a fish-eye lens, is used to capture video images, employing video analytics to digitally pan, tilt, and zoom on participants, while leveraging integrated graphics processing units for efficient feature detection and framing, reducing the need for additional processing units and minimizing computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras and complex processing units are used to automatically frame participants, then framing accuracy is improved, but system cost and device complexity increase
Solution Approach 1:
The patent combines multiple camera feeds into a single composite video frame by detecting regions of interest in each camera feed and compositing them together. This merging approach allows the system to achieve accurate participant framing using multiple camera inputs while processing them through a unified compositing algorithm, thereby improving framing accuracy without proportionally increasing system complexity
Solution Approach 2:
The patent replaces physical camera movement mechanisms with digital image processing techniques. Instead of mechanically moving cameras to track participants, the system uses video analytics and compositing algorithms to digitally frame participants within the video conference, eliminating the need for complex mechanical tracking systems while maintaining framing accuracy
2Reliability
If multiple cameras and independent processing units are used for participant tracking, then framing reliability is improved, but computational requirements and cost increase
Solution Approach 1:
The patent performs preliminary detection of regions of interest in each camera feed before compositing. By pre-identifying participants and their locations in individual camera feeds using video analytics, the system ensures reliable participant tracking is established early in the processing pipeline, allowing subsequent compositing operations to proceed efficiently without requiring redundant processing
Solution Approach 2:
The patent employs a universal compositing algorithm that can process video feeds from multiple cameras with different resolutions and orientations. This multi-functional approach allows the same processing unit to handle various camera configurations and participant arrangements, improving framing reliability across different scenarios without requiring specialized processing units for each case
3Adaptability or versatility
If conventional video conferencing systems use multiple cameras for automatic framing, then participant tracking capability is improved, but system cost increases
Solution Approach 1:
The patent creates a composite video frame that copies and combines regions of interest from multiple camera feeds. By digitally copying participant images from different camera views and compositing them into a unified frame, the system achieves versatile participant tracking capability without requiring physical camera movement or switching, thereby reducing system complexity and cost
Data Source
AI summary
Systems and methods are described for automatically framing participants in a video conference using a single camera of a video conferencing system. A camera of a video conferencing system may capture video images of a conference room. A processor of the video conferencing system may identify a potential region of interest within a video image of the captured video images, the potential region of interest including an identified participant. Feature detection may be executed on the potential region of interest, and a region of interest may be computed based on the executed feature detection. The processor may then automatically frame the identified participant within the computed region of interest, the automatic framing including at least one of cropping the video image to match the computed region of interest and rescaling the video image to a desired resolution.


