Auto-Framing Video Conference Camera Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems face challenges in accurately framing meeting participants due to difficulties in manually adjusting camera settings, with voice-tracking cameras often losing track of speakers or directing to reflections rather than actual sound sources, especially in reverberant environments.
Innovation Solution
The implementation of a video conferencing apparatus that uses a combination of stationary and adjustable cameras, employing fast and reliable people detection methods such as face, torso, and motion detection, along with audio localization, to automatically and continuously adjust the camera view to provide an optimal framing of all participants or the speaker, eliminating the need for manual initiation and improving accuracy even when participants turn away or change positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice-tracking cameras are used to direct toward speakers, then speaker tracking capability is improved, but the system may lose track of speakers or direct at reflections in reverberant environments
Solution Approach 1:
The patent combines multiple detection methods (face detection, torso detection, motion detection) with audio localization to create a hybrid tracking system. This merging of visual and audio cues allows the system to maintain reliable speaker tracking even when audio alone is ambiguous due to reflections or when participants turn away from microphones.
Solution Approach 2:
The system introduces visual detection results as an intermediary to verify and supplement audio-based speaker identification. When audio localization is uncertain (due to reverberation or speaker orientation), the visual detection of faces and torsos serves as a mediator to confirm the actual speaker location, preventing misdirection to reflection points.
2Ease of operation
If manual camera adjustment is used, then framing control is achieved, but the operation becomes cumbersome and requires repeated adjustments
Solution Approach 1:
The system implements automated framing that performs camera adjustments without user intervention. The detection and tracking modules continuously monitor participant positions and automatically control camera pan, tilt, and zoom to maintain optimal framing, eliminating the need for manual operations and repeated adjustments when participants move.
Solution Approach 2:
The system establishes a feedback loop where detection results from visual and audio sensors continuously inform camera control decisions. This real-time feedback mechanism allows the camera to automatically adapt to changing participant positions and speaker changes, providing responsive framing without manual input.
3Extent of automation
If automated detection systems are implemented, then participant tracking is improved, but the system requires manual initiation and training mode switching
Solution Approach 1:
The system performs preliminary detection and tracking setup automatically upon initialization, eliminating the need for manual training mode activation. The detection modules begin operating immediately in the background, and the system transitions smoothly to active tracking without requiring user intervention to switch modes or initiate automated functions.
4Area of stationary object
If wide-angle fixed camera view is used, then coverage of entire room is achieved, but participant size in display becomes too small
Solution Approach 1:
The system transforms the camera from a static wide-angle view to a dynamic view that automatically adjusts pan, tilt, and zoom based on detected participant positions and speaker location. This dynamic adjustment allows the camera to maintain appropriate framing and participant size in the display while still being able to cover the entire room when needed.
Data Source
AI summary
A videoconference apparatus and method coordinates a stationary view obtained with a stationary camera to an adjustable view obtained with an adjustable camera. The stationary camera can be a web camera, while the adjustable camera can be a pan-tilt-zoom camera. As the stationary camera obtains video, participants are detected and localized by establishing a static perimeter around a participant in which no motion is detected. Thereafter, if no motion is detected in the perimeter, any personage objects such as head, face, or shoulders which are detected in the region bounded by the perimeter are determined to correspond to the participant.


