PTZ Camera Automatic Framing with Depth Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing systems require manual adjustment of PTZ cameras, which is time-consuming and often results in capturing unnecessary room environments rather than focusing on participants, and face detection methods are inconsistent, especially in low light conditions or with partial views.
Innovation Solution
An automatic framing system utilizing a PTZ camera combined with depth cameras to distinguish between occupants and inanimate objects, adaptively determining their location and adjusting the camera settings to center participants in the main video frame, even in low light scenarios and with partial views, and dynamically updating the view as participants move.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual adjustment of PTZ camera is used, then the camera can be precisely positioned to frame occupants, but the process is time-consuming and tedious
Solution Approach 1:
The system enables automatic framing by having the camera system itself perform the positioning task through integrated depth cameras and automated control algorithms, eliminating the need for manual user intervention while maintaining precise framing of occupants
Solution Approach 2:
The patent replaces manual mechanical adjustment with an automated optical and computational system using depth cameras to detect occupant positions and algorithms to calculate optimal camera angles, substituting human operation with automated sensing and processing
2Quantity of substance
If the camera captures the entire room view, then all occupants are visible, but unnecessary empty surrounding environment is included
Solution Approach 1:
The system applies different capture strategies to different regions of the room by using depth information to identify occupied zones versus empty spaces, dynamically adjusting the camera field of view to focus on areas with occupants while minimizing or eliminating capture of empty surrounding environments
Solution Approach 2:
The camera framing is made dynamic and adaptive, automatically adjusting the visible area based on real-time detection of occupant positions and distributions, allowing the system to optimize the balance between capturing all occupants and minimizing unnecessary background space
3Extent of automation
If face detection methods are used for automatic framing, then the system can automatically locate occupants, but the detection becomes inconsistent in low light scenarios or with partial views
Solution Approach 1:
The patent introduces depth cameras as an intermediary sensing modality that operates independently of visible light conditions, providing reliable occupancy detection in low light scenarios where traditional face detection fails, and using this depth information to guide the framing process
Solution Approach 2:
The system changes the detection parameter from optical intensity-based face features to depth-based spatial occupancy measurements, which remain reliable under varying lighting conditions and can detect partial occlusions more effectively, thereby improving detection consistency across different environmental conditions
Data Source
AI summary
A dynamically adjustable framed view of occupants in a room is captured through an automatic framing system. The system employs a camera system, including a pan/tilt/zoom (PTZ) camera and one or more depth cameras, to automatically locate occupants in a room and adjust the PTZ camera's pan, tilt, and zoom settings to focus in on the occupants and center them in the main video frame. The depth cameras may distinguish between occupants and inanimate objects and adaptively determine the location of the occupants in the room. The PTZ camera may be calibrated with the depth cameras in order to use the location information determined by the depth cameras to automatically center the occupants in the main video frame for a framed view. Additionally, the system may track position changes in the room and may dynamically adjust and update the framed view when changes occur.


