Camera Calibration Using Human Skeleton Ground-Plane Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing camera systems with overlapping fields of view face challenges in efficiently calibrating and tracking objects across multiple cameras, requiring time-consuming manual processes and recalibration due to camera movement or vibration, which disrupt normal operation.
Innovation Solution
Utilizing depth sensors and two-dimensional image sensors to detect human skeletons, determine a ground plane by tracking skeletal representations, and compute camera transformation matrices based on overlapping views to establish spatial relationships among cameras, enabling automatic calibration and tracking without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual calibration processes are used for camera systems, then calibration accuracy can be achieved, but the process is time-consuming and requires manual intervention
Solution Approach 1:
The system performs automatic calibration using skeletal representations of humans detected in the environment. The camera system calibrates itself by tracking human skeletons and computing transformation matrices without manual intervention, eliminating the need for operators to perform time-consuming manual calibration procedures while maintaining accurate spatial relationships between cameras
Solution Approach 2:
The system pre-establishes a library of skeletal representations and ground plane models that can be quickly applied during calibration. By having these reference models prepared in advance, the system can rapidly perform calibration operations when humans are detected in the camera fields of view, significantly reducing calibration time compared to manual methods
2Area of stationary object
If cameras are positioned to cover wide areas, then monitoring coverage is improved, but camera movement or vibration causes recalibration needs
Solution Approach 1:
The system continuously monitors the environment for humans and performs recalibration automatically when skeletal representations are detected. This feedback mechanism allows the system to maintain accurate calibration despite camera movement or vibration, as long as humans are present in the fields of view to provide reference points for recalibration
Solution Approach 2:
The calibration system is designed to be dynamic rather than static, allowing automatic recalibration when needed. The system adapts to camera position changes by continuously detecting humans and updating transformation matrices, enabling wide-area coverage with movable or vibration-prone cameras without sacrificing calibration reliability
3Adaptability or versatility
If multiple cameras with overlapping fields of view are used, then object tracking across views is enabled, but correspondence establishment between tracks becomes complex
Solution Approach 1:
The system uses skeletal representations as an intermediary to establish correspondence between camera views. Instead of directly matching complex object tracks across multiple cameras, the system first detects human skeletons and uses their standardized anatomical landmarks as intermediate reference points, greatly simplifying the correspondence establishment process while enabling effective multi-camera tracking
Solution Approach 2:
The system transforms the calibration problem from matching arbitrary image features to matching standardized skeletal parameters. By changing the reference framework from general image coordinates to anatomically-defined skeletal landmarks, the system simplifies correspondence computation across multiple cameras with overlapping fields of view
Data Source
Figure 1~2
Figure 3A~3C
Figure 4
AI summary
Examples are disclosed herein that relate to automatically calibrating cameras based on human detection. One example provides a computing system comprising instructions executable to receive image data comprising depth image data and two-dimensional image data of a space from a camera, detect a person in the space via the image data, determine a skeletal representation for the person via the image data, determine over a period of time a plurality of locations at which a reference point of the skeletal representation is on a ground area in the image data, determine a ground plane of the three-dimensional representation based upon the plurality of locations at which the reference point of the skeletal representation is on the ground area in the image data, and track a location of an object within the space relative to the ground plane.