Stereo Height Map Clustering for Human Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for detecting, tracking, and counting human objects of interest are prone to issues such as background clutter, object detection drift, occlusion, and difficulty in distinguishing between multiple objects or different heights, especially under varying lighting conditions.
Innovation Solution
A computer-implemented system and method that uses stereo image pairs to generate height maps, allowing for robust detection, tracking, and counting of human objects by leveraging range information rather than intensity images, and includes features for accurate height calculation and differentiation between moving and static objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular video streams are used for object detection, then the system is simpler and cheaper, but the detection is affected by background clutter and lighting changes
Solution Approach 1:
The patent transitions from monocular 2D intensity images to stereo 3D range images, adding the depth dimension. This allows the system to detect objects based on their spatial position and height, making detection robust to lighting changes and background clutter that plague 2D intensity-based methods.
Solution Approach 2:
The patent replaces the optical intensity-based detection mechanism with a geometric range-based detection mechanism. Instead of relying on pixel intensity variations affected by lighting, the system uses measured 3D spatial coordinates that are inherently immune to illumination conditions.
2Adaptability or versatility
If adaptive template matching is used, then object detection can adapt to changes, but detections drift from true locations or get fixed to background features
Solution Approach 1:
The patent replaces template matching algorithms with a geometric approach using 3D range data. Instead of comparing 2D intensity patterns that can drift or lock onto background features, the system uses spatial coordinate geometry and height map analysis, which provides precise localization based on actual 3D object positions.
3Reliability
If range background differencing is used in stereo systems, then lighting robustness is improved, but difficulty in differentiating between multiple closely positioned objects remains
Solution Approach 1:
The patent utilizes the full 3D spatial information from stereo range images, including height (Z-dimension), to differentiate between closely positioned objects. By analyzing the height map and 3D spatial relationships, the system can separate objects that may appear overlapping in 2D projections but have distinct vertical positions or depth separations.
Solution Approach 2:
The patent segments the range image data into distinct object regions based on 3D spatial criteria, including height thresholds and spatial separation. This segmentation approach, applied to 3D range data rather than 2D intensity data, enables better separation of closely positioned objects by exploiting their spatial and vertical differences.
4Speed
If contour tracking is used, then object tracking can follow motion, but degradation by intensity gradients in the background near contours reduces performance
Solution Approach 1:
The patent replaces intensity gradient-based contour tracking with 3D spatial coordinate-based tracking. Instead of following contours defined by intensity variations that are degraded by background gradients, the system tracks objects using their 3D spatial positions derived from range data, which are unaffected by intensity gradients.
Data Source
AI summary
Techniques for evaluating height data are disclosed herein. Height values for a tracked object are obtained over a specified period of time. Clusters including at least one of the height values are generated. The cluster having the largest amount of height values is identified. A height of the tracked object is estimated from the identified cluster.


