Stereo Vision Height Maps for Human Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current object detection and tracking systems for human objects in facilities face challenges such as background clutter, lighting conditions, occlusion, and the inability to distinguish between incoming and outgoing traffic, static and moving objects, and different heights, while also requiring high maintenance and budget. Additionally, existing systems struggle to separately track employees and customers in retail environments.
Innovation Solution
A computer-implemented system using stereo cameras to capture images and generate height maps, which are then analyzed for object detection and tracking, allowing for real-time, unobtrusive, and low-maintenance human object detection, tracking, and counting, with the ability to differentiate between entering and exiting traffic and distinguish object heights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular video streams or image sequences are used for object detection and tracking, then the system is simpler and less expensive, but the system performance is affected by background clutter, lighting conditions, and shadows
Solution Approach 1:
The patent transitions from 2D intensity images to 3D range/height maps by introducing depth information through stereo vision. This dimensional change allows the system to segment objects from background based on height differences, effectively eliminating background clutter and lighting-related issues that plague monocular systems.
Solution Approach 2:
The patent introduces range maps and height maps as intermediary representations between the raw video input and object detection. These intermediate depth-based representations serve as a mediator that transforms the problematic intensity-based input into a robust depth-based input for tracking, filtering out lighting variations and background clutter.
2Reliability
If stereo or multi-sensor systems are used to generate range or height maps, then the system is more robust to lighting conditions and shadows, but the system complexity and cost increase
Solution Approach 1:
The patent uses stereo vision to add the depth dimension, transforming 2D intensity images into 3D range maps. This dimensional enhancement provides robustness to lighting conditions while maintaining a relatively simple dual-camera configuration compared to other multi-sensor systems.
3Productivity
If range background differencing is used in stereo systems, then object detection is performed, but the system still suffers from background clutter and difficulty in differentiating between multiple closely positioned objects
Solution Approach 1:
The patent applies local quality by using height information specifically for background subtraction and object segmentation. Instead of using global intensity-based methods, the system locally evaluates height differences at each pixel position to determine foreground objects, enabling accurate detection even in cluttered environments with multiple closely positioned objects.
4Productivity
If adaptive template matching is used for object detection, then the system can track objects, but detections drift from true locations or get fixed to strong background features
Solution Approach 1:
The patent introduces height maps as an intermediary that mediates between template matching and object detection. By performing matching in the height domain rather than intensity domain, the system avoids getting locked onto strong background features, as height-based segmentation naturally separates foreground objects from static background.
5Productivity
If contour tracking is used for object tracking, then the system can follow object boundaries, but it suffers from degradation by intensity gradients in the background near object contours
Solution Approach 1:
The patent uses height maps as an intermediary representation that eliminates intensity gradient problems. By tracking object contours in the height domain rather than intensity domain, the system achieves robust tracking even when intensity gradients are present in the background near object contours.
Data Source
Figure 1~2
Figure 3
Figure 4~5
AI summary
A system for counting and tracking objects of interest within a predefined area with a sensor that captures object data and a data capturing device that receives subset data to produce reports that provide information related to a time, geographic, behavioral, or demographic dimension.