Actor Localization via Camera Machine Learning and Trajectory Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In dynamic environments like materials handling facilities and financial institutions, detecting and locating large numbers of objects or actors using digital cameras is challenging due to the computational expense and data storage requirements of generating 3D models from imaging data captured by multiple cameras.
Innovation Solution
A distributed system utilizing cameras programmed with machine learning tools to detect body parts, predict positions in 3D space, and correlate trajectories based on visual descriptors, allowing for real-time detection and localization of actors even when temporarily lost or at low confidence, using a network of cameras and a central server to manage and resolve confusion between trajectories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D models are generated from imaging data captured by large numbers of digital cameras, then measurement precision and object detection capability are improved, but device complexity and computational cost increase substantially
Solution Approach 1:
The patent divides the complex task of 3D object detection into multiple stages: initial detection using simplified models, candidate generation, and refined localization. This segmentation allows the system to achieve high measurement precision without requiring all cameras to process full 3D models simultaneously, thereby reducing overall system complexity.
Solution Approach 2:
The patent introduces intermediate data structures and processing layers between the cameras and final 3D model generation. These intermediaries (such as 2D detection results, candidate lists, and projected model views) act as mediators that reduce the computational burden on individual cameras while maintaining overall detection precision.
2Measurement precision
If 3D models are generated from multiple camera imaging data, then object detection capability is improved, but data storage and processing capacity requirements increase
Solution Approach 1:
The patent extracts only the essential information needed for detection from the full imaging data. Instead of storing and processing complete 3D models from all cameras, the system extracts key features and characteristics, storing only what is necessary for accurate object detection and tracking.
Solution Approach 2:
The patent employs partial processing where not all cameras contribute full 3D model data to every detection task. Instead, only relevant camera data is processed for each object, and simplified models are used initially with full models generated only when necessary, reducing overall data storage requirements.
3Measurement precision
If 3D models are generated from multiple camera imaging data, then object detection capability is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary detection using simplified 2D analysis and projected model views before committing to full 3D model generation. This preliminary action identifies candidate objects and regions of interest, allowing the system to achieve fast initial detection with minimal processing time, and only invest heavy computational resources when actually needed.
Solution Approach 2:
The patent uses periodic updates of 3D models rather than continuous generation from all cameras. Models are updated at appropriate intervals and only when changes are detected, reducing processing time while maintaining detection precision through periodic refinement.
Data Source
AI summary
Motion of actors within a scene may be detected based on imaging data, using machine learning tools operating on cameras that captured the imaging data. The machine learning tools process images to perform a number of tasks, including detecting heads of actors, and sets of pixels corresponding to the actors, before constructing line segments from the heads of the actors to floor surfaces on which the actors stand or walk. The line segments are aligned along lines extending from locations of heads within an image to a vanishing point of a camera that captured the image. Trajectories of actors and visual data are transferred from the cameras to a central server, which links trajectories captured by multiple cameras and locates detected actors throughout the scene, even when the actors are not detected within a field of view of at least one camera.


