Actor Localization via Camera Trajectory Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In dynamic environments like materials handling facilities and financial institutions, detecting and locating large numbers of objects or actors using digital cameras is challenging due to the computational expense and data storage requirements of generating 3D models from imaging data captured by multiple cameras.
Innovation Solution
A distributed system of cameras equipped with machine learning tools that detect body parts, generate trajectories, and provide visual descriptors to a server for correlation, allowing for real-time localization of actors in 3D space without relying on extensive depth imaging data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D models are generated from imaging data captured by large numbers of digital cameras, then object detection and location accuracy is improved, but computational expense and data storage requirements increase substantially
Solution Approach 1:
The system segments the complex task of 3D object detection into separate processing stages: individual cameras capture 2D images, servers generate candidate object lists from each camera's data, and then correlate these lists across multiple cameras to identify objects visible to multiple cameras. This segmentation reduces the computational burden on any single component while maintaining detection accuracy.
Solution Approach 2:
Instead of processing all imaging data from all cameras to generate complete 3D models for every detected object, the system performs partial action by generating candidate object lists from individual cameras and then selectively correlating these candidates across cameras. This approach achieves sufficient detection accuracy without the excessive computational cost of full 3D model generation for all objects.
2Measurement precision
If 3D models are generated from imaging data captured by large numbers of digital cameras, then object detection and location accuracy is improved, but data storage and transmission capacities are consumed
Solution Approach 1:
The system extracts only the essential information needed for object detection from each camera's imaging data - specifically, generating candidate object lists containing object identities and locations - rather than storing or transmitting complete 3D models. This extraction approach maintains detection accuracy while significantly reducing data storage and transmission requirements.
3Measurement precision
If 3D models are generated for large numbers of objects in crowded environments, then comprehensive object location is achieved, but processing time increases comparatively lengthily
Solution Approach 1:
The system performs preliminary action by generating candidate object lists from each camera's imaging data before correlating these candidates across multiple cameras. This preliminary processing organizes the data in advance, enabling faster correlation and identification of objects visible to multiple cameras without requiring lengthy processing of complete 3D models.
Data Source
AI summary
Motion of actors within a scene may be detected based on imaging data, using machine learning tools operating on cameras that captured the imaging data. The machine learning tools process images to perform a number of tasks, including detecting heads of actors, and sets of pixels corresponding to the actors, before constructing line segments from the heads of the actors to floor surfaces on which the actors stand or walk. The line segments are aligned along lines extending from locations of heads within an image to a vanishing point of a camera that captured the image. Trajectories of actors and visual data are transferred from the cameras to a central server, which links trajectories captured by multiple cameras and locates detected actors throughout the scene, even when the actors are not detected within a field of view of at least one camera.


