Actor Localization via Camera Trajectory Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In dynamic environments like materials handling facilities and financial institutions, detecting and locating large numbers of objects or actors using digital cameras is challenging due to the computational expense and data storage requirements of generating 3D models from imaging data captured by multiple cameras.

Innovation Solution

A distributed system of cameras equipped with machine learning tools that detect body parts, generate trajectories, and provide visual descriptors to a server for correlation, allowing for real-time localization of actors in 3D space without relying on extensive depth imaging data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 3D models are generated from imaging data captured by large numbers of digital cameras, then object detection and location accuracy is improved, but computational expense and data storage requirements increase substantially

Engineering Contradiction:
Improveobject detection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex task of 3D object detection into separate processing stages: individual cameras capture 2D images, servers generate candidate object lists from each camera's data, and then correlate these lists across multiple cameras to identify objects visible to multiple cameras. This segmentation reduces the computational burden on any single component while maintaining detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of processing all imaging data from all cameras to generate complete 3D models for every detected object, the system performs partial action by generating candidate object lists from individual cameras and then selectively correlating these candidates across cameras. This approach achieves sufficient detection accuracy without the excessive computational cost of full 3D model generation for all objects.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If 3D models are generated from imaging data captured by large numbers of digital cameras, then object detection and location accuracy is improved, but data storage and transmission capacities are consumed

Engineering Contradiction:
Improveobject detection accuracyVSAvoiddata storage capacity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential information needed for object detection from each camera's imaging data - specifically, generating candidate object lists containing object identities and locations - rather than storing or transmitting complete 3D models. This extraction approach maintains detection accuracy while significantly reducing data storage and transmission requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If 3D models are generated for large numbers of objects in crowded environments, then comprehensive object location is achieved, but processing time increases comparatively lengthily

Engineering Contradiction:
Improveobject location completenessVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by generating candidate object lists from each camera's imaging data before correlating these candidates across multiple cameras. This preliminary processing organizes the data in advance, enabling faster correlation and identification of objects visible to multiple cameras without requiring lengthy processing of complete 3D models.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11398094B1Locally and globally locating actors by digital cameras and machine learning
Publication Date: 2022.07.26 AMAZON TECH INC
  • US11398094B1 patent drawing
  • US11398094B1 patent drawing
  • US11398094B1 patent drawing

AI summary

Motion of actors within a scene may be detected based on imaging data, using machine learning tools operating on cameras that captured the imaging data. The machine learning tools process images to perform a number of tasks, including detecting heads of actors, and sets of pixels corresponding to the actors, before constructing line segments from the heads of the actors to floor surfaces on which the actors stand or walk. The line segments are aligned along lines extending from locations of heads within an image to a vanishing point of a camera that captured the image. Trajectories of actors and visual data are transferred from the cameras to a central server, which links trajectories captured by multiple cameras and locates detected actors throughout the scene, even when the actors are not detected within a field of view of at least one camera.