Streaming Pooled Camera Features for Storage-Event Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for detecting events in dynamic environments, such as materials handling facilities, require large numbers of cameras and consume substantial data storage, processing, and transmission resources, leading to lengthy processing times.
Innovation Solution
A network of cameras captures images, processes them to determine location features, and streams pooled features to a system that generates hypotheses about interactions with storage units, using machine learning models to efficiently detect events.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If large numbers of individual cameras are used to capture imaging data from dynamic environments, then detection coverage and reliability are improved, but computational cost, data storage requirements, and processing time increase substantially
Solution Approach 1:
The patent segments the complex task of event detection by dividing it into two distinct stages: (1) cameras capture and process imaging data to extract location features and pool them locally, and (2) a separate system receives the pooled features and determines events. This segmentation reduces the computational burden on individual cameras and enables parallel processing across multiple devices, thereby improving reliability without proportionally increasing overall system complexity.
Solution Approach 2:
The patent extracts only the essential location features from imaging data at the camera level, rather than transmitting and processing complete imaging datasets. By taking out only the relevant pooled location features needed for event determination, the system reduces data storage requirements and computational costs while maintaining detection reliability through feature-based analysis.
2Measurement precision
If complete imaging data from large numbers of cameras is processed to determine events, then measurement precision is improved, but processing time and computational resources are substantially consumed
Solution Approach 1:
The patent extracts only the essential location features from imaging data at the camera level, rather than transmitting and processing complete imaging datasets. By taking out only the relevant pooled location features needed for event determination, the system reduces data storage requirements and computational costs while maintaining detection reliability through feature-based analysis.
Solution Approach 2:
The patent performs preliminary processing of imaging data at the camera level by extracting and pooling location features before transmission. This preliminary action prepares the data in advance, so that when the data reaches the event determination system, only the essential pooled features need to be analyzed, significantly reducing processing time while maintaining measurement precision.
3Measurement precision
If pooled location features are streamed from multiple cameras to a centralized system, then event detection accuracy is improved, but data transmission requirements increase
Solution Approach 1:
The patent extracts only the essential location features from imaging data at the camera level, rather than transmitting and processing complete imaging datasets. By taking out only the relevant pooled location features needed for event determination, the system reduces data storage requirements and computational costs while maintaining detection reliability through feature-based analysis.
Data Source
AI summary
Cameras having storage fixtures within their fields of view are programmed to capture images and process the images to determine whether such images depict an interaction by an actor with one of the storage fixtures. If a camera determines that an actor is present within imaging data captured thereby, the camera generates feature tensors from images captured over a predetermined period of time and pools the feature tensors into features corresponding to locations of the respective storage units of the actor within such images. Cameras provide pooled features to a multi-camera system that processes the pooled features to determine whether an interaction occurred at a storage unit, and to update a record of items associated with the actor accordingly.


