Camera Network Feature Extraction for Shopping Event Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In dynamic environments like materials handling facilities and financial institutions, detecting and locating large numbers of objects or actors using cameras is computationally expensive and resource-intensive, requiring substantial data storage, processing, and transmission capacities, and often results in lengthy processing times.
Innovation Solution
A system of networks of cameras configured to capture and process images, using machine learning models like transformers to generate hypotheses about shopping events by analyzing sequences of images, determining positions of body parts in 3D space, and predicting interactions with product spaces, thereby reducing the computational load and processing time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If large numbers of cameras are used to capture imaging data for detecting and locating objects or actors, then detection capability and coverage are improved, but computational cost, data storage requirements, and processing time increase substantially
Solution Approach 1:
The patent segments the complex task of object detection by introducing specialized hardware modules (depth estimation module, pose estimation module, interaction detection module) that process different aspects of the imaging data independently. This modular segmentation reduces the computational burden on centralized systems while maintaining comprehensive detection capabilities across multiple cameras.
Solution Approach 2:
The patent introduces intermediate processing modules that act as mediators between the camera array and the centralized system. These modules perform preliminary processing of imaging data, extracting key features and reducing data volume before transmission, thereby decreasing computational cost and data storage requirements while preserving detection reliability.
2Measurement precision
If imaging data from large numbers of cameras is processed to determine events or interactions, then detection accuracy is improved, but processing time becomes lengthy
Solution Approach 1:
The patent performs preliminary processing actions by estimating depth and pose information from imaging data before full event detection. This preliminary extraction of geometric features prepares the data in advance, allowing the centralized system to quickly determine interactions without reprocessing raw imaging data, thus reducing processing time while maintaining detection accuracy.
Solution Approach 2:
The patent replaces computationally intensive mechanical processing of raw imaging data with optimized algorithms and specialized hardware modules that estimate depth and pose more efficiently. This substitution reduces processing time by using dedicated computational pathways rather than general-purpose image processing pipelines.
3Reliability
If comprehensive imaging data is captured and stored for event detection, then event detection reliability is improved, but data storage and transmission capacities are substantially consumed
Solution Approach 1:
The patent extracts only the essential geometric features (depth estimates, pose information, body part locations) from comprehensive imaging data before storage and transmission. By taking out only the critical information needed for event detection rather than storing complete high-resolution images, the system maintains event detection reliability while substantially reducing data storage and transmission requirements.
Data Source
AI summary
Cameras having storage fixtures within their fields of view are programmed to capture images and process clips of the images to generate sets of features representing product spaces and actors depicted within such images, and to classify the clips as depicting or not depicting a shopping event. Where consecutive clips are determined to depict a shopping event, features of such clips are combined into a sequence and transferred, along with classifications of the clips and a start time and end time of the shopping event, to a multi-camera system. A shopping hypothesis is generated based on such sequences of features received from cameras, along with information regarding items detected within the hands of such actors, to determine a summary of shopping activity by an actor, and to update a record of items associated with the actor accordingly.


