Camera Network Feature Extraction for Shopping Event Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In dynamic environments like materials handling facilities and financial institutions, detecting and locating large numbers of objects or actors using cameras is computationally expensive and resource-intensive, requiring substantial data storage, processing, and transmission capacities, and often results in lengthy processing times.

Innovation Solution

A system of networks of cameras configured to capture and process images, using machine learning models like transformers to generate hypotheses about shopping events by analyzing sequences of images, determining positions of body parts in 3D space, and predicting interactions with product spaces, thereby reducing the computational load and processing time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large numbers of cameras are used to capture imaging data for detecting and locating objects or actors, then detection capability and coverage are improved, but computational cost, data storage requirements, and processing time increase substantially

Engineering Contradiction:
Improvedetection capabilityVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex task of object detection by introducing specialized hardware modules (depth estimation module, pose estimation module, interaction detection module) that process different aspects of the imaging data independently. This modular segmentation reduces the computational burden on centralized systems while maintaining comprehensive detection capabilities across multiple cameras.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing modules that act as mediators between the camera array and the centralized system. These modules perform preliminary processing of imaging data, extracting key features and reducing data volume before transmission, thereby decreasing computational cost and data storage requirements while preserving detection reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If imaging data from large numbers of cameras is processed to determine events or interactions, then detection accuracy is improved, but processing time becomes lengthy

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing actions by estimating depth and pose information from imaging data before full event detection. This preliminary extraction of geometric features prepares the data in advance, allowing the centralized system to quickly determine interactions without reprocessing raw imaging data, thus reducing processing time while maintaining detection accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces computationally intensive mechanical processing of raw imaging data with optimized algorithms and specialized hardware modules that estimate depth and pose more efficiently. This substitution reduces processing time by using dedicated computational pathways rather than general-purpose image processing pipelines.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If comprehensive imaging data is captured and stored for event detection, then event detection reliability is improved, but data storage and transmission capacities are substantially consumed

Engineering Contradiction:
Improveevent detection reliabilityVSAvoiddata storage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential geometric features (depth estimates, pose information, body part locations) from comprehensive imaging data before storage and transmission. By taking out only the critical information needed for event detection rather than storing complete high-resolution images, the system maintains event detection reliability while substantially reducing data storage and transmission requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12131539B1Detecting interactions from features determined from sequences of images captured using one or more cameras
Publication Date: 2024.10.29 AMAZON TECH INC
  • US12131539B1 patent drawing
  • US12131539B1 patent drawing
  • US12131539B1 patent drawing

AI summary

Cameras having storage fixtures within their fields of view are programmed to capture images and process clips of the images to generate sets of features representing product spaces and actors depicted within such images, and to classify the clips as depicting or not depicting a shopping event. Where consecutive clips are determined to depict a shopping event, features of such clips are combined into a sequence and transferred, along with classifications of the clips and a start time and end time of the shopping event, to a multi-camera system. A shopping hypothesis is generated based on such sequences of features received from cameras, along with information regarding items detected within the hands of such actors, to determine a summary of shopping activity by an actor, and to update a record of items associated with the actor accordingly.