Actor-Event Association via Articulated Models in Materials Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In dynamic environments like materials handling facilities, it is challenging to determine which individuals or objects are associated with specific events based on imaging data alone, especially when multiple cameras with fixed orientations capture a large number of people, objects, or machines with varying sizes and velocities.

Innovation Solution

The system uses digital imagery from multiple cameras to detect events, generate articulated models of actors by recognizing body parts, and rank these models based on features and motion to identify which actor is associated with the event, employing techniques such as deep neural networks and support vector machines to process images before, during, and after the event.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If multiple digital cameras with fixed orientations are deployed to capture imaging data in dynamic environments, then the coverage area and number of detectable actors increase, but the difficulty of distinguishing and associating actors with specific events increases significantly

Engineering Contradiction:
Improvecoverage areaVSAvoiddifficulty of associating actors with events
Core Design Contradiction:
Area of stationary objectVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the complex task of actor-event association into multiple processing stages: event detection, actor detection, feature extraction, and association ranking. Each stage handles a specific aspect of the problem, breaking down the overwhelming complexity of tracking multiple actors with varying poses and velocities across multiple camera views into manageable computational steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from two-dimensional image analysis to three-dimensional spatial reasoning by generating articulated models that represent actors in 3D space. This dimensional transformation enables the system to disambiguate actors across different camera views and temporal frames, solving the association problem that cannot be resolved within a single 2D image plane.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If digital cameras capture images of actors with varying sizes, shapes, and velocities, then the system can detect more diverse actors, but recognizing and distinguishing between their poses becomes exceptionally challenging

Engineering Contradiction:
Improveability to detect diverse actorsVSAvoidprecision of pose recognition
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter representation of actors from raw pixel data to articulated model parameters including joint positions, body orientation, and motion vectors. This parameter transformation standardizes the representation of actors regardless of their varying sizes, shapes, and velocities, enabling consistent pose recognition and distinction across diverse actor types.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system employs feedback mechanisms where detected actor features are used to refine the association between actors and events. The articulated models provide continuous feedback about actor position and pose, which is used to update tracking information and improve the accuracy of event-actor associations over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11284041B1Associating items with actors based on digital imagery
Publication Date: 2022.03.22 AMAZON TECH INC
  • US11284041B1 patent drawing
  • US11284041B1 patent drawing
  • US11284041B1 patent drawing

AI summary

In a materials handling facility, events may be associated with users based on imaging data captured from multiple fields of view. When an event is detected at a location within the fields of view of multiple cameras, two or more of the cameras may be identified as having captured images of the location at a time of the event. Users within the materials handling facility may be identified from images captured prior to, during or after the event, and visual representations of the respective actors may be generated from the images. The event may be associated with one of the users based on distances between the users' hands and the location of the event, as determined from the visual representations, or based on imaging data captured from the users' hands, which may be processed to determine which, if any, of such hands includes an item associated with the event.