Vision-Based Shelf Event Detection for Virtual Cart Updates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to accurately detect and record events, such as item picking or returning, in environments like retail facilities using image data without relying heavily on additional sensors.

Innovation Solution

Utilizing convolutional neural networks (CNNs) and computer vision algorithms to analyze image data from cameras, identifying interactions and generating event data to update virtual carts, including segmentation maps, customer-interaction scores, and direction/velocity of hands, to determine events like picking or returning items.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional sensor-based systems are used to detect events in retail facilities, then detection reliability is improved, but device complexity and cost increase due to additional sensors

Engineering Contradiction:
Improveevent detection reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical/sensor-based event detection systems with a computer vision system using cameras and CNN algorithms. The system captures images from existing cameras, processes them through trained neural networks to detect customer interactions with products, and generates event data without requiring additional physical sensors in the environment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates a virtual representation of physical events by generating segmentation maps and interaction data from image processing. Instead of directly detecting physical interactions with sensors, the system creates digital copies of the interaction through image analysis, identifying customers, products, and interaction types from visual data.

Inventive Principle:
Principle #26Copying

2Device complexity

If computer vision algorithms are used to analyze image data, then device complexity is reduced, but measurement precision of events may worsen

Engineering Contradiction:
Improvesystem complexityVSAvoidevent detection precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system performs preliminary training of CNN models using labeled interaction data before deployment. The training phase prepares the neural network to recognize specific interaction patterns, and the trained model is then used for accurate event detection. This preliminary action ensures the system can precisely measure and detect events without requiring complex real-time processing adjustments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies segmentation by dividing the image analysis into distinct components: customer identification, product identification, interaction type classification, and confidence scoring. The CNN model processes images to generate segmentation maps that separate different elements of the scene, allowing precise measurement of each component independently and improving overall detection accuracy.

Inventive Principle:
Principle #1Segmentation

3Productivity

If real-time event detection is implemented, then productivity is improved, but use of energy increases due to continuous image processing

Engineering Contradiction:
Improveinventory management efficiencyVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system processes images at periodic intervals rather than continuously analyzing every frame. The CNN model evaluates images at scheduled time points or when specific triggers occur (such as customer entry or product availability changes), reducing computational load and energy consumption while maintaining effective real-time monitoring capabilities for inventory management.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS12524503B1System and method for vision-based event detection
Publication Date: 2026.01.13 AMAZON TECH INC
  • US12524503B1 patent drawing
  • US12524503B1 patent drawing
  • US12524503B1 patent drawing

AI summary

This disclosure describes systems and techniques for identifying events that occur within an environment using image data captured at the environment. For example, one or more cameras may generate image data representative of a user interacting with an item on the shelf. This image data may be used to generate feature data associated with the user and the item, which may be analyzed by one or more classifiers for identifying an interaction between the user and the item. The systems and techniques may then generate interaction data, which in turn may be analyzed by one or more additional classifiers for identifying an event, such as the user picking a particular item from the shelf within the environment. Event data indicative of the event may then be used to update a virtual cart of the user.