Occluded item detection for vision-based self-checkouts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current item checkout methods are inefficient and error-prone when dealing with multiple items, especially when they are occluded, requiring extensive training data and manual annotation, leading to poor accuracy and long training times.
Innovation Solution
A vision-based detection system that uses two machine-learning algorithms to recognize multiple items with occluded views, reducing the need for exhaustive training data by focusing on features and angles of individual items and pairs of items, allowing for accurate identification without scanning each item individually.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple items are placed on the checkout counter for recognition, then item checkout efficiency is improved, but items partially cover or occlude full views of one another leading to poor recognition accuracy
Solution Approach 1:
The patent segments the item recognition task into two distinct machine learning algorithms: one trained on non-occluded item views and another trained on occluded item views. This segmentation allows each algorithm to specialize in specific detection scenarios, resolving the contradiction by handling multiple items (improving productivity) while maintaining accuracy through specialized recognition models for different occlusion states
Solution Approach 2:
The patent changes the training parameters and data characteristics for different machine learning algorithms. One algorithm is trained on images with clear, non-occluded item views, while another is trained specifically on images where items are partially occluded. This parameter change in training data characteristics enables accurate recognition under varying occlusion conditions, maintaining measurement precision while enabling multi-item processing
2Measurement precision
If training is performed on all possible combinations of multiple items from different positions and angles, then item recognition accuracy is improved, but the training process becomes infeasible due to the enormous size of training images required
Solution Approach 1:
The patent divides the comprehensive training task into separate, manageable segments: one training set for non-occluded views and another for occluded views. This segmentation reduces the complexity of training data by avoiding the need to create all possible combinations of multiple items from different positions and angles, making the training process feasible while maintaining recognition accuracy through specialized models
Solution Approach 2:
Instead of training on all possible combinations of multiple items (excessive action), the patent trains on representative subsets: non-occluded views and occluded views. This partial action approach captures the essential recognition patterns without requiring exhaustive training data, reducing training complexity while achieving sufficient accuracy for practical application
3Measurement precision
If exhaustive training data covering all item combinations is collected and annotated, then item recognition accuracy is improved, but the time and resources required for manual annotation increase significantly
Solution Approach 1:
The patent segments the annotation task into two distinct workflows: annotating non-occluded item views and annotating occluded item views. This segmentation reduces the total annotation time by focusing on specific occlusion scenarios rather than requiring exhaustive annotation of all possible multi-item combinations, thereby improving measurement precision while reducing the loss of time
Data Source
AI summary
Item recognition of a given item is trained on a single item from different views. The item recognition is then trained on images of the given item partially occluded by a second item having same, similar, or different shapes and features to that of the given item. General features of the item are noted and used to detect the given item when the given item is presented with multiple different items having multiple different occluded views.


