Smart Container Item Identification via Multi-Modal Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current item identification methods in the new retail industry face challenges with low accuracy, high costs, and high loss rates due to limitations in RFID technology and visual recognition systems, particularly with metal and liquid items, and poor space utilization.
Innovation Solution
An item identification method and system that uses multi-frame image processing, multi-modality fusion of position and auxiliary information, and fine-grained classification to accurately determine item categories and numbers, incorporating image pre-processing, non-maximum suppression, and depth camera data for enhanced accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If RFID electronic tags are used for item identification, then item identification capability is achieved, but cost and labor cost increase significantly
Solution Approach 1:
The patent replaces expensive RFID electronic tags with a free, disposable alternative: visual codes (barcodes or images) printed directly on item packaging. This eliminates the need for purchasing and manually attaching costly electronic tags to each item, while still enabling reliable identification through image capture and recognition
Solution Approach 2:
The patent substitutes the radio frequency identification system with an optical identification system. Instead of using RFID readers and electronic tags that communicate via radio waves, the system uses cameras to capture images and software to recognize visual codes, replacing the mechanical/electrical RFID infrastructure with an optical-based solution
2Reliability
If RFID tags are attached to items, then item identification is enabled, but tags are easily torn off resulting in high loss rate
Solution Approach 1:
The patent creates a visual copy of the identification information directly on the item packaging through printed barcodes or images. This copy is permanently integrated into the packaging design, making it impossible to separate the identification element from the item itself, thereby eliminating the loss problem associated with detachable RFID tags
3Reliability
If camera is installed on top of each layer for visual recognition, then item identification is achieved, but space utilization rate decreases due to large distance requirement
Solution Approach 1:
The patent transitions from a top-down camera viewing angle to a side-view camera angle. By positioning the camera at the side of the container rather than above, items can be stacked vertically without requiring the camera to be at a large distance, thus improving space utilization while maintaining identification capability
Solution Approach 2:
The patent implements multiple cameras at different heights along the side of the container, with each camera responsible for capturing images of items in its specific vertical zone. This localized approach allows cameras to be positioned closer to items while ensuring all items within each camera's field of view are properly captured, improving both space utilization and identification accuracy
4Volume of moving object
If camera captures images from large distance, then space utilization improves, but identification accuracy decreases due to item sheltering
Solution Approach 1:
The patent divides the container into multiple vertical zones, each monitored by a dedicated camera. This segmentation ensures that each camera focuses on a specific region where items are more visible and less likely to be completely obscured by items in front, thereby maintaining identification accuracy while allowing for better space utilization
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An item identification method, system and electronic device are provided. The method includes: acquiring multi-frame images of the item by an image capturing device; processing the multi-frame images of the item to obtain position information and category information of the item in each frame image; acquiring auxiliary information of the item by an information capturing device; performing multi-modality fusion on the position information and the auxiliary information to obtain a fusion result; and determining an identification result of the item according to the category information and the fusion result. Through at least some embodiments of the present invention, a problem of low identification accuracy when identifying an item in the related art is partially solved.