Image Item Identification Using Region-of-Interest Feature Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack an efficient method to identify specific items within images by comparing features of the items to regions-of-interest (ROIs) in a way that accurately retrieves images depicting the item, especially in varying contexts and depths.
Innovation Solution
The use of a trained convolutional neural network or Siamese network to compare features of items with defined ROIs in images, calculating ROI sizes based on item size and image depth, and filtering regions by color similarity before feature comparison.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If feature comparison is performed across entire images, then item identification completeness is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent divides the image into multiple regions-of-interest (ROIs) based on spatial location and object characteristics. Instead of comparing features across the entire image, the system performs feature comparison only within these segmented regions, significantly reducing computational complexity while maintaining identification accuracy for specific items like clothing.
Solution Approach 2:
The patent applies different processing strategies to different regions of the image based on their characteristics. ROIs are selected and processed with appropriate feature comparison methods tailored to their specific properties (e.g., color histograms for clothing items), improving efficiency without sacrificing identification reliability.
2Measurement precision
If ROI size is increased to capture more context, then item identification accuracy is improved, but false positive rate increases
Solution Approach 1:
The patent dynamically adjusts ROI size and position based on the specific item being searched for and the image characteristics. The system optimizes ROI parameters to capture sufficient context for accurate identification while minimizing inclusion of irrelevant regions that could cause false positives.
Solution Approach 2:
The patent modifies ROI parameters (size, position, shape) based on the search query and image analysis results. By changing these parameters adaptively, the system achieves optimal balance between capturing enough context for accurate identification and avoiding irrelevant regions that lead to false positives.
3Reliability
If multiple ROIs are analyzed per image, then item identification reliability is improved, but computational complexity increases
Solution Approach 1:
The patent segments the image into multiple ROIs and applies feature comparison to each region. This segmentation allows parallel processing of different regions, improving reliability through comprehensive coverage while managing computational complexity through efficient region selection and processing strategies.
Solution Approach 2:
The patent analyzes multiple ROIs per image, performing more comparisons than a single-region approach. This partial or excessive action across multiple regions improves identification reliability by reducing false negatives, while the system manages the increased computational load through efficient algorithms and selective ROI processing.
Data Source
AI summary
The techniques described herein may identify images that likely depict one or more items by comparing features of the items to features of different regions-of-interest (ROIs) of the images. For instance, some of the images may include a user, and the techniques may define multiple regions within the image corresponding to different portions of the user. The techniques may then use a trained convolutional neural network or any other type of trained classifier to determine, for each region of the image, whether the region depicts a particular item. If so, the techniques may designate the corresponding image as depicting the item and may output an indication that the image depicts the item. The techniques may perform this process for multiple images, outputting an indication of each image deemed to depict the particular item.


