Product Labeling via Image Embedding and Weighted Voting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing e-commerce product search systems face challenges in accurately and automatically flagging and correcting erroneous or outdated product attributes in product catalogs, which affects user search experience.
Innovation Solution
A system and method that utilize image classification techniques, specifically by pre-training a neural network model to generate embedding vectors for item images, and then using a search engine to find visually similar products with verified labels, followed by a weighted majority vote to determine the accurate product label.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual or automatic computer system methods are used to assign product attributes, then productivity is improved, but measurement precision deteriorates leading to incorrect product attributes
Solution Approach 1:
The patent introduces an image classification system as an intermediary between product images and attribute assignment. The system uses trained machine learning models to classify product images and generate suggested attributes, serving as a mediator that improves both efficiency and accuracy compared to direct manual assignment or simple automatic systems.
Solution Approach 2:
The patent replaces manual mechanical processes of attribute assignment with automated image classification algorithms. By substituting human manual work and simple computer automation with sophisticated machine learning-based image analysis, the system achieves both high productivity and improved measurement precision through advanced pattern recognition.
2Adaptability or versatility
If frequent re-training of classification models is performed to accommodate taxonomy changes, then adaptability is improved, but loss of time increases due to re-training requirements
Solution Approach 1:
The patent performs preliminary actions by pre-training image classification models on comprehensive datasets that cover multiple product categories and potential taxonomy variations. This advance preparation allows the models to adapt to taxonomy changes without requiring frequent re-training, as they are already equipped with generalized knowledge from the preliminary training phase.
Solution Approach 2:
The patent creates universal image classification models that can handle multiple product categories and taxonomy types simultaneously. By designing models with multi-functional capabilities that work across different product domains and taxonomy structures, the system achieves high adaptability without needing separate models or frequent re-training for each specific case.
Data Source
AI summary
A method including automatically determining, by a machine learning model trained based at least in part on sample items stored in a sample database, a query embedding vector for a query image of a query item. The method further can include determining, based on a respective embedding distance between the query image of the query item and a respective image of each of the sample items, neighboring items from among the sample items. The respective embedding distance can be calculated based on the query embedding vector for the query image and a respective embedding vector for the respective image of each of the sample items. Each of the sample items can include the respective image and at least one respective item label. The method also can include determining a respective normalized weight for each of the neighboring items based on the respective embedding distance between the query image and the respective image of the each of the neighboring items. The method additionally can include determining a query item label of the query item based on a weighted majority vote by the neighboring items via the respective normalized weight for the each of the neighboring items. The method further can include upon determining that the query item label of the query item is different from a first item label of the query item. storing the query item with the query item label in a product database. The method also can include selectively updating the sample items stored in the sample database from items in the product database. In addition, the method can include re-training the machine learning model based at least in part on the sample items in the sample database, as updated. Other embodiments are disclosed.


