Machine Learning Item Mapping Across Retail Domains
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Retailers face challenges in matching identical items across different sources of product information, leading to duplicate purchases, inadequate product offerings, and lost sales opportunities.
Innovation Solution
The use of trained machine learning processes, specifically Siamese Bidirectional Encoder Representations from Transformers (BERT) networks, to map items across various domains by generating features from textual data and determining textual similarity between items.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional manual or basic automated matching methods are used to identify items across different sources, then the system complexity remains low, but the accuracy of item matching deteriorates leading to duplicate purchases and lost sales
Solution Approach 1:
The patent replaces traditional mechanical or rule-based matching systems with machine learning models that automatically learn patterns from data. The system uses trained models to predict item matches across different sources without requiring complex manual configuration, achieving high accuracy through automated learning rather than complex deterministic rules.
Solution Approach 2:
The system changes the approach from using fixed thresholds or simple rules to using probability scores generated by machine learning models. The matching process transitions from binary match/no-match decisions based on rigid parameters to nuanced probability-based matching that adapts to different data sources and item characteristics.
2Productivity
If retailers manually verify and match items across multiple sources, then the accuracy of product offerings improves, but the time consumption and labor requirements increase significantly
Solution Approach 1:
The machine learning system performs self-learning and automatic matching without requiring manual verification. The models are trained on historical data and automatically apply their knowledge to new items, enabling the system to serve itself rather than requiring human intervention for routine matching tasks.
Solution Approach 2:
The system performs preliminary training on historical item data before actual matching occurs. By pre-learning patterns from past purchases and product information, the model is prepared to quickly and accurately match new items without requiring time-consuming manual verification during the actual matching process.
3Adaptability or versatility
If retailers use diverse data sources from multiple third-party providers, then the variety and completeness of product information increases, but the difficulty of consistent item identification across sources worsens
Solution Approach 1:
The machine learning system is designed to handle multiple data sources and item representations universally. The models can process different formats, vocabularies, and structures from various third-party providers and transform them into a common understanding, enabling consistent identification across diverse sources without requiring source-specific processing logic.
Solution Approach 2:
The trained machine learning models act as intermediaries between different data sources and the retailer's inventory system. These models translate and reconcile different item representations from various sources into a unified understanding, mediating the complexity of multi-source data integration and enabling consistent identification.
Data Source
AI summary
This application relates to employing trained machine learning processes to map items across various domains. For example, a computing device may obtain first textual data for a first item, and second textual data for a second item. The computing device generates a plurality of features based on the first textual data and the second textual data. Further, the computing device inputs the generated plurality of features to a trained machine learning process to generate output data characterizing a textual similarity between the first textual data and the second textual data. The computing device may also determine whether the first item maps to the second item based on the output data. The computing device generates mapping data based on the determination, and stores the mapping data in a data repository.


