Seeding Rule-Based Models for Duplicate Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network-accessible marketplaces face challenges in identifying and managing duplicate item description entries from different suppliers, leading to inefficiencies and potential customer-merchant relationship issues due to variations in item information formats and nomenclature.
Innovation Solution
A system and method utilizing rule-based machine learning models, specifically genetic algorithms and decision trees, to generate and apply rule sets for duplicate detection, which can identify and merge duplicate item description entries, leveraging seed rules from previously generated models to accelerate the process and improve precision and recall.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional duplicate detection methods are used to identify duplicate item description entries from different suppliers, then duplicate items can be identified, but significant time and computing resources are consumed
Solution Approach 1:
The patent applies preliminary action by pre-processing item description data into standardized features before duplicate detection. The system extracts and normalizes key attributes (title, description, specifications, images) into a consistent format, creating a prepared dataset that accelerates subsequent duplicate identification. This pre-processing step reduces the computational burden during actual duplicate detection operations.
Solution Approach 2:
The patent replaces traditional mechanical comparison methods with machine learning-based duplicate detection. Instead of using rule-based or manual comparison algorithms, the system employs trained models that automatically identify duplicate items based on learned patterns from training data. This substitution significantly reduces both time and computational resources while maintaining high detection accuracy.
2Measurement precision
If traditional duplicate detection methods are used to identify duplicate item description entries from different suppliers, then duplicate items can be identified, but significant computing resources are consumed
Solution Approach 1:
The system performs preliminary feature extraction and data normalization, converting raw item descriptions into standardized vectors. This pre-computation reduces the complexity of subsequent duplicate detection operations, requiring fewer computational resources during actual detection while maintaining high accuracy through structured data representation.
Solution Approach 2:
The patent replaces computationally intensive traditional algorithms with efficient machine learning models. The trained models leverage learned patterns to rapidly identify duplicates without requiring exhaustive comparisons, significantly reducing CPU usage and energy consumption while preserving detection precision.
3Adaptability or versatility
If item information from different suppliers is maintained without standardization, then supplier-specific information formats are preserved, but identifying duplicate items becomes complex and resource-intensive
Solution Approach 1:
The patent applies local quality by maintaining supplier-specific information formats in their original form while extracting standardized features from each. The system processes different supplier formats with appropriate extraction rules tailored to each format's characteristics, then normalizes the extracted features into a unified representation. This approach preserves format flexibility locally while achieving global standardization for duplicate detection.
Solution Approach 2:
The system introduces an intermediary layer between supplier-specific formats and the duplicate detection process. This intermediary performs format-agnostic feature extraction, converting diverse supplier data into a standardized intermediate representation that the detection algorithm can process uniformly. This mediator simplifies the overall system complexity by handling format variations in a centralized manner.
Data Source
AI summary
Embodiments of a system and method for seeding rule-based machine learning models include generating a set of seed rules including one or more rules resulting from one or more previously performed machine learning operations and one or more randomly or pseudo-randomly generated rules. Embodiments may include performing one or more machine learning operations on the set of seed rules to generate a new set of rules for determining whether a pair of information items have a specific relationship. Generating the new set of rules from the seed rules may be faster than generating a set of rules from random data. Embodiments may also include applying the new set of rules to one or more pairs of information items to identify at least one pair of information items as having the specific relationship. For instance, the rules may be applied to identify pairs of information items that are duplicate pairs.


