Seeding Rule-Based Models for Duplicate Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Network-accessible marketplaces face challenges in identifying and managing duplicate item description entries from different suppliers, leading to inefficiencies and potential customer-merchant relationship issues due to variations in item information formats and nomenclature.

Innovation Solution

A system and method utilizing rule-based machine learning models, specifically genetic algorithms and decision trees, to generate and apply rule sets for duplicate detection, which can identify and merge duplicate item description entries, leveraging seed rules from previously generated models to accelerate the process and improve precision and recall.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional duplicate detection methods are used to identify duplicate item description entries from different suppliers, then duplicate items can be identified, but significant time and computing resources are consumed

Engineering Contradiction:
Improveduplicate detection accuracyVSAvoidtime for duplicate detection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing item description data into standardized features before duplicate detection. The system extracts and normalizes key attributes (title, description, specifications, images) into a consistent format, creating a prepared dataset that accelerates subsequent duplicate identification. This pre-processing step reduces the computational burden during actual duplicate detection operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical comparison methods with machine learning-based duplicate detection. Instead of using rule-based or manual comparison algorithms, the system employs trained models that automatically identify duplicate items based on learned patterns from training data. This substitution significantly reduces both time and computational resources while maintaining high detection accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If traditional duplicate detection methods are used to identify duplicate item description entries from different suppliers, then duplicate items can be identified, but significant computing resources are consumed

Engineering Contradiction:
Improveduplicate detection accuracyVSAvoidcomputational resources for duplicate detection
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary feature extraction and data normalization, converting raw item descriptions into standardized vectors. This pre-computation reduces the complexity of subsequent duplicate detection operations, requiring fewer computational resources during actual detection while maintaining high accuracy through structured data representation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces computationally intensive traditional algorithms with efficient machine learning models. The trained models leverage learned patterns to rapidly identify duplicates without requiring exhaustive comparisons, significantly reducing CPU usage and energy consumption while preserving detection precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If item information from different suppliers is maintained without standardization, then supplier-specific information formats are preserved, but identifying duplicate items becomes complex and resource-intensive

Engineering Contradiction:
Improvesupplier information format flexibilityVSAvoidduplicate detection system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by maintaining supplier-specific information formats in their original form while extracting standardized features from each. The system processes different supplier formats with appropriate extraction rules tailored to each format's characteristics, then normalizes the extracted features into a unified representation. This approach preserves format flexibility locally while achieving global standardization for duplicate detection.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system introduces an intermediary layer between supplier-specific formats and the duplicate detection process. This intermediary performs format-agnostic feature extraction, converting diverse supplier data into a standardized intermediate representation that the detection algorithm can process uniformly. This mediator simplifies the overall system complexity by handling format variations in a centralized manner.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8793201B1System and method for seeding rule-based machine learning models
Publication Date: 2014.07.29 AMAZON TECH INC
  • US8793201B1 patent drawing
  • US8793201B1 patent drawing
  • US8793201B1 patent drawing

AI summary

Embodiments of a system and method for seeding rule-based machine learning models include generating a set of seed rules including one or more rules resulting from one or more previously performed machine learning operations and one or more randomly or pseudo-randomly generated rules. Embodiments may include performing one or more machine learning operations on the set of seed rules to generate a new set of rules for determining whether a pair of information items have a specific relationship. Generating the new set of rules from the seed rules may be faster than generating a set of rules from random data. Embodiments may also include applying the new set of rules to one or more pairs of information items to identify at least one pair of information items as having the specific relationship. For instance, the rules may be applied to identify pairs of information items that are duplicate pairs.