Relevance-Based Schema Matching for Catalog Enrichment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing schema matching approaches for large-scale Internet-accessible product catalogs struggle to unify heterogeneous product data from millions of schemas across thousands of categories and attributes, leading to inconsistent and noisy data that fails to scale effectively, resulting in poor customer experience and increased manual annotation efforts.

Innovation Solution

A catalog management system employing unsupervised domain-specific attribute representations and general attribute similarity metrics to identify and prioritize relevant attributes based on customer information, such as reviews and search queries, enabling scalable schema matching that consolidates product data into a unified and consistent schema, reducing manual annotation and improving data quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional schema matching approaches are used for large-scale product catalogs, then manual annotation efforts are required, but scalability deteriorates and data consistency worsens

Engineering Contradiction:
Improveschema matching scalabilityVSAvoiddata inconsistency
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system employs unsupervised domain-specific attribute representations that enable the schema matching process to automatically identify and unify attributes across heterogeneous schemas without requiring manual annotation. The attribute similarity metrics self-evaluate and rank candidate matches, allowing the system to serve itself in resolving data consistency issues at scale.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the schema matching problem by changing the parameters from manual comparison to automated similarity scoring. By introducing domain-specific attribute representations and computing similarity metrics based on customer information relevance, the system fundamentally alters how attributes are matched, enabling scalable processing while maintaining consistency.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If all attributes from heterogeneous schemas are retained, then data completeness is maintained, but storage size increases and data quality deteriorates

Engineering Contradiction:
Improvedata qualityVSAvoidstorage size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system extracts only the most relevant attributes for unified representation by evaluating candidate attributes against domain-specific criteria and customer information relevance. This extraction process removes redundant and noisy attributes from the heterogeneous schemas, retaining only those that contribute to data quality while reducing overall storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different selection criteria to different attribute types based on their domain-specific importance. Rather than uniformly retaining all attributes, the system applies local quality filters that prioritize attributes relevant to specific product categories and customer needs, thereby improving overall data quality while optimizing storage usage.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If manual annotation is used for schema matching, then attribute accuracy is improved, but time consumption increases and automation deteriorates

Engineering Contradiction:
Improveattribute matching accuracyVSAvoidmanual annotation requirement
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system incorporates feedback loops where attribute similarity scores are computed and ranked, then used to automatically select the best matches. The feedback mechanism continuously refines the matching process by learning from customer information patterns, thereby maintaining high accuracy while eliminating the need for manual annotation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces the mechanical process of manual annotation with an automated computational system that uses unsupervised learning and similarity metrics. This substitution maintains measurement precision by using domain-specific attribute representations that capture the essential characteristics of product attributes, while completely eliminating manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Speed

If heterogeneous schemas are unified without domain-specific approaches, then processing speed is improved, but attribute relevance to customers deteriorates

Engineering Contradiction:
Improveschema matching speedVSAvoidcustomer relevance information
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The system performs preliminary action by pre-defining domain-specific attribute representations and similarity metrics before the actual schema matching process. This preliminary preparation enables fast automated processing while ensuring that customer-relevant attributes are prioritized from the outset, preventing loss of important information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces domain-specific attribute representations as intermediaries between heterogeneous schemas and the unified catalog structure. These intermediaries preserve customer relevance information by translating diverse attribute formats into a common domain-specific framework, enabling both fast processing and information retention.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12020297B1Relevance-based schema matching for targeted catalog enrichment
Publication Date: 2024.06.25 AMAZON TECH INC
  • US12020297B1 patent drawing
  • US12020297B1 patent drawing
  • US12020297B1 patent drawing

AI summary

Methods, systems, and computer-readable media for relevance-based schema matching for targeted catalog enrichment are disclosed. A catalog management system determines textual descriptors associated with a category of items in a catalog. The system determines relevant attributes for the category based (at least in part) on analysis of the textual descriptors. The relevant attributes are selected from a larger set of candidate attributes for the category. The system modifies the catalog based (at least in part) on schema matching for a plurality of items in the category. The items are associated with descriptive terms from a plurality of data sources and expressed according to a plurality of source schemas. For an individual item, the schema matching determines a correspondence between one or more of the descriptive terms in one of the source schemas and a corresponding attribute in a target schema comprising the relevant attributes.