AI Product Catalog Generator for Automated Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

E-commerce websites face challenges in automatically collecting and organizing product information and attributes across multiple product pages, leading to inefficiencies in data collection and inconsistency in product datasets.

Innovation Solution

A method and system for crawling websites to identify product pages, extracting and normalizing product attributes, and storing them in a structured database using a product catalog generator that employs web crawling, machine learning, and computer vision to automate the process of generating product page variations and extracting attributes, and standardizing data across multiple websites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated web crawling is used to collect product data, then data collection efficiency is improved, but data consistency and quality deteriorate

Engineering Contradiction:
Improvedata collection efficiencyVSAvoiddata consistency
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system implements feedback mechanisms where extracted product attributes are validated against predefined schemas and patterns. Inconsistent data triggers re-extraction or correction processes, ensuring data quality improves over time while maintaining high collection efficiency through automated feedback loops.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent transforms unstructured product data into standardized parameters by applying normalization rules and data transformation processes. This converts variable-format web data into consistent structured parameters, resolving the consistency issue while preserving automated collection benefits.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual data collection methods are used, then data consistency is maintained, but productivity and automation level deteriorate

Engineering Contradiction:
Improvedata consistencyVSAvoiddata collection efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs self-service data validation and correction by automatically detecting inconsistencies and applying repair rules without human intervention. This maintains data consistency through automated quality control while preserving high productivity through continuous autonomous operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical data verification processes with automated computational validation systems. This substitution maintains consistency through algorithmic checking while dramatically improving productivity by eliminating manual labor bottlenecks.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If comprehensive product attributes are extracted from multiple pages, then data completeness is improved, but system complexity and processing time worsen

Engineering Contradiction:
Improvedata completenessVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the complex task of comprehensive data extraction into modular components: crawling module, extraction module, validation module, and storage module. Each handles specific aspects of data collection, reducing overall system complexity while achieving complete product attribute extraction through coordinated modular operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary data filtering and preprocessing to identify only relevant product attributes before full extraction. This preliminary action reduces the scope of subsequent processing, maintaining data completeness while simplifying the overall system architecture and reducing processing complexity.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If comprehensive product attributes are extracted from multiple pages, then data completeness is improved, but processing time worsens

Engineering Contradiction:
Improvedata completenessVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system employs periodic action by extracting data in structured batches or cycles rather than continuously processing all pages sequentially. This periodic approach maintains complete data extraction while optimizing processing time through efficient resource utilization and parallel processing capabilities.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent performs preliminary identification and prioritization of product pages before full extraction. High-priority pages with complete attributes are processed first, while lower-priority pages are handled subsequently or skipped if redundant. This preliminary sorting maintains data completeness while significantly reducing total processing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11550856B2Artificial intelligence for product data extraction
Publication Date: 2023.01.10 HEARST MAGAZINE MEDIA INC
  • US11550856B2 patent drawing
  • US11550856B2 patent drawing
  • US11550856B2 patent drawing

AI summary

A computer system and method may be used to generate a product catalog from one or more websites. One or more product pages on the websites may be identified and parsed. Attribute information may be identified in each page. Moreover, one or more automated interactions may be performed to generate page variations and identify attribute values. The attribute information and attribute values may be stored as structured data in a database.