Automated Item Data Model Generation via Web Crawling and ETL

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing item database systems are inefficient due to their static nature, requiring manual data collection and management, which limits the amount of information that can be maintained for products, making it difficult to keep records up-to-date and accommodate the diverse data needs of various product types, leading to a tedious and slow process for modifying database structures.

Innovation Solution

An automated method for developing item data models using data virtualization, analytics, ETL processes, web crawlers, and reverse engineering systems to extract, transform, and load data from multiple sources, enabling the creation and updating of item data models that can efficiently organize and normalize information, and provide relevant details to customers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data collection and management is used, then data accuracy can be maintained, but productivity is reduced and time consumption increases

Engineering Contradiction:
Improvedata accuracyVSAvoiddata collection efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables automated self-service data collection where web crawlers autonomously navigate, extract, and process product information from multiple sources without manual intervention. The automated item data model development system self-manages the entire data pipeline from source identification to normalized database output, eliminating the need for manual data entry while maintaining accuracy through systematic validation processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical data collection processes with automated computational systems. Web crawlers use algorithms to automatically scrape product data from vendor websites and social media, while ETL processes automatically transform and load data into standardized formats. This substitution of mechanical manual work with automated digital systems dramatically increases productivity while maintaining data quality through consistent application of extraction rules.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If static database structure is used, then system simplicity is maintained, but adaptability is reduced when new product information needs to be added

Engineering Contradiction:
Improvedatabase structure simplicityVSAvoiddatabase modification capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system implements dynamic database structure development where the item data model evolves automatically based on discovered product attributes. As web crawlers encounter new product types and attributes during data collection, the system dynamically generates and updates database schemas to accommodate these variations. This allows the database structure to adapt flexibly to new product information while maintaining organized, normalized relationships through automated model generation.

Inventive Principle:
Principle #15Dynamics

3Loss of information

If comprehensive product information is collected from multiple sources, then information completeness is improved, but data redundancy increases

Engineering Contradiction:
Improveproduct information completenessVSAvoiddata redundancy
Core Design Contradiction:
Loss of informationVSLoss of substance

Solution Approach 1:

The ETL process extracts only the necessary and unique product attributes from multiple data sources, filtering out redundant information. The system identifies and extracts key product attributes such as specifications, pricing, and availability from vendor websites and social media, while automatically deduplicating data across sources. This selective extraction maintains information completeness by capturing all essential product details while eliminating redundant copies of the same data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The normalized item data model serves as a universal structure that can accommodate multiple product types and data sources through standardized attribute schemas. By creating a multi-functional database model that can represent diverse product information in a consistent format, the system eliminates the need to maintain separate data structures for different sources, thereby reducing redundancy while preserving comprehensive product information across all categories.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Speed

If automated web crawling is used, then data collection speed is improved, but system complexity increases

Engineering Contradiction:
Improvedata collection speedVSAvoidautomation system complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The automated data collection system is segmented into distinct functional modules: web crawlers for data extraction, ETL processes for transformation, and automated model generation for database structuring. Each module performs a specific function independently, allowing the system to achieve high data collection speed through parallel processing while managing complexity through modular design. This segmentation enables independent optimization of each component without increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10936675B2Developing an item data model for an item
Publication Date: 2021.03.02 WALMART APOLLO LLC
  • US10936675B2 patent drawing
  • US10936675B2 patent drawing
  • US10936675B2 patent drawing

AI summary

The present invention extends to methods, systems, and computer program products for developing an item data model for an item. Aspects of the invention can automate the process of data collection of “facts” for “items” that information is needed about. Facts can be organized and normalized to eliminate redundant facts, and interpret what is found. Data requirements extraction and automated modeling using a combination of data virtualization, data analytics, extract, transform, and load (ETL), web crawlers, and reverse engineering systems, can be used along with other technologies to develop an item model. A model owner feeds a curating module with the information for locating the facts to be used, and initiating the modeling process. Existing data structures, websites, vendor input, etc. can be described to the import process, and an item model is produced. The model can be imported into existing modeling tools for viewing, or viewed as XML.