Automated Item Data Model Generation via Web Crawling and ETL
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing item database systems are inefficient due to their static nature, requiring manual data collection and management, which limits the amount of information that can be maintained for products, making it difficult to keep records up-to-date and accommodate the diverse data needs of various product types, leading to a tedious and slow process for modifying database structures.
Innovation Solution
An automated method for developing item data models using data virtualization, analytics, ETL processes, web crawlers, and reverse engineering systems to extract, transform, and load data from multiple sources, enabling the creation and updating of item data models that can efficiently organize and normalize information, and provide relevant details to customers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data collection and management is used, then data accuracy can be maintained, but productivity is reduced and time consumption increases
Solution Approach 1:
The system enables automated self-service data collection where web crawlers autonomously navigate, extract, and process product information from multiple sources without manual intervention. The automated item data model development system self-manages the entire data pipeline from source identification to normalized database output, eliminating the need for manual data entry while maintaining accuracy through systematic validation processes.
Solution Approach 2:
The patent replaces manual mechanical data collection processes with automated computational systems. Web crawlers use algorithms to automatically scrape product data from vendor websites and social media, while ETL processes automatically transform and load data into standardized formats. This substitution of mechanical manual work with automated digital systems dramatically increases productivity while maintaining data quality through consistent application of extraction rules.
2Device complexity
If static database structure is used, then system simplicity is maintained, but adaptability is reduced when new product information needs to be added
Solution Approach 1:
The system implements dynamic database structure development where the item data model evolves automatically based on discovered product attributes. As web crawlers encounter new product types and attributes during data collection, the system dynamically generates and updates database schemas to accommodate these variations. This allows the database structure to adapt flexibly to new product information while maintaining organized, normalized relationships through automated model generation.
3Loss of information
If comprehensive product information is collected from multiple sources, then information completeness is improved, but data redundancy increases
Solution Approach 1:
The ETL process extracts only the necessary and unique product attributes from multiple data sources, filtering out redundant information. The system identifies and extracts key product attributes such as specifications, pricing, and availability from vendor websites and social media, while automatically deduplicating data across sources. This selective extraction maintains information completeness by capturing all essential product details while eliminating redundant copies of the same data.
Solution Approach 2:
The normalized item data model serves as a universal structure that can accommodate multiple product types and data sources through standardized attribute schemas. By creating a multi-functional database model that can represent diverse product information in a consistent format, the system eliminates the need to maintain separate data structures for different sources, thereby reducing redundancy while preserving comprehensive product information across all categories.
4Speed
If automated web crawling is used, then data collection speed is improved, but system complexity increases
Solution Approach 1:
The automated data collection system is segmented into distinct functional modules: web crawlers for data extraction, ETL processes for transformation, and automated model generation for database structuring. Each module performs a specific function independently, allowing the system to achieve high data collection speed through parallel processing while managing complexity through modular design. This segmentation enables independent optimization of each component without increasing overall system complexity.
Data Source
AI summary
The present invention extends to methods, systems, and computer program products for developing an item data model for an item. Aspects of the invention can automate the process of data collection of “facts” for “items” that information is needed about. Facts can be organized and normalized to eliminate redundant facts, and interpret what is found. Data requirements extraction and automated modeling using a combination of data virtualization, data analytics, extract, transform, and load (ETL), web crawlers, and reverse engineering systems, can be used along with other technologies to develop an item model. A model owner feeds a curating module with the information for locating the facts to be used, and initiating the modeling process. Existing data structures, websites, vendor input, etc. can be described to the import process, and an item model is produced. The model can be imported into existing modeling tools for viewing, or viewed as XML.


