Automated Data Extraction for Multi-Vendor Product Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing search engines fail to efficiently identify specific products with unique part attributes across multiple vendor websites, leading to time-consuming and complex searches due to inconsistencies in terminology, limited search capabilities, and varying product categorization.
Innovation Solution
Systems and methods for automated data extraction and ingestion that execute search requests across multiple data sources, using intelligent algorithms to identify and extract relevant data, providing real-time pricing and availability information, and sorting results by price and availability for unique part attributes like part numbers or model numbers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual searching is performed across multiple vendor websites, then users can find product information, but the process becomes time-consuming and complex
Solution Approach 1:
The patent introduces a specialized search engine as an intermediary system that connects users to multiple vendor websites. This search engine aggregates product data from various vendors, normalizes terminology differences, and presents unified search results, thereby reducing the time users would otherwise spend manually searching multiple websites while ensuring comprehensive product information is found.
Solution Approach 2:
The search engine performs preliminary actions by pre-aggregating and indexing product data from multiple vendor websites before users initiate searches. This includes pre-processing vendor-specific terminology, categorizing products, and organizing data structures, so that when users search, they immediately receive processed results without having to navigate multiple vendor interfaces manually.
2Adaptability or versatility
If existing search engines are used to search generic product terms, then broad product categories can be found, but specific parts with unique attributes cannot be efficiently identified
Solution Approach 1:
The patent applies local quality by enabling the search engine to adapt its search strategy and data processing based on the specific search query type. When users search for generic product terms, the engine provides broad category results; when users search for specific parts with unique attributes like part numbers, the engine switches to precision mode that targets exact matches and filters results accordingly, thereby achieving both versatility and precision.
Solution Approach 2:
The search engine dynamically adjusts its behavior based on the search input. It detects whether a query is for a generic product term or a specific part number, and automatically modifies its search algorithms, data filtering criteria, and result presentation format. This dynamic adaptation allows the same system to efficiently handle both broad category searches and precise part identification.
3Adaptability or versatility
If vendors use different terminology and categorization for the same products, then each vendor can maintain their own naming conventions, but cross-vendor product comparison becomes difficult
Solution Approach 1:
The search engine implements universality by creating a standardized data model that can represent products from multiple vendors with different terminology and categorization schemes. It maintains vendor-specific data structures while translating them into a universal format for comparison, allowing users to compare products across vendors using consistent criteria while each vendor retains independence in their own naming conventions and product organization.
Data Source
AI summary
Systems, methods, and devices for data extraction and data ingestion. A method includes receiving a search request comprising a product descriptor and searching a plurality of vendor websites to identify a plurality of product listings that each comprise information matching the product descriptor. The method includes extracting data from each of the plurality of product listings, wherein the extracted data comprises unstructured data. The method includes providing at least a portion of the extracted data to a machine learning algorithm trained to identify one or more unique part attributes within the portion of the extracted data. The method includes determining whether two or more of the plurality of product listings are duplicate product listings based on the one or more unique part attributes identified by the machine learning algorithm.


