Web Crawler Virtual Browser Data Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing web crawlers face inefficiencies in extracting data from large websites with many webpages, as they take a long time to discover all pages and generate numerous requests that can impact website performance.
Innovation Solution
A retail crawler module with a crawling engine, virtual browser, and language processor that processes webpages efficiently by loading a 'thin' version, suppressing irrelevant elements, and using engine control statements to extract product data, while also implementing self-healing mechanisms to adapt to changes in webpage configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a web crawler requests all webpages to discover content, then complete data extraction is achieved, but the time required increases significantly and website performance is impacted
Solution Approach 1:
The patent applies preliminary action by using AI models to predict and prioritize which webpages are most likely to contain relevant product data before the crawler actually visits them. The system analyzes historical data, webpage patterns, and product information to pre-determine the crawling sequence, allowing the crawler to focus on high-value pages first and achieve complete data extraction more efficiently
Solution Approach 2:
The patent uses copying by creating virtual representations of webpages through AI-generated summaries and metadata before full crawling is performed. The system generates condensed versions of webpage content that capture essential product information, allowing the crawler to work with these copies initially and only visit full pages when necessary, significantly reducing crawling time while maintaining data completeness
2Measurement precision
If a web crawler requests all webpages to ensure complete data extraction, then all product information is captured, but the number of requests impacts website performance
Solution Approach 1:
The patent applies partial action by having the crawler visit only the subset of webpages that AI models predict contain relevant product data, rather than requesting all webpages. The system determines the optimal number and selection of pages to crawl based on predicted data value, achieving complete product information extraction with fewer requests and reduced impact on website performance
Solution Approach 2:
The patent introduces an intermediary layer between the crawler and the website - an AI prediction system that acts as a mediator to filter and prioritize target pages. This intermediary analyzes various signals and generates a ranked list of webpages to crawl, preventing the crawler from directly requesting all pages and thereby reducing the harmful impact on website performance while maintaining data extraction completeness
3Measurement precision
If the crawler processes all webpage elements to extract product data, then complete information is obtained, but the processing time increases
Solution Approach 1:
The patent applies extraction by using AI models to identify and extract only the specific product data elements that are relevant, rather than processing all webpage elements. The system extracts key product information such as product identifiers, prices, and descriptions directly from predicted relevant sections, obtaining complete product data while significantly reducing the time required by avoiding unnecessary processing of irrelevant content
4Measurement precision
If the crawler uses detailed processing to extract accurate product data, then data precision is improved, but the complexity of the crawling system increases
Solution Approach 1:
The patent replaces complex mechanical crawling and data extraction mechanisms with AI-based predictive models. Instead of using elaborate rule-based systems to identify and extract product data, the system uses machine learning models that automatically predict relevant pages and extract information, achieving high data accuracy while reducing system complexity through智能化 automation
Data Source
AI summary
A computer system extracts product data from a website and correlates product records from multiple sources to one another as corresponding to the same product. A website is crawled efficiently by rendering webpages using a virtual browser that ignores blacklisted elements, extracts data from objects without rendering, and suppressing retrieval of remote resources. Data is extracted according to engine control statements including a selector and extractor. A website may be crawled repeatedly and changes in extracted data may be detected and flagged. Engine control statements may be automatically changed in response to detecting a change in the configuration of the website. Images of product records may be correlated with one another by first comparing text of the product records and selecting images for comparison based on composition. Images are compared using a machine learning model. Images determined to be similar may be presented to a human for a correlation decision.


