Headless Browser Fingerprinting for Merchant Server Data Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for extracting and comparing online product and service information from merchant computer servers are often blocked by competitors, leading to inaccurate and tedious manual methods for price monitoring.
Innovation Solution
A system comprising distributed query servers with multiple proxy addresses and headless web browsers, which generate random browser fingerprint identifiers to simulate human interactions, allowing for the extraction of product and service information without being blocked.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If query servers send requests to merchant computer servers for information extraction, then product and service information can be obtained for comparison, but the requests are blocked by competitors to preserve server resources
Solution Approach 1:
The system segments the extraction process into multiple phases: initial HTTP requests for straightforward cases, and headless browser-based requests for protected sites. This segmentation allows the system to adapt to different server response types and avoid blocking by diversifying request methodologies.
Solution Approach 2:
The patent introduces headless browsers as intermediary components between the query servers and merchant computer servers. These browsers act as mediators that can bypass server-side blocking mechanisms by simulating genuine user interactions, thereby maintaining reliable information extraction even when direct HTTP requests are blocked.
2Reliability
If manual extraction methods are used to avoid blocking, then information can be obtained, but the process becomes tedious and time-consuming
Solution Approach 1:
The system creates automated copies of manual extraction processes through headless browsers that replicate human browsing behavior. These automated agents copy the interactions a human would perform (clicking, scrolling, form filling) but execute them programmatically, maintaining extraction accuracy while eliminating the time and effort required for manual operations.
Solution Approach 2:
The extraction system performs self-service by automatically detecting when responses indicate blocking attempts and autonomously switching to headless browser modes. The system monitors response patterns, identifies blocked requests, and independently initiates alternative extraction methods without human intervention, thereby maintaining both reliability and efficiency.
3Productivity
If existing query server solutions are used, then information extraction can be automated, but they are blocked by competitors identifying requests from data center IP addresses
Solution Approach 1:
The patent transitions from a single-dimension approach (HTTP requests from data center IPs) to a multi-dimensional approach by introducing headless browsers that operate in a different dimensional space. This dimensional shift allows requests to bypass traditional blocking mechanisms that target data center IP addresses, as the browser-based requests present different identifying characteristics.
Solution Approach 2:
The system dynamically changes request parameters based on server responses. When blocking is detected, the system modifies request characteristics by switching to headless browser modes with different user agents, timing patterns, and interaction sequences. These parameter changes make it difficult for servers to identify and block automated extraction attempts.
4Productivity
If multiple requests are sent en masse from data center IP addresses, then information extraction speed increases, but server resources are consumed and blocking occurs
Solution Approach 1:
The system implements periodic action by introducing variable time delays between requests when using headless browsers. Instead of sending requests en masse at constant intervals, the system employs human-like timing patterns with random variations, reducing server resource consumption while maintaining extraction productivity over extended periods.
Solution Approach 2:
The extraction system dynamically adjusts its operational characteristics based on server responses and detected blocking patterns. When high-speed extraction is feasible, the system operates in HTTP mode with faster request rates. When blocking or resource consumption becomes an issue, the system dynamically transitions to headless browser mode with adjusted timing, optimizing the balance between productivity and server resource usage.
Data Source
AI summary
The invention provides a system for extracting information accessible by queries from merchant computer servers, such system comprising query servers, configured to run instances of a headless internet browser and configured to randomly generate a browser fingerprint identifier in requests from instances of the browsers, execute a driver module of a browser with timed keyboard input simulation commands, each request including a browser fingerprint identifier combining values for parameters selected from the group constituted by the name of a browser, the version of this browser, an operating system, the language of the browser, a type of device supposed to run the operating system, plug-ins available in the browser, and retrieve requests answers. The invention also provides a system for benchmarking information, comprising a module for extracting information and a module for comparing virtual references.
