Asynchronous Procurement System for Interactive Web Scraping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern interactive supplier websites pose challenges for procurement professionals due to issues like poor data freshness, high rogue spend, savings leaks, and the inability to retrieve comprehensive product information efficiently, as traditional scraping techniques are thwarted by security measures and fail to provide real-time insights or compliance with procurement rules.
Innovation Solution
The implementation of a system using Real-Time Dual Mode Agents, Real-Time Cross-Catalog Relevance, Real-Time Guided Buying, Real-Time Universal Alternate Supplier Checking, and Real-Time Price Dispersion Analytics, which enables asynchronous progressive requests to interactively retrieve and organize product information from multiple sources, ensuring timely and relevant data presentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional scraping techniques are used to retrieve data from interactive supplier websites, then data can be obtained from static pages, but comprehensive product information cannot be retrieved efficiently due to security measures and interactive data revelation
Solution Approach 1:
The patent applies dynamics by making the scraping system adaptive and responsive to interactive website behaviors. The system dynamically adjusts its scraping strategy based on detected website responses, security measures, and data revelation patterns, allowing it to effectively retrieve information from modern interactive catalogs that thwart traditional static scraping approaches
Solution Approach 2:
The system implements feedback mechanisms by monitoring website responses, security throttling patterns, and data availability. This feedback loop allows the scraping system to adjust its request rate, timing, and methodology in real-time, optimizing both the completeness of information retrieved and the efficiency of the scraping process while avoiding detection and blocking
2Loss of information
If traditional scraping techniques are used to retrieve data from interactive websites, then basic information can be obtained, but comprehensive detail information cannot be retrieved due to interactive data revelation mechanisms
Solution Approach 1:
The patent applies preliminary action by performing initial assessments of website structures, data revelation patterns, and security measures before actual data extraction. The system pre-configures scraping strategies, identifies interactive elements, and establishes optimal retrieval sequences in advance, enabling comprehensive detail information to be obtained efficiently without excessive delays
Solution Approach 2:
The system dynamically adapts to interactive data revelation mechanisms by detecting when additional information becomes available through user interactions. It automatically adjusts its scraping approach to capture detail information that is revealed progressively, ensuring complete data retrieval while minimizing the time penalty imposed by interactive requirements
3Productivity
If more powerful hardware is used to support traditional scraping techniques, then scraping capacity increases, but security measures such as throttling and blocking still prevent comprehensive data retrieval
Solution Approach 1:
The patent applies parameter changes by dynamically adjusting scraping parameters such as request timing, data volume per request, and interaction patterns. Instead of relying on brute-force hardware power, the system optimizes software parameters to match website security thresholds, maintaining high scraping capacity while avoiding throttling and blocking mechanisms
4Productivity
If caching is used to store retrieved data locally, then processing speed may improve, but data becomes stale and memory resources are consumed
Solution Approach 1:
The patent applies periodic action by implementing scheduled data refreshes and validity checks for cached information. The system periodically updates cached data based on detected changes at source websites or based on time-based expiration policies, ensuring data freshness is maintained while still benefiting from caching speed improvements
Solution Approach 2:
The system implements feedback mechanisms to monitor data freshness and trigger updates when necessary. By detecting changes in source data or receiving feedback about data age, the system intelligently determines when to refresh cached information, balancing processing speed benefits against data freshness requirements
Data Source
AI summary
Embodiments disclosed herein provide computerized, networked procurement systems designed to interact with source sites through asynchronous, progressive scripting requests to retrieve richer data sets from websites utilizing interactive loading and multiple hyperlinked pages either with a single vendor or across a plurality of vendors. These may provide improvements on prior art systems, such as by improving the response time relative to a prior art synchronous system from 30-60 seconds to less than ten seconds.


