Web Scraping Control via Real-Time Access Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in effectively identifying and controlling web scraping activities, which can lead to data misuse and website overload, as existing methods are inadequate in distinguishing between legitimate and malicious scrapers in real time.

Innovation Solution

A centralized system that logs and compiles web access statistics in real time across multiple web servers, using a centralized database to identify and block unwanted bots by analyzing request patterns, thresholds, and user behavior, allowing for dynamic control of access to prevent web scraping.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If web scraping is allowed to extract data from websites, then data availability for indexing and search is improved, but website performance degradation and overload occur

Engineering Contradiction:
Improvedata availabilityVSAvoidwebsite performance
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent applies local quality by treating different web scrapers differently based on their characteristics. Friendly scrapers (like search engine bots) are allowed to access data while malicious scrapers are blocked. This is achieved by analyzing scraper behavior patterns, request frequencies, and data extraction methods to determine whether to permit or deny access locally for each scraper instance, thus preserving website performance while maintaining data availability for legitimate purposes

Inventive Principle:
Principle #3Local quality

2Reliability

If automated blocking tools are used to stop web scrapers, then website performance is protected, but legitimate scrapers are also blocked

Engineering Contradiction:
Improvewebsite performanceVSAvoidscraper access control
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements feedback mechanisms by continuously monitoring scraper behavior and adjusting access control decisions dynamically. The system analyzes request patterns, data extraction rates, and scraper identification information in real-time, then provides feedback by either permitting or blocking subsequent requests. This adaptive feedback loop enables the system to distinguish between friendly and malicious scrapers, protecting website performance while allowing legitimate scraping activities to continue

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies dynamics by making the blocking mechanism adaptive rather than static. Instead of using fixed blocking rules, the system dynamically adjusts access control based on real-time analysis of scraper behavior, request frequencies, and pattern recognition. This dynamic approach allows the system to respond to changing scraper tactics while maintaining protection against malicious activities and permitting legitimate scraping

Inventive Principle:
Principle #15Dynamics

3Loss of time

If real-time monitoring of web access is implemented, then web scraping activities are identified promptly, but system complexity increases

Engineering Contradiction:
Improvedetection timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-establishing monitoring infrastructure and access control mechanisms before scraping activities occur. The system is configured with predetermined rules, thresholds, and analysis algorithms that are ready to evaluate scraper behavior in real-time. This preliminary setup enables prompt identification of web scraping activities without requiring complex real-time decision-making, thus reducing detection time while managing system complexity through advance preparation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9385928B2Systems and methods to control web scraping
Publication Date: 2016.07.05 THRYV INC
  • US9385928B2 patent drawing
  • US9385928B2 patent drawing
  • US9385928B2 patent drawing

AI summary

Systems and methods to control web scraping through a plurality of web servers using real time access statistics are described. For example, in one embodiment a web request is categorized based at least in part on a type of data access. The web request can be processed based on threshold determinations. The processing can include blocking, delaying, timely replying, or prioritizing a response to the web request.