Distributed Deep Web Data Extraction System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current web crawling technologies are inadequate for efficiently retrieving and analyzing large volumes of unstructured or poorly structured data from the World Wide Web, as they lack the scalability and processing capabilities to handle the vast amounts of data generated, particularly in the 'deep web' which lacks machine-parsable descriptive tags, making it inaccessible without specialized and time-consuming programming.

Innovation Solution

A distributed system for deep web data extraction that is highly scalable, allowing multiple concurrent searches, with a customizable search agent configuration interface, and includes modules for scrape campaign control, data storage, and output processing, enabling efficient coordination of web scraping agents across multiple servers and interfaces for direct output and storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current web crawling technology is used to retrieve unstructured deep web data, then data extraction can be performed, but the system lacks scalability and cannot handle large volumes of data efficiently

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidsystem scalability
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the web crawling task into multiple independent scrape agents that can operate concurrently across distributed servers. Each agent handles specific portions of the deep web, allowing the system to scale horizontally by adding more agents and servers without increasing individual component complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from single-server monolithic crawling to a multi-dimensional distributed architecture across multiple servers and networks. This dimensional expansion allows parallel processing of deep web data, dramatically increasing productivity while maintaining manageable complexity through modular design.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If specialized programming methodology is used to access deep web data, then data retrieval is possible, but the process becomes tedious and time-consuming

Engineering Contradiction:
Improvedata retrieval speedVSAvoidprogramming complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The scrape agents are designed with autonomous capabilities to navigate, identify, and extract data from unstructured deep web sources without requiring manual programming intervention for each target. The system self-configures and adapts to different data sources, eliminating tedious specialized programming while maintaining high retrieval speeds.

Inventive Principle:
Principle #25Self-service

3Productivity

If web scraping agents are distributed across multiple servers, then processing capability increases, but coordination and monitoring become more complex

Engineering Contradiction:
Improveprocessing capabilityVSAvoidcoordination complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The distributed scrape agents implement a universal communication protocol and data format that enables seamless coordination across multiple servers. This multi-functionality allows the same agent architecture to operate independently or in concert, simplifying coordination while maximizing processing capability through flexible distributed execution.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Quantity of substance

If large volumes of unstructured data are stored, then data availability increases, but storage and management become more challenging

Engineering Contradiction:
Improvedata volumeVSAvoidstorage management
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments stored data into structured categories and metadata tags during the scraping process. This pre-organization of unstructured data into manageable segments reduces subsequent storage management complexity while preserving access to the full volume of retrieved information.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10210255B2Distributed system for large volume deep web data extraction
Publication Date: 2019.02.19 QOMPLX INC
  • US10210255B2 patent drawing
  • US10210255B2 patent drawing
  • US10210255B2 patent drawing

AI summary

A distributed system for large volume deep web data extraction that is extremely scalable, allows multiple heterogeneous concurrent searches, has power web scrape result processing capabilities and uses a well defined, highly customizable, simplified, search agent configuration interface requiring minimal specialized programming knowledge. A scrape campaign control module receives scrape control and web spider configuration parameters through either a command line interface of an HTTP based application programming interface. The control module uses those parameters to have an arbitrary plurality of web spiders created and deployed by a plurality of servers. Scrape campaign results are presented as prescribed.