Event Identifier Partitioning for Distributed AJAX Crawling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional web crawling techniques are ineffective for AJAX-based rich Internet applications (RIAs) due to the lack of a one-to-one correspondence between the document object model (DOM) state and URL, leading to inefficient crawling and variable costs in reaching page states.
Innovation Solution
A computer-implemented process that computes event identifiers for each event in a set of events, segments them into partitions, assigns partitions to nodes for execution, and notifies other nodes of new states discovered, allowing for efficient distributed crawling by minimizing communication and ensuring complete coverage of DOM states.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional URL-based partitioning is used for crawling, then crawling of conventional web applications is efficient, but crawling of AJAX-based rich Internet applications becomes ineffective due to lack of one-to-one correspondence between DOM state and URL
Solution Approach 1:
The patent changes the partitioning parameter from URL to event identifier. Instead of using URL as the basis for dividing the crawling space, the system computes event identifiers for each event in the DOM and uses these identifiers for partitioning. This parameter change enables effective crawling of AJAX applications where URL no longer uniquely identifies DOM state.
Solution Approach 2:
The patent introduces a new dimension for partitioning by using event identifiers computed from DOM events rather than relying solely on the URL dimension. This adds an event-based dimension to the crawling space partitioning, allowing nodes to be assigned specific event identifiers to process.
2Productivity
If distributed crawling is implemented with multiple nodes, then crawling coverage increases, but communication overhead increases due to nodes needing to notify each other of newly discovered URLs
Solution Approach 1:
The patent applies preliminary action by pre-computing event identifiers for all events in the DOM before distributing the crawling task to nodes. This allows the system to determine in advance which node should process which event, reducing the need for communication during the crawling process.
Solution Approach 2:
The patent segments the crawling space by dividing event identifiers into partitions and assigning each partition to a specific node. This segmentation ensures that each node processes a distinct set of events, minimizing overlap and reducing communication overhead when nodes discover new states.
3Reliability
If a crawler executes events to materialize new pages in AJAX applications, then complete DOM state coverage is achieved, but the cost of reaching a page state becomes variable and increases
Solution Approach 1:
The patent uses copying by creating a virtual representation of the DOM state through event identifiers rather than physically navigating to each state. Nodes can determine which events they need to process by comparing event identifiers, avoiding redundant execution of the same events and reducing the cost of reaching page states.
Data Source
AI summary
An illustrative embodiment of a computer-implemented process for partitioning a crawling space computes an event identifier for each event in the set of events to form an identified set of events, segments the identified set of events into a number of partitions, assigns a partition to each node in a set of nodes and executes each event in each assigned partition by a respective node. In response to a determination that a new state is discovered, other nodes are notified of the new state, in which information associated with the new state is added to a respective assigned set of event IDs at each node. In response to a determination that no more notifications exist, the computer-implemented process determines whether more events to process exist and terminates in response to a determination that no more events to process exist.


