Search Service Administration Web Service Protocol for Crawl Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current web crawling technologies are inefficient due to the vast amount of data on the web, leading to excessive network bandwidth consumption and resource wastage, as well as the inability to access valuable information due to authentication and authorization issues, resulting in missed relevant data.
Innovation Solution
A client-centric approach is implemented to configure and control web crawling functions through a Search Service application, utilizing an Application Programming Interface (API) governed by the Search Service Administration Web Service protocol, allowing users to define crawl rules, authentication, and access parameters, thereby controlling the crawling process and ensuring efficient data retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If web crawling is performed without restrictions to ensure comprehensive data retrieval, then the completeness of information is improved, but network bandwidth consumption and resource usage increase excessively
Solution Approach 1:
The patent segments the web crawling process into controlled phases with defined boundaries through crawl rules. The URL space is divided into crawlable and non-crawlable sections, allowing the crawler to systematically process only relevant portions of the web while avoiding unnecessary data retrieval and reducing network bandwidth consumption.
Solution Approach 2:
The patent implements dynamic control mechanisms where crawl rules can be modified, added, or removed during the crawling process. The system adapts to changing requirements by allowing users to dynamically adjust crawl boundaries and authentication settings, optimizing the balance between information completeness and resource consumption.
2Area of stationary object
If web crawling is performed without authentication to ensure broad access, then the coverage of accessible data is improved, but the ability to retrieve protected valuable information is lost
Solution Approach 1:
The patent implements preliminary authentication setup through crawl rules before the actual crawling process begins. Users configure authentication credentials and authorization parameters in advance, allowing the crawler to automatically present these credentials when encountering protected resources, thus ensuring both broad coverage and access to protected information.
Solution Approach 2:
The patent introduces authentication credentials as an intermediary between the crawler and protected web resources. The crawl rules serve as a mediator that manages the authentication process, allowing the system to access protected content without requiring manual intervention while maintaining security protocols.
3Measurement precision
If manual filtering of retrieved documents is performed to ensure quality information, then the relevance of retrieved data is improved, but the time and resources required increase significantly
Solution Approach 1:
The patent performs preliminary filtering by configuring crawl rules that define inclusion and exclusion criteria before the crawling process begins. This pre-filtering approach eliminates the need for extensive manual filtering afterward, as the crawler systematically retrieves only relevant documents based on pre-defined parameters, significantly reducing the time and resources required for post-processing.
4Productivity
If crawl rules and authentication settings are configured to reduce data retrieval volume, then resource efficiency is improved, but the complexity of system configuration increases
Solution Approach 1:
The patent implements a universal crawl rule configuration system that handles multiple functions through a unified interface. The same crawl rule structure manages URL patterns, authentication credentials, inclusion/exclusion criteria, and scheduling parameters, reducing configuration complexity while maintaining high resource efficiency through centralized control.
Data Source
AI summary
The embodiments described herein generally relate to a method and system for enabling a client to configure and control the crawling function available through a crawl configuration Web service. A client is able to configure and control the crawling function by defining the URL space of the crawl. Such space may be defined by configuring the starting point(s) and other properties of the crawl. The client further configures the crawling function by creating and configuring a content source and/or a crawl rule. Further, a client defines authentication information applicable to the crawl to enable the discovery and retrieval of electronic documents requiring authentication and/or authorization information for access thereof. A protocol governs the format, structure and syntax (using a Web Services Description Language schema) of messages for communicating to and from the Web crawler through an application programming interface on a server hosting the crawler application.


