Search Service Administration Web Service Protocol for Crawl Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current web crawling technologies are inefficient due to the vast amount of data on the web, leading to excessive network bandwidth consumption and resource wastage, as well as the inability to access valuable information due to authentication and authorization issues, resulting in missed relevant data.

Innovation Solution

A client-centric approach is implemented to configure and control web crawling functions through a Search Service application, utilizing an Application Programming Interface (API) governed by the Search Service Administration Web Service protocol, allowing users to define crawl rules, authentication, and access parameters, thereby controlling the crawling process and ensuring efficient data retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If web crawling is performed without restrictions to ensure comprehensive data retrieval, then the completeness of information is improved, but network bandwidth consumption and resource usage increase excessively

Engineering Contradiction:
Improvecompleteness of informationVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The patent segments the web crawling process into controlled phases with defined boundaries through crawl rules. The URL space is divided into crawlable and non-crawlable sections, allowing the crawler to systematically process only relevant portions of the web while avoiding unnecessary data retrieval and reducing network bandwidth consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic control mechanisms where crawl rules can be modified, added, or removed during the crawling process. The system adapts to changing requirements by allowing users to dynamically adjust crawl boundaries and authentication settings, optimizing the balance between information completeness and resource consumption.

Inventive Principle:
Principle #15Dynamics

2Area of stationary object

If web crawling is performed without authentication to ensure broad access, then the coverage of accessible data is improved, but the ability to retrieve protected valuable information is lost

Engineering Contradiction:
Improvecoverage of accessible dataVSAvoidaccess to protected information
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The patent implements preliminary authentication setup through crawl rules before the actual crawling process begins. Users configure authentication credentials and authorization parameters in advance, allowing the crawler to automatically present these credentials when encountering protected resources, thus ensuring both broad coverage and access to protected information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces authentication credentials as an intermediary between the crawler and protected web resources. The crawl rules serve as a mediator that manages the authentication process, allowing the system to access protected content without requiring manual intervention while maintaining security protocols.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual filtering of retrieved documents is performed to ensure quality information, then the relevance of retrieved data is improved, but the time and resources required increase significantly

Engineering Contradiction:
Improverelevance of retrieved dataVSAvoidtime for filtering documents
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary filtering by configuring crawl rules that define inclusion and exclusion criteria before the crawling process begins. This pre-filtering approach eliminates the need for extensive manual filtering afterward, as the crawler systematically retrieves only relevant documents based on pre-defined parameters, significantly reducing the time and resources required for post-processing.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If crawl rules and authentication settings are configured to reduce data retrieval volume, then resource efficiency is improved, but the complexity of system configuration increases

Engineering Contradiction:
Improveresource efficiencyVSAvoidconfiguration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal crawl rule configuration system that handles multiple functions through a unified interface. The same crawl rule structure manages URL patterns, authentication credentials, inclusion/exclusion criteria, and scheduling parameters, reducing configuration complexity while maintaining high resource efficiency through centralized control.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8364795B2Search service administration web service protocol
Publication Date: 2013.01.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8364795B2 patent drawing
  • US8364795B2 patent drawing
  • US8364795B2 patent drawing

AI summary

The embodiments described herein generally relate to a method and system for enabling a client to configure and control the crawling function available through a crawl configuration Web service. A client is able to configure and control the crawling function by defining the URL space of the crawl. Such space may be defined by configuring the starting point(s) and other properties of the crawl. The client further configures the crawling function by creating and configuring a content source and/or a crawl rule. Further, a client defines authentication information applicable to the crawl to enable the discovery and retrieval of electronic documents requiring authentication and/or authorization information for access thereof. A protocol governs the format, structure and syntax (using a Web Services Description Language schema) of messages for communicating to and from the Web crawler through an application programming interface on a server hosting the crawler application.