System and method for real-time intelligent capturing of article

An article and intelligent technology, which is applied in the field of Internet technology to capture technology, can solve the problems of inability to accurately extract articles, low usability of captured articles, and consumption of network hardware resources, so as to improve news coverage and real-time performance, and improve Coverage and real-time performance, fast approximate weight-removal effect

CN102402627AActive Publication Date: 2012-04-04凤凰在线(北京)信息技术有限公司
3 Cites 13 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Publication Date
2012-04-04

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

The invention discloses a system for real-time intelligent capturing of an article. The system comprises a real-time capturing module, a webpage extraction system, a similar document duplicate-removing module, a document automatic classification module and an article publishing module. The real-time capturing module further comprises seven modules running online: a task extraction module, a task analysis module, a task capturing time range test module, a task capturing time interval test module, a task scheduling module, a task downloading module and a task capturing frequency regulation module; and the real-time capturing module still comprises three modules running offline: a task capturing time range discovery module, a task capturing time internal discovery module and a nonprofit agent collection and authentication module.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The invention relates to the fields of crawling technology, web mining technology, information extraction technology, and natural language processing technology in Internet technology; it can be applied to Internet fields such as portal websites and search engine websites that require large-scale, accurate, and real-time crawling of articles. Background technique

[0002] Internet portal websites have a large demand for reprinting articles every day, and have high requirements for the quality of articles. Many existing crawling systems can meet this requirement, but they all suffer from the following three problems:

[0003] 1) The crawling system that uses the machine-automatically generated extraction wrapper technology can capture a large number of articles, but it cannot achieve accurate extraction of articles, and the usability of crawling articles is low;

[0004] 2) The article extraction results of the crawling system using the artificially ge...

Examples

Embodiment Construction

[0080] The grabbing system consists of 5 modules or subsystems, such as figure 1 shown. Including: real-time crawling module, web page extraction system, document approximation deduplication module, document automatic classification module, and article publishing module.

[0081] The overall data flow of the system is as follows: figure 2 As shown, the specific steps are as follows:

[0082] Step 1, submit a job or a bunch of jobs to the real-time capture module of the system; the real-time capture module can be mainly divided into two main steps: a jobs analysis scheduling module and a crawler download module (task download module);

[0083] Step 2, the jobs parsing and scheduling module of the real-time crawling module is responsible for explaining each job to several rules stipulated by the cost system. These rules specify the specific crawling logic of the crawler module in the next step; A job schedule is distributed to a suitable server to achieve faster job capture ...