System and method for real-time intelligent capturing of article
An article and intelligent technology, which is applied in the field of Internet technology to capture technology, can solve the problems of inability to accurately extract articles, low usability of captured articles, and consumption of network hardware resources, so as to improve news coverage and real-time performance, and improve Coverage and real-time performance, fast approximate weight-removal effect
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Publication Date
- 2012-04-04
Smart Images
Figure 1 Figure 2 Figure 3
Abstract
Description
technical field
[0001] The invention relates to the fields of crawling technology, web mining technology, information extraction technology, and natural language processing technology in Internet technology; it can be applied to Internet fields such as portal websites and search engine websites that require large-scale, accurate, and real-time crawling of articles. Background technique
[0002] Internet portal websites have a large demand for reprinting articles every day, and have high requirements for the quality of articles. Many existing crawling systems can meet this requirement, but they all suffer from the following three problems:
[0003] 1) The crawling system that uses the machine-automatically generated extraction wrapper technology can capture a large number of articles, but it cannot achieve accurate extraction of articles, and the usability of crawling articles is low;
[0004] 2) The article extraction results of the crawling system using the artificially ge...
Examples
Embodiment Construction
[0080] The grabbing system consists of 5 modules or subsystems, such as figure 1 shown. Including: real-time crawling module, web page extraction system, document approximation deduplication module, document automatic classification module, and article publishing module.
[0081] The overall data flow of the system is as follows: figure 2 As shown, the specific steps are as follows:
[0082] Step 1, submit a job or a bunch of jobs to the real-time capture module of the system; the real-time capture module can be mainly divided into two main steps: a jobs analysis scheduling module and a crawler download module (task download module);
[0083] Step 2, the jobs parsing and scheduling module of the real-time crawling module is responsible for explaining each job to several rules stipulated by the cost system. These rules specify the specific crawling logic of the crawler module in the next step; A job schedule is distributed to a suitable server to achieve faster job capture ...