A Timely and Efficient Internet Information Crawling Method
A technology for Internet information and web page information, applied in the information field to simplify resource allocation, simplify the scope and complexity, and reduce misjudgments
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Publication Date
- 2016-06-29
Smart Images
Figure 1 Figure 2 Figure 3
Abstract
Description
technical field
[0001] The invention belongs to the field of information technology, and in particular relates to a timely and efficient Internet information crawling method. Background technique
[0002] With the rapid development of the Internet, it has become the largest public data source in the world, and its scale is still growing. Judging from the content contained therein, there are many webpage information linked together by hyperlinks on the Internet, and a considerable part of them has the characteristics of dynamic changes; based on this, many services can be provided on the Internet, and through The communication between people and organizations forms a virtual society that has a certain correspondence and relationship with the real society. For this reason, Web data mining, which aims to find useful knowledge from the structure, content, and logs of the Internet, has received great attention and development, especially the content mining that takes the content...
Examples
Embodiment Construction
[0037] The specific embodiment of the present invention is as figure 1 shown. The steps are detailed below.
[0038] 1. Information collection and collation (such as figure 2 shown)
[0039] 1. Collect relevant information Url address
[0040] According to the pre-determined topic meaning, first select a certain part (such as 3-5) topic keywords; enter these topic keywords on a general search engine to get a list of query results; organize the query results and extract Url to get some relevant information URL address.
[0041] 2. Initial Url setting and web page information crawling
[0042] Select Internet information crawler software (such as Heritrix, Nutch, etc.), and set these Url addresses obtained in steps 1 and 1 as seed Url addresses in the software. Parameters such as the number of pages (determined in advance) are set in the software, and then the general Internet information crawling method (without subject-related judgment and timeliness prediction) is used...