Web page text extraction method and device based on aggregated text density

A webpage text extraction and text technology, which is applied in the direction of text database indexing, unstructured text data retrieval, website content management, etc., can solve the problems of cumbersome and complicated extraction of text, complicated simple problems, and unfavorable wide application, so as to achieve accurate extraction High efficiency, avoid low efficiency, and strong versatility

CN105740355BActive Publication Date: 2019-03-26NAT UNIV OF DEFENSE TECH
5 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Publication Date
2019-03-26

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
Patent Text Reader

Abstract

The present invention provides an aggregated text density based webpage body text extraction method and apparatus. In the method, webpage text content is segmented by a method of separating a webpage HTML according to a tag, so as to effectively separate various types of texts in the content. A special website extraction rule does not need to be customized, so that the method is high in generality; a complex text mining means is not required, so that the method is simple and efficient and accurate for extraction of various types of webpage body texts.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The present invention relates to the technical field of webpage reptiles, in particular to a method and device for extracting webpage text based on aggregated text density. Background technique

[0002] With the rapid development of social informatization, the Internet has become an important source of information for people. Netizens usually use browsers to directly view the content of web pages. In addition, there are many Internet-based information processing tasks (such as information retrieval, data mining, machine translation, etc.) that are also based on the information content of web pages. The body of the web page is processed. But besides useful information (such as body content), most web pages also contain a lot of noise information, such as website navigation information, related links and advertisements, copyright information, and some scripting languages. How to extract the text information of web pages accurately and efficiently, so a...

Examples

Embodiment Construction

[0027]

[0028] figure 1

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043] i+2 i+2

[0044]

[0045]

[0046] figure 2

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056] i+2 i+2 i+2

[0057]

[0058]

[0059] Tags are parsed and stored as units; paragraphs are clustered using a text clustering algorithm and the text is finally generated. Existing problems: simple problems are complicated, which makes extracting the text cumbersome and complicated, which is not conducive to wide application. SUMMARY OF THE INVENTION The purpose of the present invention is to provide a method and device for extracting webpage text based on aggregated text density in order to solve the technical problems in the prior art mentioned in the background art above. The ...