Web page text extraction method and device based on aggregated text density
A webpage text extraction and text technology, which is applied in the direction of text database indexing, unstructured text data retrieval, website content management, etc., can solve the problems of cumbersome and complicated extraction of text, complicated simple problems, and unfavorable wide application, so as to achieve accurate extraction High efficiency, avoid low efficiency, and strong versatility
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Publication Date
- 2019-03-26
Smart Images

Figure 1 
Figure 2
Abstract
Description
technical field
[0001] The present invention relates to the technical field of webpage reptiles, in particular to a method and device for extracting webpage text based on aggregated text density. Background technique
[0002] With the rapid development of social informatization, the Internet has become an important source of information for people. Netizens usually use browsers to directly view the content of web pages. In addition, there are many Internet-based information processing tasks (such as information retrieval, data mining, machine translation, etc.) that are also based on the information content of web pages. The body of the web page is processed. But besides useful information (such as body content), most web pages also contain a lot of noise information, such as website navigation information, related links and advertisements, copyright information, and some scripting languages. How to extract the text information of web pages accurately and efficiently, so a...
Examples
Embodiment Construction
[0027]
[0028] figure 1
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043] i+2 i+2
[0044]
[0045]
[0046] figure 2
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056] i+2 i+2 i+2
[0057]
[0058]
[0059] Tags are parsed and stored as units; paragraphs are clustered using a text clustering algorithm and the text is finally generated. Existing problems: simple problems are complicated, which makes extracting the text cumbersome and complicated, which is not conducive to wide application. SUMMARY OF THE INVENTION The purpose of the present invention is to provide a method and device for extracting webpage text based on aggregated text density in order to solve the technical problems in the prior art mentioned in the background art above. The ...