A method for generating web document summaries based on text structure analysis
A technology for text structure and document summarization, applied in the fields of Chinese automatic summarization, natural language processing, and web page text extraction, it can solve the problems of short research history of automatic summarization technology, inability to truly realize automatic summarization, and less rigorous title naming. High abstract coverage, fast and accurate information search, and smooth abstract effect
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Publication Date
- 2017-02-08
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention relates to the technical fields of web page text extraction, natural language processing, and Chinese automatic summarization, in particular to a generation method of web document summaries based on text structure analysis. Background technique
[0002] At present, the Internet has become the main source for people to obtain information. Especially with the rapid development of User Generated Content (UGC) in recent years, the information on the Internet is growing explosively. Although search engines can return search results according to user requirements. But users still need to find the most suitable webpage for their needs from the search list, especially because there are a large number of search engine optimization and reposting phenomena on the Internet, which brings great difficulties to users to quickly and accurately find information.
[0003] The automatic summarization system uses computers to quickly process web documents,...
Examples
Embodiment Construction
[0049] The invention discloses a search engine-oriented method for generating web document abstracts, which can automatically analyze a web page and generate text abstracts reflecting the theme of the web page.
[0050] The invention includes a webpage body text extraction that combines visual features and text features and an automatic text summary based on subtopic division through text structure analysis.
[0051] The invention takes a URL as input, and finally generates a text summary through two stages of web page text extraction and automatic summary.
[0052] The following is a further description of the specific algorithms of the two stages, combined with an example of summarizing a news web page:
[0053] figure 1 Describes the overall process from the URL to be summarized to generating the summary, including the web page preprocessing process and the automatic summarization process.
[0054] Specifically, in an embodiment, the present invention is in the web page p...