Automatic Web Page Summarization Using Information Element Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automatic summarization methods for web page and text information lack quality due to incomplete information elements and the absence of organization structure words, making it difficult to extract essential information elements, leading to low-quality summaries.
Innovation Solution
A method that utilizes a double-ten law to categorize and tag information elements, extracting content keywords and contexts by matching them with high-frequency organization structure words, and adjusting the summarization process based on quality indicators to improve summary quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full-text keyword search is used to organize web page information, then information can be retrieved, but summary quality remains low and users must browse multiple pages
Solution Approach 1:
The patent segments web page information into distinct information elements (title, author, source, time, keywords, summary, etc.) and extracts them separately using structured templates. This segmentation allows the system to identify and organize key information elements systematically, improving summary quality without requiring users to browse multiple pages.
Solution Approach 2:
The patent introduces an intermediary information element extraction layer between full-text search and summary generation. This intermediary process uses structured templates and extraction rules to identify and organize key information elements from web pages, transforming unstructured content into structured summaries that directly answer user queries.
2Measurement precision
If organization structure words are not present in original text, then information elements cannot be easily identified, but adding them increases processing complexity
Solution Approach 1:
The patent performs preliminary action by pre-defining information element templates and extraction rules before processing web page content. These templates specify the structure and characteristics of different information elements (title, author, time, etc.), allowing the system to systematically identify and extract elements without adding complex processing steps during actual summarization.
Solution Approach 2:
The patent changes parameters by transforming unstructured web page text into structured information elements with defined attributes. By converting content into standardized formats with specific parameters (element type, position, content), the system achieves high extraction accuracy while maintaining manageable process complexity through consistent transformation rules.
3Measurement precision
If information elements are incomplete in original text, then high-quality summary cannot be generated, but requiring complete elements reduces the scope of summarizable content
Solution Approach 1:
The patent applies partial action by extracting and utilizing only the information elements that are present in the source text, rather than requiring all possible elements to be complete. The system generates high-quality summaries using available elements and omits or marks missing elements appropriately, maintaining summary quality while preserving adaptability to various text types with different levels of information completeness.
Data Source
AI summary
Based on a double-ten law of an Internet information organization structure, the present invention provides a method for automatically summarizing web page and text information, where matching information elements by category is added to an existing method, and content keywords and related contexts of summary (or abstract) information are directly extracted by using various successfully matched information element organization structure words. On this basis, the present invention further provides a method for supplementing summary information elements based on title information, and a method for superposing title information element organization structure words. Therefore, a new method for automatically summarizing web page and text information is comprehensively and systematically provided, and automatic summarization quality of the web page and text information can be greatly improved.


