Automatic Web Page Summarization Using Information Element Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic summarization methods for web page and text information lack quality due to incomplete information elements and the absence of organization structure words, making it difficult to extract essential information elements, leading to low-quality summaries.

Innovation Solution

A method that utilizes a double-ten law to categorize and tag information elements, extracting content keywords and contexts by matching them with high-frequency organization structure words, and adjusting the summarization process based on quality indicators to improve summary quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If full-text keyword search is used to organize web page information, then information can be retrieved, but summary quality remains low and users must browse multiple pages

Engineering Contradiction:
Improvesummary qualityVSAvoidtime to find needed information
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments web page information into distinct information elements (title, author, source, time, keywords, summary, etc.) and extracts them separately using structured templates. This segmentation allows the system to identify and organize key information elements systematically, improving summary quality without requiring users to browse multiple pages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary information element extraction layer between full-text search and summary generation. This intermediary process uses structured templates and extraction rules to identify and organize key information elements from web pages, transforming unstructured content into structured summaries that directly answer user queries.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If organization structure words are not present in original text, then information elements cannot be easily identified, but adding them increases processing complexity

Engineering Contradiction:
Improveinformation element extraction accuracyVSAvoidsummarization process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-defining information element templates and extraction rules before processing web page content. These templates specify the structure and characteristics of different information elements (title, author, time, etc.), allowing the system to systematically identify and extract elements without adding complex processing steps during actual summarization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by transforming unstructured web page text into structured information elements with defined attributes. By converting content into standardized formats with specific parameters (element type, position, content), the system achieves high extraction accuracy while maintaining manageable process complexity through consistent transformation rules.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If information elements are incomplete in original text, then high-quality summary cannot be generated, but requiring complete elements reduces the scope of summarizable content

Engineering Contradiction:
Improvesummary qualityVSAvoidrange of summarizable information
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies partial action by extracting and utilizing only the information elements that are present in the source text, rather than requiring all possible elements to be complete. The system generates high-quality summaries using available elements and omits or marks missing elements appropriately, maintaining summary quality while preserving adaptability to various text types with different levels of information completeness.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11514242B2Method for automatically summarizing internet web page and text information
Publication Date: 2022.11.29 CHONGQING SIZAI INFORMATION TECH CO LTD
  • US11514242B2 patent drawing
  • US11514242B2 patent drawing
  • US11514242B2 patent drawing

AI summary

Based on a double-ten law of an Internet information organization structure, the present invention provides a method for automatically summarizing web page and text information, where matching information elements by category is added to an existing method, and content keywords and related contexts of summary (or abstract) information are directly extracted by using various successfully matched information element organization structure words. On this basis, the present invention further provides a method for supplementing summary information elements based on title information, and a method for superposing title information element organization structure words. Therefore, a new method for automatically summarizing web page and text information is comprehensively and systematically provided, and automatic summarization quality of the web page and text information can be greatly improved.