Web Page Characteristic Content Extraction via Frequency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for extracting characteristic content from Web pages are inefficient, requiring manual user designation and relying on specific content types like images or texts, leading to low reliability and excessive user effort, especially when dealing with multiple target Web pages.
Innovation Solution
A device and method that calculates the frequency of appearance of content within a designated Web page and across other pages to determine characteristic content, using extraction units, calculation units, and determination units to identify and generate new content based on these frequencies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual user designation is used to extract characteristic content, then user control over content selection is improved, but user effort and time consumption increase enormously
Solution Approach 1:
The system performs automatic extraction of characteristic content by itself without requiring manual user designation. The extraction unit automatically identifies and extracts characteristic content based on frequency analysis, making the system self-sufficient and eliminating the need for extensive user intervention.
Solution Approach 2:
The patent replaces manual mechanical selection processes with automated computational analysis. Frequency calculation units compute occurrence frequencies programmatically, and determination units automatically identify characteristic content based on these calculations, substituting manual effort with automated information processing.
2Ease of manufacture
If only specific content types are extracted based on HTML tags, then extraction process is simplified, but extraction reliability decreases due to inclusion of non-characteristic content
Solution Approach 1:
The patent introduces frequency of appearance as a new parameter to distinguish characteristic content from non-characteristic content. By calculating and comparing frequency parameters across different content elements, the system reliably identifies characteristic content while maintaining an automated extraction process.
Solution Approach 2:
The extraction process is segmented into distinct functional units: an extraction unit that retrieves content, frequency calculation units that compute occurrence frequencies, and determination units that identify characteristic content. This segmentation allows each unit to perform its specific function efficiently while maintaining overall reliability.
3Measurement precision
If frequency calculation is performed across multiple Web pages, then identification accuracy of characteristic content is improved, but computational complexity and processing time increase
Solution Approach 1:
The computational process is divided into separate frequency calculation units, each responsible for calculating frequencies of specific content elements. This segmentation reduces the complexity burden on any single unit and enables parallel processing of frequency calculations across multiple Web pages.
Solution Approach 2:
The frequency calculation mechanism is designed as a universal process that can be applied to any content element across any Web page. The same calculation logic and determination criteria are used consistently throughout the system, simplifying the overall computational framework despite processing multiple pages.
Data Source
AI summary
A characteristic content determination device extracts a content constituting a designated Web page. The characteristic content determination device calculates a first frequency of appearance of each content constituting the designated Web page in the designated Web page. The characteristic content determination device calculates a second frequency of appearance of each content constituting the designated Web page in other Web pages. Then, the characteristic content determination device determines a characteristic content of the designated Web page among contents constituting the designated Web page based on the calculated first frequency of appearance and the calculated second frequency of appearance.


