Non-Body Text Recognition via DOM Tree Frequency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for recognizing non-body text in webpages, such as garbage keyword density, are inefficient due to the need for constant dictionary updates and lag when processing large amounts of data, particularly in search engines and mobile reading applications.
Innovation Solution
A system and method utilizing a DOM tree construction unit, text statistics unit, and text recognition unit to identify non-body text by constructing a DOM tree for each webpage, analyzing unit text sections, and determining their occurrence frequency to classify them as non-body text based on a predetermined threshold.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If garbage keyword density method is used for non-body text recognition, then recognition capability is provided, but recognition accuracy deteriorates due to lag in dictionary updates
Solution Approach 1:
The patent performs preliminary actions by constructing a complete DOM tree structure before analyzing text sections. The system pre-processes the webpage by establishing the hierarchical DOM structure, then uses this pre-built structure to efficiently locate and analyze text sections without lag. This preliminary structuring enables accurate identification of non-body text sections even when new garbage keywords appear, eliminating the need for reactive dictionary updates.
2Reliability
If garbage keyword dictionary is constantly updated, then recognition capability is maintained, but processing time increases due to update operations
Solution Approach 1:
The system employs self-service by automatically analyzing the DOM tree structure to identify non-body text sections without requiring external dictionary updates. The DOM tree analysis unit autonomously determines text section boundaries and characteristics based on the webpage's own structural information, eliminating the need for time-consuming dictionary maintenance operations while maintaining accurate recognition capability.
3Measurement precision
If DOM tree construction is performed for each webpage, then text section analysis accuracy is improved, but system complexity increases
Solution Approach 1:
The patent applies universality by using the DOM tree structure for multiple purposes: it serves as the foundation for text section analysis, provides the framework for identifying non-body text, and enables efficient navigation through webpage content. This single DOM tree construction performs multiple functions that would otherwise require separate processing mechanisms, maintaining high analysis accuracy while avoiding additional system complexity.
Data Source
AI summary
The invention discloses a system and method for recognizing the non-body text in a webpage, and relates to the field of main body extraction. The system comprises: a webpage grabber configured to grab data of all the webpages of a target website; a DOM tree construction unit configured to construct a DOM tree corresponding to each webpage of the target website; a DOM tree analysis unit configured to find out a unit text section in the webpage according to the DOM tree; a text statistics unit configured to conduct statistics on the number of occurrence of the unit text section in all the webpages of the target website; and a text recognition unit configured to recognize the unit text section as a non-body text when the number of occurrence is greater than a predetermined threshold. The system and the method overcome the problem of lag of recognition of a non-body text in the prior art method, and have a high recognition accuracy.


