Non-Body Text Recognition via DOM Tree Frequency Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for recognizing non-body text in webpages, such as garbage keyword density, are inefficient due to the need for constant dictionary updates and lag when processing large amounts of data, particularly in search engines and mobile reading applications.

Innovation Solution

A system and method utilizing a DOM tree construction unit, text statistics unit, and text recognition unit to identify non-body text by constructing a DOM tree for each webpage, analyzing unit text sections, and determining their occurrence frequency to classify them as non-body text based on a predetermined threshold.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If garbage keyword density method is used for non-body text recognition, then recognition capability is provided, but recognition accuracy deteriorates due to lag in dictionary updates

Engineering Contradiction:
Improverecognition capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by constructing a complete DOM tree structure before analyzing text sections. The system pre-processes the webpage by establishing the hierarchical DOM structure, then uses this pre-built structure to efficiently locate and analyze text sections without lag. This preliminary structuring enables accurate identification of non-body text sections even when new garbage keywords appear, eliminating the need for reactive dictionary updates.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If garbage keyword dictionary is constantly updated, then recognition capability is maintained, but processing time increases due to update operations

Engineering Contradiction:
Improverecognition capabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system employs self-service by automatically analyzing the DOM tree structure to identify non-body text sections without requiring external dictionary updates. The DOM tree analysis unit autonomously determines text section boundaries and characteristics based on the webpage's own structural information, eliminating the need for time-consuming dictionary maintenance operations while maintaining accurate recognition capability.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If DOM tree construction is performed for each webpage, then text section analysis accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvetext section analysis accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by using the DOM tree structure for multiple purposes: it serves as the foundation for text section analysis, provides the framework for identifying non-body text, and enables efficient navigation through webpage content. This single DOM tree construction performs multiple functions that would otherwise require separate processing mechanisms, maintaining high analysis accuracy while avoiding additional system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10042827B2System and method for recognizing non-body text in webpage
Publication Date: 2018.08.07 BEIJING QIHOOD TECHNOLOGY CO LTD
  • US10042827B2 patent drawing
  • US10042827B2 patent drawing
  • US10042827B2 patent drawing

AI summary

The invention discloses a system and method for recognizing the non-body text in a webpage, and relates to the field of main body extraction. The system comprises: a webpage grabber configured to grab data of all the webpages of a target website; a DOM tree construction unit configured to construct a DOM tree corresponding to each webpage of the target website; a DOM tree analysis unit configured to find out a unit text section in the webpage according to the DOM tree; a text statistics unit configured to conduct statistics on the number of occurrence of the unit text section in all the webpages of the target website; and a text recognition unit configured to recognize the unit text section as a non-body text when the number of occurrence is greater than a predetermined threshold. The system and the method overcome the problem of lag of recognition of a non-body text in the prior art method, and have a high recognition accuracy.