Boilerplate Text Discrimination Using Local Language Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying boilerplate text in web pages are inadequate, as they rely on non-language characteristics like shape and pixel count, which fail to distinguish boilerplate from primary content, especially when boilerplate text varies across pages or is similar in length to main content.

Innovation Solution

A computer-implemented method that generates local language models for text elements with the same label, compares these models to derive similarity scores, and uses these scores to determine and filter out boilerplate text, independent of non-language characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If non-language characteristics like shape and pixel count are used to identify boilerplate text, then the identification process is simple, but the accuracy deteriorates when boilerplate text varies across pages or is similar in length to main content

Engineering Contradiction:
Improveidentification process complexityVSAvoidboilerplate text identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters used for identification from non-language characteristics (shape, pixel count) to language-based characteristics (local language models, word frequency, n-gram patterns). This allows the system to accurately distinguish boilerplate text even when it varies across pages or is similar in length to main content, while maintaining computational feasibility through efficient language modeling techniques.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical/visual analysis system (based on shape and pixel count) with a linguistic analysis system (based on local language models and word frequency). This substitution enables the system to capture the semantic and stylistic characteristics of boilerplate text, improving identification accuracy without excessive complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If language-based local models are generated and compared for text elements, then the identification accuracy improves, but the computational complexity increases

Engineering Contradiction:
Improveboilerplate text identification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the text analysis process by generating local language models for individual text elements (nodes) rather than analyzing entire pages. This segmentation allows for efficient computation on smaller, manageable units while maintaining high identification accuracy through comparative analysis of local characteristics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by focusing language model generation and comparison only on text elements with specific labels that are likely to contain boilerplate text. This selective approach reduces overall computational complexity while maintaining high accuracy for the target identification task.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If frequency-based methods are used to identify boilerplate text, then commonly occurring boilerplate is filtered out, but the method fails to identify boilerplate that occurs on only a few pages

Engineering Contradiction:
Improveboilerplate text identification accuracyVSAvoidmethod applicability to varied boilerplate
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the identification parameter from occurrence frequency to local language model similarity. This allows the system to identify boilerplate text based on its linguistic characteristics rather than how often it appears, enabling detection of boilerplate that occurs on only a few pages while maintaining the ability to filter commonly occurring boilerplate.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces local language models as an intermediary representation that captures the linguistic essence of text elements. This intermediary allows for flexible comparison and identification of boilerplate text regardless of its frequency of occurrence, bridging the gap between exact matching and frequency-based approaches.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11170759B2System and method for discriminating removing boilerplate text in documents comprising structured labelled text elements
Publication Date: 2021.11.09 VERINT SYST UK LTD
  • US11170759B2 patent drawing
  • US11170759B2 patent drawing
  • US11170759B2 patent drawing

AI summary

A method, system, and computer program product for discriminating boilerplate text in documents, such as web pages. An example method includes: receiving documents structured as labelled text elements; generating a local language model for each labelled text element of the received documents; comparing local language models for different labelled text elements that have the same label; for each comparison of local language models, deriving a similarity indicator, and using the similarity indicators of all the comparisons to derive a similarity score for that label; using the similarity scores to determine labels associated with text elements comprising boilerplate text; and providing the textual content of the labelled text elements to a receiving computer system; and identifying the textual content of labelled text elements that include boilerplate text.