Hierarchical Content Extraction System for Web Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for monitoring social networks and user-generated content are inefficient in extracting relevant information due to reliance on keyword-based classification and inability to traverse hierarchical web structures, making it difficult for businesses to respond to unfavorable comments and reviews across multiple online platforms.

Innovation Solution

A method for automatically extracting content from hierarchical data resources using a training phase to define and identify relevant entities and properties, followed by a content extraction phase that compares data resources with composite schemas to extract targeted information, allowing for the identification and extraction of specific content across multiple levels of a web site hierarchy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If keyword-based classification is used to monitor social networks, then the system can categorize and prioritize posts, but it cannot accurately extract relevant information from hierarchical web structures and multi-level content

Engineering Contradiction:
Improvecontent extraction efficiencyVSAvoidrelevance detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the web page structure into hierarchical levels (first level, second level, third level content) and processes each level separately with dedicated extraction rules. This allows the system to handle complex hierarchical structures systematically, improving both efficiency and accuracy by addressing each structural layer with appropriate processing techniques.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from flat keyword-based search to a multi-dimensional hierarchical structure analysis. By organizing content extraction across multiple levels (first level elements, second level nested elements, third level nested elements) and processing them in sequence with increasing specificity, the system achieves precise relevance detection while maintaining productivity through structured processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If Dapper extracts content from web pages at the same hierarchical level, then it can identify relevant content, but it cannot traverse the hierarchy to extract content from multiple levels

Engineering Contradiction:
Improvehierarchy traversal capabilityVSAvoidcontent extraction system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the content extraction process into distinct phases for different hierarchical levels. First level content is extracted using initial rules, then second level nested content is extracted using refined rules, followed by third level nested content extraction. This segmentation allows the system to traverse multiple hierarchy levels while managing complexity through structured, phased processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary identification and extraction of first level content before proceeding to extract second and third level nested content. This preliminary action establishes the hierarchical structure and context before extracting deeper nested elements, enabling the system to traverse complex hierarchies systematically without overwhelming complexity.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If businesses monitor all comments and user-generated content across multiple platforms, then they can respond to unfavorable information, but the volume of data makes practical monitoring impossible

Engineering Contradiction:
Improvemonitoring coverageVSAvoiddata processing capacity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts only the relevant content from the vast amount of user-generated data by applying hierarchical extraction rules that identify and isolate meaningful elements (first level, second level, third level content) from the overall data stream. This extraction process enables reliable monitoring coverage by focusing processing capacity on extracted relevant content rather than processing all data uniformly.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different extraction and processing qualities to different hierarchical levels. First level content receives initial processing, second level nested content receives refined processing, and third level nested content receives specialized processing. This local quality approach optimizes productivity by applying appropriate processing intensity to appropriate data layers, enabling effective monitoring without overwhelming resources.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10545928B2Textual analysis system for automatic content extaction
Publication Date: 2020.01.28 BRITISH TELECOM PLC
  • US10545928B2 patent drawing
  • US10545928B2 patent drawing
  • US10545928B2 patent drawing

AI summary

The present invention provides a method, and an associated apparatus configured to implement such a method, for analysing mark-up language text content, such as might be found on a website or within online user generated content. The method comprises a training phase, in which plurality of schemas are automatically generated from a specified text and a final schema is compiled. This final schema can then be used to compare with other online text content such that content which matched the final schema can be identified, for example for further analysis and comparison.