Hierarchical Content Extraction System for Web Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for monitoring social networks and user-generated content are inefficient in extracting relevant information due to reliance on keyword-based classification and inability to traverse hierarchical web structures, making it difficult for businesses to respond to unfavorable comments and reviews across multiple online platforms.
Innovation Solution
A method for automatically extracting content from hierarchical data resources using a training phase to define and identify relevant entities and properties, followed by a content extraction phase that compares data resources with composite schemas to extract targeted information, allowing for the identification and extraction of specific content across multiple levels of a web site hierarchy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If keyword-based classification is used to monitor social networks, then the system can categorize and prioritize posts, but it cannot accurately extract relevant information from hierarchical web structures and multi-level content
Solution Approach 1:
The patent segments the web page structure into hierarchical levels (first level, second level, third level content) and processes each level separately with dedicated extraction rules. This allows the system to handle complex hierarchical structures systematically, improving both efficiency and accuracy by addressing each structural layer with appropriate processing techniques.
Solution Approach 2:
The patent transitions from flat keyword-based search to a multi-dimensional hierarchical structure analysis. By organizing content extraction across multiple levels (first level elements, second level nested elements, third level nested elements) and processing them in sequence with increasing specificity, the system achieves precise relevance detection while maintaining productivity through structured processing.
2Adaptability or versatility
If Dapper extracts content from web pages at the same hierarchical level, then it can identify relevant content, but it cannot traverse the hierarchy to extract content from multiple levels
Solution Approach 1:
The patent divides the content extraction process into distinct phases for different hierarchical levels. First level content is extracted using initial rules, then second level nested content is extracted using refined rules, followed by third level nested content extraction. This segmentation allows the system to traverse multiple hierarchy levels while managing complexity through structured, phased processing.
Solution Approach 2:
The patent performs preliminary identification and extraction of first level content before proceeding to extract second and third level nested content. This preliminary action establishes the hierarchical structure and context before extracting deeper nested elements, enabling the system to traverse complex hierarchies systematically without overwhelming complexity.
3Reliability
If businesses monitor all comments and user-generated content across multiple platforms, then they can respond to unfavorable information, but the volume of data makes practical monitoring impossible
Solution Approach 1:
The patent extracts only the relevant content from the vast amount of user-generated data by applying hierarchical extraction rules that identify and isolate meaningful elements (first level, second level, third level content) from the overall data stream. This extraction process enables reliable monitoring coverage by focusing processing capacity on extracted relevant content rather than processing all data uniformly.
Solution Approach 2:
The patent applies different extraction and processing qualities to different hierarchical levels. First level content receives initial processing, second level nested content receives refined processing, and third level nested content receives specialized processing. This local quality approach optimizes productivity by applying appropriate processing intensity to appropriate data layers, enabling effective monitoring without overwhelming resources.
Data Source
AI summary
The present invention provides a method, and an associated apparatus configured to implement such a method, for analysing mark-up language text content, such as might be found on a website or within online user generated content. The method comprises a training phase, in which plurality of schemas are automatically generated from a specified text and a final schema is compiled. This final schema can then be used to compare with other online text content such that content which matched the final schema can be identified, for example for further analysis and comparison.


