HTML Data Extraction via Semantic Clustering for Malformed Pages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Web search engines lack an efficient method to programmatically compare information across multiple HTML pages, leading to time-consuming manual searches and inaccurate results due to the unstructured nature of HTML documents and high rates of malformed pages.

Innovation Solution

A computer-implemented system and method that automatically extracts and consolidates data from multiple HTML pages into a tabular format, resistant to malformed HTML and not requiring specific tag names or spatial reasoning, allowing for real-time comparison of people, places, or things by creating and merging tables, and identifying relevant information without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review of Web search results is used, then accuracy of information comparison can be maintained, but time consumption increases significantly

Engineering Contradiction:
Improveaccuracy of information comparisonVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically extracts, consolidates, and formats data from multiple HTML pages into comparison tables without requiring manual user intervention. The automated algorithm processes web results, identifies relevant information, and generates structured comparison data, allowing the system to serve itself rather than requiring continuous manual review by users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical review processes with an automated computer-based extraction and consolidation system. The algorithm automatically parses HTML documents, identifies tabular data, merges information from multiple sources, and presents consolidated results, substituting the manual mechanical task of reviewing each web result individually with an automated digital process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If traditional HTML data extraction methods are used, then data can be obtained from Web pages, but reliability decreases due to malformed HTML and lack of structure

Engineering Contradiction:
Improvedata extraction capabilityVSAvoiddata extraction accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system changes the approach to data extraction by transitioning from relying on fixed HTML tag structures to using semantic meaning and contextual relationships. The algorithm identifies data based on the meaning of content and its relationships rather than depending on specific HTML syntax, making the extraction process robust against malformed or non-standard HTML while maintaining high data quality.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary processing layer between the raw HTML data and the final extracted information. This intermediate algorithmic layer analyzes the semantic content, identifies relationships between elements, and consolidates data regardless of the original HTML structure, acting as a mediator that transforms unreliable raw HTML into reliable structured data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated data extraction algorithms are implemented, then productivity increases, but device complexity increases

Engineering Contradiction:
Improvedata extraction and consolidation speedVSAvoidalgorithm complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The complex data extraction and consolidation process is segmented into distinct manageable stages: HTML parsing, semantic analysis, data identification, information consolidation, and table generation. Each stage handles a specific aspect of the processing pipeline, making the overall complex system more manageable and maintainable while achieving high productivity through automated execution of these segmented tasks.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8868621B2Data extraction from HTML documents into tables for user comparison
Publication Date: 2014.10.21 RILLIP
  • US8868621B2 patent drawing
  • US8868621B2 patent drawing
  • US8868621B2 patent drawing

AI summary

The Computer-implemented system, method or computer program that creates a data table of rows and columns from an HTML Web page or document independent of the HTML markup tags. Data embedded in the HTML is identified using clustering of text and extracted into a data table. The generation of data tables can be performed in real-time and is not subject to problems with malformed or poorly created HTML.