HTML Data Extraction via Semantic Clustering for Malformed Pages
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Web search engines lack an efficient method to programmatically compare information across multiple HTML pages, leading to time-consuming manual searches and inaccurate results due to the unstructured nature of HTML documents and high rates of malformed pages.
Innovation Solution
A computer-implemented system and method that automatically extracts and consolidates data from multiple HTML pages into a tabular format, resistant to malformed HTML and not requiring specific tag names or spatial reasoning, allowing for real-time comparison of people, places, or things by creating and merging tables, and identifying relevant information without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of Web search results is used, then accuracy of information comparison can be maintained, but time consumption increases significantly
Solution Approach 1:
The system automatically extracts, consolidates, and formats data from multiple HTML pages into comparison tables without requiring manual user intervention. The automated algorithm processes web results, identifies relevant information, and generates structured comparison data, allowing the system to serve itself rather than requiring continuous manual review by users.
Solution Approach 2:
The patent replaces manual mechanical review processes with an automated computer-based extraction and consolidation system. The algorithm automatically parses HTML documents, identifies tabular data, merges information from multiple sources, and presents consolidated results, substituting the manual mechanical task of reviewing each web result individually with an automated digital process.
2Quantity of substance
If traditional HTML data extraction methods are used, then data can be obtained from Web pages, but reliability decreases due to malformed HTML and lack of structure
Solution Approach 1:
The system changes the approach to data extraction by transitioning from relying on fixed HTML tag structures to using semantic meaning and contextual relationships. The algorithm identifies data based on the meaning of content and its relationships rather than depending on specific HTML syntax, making the extraction process robust against malformed or non-standard HTML while maintaining high data quality.
Solution Approach 2:
The patent introduces an intermediary processing layer between the raw HTML data and the final extracted information. This intermediate algorithmic layer analyzes the semantic content, identifies relationships between elements, and consolidates data regardless of the original HTML structure, acting as a mediator that transforms unreliable raw HTML into reliable structured data.
3Productivity
If automated data extraction algorithms are implemented, then productivity increases, but device complexity increases
Solution Approach 1:
The complex data extraction and consolidation process is segmented into distinct manageable stages: HTML parsing, semantic analysis, data identification, information consolidation, and table generation. Each stage handles a specific aspect of the processing pipeline, making the overall complex system more manageable and maintainable while achieving high productivity through automated execution of these segmented tasks.
Data Source
AI summary
The Computer-implemented system, method or computer program that creates a data table of rows and columns from an HTML Web page or document independent of the HTML markup tags. Data embedded in the HTML is identified using clustering of text and extracted into a data table. The generation of data tables can be performed in real-time and is not subject to problems with malformed or poorly created HTML.


