Ontology-Directed Data Structuring for Heterogeneous Web Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in efficiently retrieving and organizing vast amounts of information from diverse sources on the internet and integrating it with existing data systems, as current methods are costly, time-consuming, and often inaccurate due to the lack of standard formatting and varying lexicons.

Innovation Solution

The solution involves creating a system with tools like web agent creators, text extractors, ontology management systems, and validation components to acquire, structure, and categorize data from multiple sources, using algorithms to infer meanings and reformat data into a uniform format for use in queriable databases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual extraction, classification, structuring and categorization of data is performed by skilled experts, then data accuracy and quality are improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service data extraction and structuring through AI algorithms that autonomously navigate web pages, extract relevant information, and organize data without continuous human intervention. The automated agent performs classification and structuring tasks that would otherwise require skilled manual labor, thereby reducing time consumption while maintaining data quality through algorithmic precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of data extraction and classification with an automated software agent that uses AI and machine learning techniques. This substitution transforms the manual expert system into an automated intelligent system, eliminating the need for continuous human involvement while preserving and enhancing data accuracy through consistent algorithmic application.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If information is retrieved from diverse internet sources with varying formats and lexicons, then information availability increases, but data organization and integration become more difficult

Engineering Contradiction:
Improveinformation availabilityVSAvoiddata organization complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The automated agent is designed with universal capabilities to handle multiple data formats and sources. It can navigate various web page structures, extract information from different formats (HTML, PDF, text), and adapt to diverse lexicons through AI-powered semantic understanding. This multi-functionality allows the system to retrieve information from diverse sources while maintaining consistent organization and integration processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts its extraction and classification parameters based on the characteristics of each data source. By using AI algorithms that can recognize and adapt to different formats, structures, and lexicons, the agent automatically modifies its processing approach for each source, thereby managing the complexity of organizing diverse information while maximizing information availability.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If existing search engines are used to organize web content, then casual user needs are met, but accuracy and completeness of information retrieval are insufficient

Engineering Contradiction:
Improveuser accessibilityVSAvoidinformation accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an automated agent as an intermediary between users and web information. This agent acts as a specialized mediator that goes beyond traditional search engines by autonomously navigating, extracting, and organizing information according to specific criteria. The intermediary maintains ease of operation through automated processes while significantly improving information accuracy and completeness through intelligent selection and verification of retrieved data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by proactively retrieving, extracting, and organizing information before users need it. The automated agent continuously navigates web sources, extracts relevant data, and structures it in advance, making accurate and complete information readily available when needed, rather than relying on users to manually search and evaluate sources.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7542958B1Methods for determining the similarity of content and structuring unstructured content from heterogeneous sources
Publication Date: 2009.06.02 XSB
  • US7542958B1 patent drawing
  • US7542958B1 patent drawing
  • US7542958B1 patent drawing

AI summary

The invention includes methods and software tools for acquiring data from diverse sources, and structuring the data in a form that may be used to determine object equivalence. Practice of the invention includes one or more of the following tools: a data acquisition web agent creator, a web agent created by the web agent creator, an agent manager for deploying said web agent, and ontology-directed classifier, an ontology-directed extractor, and an ontology-directed matcher. The tools are example driven through a graphical user interface.