Semantic Data Ingestion System for Distributed Knowledge Bases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in efficiently ingesting, modeling, and querying large knowledge bases across multiple disciplines, requiring significant time and effort, especially when dealing with terabytes or petabytes of data, which complicates research processes.

Innovation Solution

A system and method utilizing a cluster of computers for ingesting and analyzing large data sets, employing semantic processing, ontology alignment, and parallelized ontology reasoning to facilitate rapid ingestion, modeling, and querying of data, allowing for automated access to distributed knowledge bases and reducing the complexity of data integration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data integration methods are used to ingest and model large knowledge bases, then data integration can be achieved, but the process requires thousands of man hours and large numbers of computers, significantly increasing time consumption and system complexity

Engineering Contradiction:
Improvedata ingestion speedVSAvoidtime required for data integration
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical data integration methods with semantic processing and ontology-based automated reasoning. The system uses semantic annotations, ontology alignment algorithms, and automated transformation rules to ingest and integrate data from multiple sources simultaneously, eliminating the need for manual, sequential processing that previously required thousands of man hours.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If multiple consultations are performed across different knowledge bases, then comprehensive research coverage is achieved, but the complexity of the inquiry and time required to execute it increases significantly

Engineering Contradiction:
Improveresearch coverageVSAvoidinquiry complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple distributed knowledge bases into a unified semantic framework using ontology alignment and integration. The system combines data from diverse sources (biomedical research, genetic databases, clinical trials) into a single coherent model that allows researchers to conduct comprehensive inquiries across all knowledge bases simultaneously through a single interface, rather than performing multiple separate consultations.

Inventive Principle:
Principle #5Merging (Combining)

3Quantity of substance

If traditional data processing systems are used, then data can be stored and accessed, but the system cannot efficiently handle terabytes, petabytes or exabytes of data across multiple disciplines

Engineering Contradiction:
Improvedata volumeVSAvoiddata analysis efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments massive data volumes into manageable semantic units organized around ontological concepts and relationships. The system divides terabytes or petabytes of data into structured knowledge representations that can be processed in parallel across distributed computing resources, enabling efficient handling of large-scale multi-disciplinary data while maintaining semantic coherence across all segments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11003661B2System for rapid ingestion, semantic modeling and semantic querying over computer clusters
Publication Date: 2021.05.11 INFOTECH SOFT
  • US11003661B2 patent drawing
  • US11003661B2 patent drawing
  • US11003661B2 patent drawing

AI summary

A computer-implemented system within a computational cluster for aggregating data from a plurality of heterogeneous data sets into a homogeneous representation within a distributed, fault tolerant data source for the subsequent high-throughput retrieval, integration, and analysis of constituent data regardless of original data source location, data source format, and data encoding is configured for reading an input data set, generating a source data model based on the input data set, generating a vocabulary annotation and a profile of said source data model, executing semantic processing on said source data model so as to produce a normalized model based on standard ontology axioms and expanding the normalized model to identify and make explicit all implicit ontology axioms, and executing semantic querying on said data model using parallelized ontology reasoning based on an Enhance Most-Specific-Concept (MSC) algorithm.