Knowledge Graph Linking Layer for Heterogeneous Data Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data systems face challenges in correlating and connecting disparate data sets from multiple sources, leading to incomplete and uncertain inferences due to lack of context, especially when dealing with heterogeneous and unstructured data.

Innovation Solution

A knowledge graph-based data storage and retrieval system that employs a linking layer to interconnect multiple subgraphs representing diverse datasets, providing a unified framework for understanding and analyzing complex relationships within and between datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional data systems organize data into discrete sets with standardized identifiers, then data searching within single entities becomes easier and faster, but correlating data sets across multiple entities becomes difficult

Engineering Contradiction:
Improvesearch speedVSAvoiddata correlation capability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system segments data into multiple subgraphs representing different data sets from various sources, while using a linking layer to connect these subgraphs. This allows efficient searching within each subgraph while enabling correlation across subgraphs through standardized linking nodes, resolving the contradiction between fast single-entity search and cross-entity data correlation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The linking layer acts as an intermediary between multiple subgraphs, providing standardized linking nodes that mediate connections between different data sets. This intermediary structure enables correlation across heterogeneous data sources without compromising the integrity or search efficiency of individual data sets.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If data systems retrieve documents from multiple data sets, then broader information retrieval is achieved, but documents are provided with no context or relation to other documents

Engineering Contradiction:
Improveinformation quantityVSAvoidcontext information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The system segments information into subgraphs while preserving contextual relationships through the linking layer. When retrieving documents from multiple data sets, the context information is maintained through the standardized linking nodes that show relationships between documents across subgraphs, preventing context loss while expanding information quantity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The linking layer provides universal connection points that work across multiple subgraphs and data types. This multi-functional linking mechanism enables retrieval of documents from diverse sources while simultaneously providing contextual relationships, as the same linking infrastructure serves both broad retrieval and context preservation functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If conventional systems attempt to draw inferences between data sets, then cross-data insights are attempted, but results present unacceptable degrees of uncertainty

Engineering Contradiction:
Improveinference capabilityVSAvoidinference accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The standardized linking nodes in the linking layer serve as reliable intermediaries that provide consistent connection semantics across different subgraphs. This standardization reduces inference uncertainty by ensuring that links between data sets have well-defined meanings and relationships, improving both the versatility and reliability of cross-data inferences.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system applies homogeneity through standardized linking nodes and consistent data organization structures across different subgraphs. This homogeneity in the linking layer enables more reliable inferences by ensuring that connection patterns are consistent and predictable across different data sets, reducing the uncertainty that would otherwise result from heterogeneous data organization.

Inventive Principle:
Principle #33Homogeneity

Data Source

PatentUS12131262B2Data storage and retrieval system including a knowledge graph employing multiple subgraphs and a linking layer including multiple linking nodes, and methods, apparatus and systems for constructing and using same
Publication Date: 2024.10.29 PAREXEL INTERNATIONAL LLC
  • US12131262B2 patent drawing
  • US12131262B2 patent drawing
  • US12131262B2 patent drawing

AI summary

A graph-based data storage and retrieval system in which multiple subgraphs representing respective datasets in different namespaces are interconnected via a linking or “canonical” layer. Datasets represented by subgraphs in different namespaces may pertain to a particular information domain (e.g., the health care domain), and may include heterogeneous datasets. The canonical layer provides for a substantial reduction of graph complexity required to interconnect corresponding nodes in different subgraphs, which in turn offers advantages as the number of subgraphs (and the number of corresponding nodes in different subgraphs) increases for the particular domain(s) of interest. Examples of such advantages include reductions in data storage and retrieval times, and enhanced query/search efficacy, discovery of relationships in different parts of the system, ability to infer relationships in different parts of the system, and ability to train data models for natural language processing (NLP) and other purposes based on information extracted from the system.