Knowledge Graph Link Prediction via Neural Network Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing analytical applications and data warehousing systems fail to fully utilize large volumes of genetic and molecular data due to lack of proper data quality screening and contextual information, making it difficult to identify gene-disease associations efficiently.

Innovation Solution

A system and method for predicting node-to-node links in knowledge graphs using neural networks, which involves determining positive and negative structural scores based on significance parameters, generating synthetic negative datasets, and calculating likelihood scores to identify links between genes and diseases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is aggregated into large data warehouses without proper data quality screening and contextual information, then data volume increases, but data usefulness decreases

Engineering Contradiction:
Improvedata volumeVSAvoidcontextual information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the knowledge graph data into structured triples (subject, predicate, object) with associated metadata, allowing systematic processing and quality assessment of individual data elements rather than treating data as an undifferentiated aggregate

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces knowledge graphs as an intermediary layer between raw data warehouses and analytical applications, adding contextual information and relationships that bridge the gap between voluminous but unstructured data and useful insights

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If conventional string matching mechanisms are used to query data without context, then query simplicity is maintained, but identification accuracy decreases

Engineering Contradiction:
Improvequery simplicityVSAvoididentification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces conventional string matching mechanisms with neural network-based semantic analysis, enabling context-aware query processing that maintains ease of use while dramatically improving identification accuracy through learned representations of data relationships

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If large amounts of computing resources are used to transform information into searchable data, then data accessibility improves, but computational efficiency decreases

Engineering Contradiction:
Improvedata accessibilityVSAvoidcomputational efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent performs preliminary transformation of data into knowledge graph format with pre-computed embeddings and structured relationships, enabling efficient querying without requiring intensive computational resources at query time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates compressed vector representations (embeddings) of knowledge graph entities and relationships, which serve as efficient copies that enable fast similarity search and querying without requiring access to the full original data structures

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11593665B2Systems and methods driven by link-specific numeric information for predicting associations based on predicate types
Publication Date: 2023.02.28 ACCENTURE GLOBAL SOLUTIONS LTD
  • US11593665B2 patent drawing
  • US11593665B2 patent drawing
  • US11593665B2 patent drawing

AI summary

The present disclosure describes methods and systems to predict predicate metadata parameters in knowledge graphs via neural networks. The method includes receiving a knowledge graph based on a knowledge base including a graph-based dataset. The knowledge graph includes a predicate between two nodes and a set of predicate metadata. The method also includes determining a positive structural score, adjusting each positive structural score based on each corresponding significance parameter, generating a synthetic negative graph-based dataset, determining a negative structural score for each synthetic negative triple of the synthetic negative graph-based dataset, adjusting each negative structural score based on each corresponding significance parameter, determining a significance loss value based on the adjusted positive structural scores and the adjusted negative structural scores, and determining a likelihood score of a link between a third node and a fourth node in the knowledge graph based on the significance loss value.