Probabilistic Knowledge Graph Completion for Sparse Entity Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional knowledge graph systems struggle to complete knowledge graphs for entities like humans or new materials due to the scarcity of accurate data, leading to incomplete or inaccurate representations that reduce their utility in applications such as health analytics and material selection.

Innovation Solution

A controller is used to crawl open-world data sources, identify relevant data, and assign probabilistic confidence scores to nodes and edges, generating a probabilistic knowledge graph that includes both factual and probabilistic entities and relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional knowledge graph systems are used to store and evaluate data, then the data structure allows for logical inference and access to implicit information, but the systems struggle to complete knowledge graphs for entities like humans or new materials due to scarcity of accurate data, leading to incomplete or inaccurate representations

Engineering Contradiction:
Improveaccuracy of knowledge graph representationVSAvoidcompleteness of knowledge graph
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent transforms the binary existence/non-existence of knowledge graph entities into a probabilistic framework where each entity and relationship is associated with a confidence score. This parameter change allows the system to represent uncertain or incomplete information quantitatively, enabling more nuanced handling of scarce data while maintaining logical inference capabilities.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The knowledge graph system dynamically adjusts confidence scores based on multiple data sources and their reliability. As new data becomes available or existing data is re-evaluated, the confidence scores are updated, allowing the system to adapt to changing information quality and quantity without requiring complete data availability.

Inventive Principle:
Principle #15Dynamics

2Loss of information

If probabilistic confidence scores are assigned to nodes and edges in the knowledge graph, then the completeness and utility of the knowledge graph is enhanced, but the complexity of data processing and score assignment increases

Engineering Contradiction:
Improvecompleteness of knowledge graphVSAvoidcomplexity of probabilistic data processing
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the confidence scoring process into distinct components: data source reliability assessment, entity existence probability, and relationship probability. Each component can be processed and evaluated independently, then combined to form overall confidence scores. This segmentation reduces the computational complexity of the overall probabilistic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary confidence scores that act as mediators between raw data and final knowledge graph representations. These intermediate probabilistic values allow for gradual refinement of information quality without requiring direct, complex processing of all underlying data sources simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12547905B2Probabilistic entity-centric knowledge graph completion
Publication Date: 2026.02.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12547905B2 patent drawing
  • US12547905B2 patent drawing
  • US12547905B2 patent drawing

AI summary

A first set of an entity is received, where the first set of data includes distinct characteristics of the entity. A second set of data on one or more domains of the entity is received. Using the first and second set of data, a probabilistic knowledge graph for the entity is generated that includes an entity node, a first plurality of nodes, and a second plurality of nodes. The first plurality of nodes are connected to the entity node and represent each of the distinct characteristics. The second plurality of nodes are connected via probabilistic edges, where each of these probabilistic edges has an associated confidence score. This confidence score is determined using the second set of data.