Probabilistic Knowledge Graph Completion for Sparse Entity Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional knowledge graph systems struggle to complete knowledge graphs for entities like humans or new materials due to the scarcity of accurate data, leading to incomplete or inaccurate representations that reduce their utility in applications such as health analytics and material selection.
Innovation Solution
A controller is used to crawl open-world data sources, identify relevant data, and assign probabilistic confidence scores to nodes and edges, generating a probabilistic knowledge graph that includes both factual and probabilistic entities and relationships.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional knowledge graph systems are used to store and evaluate data, then the data structure allows for logical inference and access to implicit information, but the systems struggle to complete knowledge graphs for entities like humans or new materials due to scarcity of accurate data, leading to incomplete or inaccurate representations
Solution Approach 1:
The patent transforms the binary existence/non-existence of knowledge graph entities into a probabilistic framework where each entity and relationship is associated with a confidence score. This parameter change allows the system to represent uncertain or incomplete information quantitatively, enabling more nuanced handling of scarce data while maintaining logical inference capabilities.
Solution Approach 2:
The knowledge graph system dynamically adjusts confidence scores based on multiple data sources and their reliability. As new data becomes available or existing data is re-evaluated, the confidence scores are updated, allowing the system to adapt to changing information quality and quantity without requiring complete data availability.
2Loss of information
If probabilistic confidence scores are assigned to nodes and edges in the knowledge graph, then the completeness and utility of the knowledge graph is enhanced, but the complexity of data processing and score assignment increases
Solution Approach 1:
The patent segments the confidence scoring process into distinct components: data source reliability assessment, entity existence probability, and relationship probability. Each component can be processed and evaluated independently, then combined to form overall confidence scores. This segmentation reduces the computational complexity of the overall probabilistic processing.
Solution Approach 2:
The system introduces intermediary confidence scores that act as mediators between raw data and final knowledge graph representations. These intermediate probabilistic values allow for gradual refinement of information quality without requiring direct, complex processing of all underlying data sources simultaneously.
Data Source
AI summary
A first set of an entity is received, where the first set of data includes distinct characteristics of the entity. A second set of data on one or more domains of the entity is received. Using the first and second set of data, a probabilistic knowledge graph for the entity is generated that includes an entity node, a first plurality of nodes, and a second plurality of nodes. The first plurality of nodes are connected to the entity node and represent each of the distinct characteristics. The second plurality of nodes are connected via probabilistic edges, where each of these probabilistic edges has an associated confidence score. This confidence score is determined using the second set of data.


