Knowledge Graph Entity Resolution Using Blocking and ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The creation of large-scale knowledge graphs is hindered by the time and resource-intensive process of manually identifying and connecting entities from various data sources, which often contain errors and inconsistencies, making it difficult to aggregate data effectively.

Innovation Solution

A method utilizing a distributed compute system and machine learning algorithms, including adaptive blocking algorithms and entity/edge resolution components, to efficiently process and resolve entities and connections in a knowledge graph, reducing the number of potential record groups and improving data accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual methods are used to create knowledge graphs, then data accuracy can be maintained, but the time and resources required increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidtime and resources
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the knowledge graph creation process into distinct components: entity identification, relationship extraction, and graph construction. This segmentation allows automated systems to handle specific tasks while maintaining accuracy through specialized processing for each component.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing layers including data cleaning modules, entity resolution components, and validation mechanisms that act as intermediaries between raw data and the final knowledge graph. These intermediaries ensure data accuracy while enabling automated processing at scale.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated systems are used to create knowledge graphs, then time and resources are reduced, but the ability to discern entities and connections is limited

Engineering Contradiction:
Improvecreation speedVSAvoidentity recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements multi-functional automated systems that can handle various data types, entity categories, and relationship kinds through a unified framework. This universality enables the system to process diverse data while maintaining discernment capabilities through adaptable processing rules and machine learning models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs parameter adjustments and optimization techniques to enhance automated entity recognition. By tuning parameters such as confidence thresholds, matching criteria, and processing depth, the system achieves both high productivity and accurate entity discernment across different data contexts.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If data from multiple sources is aggregated, then the scale and completeness of the knowledge graph increases, but errors and inconsistencies increase

Engineering Contradiction:
Improvedata volumeVSAvoiddata consistency
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements feedback mechanisms including validation rules, consistency checks, and error detection algorithms that continuously monitor aggregated data from multiple sources. This feedback loop identifies and corrects errors and inconsistencies, maintaining reliability while enabling large-scale data aggregation across diverse sources.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12566975B1Systems and methods for creating a knowledge graph
Publication Date: 2026.03.03 SCRYER INC
  • US12566975B1 patent drawing
  • US12566975B1 patent drawing
  • US12566975B1 patent drawing

AI summary

Systems, methods, and a computer readable storage medium for producing a knowledge graph are disclosed. The method includes resolving one or more vertices from one or more record sources on a graph where each vertex of the one or more vertices represents one or more records that contain information about an entity. The resolving of the one or more vertices includes reducing a possible number of records that are represented by each of the one or more vertices with a function and processing, by a distributed compute system, the reduced possible number of records with a machine learning algorithm. The method includes resolving one or more edges that comprise a connection between two vertices. The resolving of the one or more edges includes reducing a possible number of edges that are connected to each vertex with a function and processing, by the distributed compute system, the reduced possible number of edges with a machine learning algorithm.