Graph Database Entity Disambiguation via Blocking and Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for disambiguating entities in graph databases are not scalable and fail to effectively handle variations in user-entered data, leading to redundant entities and obscured relationships, which hampers the ability to provide rich features and services based on graph databases.

Innovation Solution

A multi-stage disambiguation pipeline that includes data blocking, matching, merging, and ID generation to standardize and consolidate entities, reducing redundant entries and improving data consistency within the graph database.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users are given wide latitude to store data in free-form formats, then ease of operation for users is improved, but data consistency and entity disambiguation deteriorate

Engineering Contradiction:
Improveuser data entry flexibilityVSAvoiddata consistency
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs preliminary actions by implementing data blocking and matching algorithms before entities are fully processed in the graph database. By pre-processing entity data to identify and consolidate duplicates based on blocking parameters (such as name, location, or other attributes), the system prepares data in advance to prevent consistency issues downstream while preserving user flexibility during entry.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary disambiguation layer between user data entry and the final graph database structure. This intermediary processing stage applies matching algorithms and blocking parameters to reconcile variations in user-entered data, acting as a mediator that maintains both user freedom in data entry and data consistency in the final database.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual review methods are used to disambiguate entities, then measurement precision of entity identity is improved, but productivity and scalability deteriorate

Engineering Contradiction:
Improveentity identification accuracyVSAvoiddisambiguation processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the disambiguation process into distinct stages: blocking (grouping potential duplicates by key attributes), matching (comparing entities within blocks using similarity algorithms), and merging (consolidating matched entities). This segmentation allows automated processing at each stage while maintaining high accuracy through progressive filtering, eliminating the need for complete manual review of all entities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial automation by using blocking parameters to pre-filter and group entities before applying more computationally intensive matching algorithms. This partial action approach processes only relevant subsets of data with high-precision methods, achieving near-manual-review accuracy at automated processing speeds by avoiding excessive computation on unrelated entities.

Inventive Principle:
Principle #16Partial or excessive action

3Device complexity

If simple matching algorithms are used for entity disambiguation, then device complexity is reduced, but measurement precision of entity matching deteriorates

Engineering Contradiction:
Improvedisambiguation system complexityVSAvoidentity matching accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The matching system is segmented into multiple components: blocking based on simple attribute comparisons, followed by more sophisticated matching algorithms applied only to blocked groups. This segmentation allows the use of complex matching logic where needed while keeping the overall system manageable through modular design, improving accuracy without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different matching strategies locally to different data contexts. Blocking parameters use simple exact or fuzzy matching, while the subsequent matching stage applies more sophisticated algorithms tailored to specific entity types and attributes. This local quality approach optimizes matching precision for each data subset without requiring uniformly complex processing across all entities.

Inventive Principle:
Principle #3Local quality

4Quantity of substance

If redundant entities are not consolidated, then data volume is preserved, but processing speed and resource efficiency deteriorate

Engineering Contradiction:
Improvedata volumeVSAvoidgraph database processing speed
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent implements entity merging by consolidating duplicate or redundant entities into single canonical representations based on blocking and matching results. This merging process reduces the total number of entities in the graph database while preserving all unique information, directly improving processing speed and resource efficiency by eliminating redundant data without significant data loss.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system discards redundant entity duplicates after identification through blocking and matching, then recovers their unique information by merging it into canonical entity representations. This process eliminates harmful redundancy that slows processing while preserving valuable data content, achieving both data volume optimization and processing efficiency improvement.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS11379526B2Disambiguation of massive graph databases
Publication Date: 2022.07.05 INTUIT INC
  • US11379526B2 patent drawing
  • US11379526B2 patent drawing
  • US11379526B2 patent drawing

AI summary

Certain aspects provide techniques for disambiguating graph data. In one example, a method includes receiving entity data from a data source in a first format; converting the entity data in the first format to a second format, wherein the second format is a standardized input format for a disambiguation pipeline; determining a blocked data set from the entity data in the second format based on a blocking parameter, wherein: the blocked data set comprises data regarding a first plurality of entities, and the first plurality of entities is a subset of a second plurality of entities represented in the entity data from the data source; matching at least two entities in the first plurality of entities in the blocked data set; merging the at least two entities into a single entity; generating a unique ID for the single entity; and importing the single entity into a graph database.