Graph Database Entity Disambiguation via Blocking and Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for disambiguating entities in graph databases are not scalable and fail to effectively handle variations in user-entered data, leading to redundant entities and obscured relationships, which hampers the ability to provide rich features and services based on graph databases.
Innovation Solution
A multi-stage disambiguation pipeline that includes data blocking, matching, merging, and ID generation to standardize and consolidate entities, reducing redundant entries and improving data consistency within the graph database.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If users are given wide latitude to store data in free-form formats, then ease of operation for users is improved, but data consistency and entity disambiguation deteriorate
Solution Approach 1:
The system performs preliminary actions by implementing data blocking and matching algorithms before entities are fully processed in the graph database. By pre-processing entity data to identify and consolidate duplicates based on blocking parameters (such as name, location, or other attributes), the system prepares data in advance to prevent consistency issues downstream while preserving user flexibility during entry.
Solution Approach 2:
The patent introduces an intermediary disambiguation layer between user data entry and the final graph database structure. This intermediary processing stage applies matching algorithms and blocking parameters to reconcile variations in user-entered data, acting as a mediator that maintains both user freedom in data entry and data consistency in the final database.
2Measurement precision
If manual review methods are used to disambiguate entities, then measurement precision of entity identity is improved, but productivity and scalability deteriorate
Solution Approach 1:
The patent segments the disambiguation process into distinct stages: blocking (grouping potential duplicates by key attributes), matching (comparing entities within blocks using similarity algorithms), and merging (consolidating matched entities). This segmentation allows automated processing at each stage while maintaining high accuracy through progressive filtering, eliminating the need for complete manual review of all entities.
Solution Approach 2:
The system applies partial automation by using blocking parameters to pre-filter and group entities before applying more computationally intensive matching algorithms. This partial action approach processes only relevant subsets of data with high-precision methods, achieving near-manual-review accuracy at automated processing speeds by avoiding excessive computation on unrelated entities.
3Device complexity
If simple matching algorithms are used for entity disambiguation, then device complexity is reduced, but measurement precision of entity matching deteriorates
Solution Approach 1:
The matching system is segmented into multiple components: blocking based on simple attribute comparisons, followed by more sophisticated matching algorithms applied only to blocked groups. This segmentation allows the use of complex matching logic where needed while keeping the overall system manageable through modular design, improving accuracy without overwhelming complexity.
Solution Approach 2:
The patent applies different matching strategies locally to different data contexts. Blocking parameters use simple exact or fuzzy matching, while the subsequent matching stage applies more sophisticated algorithms tailored to specific entity types and attributes. This local quality approach optimizes matching precision for each data subset without requiring uniformly complex processing across all entities.
4Quantity of substance
If redundant entities are not consolidated, then data volume is preserved, but processing speed and resource efficiency deteriorate
Solution Approach 1:
The patent implements entity merging by consolidating duplicate or redundant entities into single canonical representations based on blocking and matching results. This merging process reduces the total number of entities in the graph database while preserving all unique information, directly improving processing speed and resource efficiency by eliminating redundant data without significant data loss.
Solution Approach 2:
The system discards redundant entity duplicates after identification through blocking and matching, then recovers their unique information by merging it into canonical entity representations. This process eliminates harmful redundancy that slows processing while preserving valuable data content, achieving both data volume optimization and processing efficiency improvement.
Data Source
AI summary
Certain aspects provide techniques for disambiguating graph data. In one example, a method includes receiving entity data from a data source in a first format; converting the entity data in the first format to a second format, wherein the second format is a standardized input format for a disambiguation pipeline; determining a blocked data set from the entity data in the second format based on a blocking parameter, wherein: the blocked data set comprises data regarding a first plurality of entities, and the first plurality of entities is a subset of a second plurality of entities represented in the entity data from the data source; matching at least two entities in the first plurality of entities in the blocked data set; merging the at least two entities into a single entity; generating a unique ID for the single entity; and importing the single entity into a graph database.


