Entity Resolution via Bayesian Joint Probability Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional entity resolution techniques are cumbersome, inaccurate, and inflexible, particularly when dealing with incomplete or inconsistent data, making it difficult to efficiently compile an accurate description of an entity across multiple data sources.
Innovation Solution
The use of statistical inference techniques, such as Bayesian inference, to estimate the joint probability of descriptor values in data sets, allowing for more accurate and automated entity resolution, with the ability to adjust error rates based on the specific application's tolerance for false positives and negatives.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional entity resolution techniques are used, then the process can be performed with simple methods, but the accuracy and automation level are insufficient
Solution Approach 1:
The patent replaces manual mechanical entity resolution processes with automated statistical inference systems. Bayesian inference algorithms automatically compute probabilities of entity matches, substituting human judgment with mathematical models that process descriptor values and compute match probabilities systematically, thereby improving accuracy while maintaining manageable complexity through algorithmic standardization.
Solution Approach 2:
The patent transforms entity resolution from a qualitative manual process to a quantitative automated process by introducing probability parameters. The system computes numerical probabilities for entity matches using Bayesian inference, allowing precise control over resolution accuracy through parameter adjustment while automating the resolution process itself.
2Productivity
If manual entity resolution is performed, then flexibility in handling edge cases is maintained, but productivity and automation are reduced
Solution Approach 1:
The patent implements self-service entity resolution through automated Bayesian inference systems that independently process data sets and determine entity matches without requiring manual intervention. The system serves itself by automatically computing probabilities, making decisions, and resolving entities through algorithmic processes, thereby maximizing productivity and automation extent simultaneously.
3Adaptability or versatility
If fixed error rate thresholds are used, then the system is simple to operate, but adaptability to different application requirements is limited
Solution Approach 1:
The patent implements dynamic error rate thresholds that can be adjusted based on specific application requirements. The system allows operators to modify probability thresholds and inference parameters dynamically, enabling adaptation to different tolerance levels for false positives and negatives while maintaining ease of operation through configurable interfaces that simplify the adjustment process.
4Reliability
If heuristic rules are used for entity resolution, then the system is easy to implement, but reliability and accuracy deteriorate with incomplete or inconsistent data
Solution Approach 1:
The patent replaces unreliable heuristic rules with robust Bayesian inference mechanisms that systematically process descriptor values and compute match probabilities. The statistical framework provides reliable handling of incomplete and inconsistent data by using probability theory to account for uncertainty, thereby improving reliability while managing complexity through well-established mathematical methods.
Data Source
AI summary
Entity resolution techniques and systems are described. An entity resolution method may include estimating a joint probability of occurrence of a plurality of values of a respective plurality of descriptors of an entity. The plurality of descriptor values may be included in a first data set. The method may further include determining that the joint probability of occurrence of the plurality of descriptor values is less than a threshold probability, identifying a second data set including the same plurality of values of the same respective plurality of descriptors, and determining, based at least in part on the joint probability of occurrence of the plurality of descriptor values being less than the threshold probability and on the first and second data sets including the same plurality of descriptor values, that the first and second data sets describe the same entity.


