Machine Learning Entity Resolution Parameter Tuning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data management systems for entity resolution rely heavily on expert users to configure complex algorithm parameters and require extensive statistical analysis, making the process manual, iterative, and resource-intensive, especially when dealing with multi-dimensional comparisons and false positives.

Innovation Solution

The implementation of a machine learning-based approach that generates comparison vectors from data record pairs, trains models using these vectors and user feedback, and autonomously refines configuration parameters to improve matching accuracy and reduce manual effort.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If expert users configure complex algorithm parameters manually, then matching accuracy can be improved, but the process becomes resource-intensive and time-consuming

Engineering Contradiction:
Improvematching accuracyVSAvoidconfiguration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-configuration by automatically learning optimal algorithm parameters through machine learning models trained on comparison vectors, eliminating the need for manual expert configuration while maintaining high matching accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual expert configuration (mechanical process) with automated machine learning-based parameter tuning, substituting human expertise with computational algorithms that learn optimal parameters from data

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If statistical analysis is performed extensively to tune parameters, then reliability of matching is improved, but device complexity increases

Engineering Contradiction:
Improvematching reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex statistical analysis procedures with machine learning models that automatically learn from comparison vectors, maintaining reliability while simplifying the overall system architecture through automated pattern recognition

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If manual iterative tuning is performed, then accuracy is improved, but productivity decreases

Engineering Contradiction:
Improvematching accuracyVSAvoidsystem throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary learning by training machine learning models on comparison vectors before actual matching operations, so that parameter optimization is completed in advance rather than through iterative tuning during operation, thereby improving throughput

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system autonomously optimizes its own parameters through self-learning from training data, eliminating the need for manual iterative tuning and enabling parallel processing that increases overall productivity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11720807B2Machine learning to tune probabilistic matching in entity resolution systems
Publication Date: 2023.08.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11720807B2 patent drawing
  • US11720807B2 patent drawing
  • US11720807B2 patent drawing

AI summary

Techniques for data evaluation are provided. A plurality of data records is received, and a first comparison vector is generated by comparing a first and a second data record of the plurality of data records, where the first comparison vector indicates differences between the first and second data records. A machine learning model is trained based at least in part on the first comparison vector. The plurality of data records is evaluated using the machine learning model, and at least two of the plurality of data records are linked based on the evaluation.