Entity Matching with Iterative Demographic Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing entity matching systems face challenges in accurately linking records of individuals who have changed their names or demographics, due to reliance on unreliable identifiers like social security numbers and data entry errors, which are exacerbated by security and privacy concerns.
Innovation Solution
A method and system for entity matching that assigns comparison weights to demographic fields, including agreement, disagreement, and null weights, to determine the similarity between records, using an iterative process to optimize weight determination and improve matching accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If social security numbers or governmental identification numbers are used to match entities, then matching reliability is improved, but security and privacy concerns increase and availability of identifiers decreases
Solution Approach 1:
The patent extracts and removes reliance on governmental identification numbers (social security numbers) from the entity matching process. Instead of using these sensitive identifiers, the system uses demographic fields (name, address, date of birth) combined with iterative optimization to achieve matching without compromising security or privacy.
Solution Approach 2:
The patent changes the parameters used for entity matching from governmental identifiers to demographic fields. By transforming the matching approach to use multiple demographic parameters with adjustable weights, the system achieves reliable matching while eliminating the need for sensitive identification numbers.
2Measurement precision
If multiple demographic fields are compared to improve matching accuracy, then measurement precision is improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent introduces dynamic weight adjustment for different demographic fields based on their relevance and reliability in specific matching contexts. The iterative optimization process dynamically determines the importance of each field, allowing the system to focus computational resources on the most informative fields while maintaining high accuracy.
Solution Approach 2:
The patent optimizes the parameters (weights) of demographic fields through iterative processes. By adjusting the weights of different fields based on empirical data and matching outcomes, the system achieves high precision without requiring all fields to be processed equally, thus reducing overall complexity.
3Reliability
If data entry errors are present in identifiers, then reliability of identifiers decreases, but correcting these errors increases processing complexity
Solution Approach 1:
The patent uses demographic fields as intermediary elements to verify and correct identifier data. By cross-referencing multiple demographic fields (name, address, date of birth) with each other and with the identifiers, the system can detect and correct data entry errors without requiring complex validation processes.
Solution Approach 2:
The iterative optimization process provides feedback about the consistency and reliability of different data fields. This feedback mechanism allows the system to automatically adjust its understanding of which fields are most reliable, effectively correcting errors through multiple rounds of comparison and validation.
Data Source
AI summary
Entity matching is provided. A request to determine a second entity matching with a first entity may be received. Demographic fields to be compared to determine the match may be determined. Comparison weights including an agreement weight, a disagreement weight, and a null weight may be assigned to each of the plurality of demographic fields. The received request may be parsed to determine demographic fields data related to the second entity. The demographic fields data may be compared with indexed demographic data which may include a plurality of records. The demographic fields data may be compared for the determined plurality of demographic fields. A comparison weight for each of the plurality of demographic fields may be determined based on the comparison. The first entity matching with the second entity may be determined from the indexed demographic data based on determined comparison weights for the plurality of demographic fields.


