Duplicate Record Reconciliation Using Probability Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In healthcare, the integration of data from disparate sources leads to duplication of records across systems, causing inefficiencies and increased storage needs, as different systems use varying standards and formats, making interoperability and record reconciliation challenging.
Innovation Solution
A computerized method and system that utilizes rules to identify and reconcile duplicate records by calculating a probability of duplication, weighting records, and ranking them to generate an updated set of records free from duplicates, using HTTP PATCH logic to consolidate information from duplicates into a single record.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is integrated from multiple disparate sources, then data completeness and continuity of care are improved, but record duplication and storage burden increase
Solution Approach 1:
The patent merges multiple records from disparate sources by identifying duplicates through probability calculations and consolidating them into single unified records. This combining process eliminates redundant data while preserving complete information from all sources, directly addressing the contradiction between data completeness and storage burden.
Solution Approach 2:
The system discards duplicate records after extracting and recovering their unique valuable information into consolidated records. By selectively removing redundant data while preserving essential information, the system reduces storage requirements without losing important data for continuity of care.
2Measurement precision
If records from disparate sources are reconciled, then data accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent transforms the reconciliation problem by changing parameters - converting duplicate detection from exact matching to probability-based duplication scoring. This parameter change allows the system to handle variable data formats and structures from disparate sources with consistent processing logic, improving data accuracy while managing complexity through mathematical transformation.
Solution Approach 2:
The system introduces an intermediary normalization layer that standardizes records from disparate sources before comparison. This intermediary processing stage handles format variations and structural differences, enabling accurate reconciliation without requiring complex source-specific processing for each record type.
3Measurement precision
If duplicate detection probability thresholds are lowered, then duplicate identification accuracy is improved, but false positive rate increases
Solution Approach 1:
The patent implements dynamic threshold adjustment where the duplication probability threshold adapts based on record characteristics and context. Rather than using a fixed threshold that causes false positives, the system dynamically modifies thresholds to maintain optimal balance between detection accuracy and false positive rates across different data scenarios.
Solution Approach 2:
The system incorporates feedback mechanisms where reconciliation results are analyzed and used to refine duplication probability calculations. This feedback loop allows the system to learn from previous reconciliations and adjust its probability assessments, improving duplicate identification accuracy while reducing false positives through iterative optimization.
Data Source
AI summary
Methods, systems, and computer-readable media are disclosed herein to provide rule-based reconciliation of records. Specifically, rules are utilized to reconcile one or more records and identify duplicates therein. Once duplicate records are identified, one or more ranking sets can be utilized to identify which of the duplicate records to write to the system.


