Secure Data Matching Across Independent Systems With Reference Records
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data matching algorithms in heterogeneous systems often result in false negatives or false positives, leading to incorrect data transactions or duplicate records due to variations in attribute descriptions across different systems.
Innovation Solution
A system and method for securely linking and matching data attributes across independent data systems by using private and reference databases, where private databases store unique attributes and reference databases contain comprehensive attribute information, enabling deterministic matching and secure data record association.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional matching algorithms are used to compare attributes between data systems, then the matching process is simple and fast, but false negatives and false positives occur resulting in incorrect rejection or acceptance of data transactions
Solution Approach 1:
The patent introduces an intermediary matching system that receives attributes from both the first and second data systems, performs comprehensive comparison using multiple algorithms including deterministic and probabilistic matching, and returns match results. This intermediary structure isolates the complexity from the original data systems while improving identification accuracy through sophisticated attribute comparison.
2Adaptability or versatility
If attribute variations are allowed across different enterprises' data systems, then adaptability and versatility are improved, but matching accuracy deteriorates due to inconsistencies in attribute formats and values
Solution Approach 1:
The patent transforms attributes from different data systems into a standardized format suitable for comparison. The matching system normalizes attribute values, handles different data types, and converts varied attribute representations into a common structure, enabling accurate matching while maintaining adaptability to diverse input formats from different enterprises.
Solution Approach 2:
The patent replaces simple exact-match mechanical comparison with sophisticated matching algorithms that include deterministic matching for exact matches and probabilistic matching for approximate matches. This substitution enables the system to handle attribute variations while maintaining high matching precision through intelligent algorithmic approaches.
3Speed
If deterministic matching is used for exact attribute matches, then matching speed is fast and results are certain, but false negatives occur when attributes have minor variations
Solution Approach 1:
The patent segments the matching process into two distinct phases: deterministic matching for rapid exact-match detection and probabilistic matching for comprehensive similarity assessment. This segmentation allows the system to first quickly identify obvious matches using deterministic algorithms, then apply more sophisticated probabilistic algorithms to detect matches with minor attribute variations, thereby improving both speed and reliability.
Solution Approach 2:
The patent applies partial matching where not all attributes need to match exactly for a successful match. The probabilistic matching algorithm allows for partial attribute matches and weighs different attributes differently, enabling the system to identify subjects even when some attributes have minor variations or are missing, thus reducing false negatives while maintaining reasonable processing speed.
4Reliability
If probabilistic matching is used to handle attribute variations, then false negatives are reduced, but processing time and computational resources increase
Solution Approach 1:
The patent segments the matching process into deterministic matching executed first for rapid exact-match detection, followed by probabilistic matching only for cases where deterministic matching fails or attributes show partial similarity. This segmentation reduces the overall processing time by applying computationally intensive probabilistic algorithms only when necessary, rather than to all matching cases.
Solution Approach 2:
The patent implements a tiered matching approach where probabilistic matching is applied partially - only to records that did not match through deterministic methods or show partial attribute similarity. This selective application of probabilistic matching maintains high detection accuracy for borderline cases while minimizing the computational overhead and processing time associated with applying probabilistic algorithms to all records.
Data Source
AI summary
A methods, systems, and devices for secure linking and matching of data elements across independent data systems of a first and a second enterprise are disclosed. The system includes a first private data system having a first private database that may include all attributes known to the first enterprise A reference data system having a reference database and is also part of the first enterprise. The reference database may contain a plurality of identity data records that only contain identity attributes that are used for resolving and verifying identities of unknown data records that are external to the reference data system. A first set of query attributes may be used determine whether a subject represented by the first set of query attributes associates uniquely with a subject of any private data record stored in the first private database. A similar comparison may be made in the reference database.


