Longitudinal Patient Data Linking via Hierarchical Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in assembling longitudinal databases from multi-source patient health care transaction records while ensuring patient privacy is maintained, as existing methods struggle to effectively de-identify and link records across different data sources without compromising personal information.
Innovation Solution
The solution involves a hierarchical matching process and algorithm that compares select sets of data attributes from incoming records with existing reference data records, using a combination of alphanumeric and demographic attributes to assign longitudinal linking tags, ensuring records can be linked without revealing personal information, and employing a two-step encryption procedure with multiple encryption keys to maintain privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all patient identification information is stripped to completely anonymize records, then patient privacy is maximally protected, but the ability to link records individual patient-by-patient into a longitudinal database is lost
Solution Approach 1:
The patent segments patient identification information into two distinct components: encrypted identifiers that preserve linkability and demographic attributes that can be used for matching. This segmentation allows the system to maintain both privacy protection and record linkability simultaneously by handling these components separately through different processing mechanisms.
Solution Approach 2:
The patent introduces encrypted patient identifiers as an intermediary element between complete anonymity and full identifiability. These encrypted identifiers serve as mediators that allow records to be linked without revealing actual patient identities, thus resolving the contradiction by providing a middle ground that satisfies both privacy and linkability requirements.
2Reliability
If patient records are encrypted to preserve privacy, then personal information is protected, but the complexity of the data processing system increases
Solution Approach 1:
The patent applies preliminary action by pre-encrypting patient identifiers at the data source before transmission to the longitudinal database facility. This advance encryption step simplifies subsequent processing by eliminating the need for complex real-time encryption operations at the database level, thereby reducing overall system complexity while maintaining strong privacy protection.
3Loss of information
If multiple data sources are integrated into a longitudinal database, then the information content and utility of the database is enhanced, but the difficulty of matching and linking de-identified records from different sources increases
Solution Approach 1:
The patent implements universality by creating a standardized encrypted identifier format that works across multiple different data sources and suppliers. This universal identifier system enables consistent matching and linking of records from diverse sources without requiring source-specific matching logic, thereby reducing matching difficulty while enhancing data integration value.
Data Source
AI summary
A longitudinal database of de-identified patient healthcare transaction data records linked by longitudinal linking tags (IDs) is provided. A new healthcare transaction data record, which may include alphanumeric identification code attributes, third party attributes and/or demographic attributes, is assigned an linking ID associated with a previous healthcare transaction data record based upon successful comparison of either a designated set of identification code attributes or a designated set of demographic attributes. The longitudinal data base is assembled by a matching process in which a new data record is compared level by level with previous healthcare transaction data records through a hierarchy of a first series of matching levels each defined by a designated set of alphanumeric identification code attributes and a second series of matching levels each defined by a designated set of attributes including demographic attributes and then assigned the ID associated with a successfully matched reference data record.


