De-identified Transaction Data Linkage via Probabilistic Engine
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current transaction processing systems face challenges in analyzing large volumes of transaction data while ensuring that personally identifiable information (PII) is not exposed or accessible, particularly when linking data from different sources such as payment networks and merchant sales ledgers.
Innovation Solution
The system employs a probabilistic engine that processes de-identified data from both payment networks and merchant transaction sources, using de-identified unique identifiers and pattern analysis modules to create anonymized data extracts, ensuring that PII is removed and maintaining anonymity throughout the analysis process, thereby allowing for the generation of reports and analyses without revealing sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If transaction data from different sources is linked and analyzed, then the ability to construct analyses and reports is improved, but the risk of exposing personally identifiable information increases
Solution Approach 1:
The patent extracts and removes personally identifiable information from transaction data before analysis. De-identification modules systematically identify and extract PII elements such as names, addresses, and social security numbers, separating them from the transaction data to enable analysis without exposing sensitive information.
Solution Approach 2:
The patent introduces de-identification technology as an intermediary layer between data collection and analysis. This intermediary process transforms identifiable transaction data into anonymized data sets that retain analytical value while eliminating PII exposure risks, allowing linkage between different data sources without direct access to sensitive information.
2Reliability
If de-identified data is used for analysis, then data security is improved, but the precision of matching and analysis may deteriorate
Solution Approach 1:
The patent changes the parameters of data representation by transforming identifiable information into anonymized forms while preserving the essential characteristics needed for matching. De-identification techniques modify data parameters such as removing direct identifiers but retaining transaction patterns, amounts, and temporal information that enable precise analysis without compromising security.
Solution Approach 2:
The patent performs de-identification as a preliminary action before data linkage and analysis. By pre-processing transaction data to remove PII while preserving analytical integrity, the system establishes a foundation for secure yet precise matching between different data sources, ensuring both security and analytical accuracy from the outset.
3Object-affected harmful factors
If strict limitations on data access are maintained, then PII protection is improved, but the ability to share and analyze data across multiple sources is reduced
Solution Approach 1:
The patent creates anonymized copies of transaction data that can be freely shared and analyzed without exposing original PII. De-identification modules generate replicated data sets containing all necessary transaction information except personally identifiable elements, enabling widespread data sharing and cross-source analysis while maintaining strict protection of the original sensitive information.
Solution Approach 2:
The patent segments transaction data into identifiable and non-identifiable components. By separating PII from transaction details through systematic data segmentation, the system allows the non-identifiable portions to be shared and analyzed across multiple sources while the identifiable portions remain protected, thus enhancing both PII protection and data sharing versatility.
Data Source
AI summary
Systems, methods, means, computer program code and computerized processes include receiving a first set of de-identified transaction data from a first transaction data source, receiving a second set of de-identified transaction data from a second transaction data source, filtering the first and second sets of de-identified transaction data to identify transactions associated with at least a first entity and to create first and second filtered data sets, removing data associated with an identifier field for each of the transactions in the first filtered data set to created a de-identified first data set, removing data associated with an identifier field for each of the transactions in the second filtered data set to create a de-identified second data set, and processing the first and second de-identified data sets using a probabilistic engine to establish a linkage between data in each data set.


