De-Identified Record Merging for Privacy-Safe Media Reach Estimates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning systems face challenges in protecting sensitive data during training and validation, particularly in digital advertising technology, where healthcare data is involved, leading to issues with data privacy and accuracy in targeting relevant audiences.
Innovation Solution
A system is implemented that stores training data in a protected environment, using de-identified records and encrypted tokens, and applies criteria to ensure that only valid and accurate machine learning systems are generated and deployed, ensuring data privacy and effectiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If machine learning systems use complex algorithms and continuous learning to generate patterns from training data, then the system's ability to learn and adapt is improved, but determining why a particular result was produced becomes difficult, leading to lack of accountability
Solution Approach 1:
The patent introduces an intermediary validation system that acts as a mediator between the complex machine learning model and the user. This validation system provides explanations and justifications for model predictions without requiring direct access to or understanding of the complex internal algorithms, thereby maintaining both model complexity and interpretability.
2Measurement precision
If machine learning systems are trained on protected sensitive data such as healthcare information, then the accuracy and relevance of targeting is improved, but data privacy and security are compromised during training and validation
Solution Approach 1:
The patent extracts only the necessary patterns and relationships from protected sensitive data during training, while removing or masking the actual sensitive information. The trained model then operates on non-sensitive inputs, producing accurate predictions without requiring access to or storage of the original protected data, thus maintaining both accuracy and privacy.
Solution Approach 2:
The patent creates copies or representations of the training data that preserve the statistical relationships and patterns needed for accurate modeling, but replace actual sensitive identifiers with synthetic or masked values. This allows the model to learn from data-like structures without exposure to real protected information.
3Measurement precision
If machine learning systems memorize vast amounts of personal data to provide one-to-one recognition, then the recognition accuracy is improved, but the system provides protected information to viewers, failing to maintain data protection
Solution Approach 1:
The patent applies local quality by providing different levels of detail and protection for different aspects of the system. The model maintains high recognition accuracy for its functional purpose while simultaneously applying strong privacy protection and masking for data access and storage. Different parts of the system have different qualities: the model layer prioritizes accuracy, while the data layer prioritizes protection.
Data Source
AI summary
In one embodiment, a computer implemented method comprises receiving and storing in relational database tables in a secure data processing environment comprising one or more first virtual machine instances coupled to one or more first data stores, master data comprising records having first de-identified token values associated with health data and second data comprising records having second de-identified token values associated with historical media delivery data, wherein each of the master data and the second data comprise references to healthcare providers (HCPs) in addition to de-identified token values based on patients; in the secure data processing environment, executing one or more database table join operations based on the references to the HCPs to merge the master data and the second data to produce a joined table having records comprising third de-identified token values associated with the health data and the second data; receiving, using one or more virtual computing instances of a service provider environment, one or more filter specifications that define a target audience and a forecast request, and in real time in response to the forecast request: based on the one or more filter specifications, executing one or more queries to the joined table in the secure data processing environment; receiving, in the service provider environment, de-identified aggregated data that the secure data processing environment has generated based upon the one or more queries to the joined table; based on the de-identified aggregated data and second data, generating an estimate of media delivery reach; presenting the estimate of the media delivery reach to a user computer that is communicatively coupled to the service provider environment.


