Data Deidentification System for Privacy Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Service providers face challenges in extracting meaningful insights from user data while complying with privacy and security standards that protect personal identifiable information (PII).
Innovation Solution
A data deidentification system that receives user data, analyzes it to extract insights that do not uniquely identify individuals, and processes the data to remove or obscure PII, allowing for the retention and use of deidentified data while complying with regulatory standards.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If service providers retain and analyze user data to extract insights, then productivity and business value are improved, but compliance with privacy standards becomes difficult to maintain
Solution Approach 1:
The patent segments user data into two distinct categories: PII (personal identifiable information) which is subject to privacy restrictions, and non-PII data which can be freely analyzed and retained. This segmentation allows the system to extract insights from non-PII data while automatically discarding or protecting PII data, thereby maintaining both productivity and compliance simultaneously
Solution Approach 2:
The system extracts and removes PII data from the user data set before analysis and retention. By taking out the problematic PII elements while retaining the valuable non-PII insights, the system resolves the contradiction between maintaining productivity through data analysis and ensuring compliance with privacy standards
2Reliability
If service providers obfuscate user data to comply with privacy standards, then privacy protection is improved, but the ability to extract meaningful insights deteriorates
Solution Approach 1:
Rather than obfuscating all user data, the system selectively extracts and removes only the PII portions that pose privacy risks, while preserving the complete non-PII data for analysis. This targeted extraction maintains insight extraction quality by leaving valuable non-PII information intact while still achieving privacy protection
Solution Approach 2:
Instead of starting with full data and obfuscating it (which loses information), the system inverts the approach by starting with full data, extracting the harmful PII elements, and retaining the useful non-PII portions. This inversion allows insight extraction to proceed on high-quality non-PII data while privacy protection is achieved through selective removal
3Reliability
If service providers are restricted in data retention period, then privacy compliance is improved, but the ability to retain data for future use deteriorates
Solution Approach 1:
The system applies different retention policies to different data segments: PII data is retained only for the regulatory minimum period required by privacy standards, while non-PII data is retained indefinitely for future analysis and business use. This segmented retention strategy simultaneously achieves privacy compliance and long-term data utility
Solution Approach 2:
By extracting and separating non-PII data from PII data, the system enables indefinite retention of the valuable non-PII insights without being constrained by privacy-based retention limits that apply only to PII data
4Productivity
If service providers export data across jurisdictions, then data utility and analysis capability are improved, but compliance with geographic data restrictions deteriorates
Solution Approach 1:
The system extracts and removes PII data that is subject to geographic export restrictions, allowing the remaining non-PII data to be freely exported and analyzed across different jurisdictions without compliance concerns
Data Source
AI summary
A data deidentification system that extracts insights from user data, and retains both the insights and user data in a form that complies with applicable data privacy and related standards. The system receives user data, which can include personal identifying information and other sensitive data governed by one or more standards, including standards specifying how the data can be used and how long it can be retained. From the data, the system extracts insights characterizing various aspects of the associated users. The system also selectively hashes portions of the data, obscuring the identity of associated users. Neither the insights nor the selectively hashed data identify individual users, and therefore they are not subject to the same standards and can be retained indefinitely. Later, after the standards-protected data has been discarded, the system can provide insight information in response to a request.


