Pseudonymous Data Marking for Privacy and Manageability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in managing pseudonymous or anonymous user data while preserving privacy and fulfilling user data management requests, as required by data protection regulations like GDPR.
Innovation Solution
A method involving converting numerical features of data points into categorical forms, determining and concatenating a data contributor identifier, and cryptographically hashing the combination to generate a mark, which is then associated with the data point to create marked pseudonymous-anonymous data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If pseudonymous or anonymous data is used to protect user privacy, then user privacy is improved, but the ability to locate and manage user-specific data deteriorates
Solution Approach 1:
The patent introduces a mark as an intermediary element that bridges user privacy protection and data manageability. The mark is generated by hashing the concatenation of the user identifier and categorical feature values, serving as a mediator that allows data platforms to locate and manage user-specific data without directly storing or accessing the user identifier itself, thus maintaining privacy while enabling operational capability
Solution Approach 2:
The patent transforms numerical feature values into categorical forms (e.g., age groups, income brackets) before hashing with the user identifier. This parameter transformation ensures that even if the user identifier is compromised, the categorical features cannot be used to re-identify the user, thereby strengthening privacy protection while still allowing data retrieval through the generated mark
2Reliability
If user identifiers are removed from data to ensure anonymity, then data anonymity is improved, but the ability to fulfill user data management requests deteriorates
Solution Approach 1:
The patent performs preliminary action by generating the mark before data storage. The mark is created by hashing the concatenation of the user identifier and categorical feature values at the time of data collection. This preliminary generation of the mark allows the data platform to later retrieve user-specific data using the mark without needing to store or access the original user identifier, thus maintaining anonymity while preserving locateability
Solution Approach 2:
The patent creates a copy of the user identifier combined with categorical features and transforms it into a hash mark. This copy approach allows the data platform to work with the hashed mark instead of the original user identifier, enabling data retrieval and management operations without exposing or storing the actual user identifier, thereby maintaining anonymity while preserving the ability to fulfill data management requests
Data Source
AI summary
An approach is provided for managing pseudonymous or anonymous user data and relevant data management requests. The approach involves, for example, converting a numerical feature of a data point into a categorical form. The categorical form represents a value range into which a numerical value of the numerical feature falls. The approach also involves determining an identifier of a data contributor associated with the data point. The approach further involves concatenating the identifier with the categorical form. The approach further involves cryptographically hashing the identifier concatenated with the categorical form to generate a mark. The approach further involves associating the mark with the data point to generate marked pseudonymous-anonymous data. The approach further involves transmitting the pseudonymous-anonymous data to a data platform.


