Noise Propagation Data Anonymization via Vector Space Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization techniques face challenges in effectively anonymizing data, particularly with categorical attributes, as they require manual specification of distance and equivalence class definitions, leading to biased representations and reduced utility and reliability of the data.
Innovation Solution
The use of word embedding vector models and noise propagation modules to map data into vector space, adding noise to quasi-identifier values while preserving sensitive information, thereby anonymizing data by suppressing explicit identifiers and maintaining the utility of the data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual specification of distance and equivalence class definitions is used for data anonymization, then the anonymization process can be performed, but the representation becomes biased and utility is reduced
Solution Approach 1:
The system uses word embedding vector models to automatically determine distances and equivalence classes without manual intervention. The model self-organizes the data in vector space, automatically identifying quasi-identifier values and their relationships, thereby eliminating biased manual specifications while preserving data utility.
Solution Approach 2:
The patent replaces manual mechanical specification of distance metrics and equivalence classes with an automated word embedding vector model. This substitution uses machine learning to automatically map data to vector space and determine relationships, eliminating the need for manual parameter tuning and improving both reliability and utility.
2Reliability
If manual specification of distance and equivalence class definitions is required, then anonymization can be achieved, but the process complexity increases
Solution Approach 1:
The word embedding vector model automatically performs all steps of anonymization including identifying quasi-identifiers, determining distances, and creating equivalence classes without manual intervention. This self-service approach simplifies the process while maintaining effectiveness.
Solution Approach 2:
The word embedding vector model serves multiple functions simultaneously: it maps data to vector space, identifies quasi-identifier values, determines distances between records, and defines equivalence classes. This multi-functionality reduces process complexity while maintaining anonymization effectiveness.
3Object-affected harmful factors
If explicit identifiers are suppressed for anonymization, then individual identities are protected, but data reliability for research purposes is reduced
Solution Approach 1:
The patent applies different processing to different attributes: explicit identifiers are suppressed while sensitive attributes are preserved. The word embedding model identifies which attributes are quasi-identifiers and applies noise propagation selectively, maintaining local quality of different data fields to protect identities while preserving research utility.
Solution Approach 2:
The system changes parameters selectively: it adds noise to quasi-identifier values in vector space while leaving sensitive attribute values unchanged. This parameter change approach allows identity protection through suppression of explicit identifiers while maintaining the reliability of sensitive information for research purposes.
Data Source
AI summary
A first set of data associated with one or more data stores is received. A distance from a representation of a subset of the first set of data to at least a second representation of another set of data in vector space is identified. In response to the identifying of the distance, the first set of data is anonymized. The anonymizing includes adding noise to at least some of the first set of data.


