Utility-Aware Anonymization of Location Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization solutions for location data either overprotect by concealing entire user trajectories or discarding temporal information, leading to significant data distortion and requiring extensive parameterization, which is not utility-aware and fails to provide privacy guarantees.
Innovation Solution
A utility-aware anonymization method that identifies privacy vulnerabilities in sequential and location datasets, generating privacy constraints to anonymize the data while adhering to utility constraints, ensuring the preservation of temporal information and providing privacy guarantees.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If entire user trajectories are anonymized, then privacy protection is improved, but data utility deteriorates due to significant data distortion
Solution Approach 1:
The patent segments the anonymization process by identifying and protecting only specific privacy-vulnerable patterns (frequent location sequences) rather than anonymizing entire trajectories. This selective approach divides the data into protected portions (specific location sequences) and preserved portions (other trajectory information), resolving the contradiction between privacy protection and data utility.
Solution Approach 2:
The patent applies different levels of protection to different parts of the trajectory data based on their privacy vulnerability. High-protection areas are applied only to frequently occurring location sequences that could identify users, while other parts of the trajectory maintain higher quality and utility. This local differentiation resolves the contradiction by protecting only where necessary.
2Device complexity
If temporal information is discarded to simplify anonymization, then processing complexity is reduced, but data utility deteriorates
Solution Approach 1:
The patent performs preliminary analysis to identify privacy-vulnerable patterns before applying anonymization. By pre-identifying which location sequences are frequent and potentially identifying, the system can target only those specific patterns for protection while preserving temporal information elsewhere. This preliminary action reduces processing complexity without sacrificing useful temporal data.
3Reliability
If extensive parameterization is applied to protect user sequences, then privacy protection is improved, but ease of operation deteriorates
Solution Approach 1:
The patent implements self-service by automatically discovering privacy-vulnerable patterns directly from the data without requiring manual parameter configuration. The system autonomously identifies frequent location sequences and determines appropriate protection levels, eliminating the need for operators to manually set QIDs, m values, taxonomies, or sensitive location definitions. This automation resolves the contradiction between strong privacy protection and ease of operation.
Data Source
AI summary
A mechanism is provided for anonymizing sequential and location datasets. Responsive to receiving the sequential and location datasets from an enterprise, the sequential and location datasets are scanned to expose a set of privacy vulnerabilities. A set of privacy constraints P is generated based on the set of discovered privacy vulnerabilities and a set of utility constraints U is identified. The sequential and location datasets is anonymized using the set of privacy constraints P and the set of utility constraints U thereby forming an anonymized dataset. The anonymized dataset is then returned to the enterprise.


