Utility-Aware Anonymization of Location Datasets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data anonymization solutions for location data either overprotect by concealing entire user trajectories or discarding temporal information, leading to significant data distortion and requiring extensive parameterization, which is not utility-aware and fails to provide privacy guarantees.

Innovation Solution

A utility-aware anonymization method that identifies privacy vulnerabilities in sequential and location datasets, generating privacy constraints to anonymize the data while adhering to utility constraints, ensuring the preservation of temporal information and providing privacy guarantees.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If entire user trajectories are anonymized, then privacy protection is improved, but data utility deteriorates due to significant data distortion

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata utility
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments the anonymization process by identifying and protecting only specific privacy-vulnerable patterns (frequent location sequences) rather than anonymizing entire trajectories. This selective approach divides the data into protected portions (specific location sequences) and preserved portions (other trajectory information), resolving the contradiction between privacy protection and data utility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different levels of protection to different parts of the trajectory data based on their privacy vulnerability. High-protection areas are applied only to frequently occurring location sequences that could identify users, while other parts of the trajectory maintain higher quality and utility. This local differentiation resolves the contradiction by protecting only where necessary.

Inventive Principle:
Principle #3Local quality

2Device complexity

If temporal information is discarded to simplify anonymization, then processing complexity is reduced, but data utility deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidtemporal information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent performs preliminary analysis to identify privacy-vulnerable patterns before applying anonymization. By pre-identifying which location sequences are frequent and potentially identifying, the system can target only those specific patterns for protection while preserving temporal information elsewhere. This preliminary action reduces processing complexity without sacrificing useful temporal data.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If extensive parameterization is applied to protect user sequences, then privacy protection is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidparameter configuration
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements self-service by automatically discovering privacy-vulnerable patterns directly from the data without requiring manual parameter configuration. The system autonomously identifies frequent location sequences and determines appropriate protection levels, eliminating the need for operators to manually set QIDs, m values, taxonomies, or sensitive location definitions. This automation resolves the contradiction between strong privacy protection and ease of operation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9760718B2Utility-aware anonymization of sequential and location datasets
Publication Date: 2017.09.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9760718B2 patent drawing
  • US9760718B2 patent drawing
  • US9760718B2 patent drawing

AI summary

A mechanism is provided for anonymizing sequential and location datasets. Responsive to receiving the sequential and location datasets from an enterprise, the sequential and location datasets are scanned to expose a set of privacy vulnerabilities. A set of privacy constraints P is generated based on the set of discovered privacy vulnerabilities and a set of utility constraints U is identified. The sequential and location datasets is anonymized using the set of privacy constraints P and the set of utility constraints U thereby forming an anonymized dataset. The anonymized dataset is then returned to the enterprise.