Dataset Perturbation Selection for Linkage-Resistant Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in sharing datasets while ensuring that the shared data remains protected from linking with other datasets, which could reveal sensitive information, particularly in scenarios where data is de-identified but can be re-identified through linkage with other datasets.
Innovation Solution
A method involving perturbing an input dataset multiple times with different perturbation parameters to generate multiple derived datasets, calculating utility scores, and selecting the most useful dataset under given protection against linking, thereby reducing computational resources and enhancing privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is de-identified by removing identifier fields, then data sharing capability is improved, but data remains vulnerable to re-identification through linking with other datasets
Solution Approach 1:
The patent applies parameter changes by introducing perturbation parameters (epsilon and delta) that control the degree of data modification. By adjusting these parameters, the system can dynamically balance between data utility for sharing and protection against linking, transforming the data to satisfy specific privacy requirements while maintaining usability.
Solution Approach 2:
The patent introduces a perturbation function as an intermediary between the original dataset and the shared dataset. This function applies controlled randomization to create a transformed version that preserves statistical properties for analysis while breaking direct linkability to individual records, thus mediating between sharing and protection needs.
2Reliability
If high degree of randomisation is applied to protect against linking, then privacy protection is improved, but computational resources required increase significantly
Solution Approach 1:
The patent uses parameter changes to control the intensity of randomisation through the epsilon parameter. By selecting appropriate parameter values, the system achieves adequate privacy protection with minimal computational overhead, avoiding excessive randomisation that would waste resources while still providing sufficient protection against linking.
Solution Approach 2:
The patent applies partial action by implementing the minimum necessary randomisation required to achieve privacy goals. Rather than applying maximum randomisation uniformly, the system applies just enough perturbation to satisfy privacy requirements, thereby reducing unnecessary computational expenditure while maintaining effective protection.
3Loss of information
If multiple perturbation parameters are tested to find optimal utility, then data utility for specific purposes is improved, but the number of randomisations and computational effort increases
Solution Approach 1:
The patent applies preliminary action by pre-defining a finite set of candidate perturbation parameters before the data sharing process. This allows the system to evaluate multiple utility scenarios in advance and select the optimal parameter set, avoiding the need for extensive real-time randomisation trials and improving computational efficiency.
Data Source
AI summary
This disclosure relates to protecting an input dataset against linking with further datasets. A processor of a computer system calculates multiple values of one or more parameters of a perturbation function, the perturbation function being configured to perturb the input dataset to protect the input dataset against linking with further datasets, each of the multiple values of the one or more parameters of the perturbation function indicating a level of protection against linking with further datasets. The processor then generates multiple derived datasets from the input dataset and calculates, for each of the multiple derived datasets, a utility score that is indicative of a utility of the derived dataset for a desired data analysis. The processor then outputs one of the multiple derived datasets that has the highest utility score.


