Anonymized Dataset Watermarking With Minimal Utility Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization techniques in raw datasets often result in significant information loss when digital watermarking is applied, as these datasets lack metadata or redundant data, necessitating manipulation of the data itself, which compromises their utility.
Innovation Solution
A computer-implemented process that incorporates digital watermarking into anonymized datasets by extending anonymization techniques such as tokenization, generalization, data blurring, and synthetic record insertion, ensuring minimal additional information loss by embedding watermarks probabilistically or through non-destructive methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If digital watermarking is applied to raw datasets by manipulating the data itself, then watermarking capability is achieved, but information loss and utility degradation occur
Solution Approach 1:
The patent applies anonymization techniques (tokenization, generalization, blurring, synthetic record insertion) to the raw dataset before watermarking. This preliminary transformation creates a modified dataset structure that allows watermark embedding without requiring direct manipulation of the original data values, thereby preserving data utility while enabling watermarking capability.
2Object-affected harmful factors
If anonymization techniques are applied to protect privacy, then privacy protection is improved, but information loss occurs
Solution Approach 1:
The patent employs multiple anonymization techniques that transform data parameters: tokenization replaces identifiers with tokens, generalization aggregates specific values into broader categories, blurring introduces noise to numerical values, and synthetic record insertion adds artificial data. These parameter changes protect privacy by removing or obscuring personally identifiable information while maintaining the statistical and structural properties of the dataset, thereby preserving data utility for analysis.
3Reliability
If multiple anonymization techniques are combined with watermarking, then privacy protection and watermarking are both achieved, but process complexity increases
Solution Approach 1:
The patent integrates multiple anonymization techniques (tokenization, generalization, blurring, synthetic record insertion) into a unified preprocessing pipeline that operates on the raw dataset before watermarking. This merging of techniques into a sequential processing flow achieves comprehensive privacy protection through multiple layers of obfuscation while managing complexity through systematic organization of the processing steps.
Data Source
AI summary
A computer-implemented process of altering original data in a dataset, in which original data is anonymised and a digital watermark is included in the anonymised data. Anonymising the original data incurs information loss, and the process of including the digital watermark does not add significant further information loss. The original data can be a tabular file, a relational or a non-relational database, or the results of interactive database queries. Anonymising the data is achieved using one or more techniques that perturb the original data, such as tokenisation, generalisation; data blurring, synthetic record insertion, record removal or re-ordering.


