Anonymized Dataset Watermarking With Minimal Utility Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data anonymization techniques in raw datasets often result in significant information loss when digital watermarking is applied, as these datasets lack metadata or redundant data, necessitating manipulation of the data itself, which compromises their utility.

Innovation Solution

A computer-implemented process that incorporates digital watermarking into anonymized datasets by extending anonymization techniques such as tokenization, generalization, data blurring, and synthetic record insertion, ensuring minimal additional information loss by embedding watermarks probabilistically or through non-destructive methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If digital watermarking is applied to raw datasets by manipulating the data itself, then watermarking capability is achieved, but information loss and utility degradation occur

Engineering Contradiction:
Improvewatermarking capabilityVSAvoiddata utility
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies anonymization techniques (tokenization, generalization, blurring, synthetic record insertion) to the raw dataset before watermarking. This preliminary transformation creates a modified dataset structure that allows watermark embedding without requiring direct manipulation of the original data values, thereby preserving data utility while enabling watermarking capability.

Inventive Principle:
Principle #10Preliminary action

2Object-affected harmful factors

If anonymization techniques are applied to protect privacy, then privacy protection is improved, but information loss occurs

Engineering Contradiction:
Improveprivacy riskVSAvoiddata utility
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent employs multiple anonymization techniques that transform data parameters: tokenization replaces identifiers with tokens, generalization aggregates specific values into broader categories, blurring introduces noise to numerical values, and synthetic record insertion adds artificial data. These parameter changes protect privacy by removing or obscuring personally identifiable information while maintaining the statistical and structural properties of the dataset, thereby preserving data utility for analysis.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If multiple anonymization techniques are combined with watermarking, then privacy protection and watermarking are both achieved, but process complexity increases

Engineering Contradiction:
Improveprivacy protectionVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent integrates multiple anonymization techniques (tokenization, generalization, blurring, synthetic record insertion) into a unified preprocessing pipeline that operates on the raw dataset before watermarking. This merging of techniques into a sequential processing flow achieves comprehensive privacy protection through multiple layers of obfuscation while managing complexity through systematic organization of the processing steps.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250328691A1Digital watermarking without significant information loss in anonymized datasets
Publication Date: 2025.10.23 PRIVITAR LTD
  • US20250328691A1 patent drawing
  • US20250328691A1 patent drawing
  • US20250328691A1 patent drawing

AI summary

A computer-implemented process of altering original data in a dataset, in which original data is anonymised and a digital watermark is included in the anonymised data. Anonymising the original data incurs information loss, and the process of including the digital watermark does not add significant further information loss. The original data can be a tabular file, a relational or a non-relational database, or the results of interactive database queries. Anonymising the data is achieved using one or more techniques that perturb the original data, such as tokenisation, generalisation; data blurring, synthetic record insertion, record removal or re-ordering.