Text Watermarking via Hash-Based Group Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text watermarking techniques are not well-suited for protecting proprietary data, particularly reference data with multiple small units of text, as they require a larger corpus for encoding and negatively affect the quality of the protected data, making them ineffective against unauthorized copying and redistribution.
Innovation Solution
The method involves assigning text data to groups, identifying symbols to be added based on these groups, and generating encoded data by including text characters representative of these symbols, using a hash function to encode auxiliary information robustly, even in small subsets of data, and inserting non-breaking whitespace characters to preserve data quality and searchability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing text watermarking techniques are used to embed auxiliary information, then data protection capability is improved, but data quality and searchability deteriorate
Solution Approach 1:
The patent divides text data into multiple groups based on hash function outputs, and embeds different auxiliary information into different groups. This segmentation allows the watermarking process to be distributed across multiple data units, reducing the impact on any single unit's quality while maintaining overall protection capability.
Solution Approach 2:
The patent applies different watermarking strategies to different groups of text data based on their hash values. By selectively modifying only certain characters in specific groups and preserving others, the technique maintains local data quality where modifications are not needed while still achieving global protection.
2Reliability
If existing text watermarking techniques are used to embed auxiliary information, then data protection capability is improved, but data searchability deteriorates
Solution Approach 1:
By segmenting text data into groups and applying watermarking selectively, the patent ensures that search operations can still efficiently process unmodified or minimally modified portions of the data, maintaining searchability while achieving protection.
Solution Approach 2:
The patent uses hash functions as an intermediary mechanism to determine which data groups require watermarking. This intermediary layer allows the system to intelligently select modification targets based on data characteristics, preserving searchability in non-critical areas while protecting sensitive portions.
3Extent of automation
If existing text watermarking techniques are used, then watermark embedding capability is improved, but effectiveness against unauthorized copying deteriorates
Solution Approach 1:
The patent performs preliminary hashing and group assignment before watermark embedding, which allows for strategic placement of auxiliary information in advance. This preliminary action ensures that watermarks are positioned optimally to withstand copying and redistribution attempts.
Solution Approach 2:
The patent changes multiple parameters including hash function selection, group distribution strategies, and character modification patterns. By varying these parameters dynamically, the system adapts to different data types and threat scenarios, enhancing resistance against unauthorized copying.
4Productivity
If small subsets of text data are used for watermarking, then encoding efficiency is improved, but reliability of watermark recovery deteriorates
Solution Approach 1:
The patent segments auxiliary information across multiple text data groups rather than concentrating it in a single location. This distribution ensures that even if some data is lost or corrupted, the watermark can still be recovered from remaining groups, maintaining reliability with efficient encoding.
Solution Approach 2:
The patent dynamically adjusts encoding parameters based on the size and characteristics of available text data subsets. By adapting the watermarking density and distribution strategy to match data availability, the system maintains reliable watermark recovery even when working with limited data portions.
Data Source
AI summary
Methods, apparatus, and articles of manufacture to encode auxiliary data into text data and methods, apparatus, and articles of manufacture to obtain encoded data from text data are disclosed. An example method to embed auxiliary data into text data includes assigning source data to one of a plurality of groups, the source data comprising text data, identifying a symbol to be added to the source data based on an assigned group of the source data, and generating encoded data by including in the source data a text character representative of the symbol.


