Invisible Unicode Whitespace Encoding for Short Text Watermarking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text watermarking techniques are not well-suited for short text data units, such as those 50 characters or less, as they require a larger corpus and negatively affect the quality of the protected data, and are not robust against unauthorized copying or redistribution.
Innovation Solution
The method involves encoding auxiliary data into alphanumeric strings organized as words separated by white spaces, using combinations of non-zero-width and flow control characters to achieve high bit rates and resilience against data shuffling and reordering, allowing for robust watermarking with minimal processing resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing text watermarking techniques are used, then watermark embedding is achieved, but the text quality deteriorates and the method is not suitable for short text data
Solution Approach 1:
The patent changes the parameter of whitespace character selection by using specific Unicode whitespace characters (U+200B, U+200C, U+200D, U+2060) that are invisible or imperceptible to humans, rather than traditional visible whitespace. This allows watermark embedding without degrading text quality while maintaining robustness for short text data units of 50 characters or less.
2Reliability
If existing text watermarking techniques are used, then watermark embedding is achieved, but the method requires a larger corpus and is not suitable for short text
Solution Approach 1:
The patent changes the encoding efficiency parameter by using combinations of invisible Unicode whitespace characters that can encode multiple bits per character position. This allows effective watermark embedding in short text data units of 50 characters or less, eliminating the requirement for larger text corpora needed by existing techniques.
3Reliability
If traditional watermarking methods are used, then data protection is achieved, but the visual integrity of the source text is compromised
Solution Approach 1:
The patent changes the character selection parameter to use invisible Unicode whitespace characters (U+200B, U+200C, U+200D, U+2060) instead of visible whitespace. These characters are either not rendered or rendered as imperceptible spaces, thereby maintaining the visual integrity of the source text while embedding protective watermarks.
4Reliability
If watermark encoding is performed, then data security is improved, but processing complexity increases
Solution Approach 1:
The patent changes the encoding scheme parameter to use predefined mappings of invisible Unicode whitespace characters to bit values. This simplified parameter-based approach reduces processing complexity compared to existing methods while maintaining strong data security through robust watermark embedding that can withstand shuffling and reordering.
Data Source
AI summary
Methods, apparatus, and articles of manufacture to encode auxiliary data into text data and methods, apparatus, and articles of manufacture to obtain encoded data from text data are disclosed. An example method includes detecting, using a processor, a first symbol present in first text data, the first symbol including a white space character; mapping the first symbol to first data using the processor; detecting, using the processor, a second symbol present in the first text data by detecting a white space character or a flow control character; mapping, using the processor, the second symbol to a first bit position of the first data in encoded data; and determining the encoded data, using the processor, based on placing the first data in the first bit position.


