Text Watermarking via Hash-Based Group Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text watermarking techniques are not well-suited for protecting proprietary data, particularly reference data with multiple small units of text, as they require a larger corpus for encoding and negatively affect the quality of the protected data, making them ineffective against unauthorized copying and redistribution.

Innovation Solution

The method involves assigning text data to groups, identifying symbols to be added based on these groups, and generating encoded data by including text characters representative of these symbols, using a hash function to encode auxiliary information robustly, even in small subsets of data, and inserting non-breaking whitespace characters to preserve data quality and searchability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing text watermarking techniques are used to embed auxiliary information, then data protection capability is improved, but data quality and searchability deteriorate

Engineering Contradiction:
Improvedata protection capabilityVSAvoiddata quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent divides text data into multiple groups based on hash function outputs, and embeds different auxiliary information into different groups. This segmentation allows the watermarking process to be distributed across multiple data units, reducing the impact on any single unit's quality while maintaining overall protection capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different watermarking strategies to different groups of text data based on their hash values. By selectively modifying only certain characters in specific groups and preserving others, the technique maintains local data quality where modifications are not needed while still achieving global protection.

Inventive Principle:
Principle #3Local quality

2Reliability

If existing text watermarking techniques are used to embed auxiliary information, then data protection capability is improved, but data searchability deteriorates

Engineering Contradiction:
Improvedata protection capabilityVSAvoiddata searchability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

By segmenting text data into groups and applying watermarking selectively, the patent ensures that search operations can still efficiently process unmodified or minimally modified portions of the data, maintaining searchability while achieving protection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses hash functions as an intermediary mechanism to determine which data groups require watermarking. This intermediary layer allows the system to intelligently select modification targets based on data characteristics, preserving searchability in non-critical areas while protecting sensitive portions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Extent of automation

If existing text watermarking techniques are used, then watermark embedding capability is improved, but effectiveness against unauthorized copying deteriorates

Engineering Contradiction:
Improvewatermark embedding capabilityVSAvoideffectiveness against unauthorized copying
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent performs preliminary hashing and group assignment before watermark embedding, which allows for strategic placement of auxiliary information in advance. This preliminary action ensures that watermarks are positioned optimally to withstand copying and redistribution attempts.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes multiple parameters including hash function selection, group distribution strategies, and character modification patterns. By varying these parameters dynamically, the system adapts to different data types and threat scenarios, enhancing resistance against unauthorized copying.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If small subsets of text data are used for watermarking, then encoding efficiency is improved, but reliability of watermark recovery deteriorates

Engineering Contradiction:
Improveencoding efficiencyVSAvoidwatermark recovery reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments auxiliary information across multiple text data groups rather than concentrating it in a single location. This distribution ensures that even if some data is lost or corrupted, the watermark can still be recovered from remaining groups, maintaining reliability with efficient encoding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts encoding parameters based on the size and characteristics of available text data subsets. By adapting the watermarking density and distribution strategy to match data availability, the system maintains reliable watermark recovery even when working with limited data portions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9042554B2Methods, apparatus, and articles of manufacture to encode auxilary data into text data and methods, apparatus, and articles of manufacture to obtain encoded data from text data
Publication Date: 2015.05.26 THE NIELSEN CO (US) LLC
  • US9042554B2 patent drawing
  • US9042554B2 patent drawing
  • US9042554B2 patent drawing

AI summary

Methods, apparatus, and articles of manufacture to encode auxiliary data into text data and methods, apparatus, and articles of manufacture to obtain encoded data from text data are disclosed. An example method to embed auxiliary data into text data includes assigning source data to one of a plurality of groups, the source data comprising text data, identifying a symbol to be added to the source data based on an assigned group of the source data, and generating encoded data by including in the source data a text character representative of the symbol.