Document Text Watermarking Using Encodable Font Variants

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital watermarking technologies face challenges in effectively encoding and decoding information within the textual part of documents without altering their aesthetic appearance, especially when the document is photocopied or digitized, due to issues with separating added information from natural digitization noise.

Innovation Solution

A method using a specific font with encodable characters and their variants, where each character can represent different values through Optical Character Recognition (OCR) processes, allowing for encoding and decoding of information within the textual part of documents without altering their visual appearance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If digital watermarking is applied to the textual part of a document, then information can be encoded and decoded, but the visual appearance of the document is altered

Engineering Contradiction:
Improveinformation encoding capabilityVSAvoidvisual appearance
Core Design Contradiction:
Loss of informationVSShape

Solution Approach 1:

The patent applies local quality by making different parts of the text have different functions. Specific characters are selected as 'encodable characters' that carry both their original semantic meaning and additional encoded information through their font variant. This allows the text to maintain its primary readability while locally embedding watermark data in selected characters without altering the overall visual appearance of the document.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses parameter changes by varying the font variant parameter of encodable characters to encode information. Each encodable character can be rendered in different font variants (e.g., different weights, italics, underlines) that are imperceptible to human readers but can be detected by OCR software. This changes the physical parameter of the character representation while maintaining the same visual appearance, enabling information encoding without aesthetic alteration.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If gray level values are assigned to text points for watermarking, then information can be encoded, but the reliability is low due to printing and digitization noise

Engineering Contradiction:
Improvewatermark encodingVSAvoidencoding reliability
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent replaces the mechanical/grayscale approach with a categorical/discrete approach. Instead of using continuous gray level values that are susceptible to noise, the invention uses discrete font variants of specific characters. The OCR system detects which font variant is present (a categorical classification problem) rather than measuring gray levels (a continuous measurement problem). This substitution fundamentally improves reliability because categorical classification is more robust to printing and digitization noise than gray level measurement.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent uses copying by creating multiple versions (font variants) of the same character that are visually identical or nearly identical but have detectable differences. The OCR system is trained to recognize these variants and decode the embedded information. This copying approach allows the watermark to be reproduced faithfully through standard printing and digitization processes without degrading the reliability of the encoded information.

Inventive Principle:
Principle #26Copying

3Loss of information

If complex watermarking technologies are used, then information can be encoded, but the implementation complexity and computing power requirements increase

Engineering Contradiction:
Improvewatermark encoding capabilityVSAvoidimplementation complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies universality by using standard OCR technology, which is already widely deployed and well-understood, for the dual purpose of both reading the document's text content and detecting the encoded watermark information. The same OCR engine that recognizes characters for their semantic meaning is also configured to recognize specific font variants as encoded data. This eliminates the need for separate, complex watermark detection systems and reduces implementation complexity while maintaining encoding capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP2974260B1Method for watermarking the text portion of a document
Publication Date: 2021.05.05 SEND ONLY OKED DOCUMENTS SOOD
  • EP2974260B1 patent drawingFigure 1
  • EP2974260B1 patent drawingFigure 2A
  • EP2974260B1 patent drawingFigure 2B

AI summary

A method for watermarking a document containing at least one text portion comprising the following steps: - determining a specific character font comprising, for at least one character, an original graphic and at least one variation, each of the variations being associated with a different value, said character being termed encodable characters; - using the specific character font to encode an item of information in the text portion of the document, by replacing at least one original graphic with a variation, the original graphic and the variation or variations being identified as a single character by a first optical character recognition process referred to as standard OCR and identified as a plurality of characters by a second optical character recognition process referred to as specific OCR that is capable of determining if the represented character is the original graphic or one of the variations of same and, if so, making it possible to determine the variation that is represented, a strict order relationship being defined on the encodable characters in order to establish the order in which the encodable characters are to be processed during the decoding phase.