Text Data Embedding via Inter-Character Spacing Modulation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing steganography techniques for text documents are not robust enough to handle flexible graphical representations and noisy environments, particularly when the text is displayed on varying screen sizes or distorted images.
Innovation Solution
A method for embedding a bit sequence in a text string by identifying carrier strings with at least two spaces, and altering these spaces to distinguish between different bit values, ensuring robustness across different graphical representations and noise levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional steganography techniques (dot encoding, space modulation) are used to embed data in text, then data hiding capability is achieved, but robustness against noise and graphical representation variations is poor
Solution Approach 1:
The patent changes the parameter being modified from individual character properties (conventional approach) to inter-character distance (spacing). By embedding data in the spacing between characters rather than modifying character appearance, the method achieves robustness against noise and graphical variations while maintaining adaptability to different display formats.
Solution Approach 2:
The patent replaces conventional mechanical/visual modification methods (altering character shape, size, or position individually) with a spacing-based modulation system. This substitution enables the embedded data to survive graphical transformations, compression, and noise better than traditional character-level modifications.
2Adaptability or versatility
If text areas and fonts are made flexible to adapt to different devices, then ease of operation and adaptability improve, but conventional embedding methods fail because they depend on fixed text dimensions
Solution Approach 1:
The spacing-based embedding method serves multiple functions: it works across different screen sizes, font types, and graphical representations while maintaining data integrity. The inter-character distance metric is universal and invariant to the specific display medium, enabling the same embedding technique to function reliably across diverse devices and formats.
3Ease of operation
If the image is transcoded, compressed, or distorted for transmission, then ease of operation and adaptability improve, but decoding accuracy deteriorates with conventional methods
Solution Approach 1:
The patent embeds data in inter-character spacing, which creates a buffer or cushion against the effects of compression and distortion. The spacing between characters is more resilient to graphical transformations than character-level modifications, providing prior protection against degradation during transmission and processing.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
The invention provides, amongst other aspects, a method for embedding a bit sequence, comprising: providing the bit sequence and a text string comprising words; identifying, among the words comprised in the text string, carrier strings suitable for embedding; generating formatting data associated with said text string, said formatting data comprising instructions for embedding, within the respective identified carrier strings, respective bits of the bit sequence; wherein the identifying of carrier strings comprises identifying, for each carrier string, one or more carrier words such that the carrier string comprises at least two spaces, the spaces being letter spaces or word spaces; wherein the embedding of a respective bit relates to altering each of the letter spaces and/or each of the word spaces of the respective carrier string so that the sum of altered spaces if the value of the bit is zero is distinguishable from the sum of altered spaces if the value of the bit is one.