Asian Document Image Text Shuffling for Secure Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Asian-language documents are challenging to encode securely, as existing OCR software often fails to correctly recognize individual characters, and there is a need for a method that maintains character recognizability while rendering the document unreadable.
Innovation Solution
A method involving the division of a document image into text and non-text portions, structuring the text into a multiple resolution-level pyramid, extracting and shuffling character images, and reshuffling them to create a scrambled image that preserves character recognizability but loses meaning, using techniques like skew correction, noise removal, and tree structure analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If OCR software is used to recognize characters in Asian documents, then character recognition is attempted, but recognition accuracy deteriorates due to limitations in OCR software
Solution Approach 1:
The patent creates a shuffled copy of the original document that preserves visual character information while disrupting semantic meaning. This copy can be manually corrected and then restored to the original, providing a reliable method to handle OCR errors without directly relying on flawed OCR recognition of the original document.
2Reliability
If a document is encoded to become unreadable for security, then security is improved, but character recognizability is lost
Solution Approach 1:
The patent segments the document into individual character images that can be independently shuffled. This segmentation allows the document to be encoded in a way that preserves the visual integrity of each character while disrupting the overall meaning, maintaining character recognizability while achieving security.
Solution Approach 2:
The patent transforms the document from a semantic structure (where character position conveys meaning) to a visual structure (where character appearance is preserved but position is randomized). This dimensional change allows the document to be secure while maintaining character recognizability for manual correction.
3Reliability
If character positions are shuffled to encode the document, then security is improved, but the ability to identify individual characters deteriorates
Solution Approach 1:
The patent creates a shuffled copy that maintains the visual appearance of individual characters while randomizing their positions. This copy preserves character identifiability for manual correction purposes while achieving encoding security, as the visual characteristics of each character remain recognizable despite position changes.
Data Source
AI summary
A method, system, and computer-readable medium containing computer-executable instructions are provided, for randomly relocating text character images of a scanned-in Asian character document to produce a shuffled image, wherein the meaning of text in the shuffled image is not understandable although individual characters forming the text in the shuffled image are recognizable. In one embodiment, the method includes generally four steps: (1) dividing an Asian character document image into a text image portion and a non-text image portion; (2) structuring the text image portion into a multiple resolution-level pyramid; (3) extracting shuffleable character images by analyzing the multiple-resolution-level pyramid; and (4) shuffling some or all of the extracted shuffleable character images to create a shuffled image. The shuffled (e.g., encoded) image can be reshuffled (e.g., decoded) back to the original text image portion of the Asian character document image, and combined with the non-text image portion to restore the Asian character document image.


