Asian Document Image Text Shuffling for Secure Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Asian-language documents are challenging to encode securely, as existing OCR software often fails to correctly recognize individual characters, and there is a need for a method that maintains character recognizability while rendering the document unreadable.

Innovation Solution

A method involving the division of a document image into text and non-text portions, structuring the text into a multiple resolution-level pyramid, extracting and shuffling character images, and reshuffling them to create a scrambled image that preserves character recognizability but loses meaning, using techniques like skew correction, noise removal, and tree structure analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If OCR software is used to recognize characters in Asian documents, then character recognition is attempted, but recognition accuracy deteriorates due to limitations in OCR software

Engineering Contradiction:
Improvecharacter recognition accuracyVSAvoidOCR software capability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent creates a shuffled copy of the original document that preserves visual character information while disrupting semantic meaning. This copy can be manually corrected and then restored to the original, providing a reliable method to handle OCR errors without directly relying on flawed OCR recognition of the original document.

Inventive Principle:
Principle #26Copying

2Reliability

If a document is encoded to become unreadable for security, then security is improved, but character recognizability is lost

Engineering Contradiction:
Improvedocument securityVSAvoidcharacter recognizability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments the document into individual character images that can be independently shuffled. This segmentation allows the document to be encoded in a way that preserves the visual integrity of each character while disrupting the overall meaning, maintaining character recognizability while achieving security.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the document from a semantic structure (where character position conveys meaning) to a visual structure (where character appearance is preserved but position is randomized). This dimensional change allows the document to be secure while maintaining character recognizability for manual correction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If character positions are shuffled to encode the document, then security is improved, but the ability to identify individual characters deteriorates

Engineering Contradiction:
Improvedocument encoding securityVSAvoidcharacter identifiability
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent creates a shuffled copy that maintains the visual appearance of individual characters while randomizing their positions. This copy preserves character identifiability for manual correction purposes while achieving encoding security, as the visual characteristics of each character remain recognizable despite position changes.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS7596270B2Method of shuffling text in an Asian document image
Publication Date: 2009.09.29 DYNACOMWARE INC
  • US7596270B2 patent drawing
  • US7596270B2 patent drawing
  • US7596270B2 patent drawing

AI summary

A method, system, and computer-readable medium containing computer-executable instructions are provided, for randomly relocating text character images of a scanned-in Asian character document to produce a shuffled image, wherein the meaning of text in the shuffled image is not understandable although individual characters forming the text in the shuffled image are recognizable. In one embodiment, the method includes generally four steps: (1) dividing an Asian character document image into a text image portion and a non-text image portion; (2) structuring the text image portion into a multiple resolution-level pyramid; (3) extracting shuffleable character images by analyzing the multiple-resolution-level pyramid; and (4) shuffling some or all of the extracted shuffleable character images to create a shuffled image. The shuffled (e.g., encoded) image can be reshuffled (e.g., decoded) back to the original text image portion of the Asian character document image, and combined with the non-text image portion to restore the Asian character document image.