Training Data Generation for Hand-Printed Text Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for automating the processing of documents with handwritten or mixed handwritten and typed text face challenges due to the lack of structural and contextual relationships in available hand-printed character databases, which are slow and error-prone to label manually.
Innovation Solution
A method that generates training data for hand-printed text recognition by iteratively processing document page images, replacing typeface characters with hand-printed characters from a database, scaling, and inserting them while preserving structural relationships, allowing for the creation of realistic training files with both typeface and hand-printed characters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human-based labeling of hand-printed documents is used to create training data, then the training data can be manually labeled with high precision, but the process is slow and error-prone
Solution Approach 1:
The patent uses optical character recognition (OCR) to create a copy of the document text with predicted character labels. This automated copy serves as a template that guides the subsequent hand-printed character generation process, eliminating the need for manual labeling while maintaining high accuracy through the OCR-based prediction framework
Solution Approach 2:
The patent replaces the mechanical process of manual human labeling with an automated computational system. The OCR system performs the labeling function that would otherwise require human operators, substituting mechanical manual work with automated digital processing to achieve both speed and precision
2Measurement precision
If existing hand-printed character databases are used for training, then character-level recognition can be achieved, but the structural and contextual relationships between characters in documents are lost
Solution Approach 1:
The patent segments the document processing into distinct stages: first processing individual characters through OCR to obtain labeled character images, then reassembling them into complete document pages while preserving the original spatial relationships and structural context. This segmentation allows each character to be accurately labeled while maintaining the overall document structure
Solution Approach 2:
The patent embeds the hand-printed character generation process within the existing document structure. The generated hand-printed characters are nested into the original document layout, maintaining the same positional relationships, spacing, and structural organization, thereby preserving contextual information while introducing hand-printed characteristics
3Reliability
If manual creation of training documents is performed, then high-quality training data can be produced, but the process is time-consuming
Solution Approach 1:
The patent performs preliminary OCR processing on the document to extract and label all characters before the hand-printed character generation step. This preliminary action prepares the character labels and images in advance, allowing the subsequent hand-printed character creation to proceed rapidly without time-consuming manual intervention
Solution Approach 2:
The system performs self-service by automatically generating the training data through the combination of OCR processing and automated hand-printed character generation. The system serves its own purpose of creating training data without requiring external manual labeling, thereby eliminating time loss while maintaining quality through the automated intelligent processing
Data Source
AI summary
A method for generating training data for hand-printed text recognition includes obtaining a structured document, obtaining a set of hand-printed character images and database metadata from a database, generating a modified document page image, and outputting a training file. The structured document includes a document page image that includes text characters and document metadata that associates each of the text characters to a document character label. The database metadata associates each of the set of hand-printed character images to a database character label. The modified document page image is generated by iteratively processing each of the text characters. The iterative processing includes determining whether an individual text character should be replaced, selecting a replacement hand-printed character image from the set of hand-printed character images, scaling the replacement hand-printed character image, and inserting the replacement hand-printed character image into the modified document page image.


