Language Model Training with Simulated OCR Defects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for training language models for optical character recognition (OCR) are ineffective due to the poor quality and incorrect positioning of synthetic OCR errors, leading to poor quality results in recognizing text from images.
Innovation Solution
A method is introduced to generate realistic OCR errors by simulating defects such as lines, spots, and other imperfections on images based on an input text corpus, creating an augmented set of images that include context-dependent information, which are then used to train language models for improved OCR performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If synthetic OCR errors are used to train language models, then training data can be generated, but the quality of OCR errors is poor and positioning is incorrect
Solution Approach 1:
The patent copies real OCR errors from actual document images into synthetic training data by overlaying defect patterns (lines, spots, noise) that replicate authentic OCR failure modes. This copying approach preserves the statistical characteristics and spatial distribution of real errors while generating abundant training examples.
Solution Approach 2:
The patent introduces an intermediary layer of simulated defect patterns that mediate between clean synthetic text images and real-world OCR challenges. These defect patterns act as a bridge, translating artificial training data into realistic error scenarios without requiring actual annotated error datasets.
2Adaptability or versatility
If simple synthetic errors are added to training images, then data augmentation is achieved, but the errors do not accurately represent real document OCR errors
Solution Approach 1:
The patent applies local quality by introducing specific defect patterns (horizontal lines, vertical lines, spots, noise) at localized positions within text regions. Each defect type targets specific OCR vulnerability zones, such as line defects affecting character continuity or spot defects affecting character recognition, thereby creating realistic error scenarios.
Solution Approach 2:
The patent employs parameter changes by varying defect characteristics including position, size, density, and type across training samples. By dynamically adjusting these parameters, the system generates diverse realistic error patterns that reflect the variability found in actual document imaging conditions.
Data Source
AI summary
Systems and methods for generating text corpora comprising realistic optical character recognition (OCR) errors and training language models using the text corpora are provided. An example method comprises: generating, by a computer system, an initial set of images based on an input text corpus comprising text; overlaying, by the computer system, one or more simulated defects over the initial set of images to generate an augmented set of images; generating an output text corpus based on the augmented set of image; and training, using the output text corpus, a language model for optical character recognition.


