Learning Image Generation for Character Recognition Robustness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional character recognition using AI faces challenges in improving accuracy due to the need for a large number of learning images, especially when dealing with ruled lines, frame lines, seal impressions, and varying fonts, sizes, thicknesses, and densities, which requires significant time and effort to prepare, and is prone to errors and decreased robustness when encountering unknown noise.
Innovation Solution
A learning image generation apparatus and method that generates a large number of learning images by combining a first image with a second image containing disturbance elements, such as ruled lines, frame lines, and seal impressions, to improve character recognition accuracy, allowing for efficient machine learning and robustness against noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large number of learning images including ruled lines, frame lines, seal impressions, and various character variations are prepared manually, then character recognition accuracy using AI is improved, but the time and effort required for preparation increases significantly
Solution Approach 1:
The system performs preliminary actions by automatically generating diverse learning images with various disturbances (ruled lines, frame lines, seal impressions, noise) and character variations (fonts, sizes, thicknesses, densities, inclinations) before the actual character recognition task. This pre-generation of training data eliminates the need for manual preparation of numerous learning images, thereby improving character recognition accuracy while reducing preparation time and effort.
2Reliability
If image processing is performed to erase ruled lines and frame lines as preprocessing, then erroneous recognition is prevented, but characters such as 'I', 'L
Solution Approach 1:
Instead of erasing ruled lines and frame lines through image processing, the system changes the approach by incorporating these elements as disturbance components in the learning images during the training phase. The AI model learns to distinguish between actual characters and disturbance elements through exposure to varied parameters including different fonts, sizes, thicknesses, densities, and orientations, thereby improving reliability without deleting valid characters.
3Measurement precision
If the learning data set is updated with reduced bias by 3DCG, then accuracy at the time of image recognition is improved, but the complexity of data preparation and processing increases
Solution Approach 1:
The system creates a universal learning image generation framework that can produce diverse training data with multiple disturbance types (ruled lines, frame lines, seal impressions, noise) and character variations (fonts, sizes, thicknesses, densities, inclinations) through a unified process. This multi-functional approach improves image recognition accuracy across various scenarios while managing data preparation complexity through systematic generation rather than separate handling of each disturbance type.
Data Source
AI summary
A learning image generation apparatus 5 includes a first image generating unit 21 that receives the known character string Dt and generates a first image G1 including the character string Dt, a second image generating unit 22 that generates a second image G2 to be combined with the first image G1, an image combining unit 23 that generates a learning image G3 by combining the first image G1 and the second image G2, and an outputting unit 24 that outputs the learning image G3 and correct answer data Da of the character string Dt.


