Learning Image Generation for Character Recognition Robustness

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional character recognition using AI faces challenges in improving accuracy due to the need for a large number of learning images, especially when dealing with ruled lines, frame lines, seal impressions, and varying fonts, sizes, thicknesses, and densities, which requires significant time and effort to prepare, and is prone to errors and decreased robustness when encountering unknown noise.

Innovation Solution

A learning image generation apparatus and method that generates a large number of learning images by combining a first image with a second image containing disturbance elements, such as ruled lines, frame lines, and seal impressions, to improve character recognition accuracy, allowing for efficient machine learning and robustness against noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large number of learning images including ruled lines, frame lines, seal impressions, and various character variations are prepared manually, then character recognition accuracy using AI is improved, but the time and effort required for preparation increases significantly

Engineering Contradiction:
Improvecharacter recognition accuracyVSAvoidtime and effort for preparing learning images
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating diverse learning images with various disturbances (ruled lines, frame lines, seal impressions, noise) and character variations (fonts, sizes, thicknesses, densities, inclinations) before the actual character recognition task. This pre-generation of training data eliminates the need for manual preparation of numerous learning images, thereby improving character recognition accuracy while reducing preparation time and effort.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If image processing is performed to erase ruled lines and frame lines as preprocessing, then erroneous recognition is prevented, but characters such as 'I', 'L

Engineering Contradiction:
Improvecharacter recognition reliabilityVSAvoiddeletion of valid characters
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

Instead of erasing ruled lines and frame lines through image processing, the system changes the approach by incorporating these elements as disturbance components in the learning images during the training phase. The AI model learns to distinguish between actual characters and disturbance elements through exposure to varied parameters including different fonts, sizes, thicknesses, densities, and orientations, thereby improving reliability without deleting valid characters.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If the learning data set is updated with reduced bias by 3DCG, then accuracy at the time of image recognition is improved, but the complexity of data preparation and processing increases

Engineering Contradiction:
Improveimage recognition accuracyVSAvoidcomplexity of data preparation and processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system creates a universal learning image generation framework that can produce diverse training data with multiple disturbance types (ruled lines, frame lines, seal impressions, noise) and character variations (fonts, sizes, thicknesses, densities, inclinations) through a unified process. This multi-functional approach improves image recognition accuracy across various scenarios while managing data preparation complexity through systematic generation rather than separate handling of each disturbance type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240062567A1Learning Image Generation Apparatus, Learning Image Generation Method, And Non-Transitory Computer-Readable Recording Medium
Publication Date: 2024.02.22 KONICA MINOLTA INC
  • US20240062567A1 patent drawing
  • US20240062567A1 patent drawing
  • US20240062567A1 patent drawing

AI summary

A learning image generation apparatus 5 includes a first image generating unit 21 that receives the known character string Dt and generates a first image G1 including the character string Dt, a second image generating unit 22 that generates a second image G2 to be combined with the first image G1, an image combining unit 23 that generates a learning image G3 by combining the first image G1 and the second image G2, and an outputting unit 24 that outputs the learning image G3 and correct answer data Da of the character string Dt.