Image generation device for learning, method for generating images for learning, and program

The learning image generation device improves AI-based character recognition accuracy by synthesizing training images with diverse disturbances, addressing the challenges of ruled lines and seal shadows, and reducing the time and effort required for training image preparation.

JP7861576B2Active Publication Date: 2026-05-19KONICA MINOLTA INC
View PDF 13 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KONICA MINOLTA INC
Filing Date
2022-08-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing AI-based character recognition technologies face challenges in accurately recognizing characters due to the presence of ruled lines, frame lines, seal shadows, and back reflections, and require a large number of training images to improve robustness, which is time-consuming and effort-intensive.

Method used

A learning image generation device that generates training images by synthesizing first images containing characters with second images including disturbances such as stamp impressions, grid lines, and noise components, using image synthesis techniques to create a variety of training scenarios.

Benefits of technology

Facilitates the easy preparation of a large number of training images, enhancing the accuracy and robustness of AI-based character recognition by simulating diverse text types and noise conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861576000001
    Figure 0007861576000001
  • Figure 0007861576000002
    Figure 0007861576000002
  • Figure 0007861576000003
    Figure 0007861576000003
Patent Text Reader

Abstract

To provide a learning image generation device, a learning image generation method, and a program capable of easily preparing a lot of learning images for improving accuracy of character recognition using AI.SOLUTION: In a character recognition system, a learning image generation device 5 includes: a first image generation unit 21 which inputs a known character string Dt and generates a first image G1 containing the character string Dt; a second image generation unit 22 which generates a second image G2 to be combined with the first image G1; an image combining unit 23 which generates a learning image G3 by combining the first image G1 and the second image G2; and an output unit 24 which outputs the learning image G3 and correct answer data Da of the character string Dt.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning image generation device, a learning image generation method, and a program, and particularly relates to a technique for generating a learning image for improving the character recognition accuracy by artificial intelligence (AI).

Background Art

[0002] In recent years, there has been an increasing need to read paper documents with a scanner or the like and save them in electronic data form. When converting paper documents into electronic data, the convenience of the electronically converted paper documents is improved by converting the characters included in the paper documents into text data. Conventionally, as one method of converting characters into text data, there is OCR (Optical Character Recognition / Reader). That is, an image area in which characters are shown is cut out from an image read by a scanner, and the characters included in the image are converted into text data by performing character recognition processing on the image area.

[0003] However, in the conventional character recognition processing, when the cut-out image area includes ruled lines or frame lines, the ruled line or frame line portions may be misrecognized as characters such as "I", "L", "1", or when there is a seal shadow or back reflection, the characters in that portion may be misrecognized. To prevent these misrecognitions, it is conceivable to perform image processing for erasing ruled lines and frame lines as a preprocessing before performing character recognition processing. However, when such image processing is performed, there is a possibility that characters such as "I", "L", "1" that should originally be recognized as characters may be erased, and there is a problem that the characters cannot be accurately recognized in the subsequent character recognition processing.

[0004] On the other hand, in recent years, technologies for recognizing images using AI (Artificial Intelligence) technology have been proposed (Patent Document 1). In this conventional technology, if there is a bias in the training dataset used for AI machine learning, the accuracy of image recognition is improved by updating the training dataset with one whose bias has been reduced using 3DCG and then performing machine learning.

[0005] AI-based character recognition can also be applied to recognize text from images of paper documents scanned using a scanner. By pre-training the AI ​​with images that include lines, borders, and seal impressions, it can appropriately separate and recognize text from these image components, thereby improving the character recognition rate without requiring any pre-processing for character recognition. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2021-111101 [Overview of the project] [Problems that the invention aims to solve]

[0007] However, in order to further improve the accuracy of character recognition in AI-based character recognition processing, a large number of training images containing lines, borders, and seal impressions are required for machine learning, and preparing such training images requires a great deal of time and effort.

[0008] Furthermore, if AI lacks robustness, its character recognition accuracy may decrease if unknown noise is introduced. For example, text in images read by scanners may be written in a variety of fonts, sizes, thicknesses, and densities, and may even be handwritten. Additionally, text within images may be slanted. Therefore, to ensure robustness in AI-based character recognition processing, it is desirable to enable the AI ​​to appropriately recognize these diverse text types.

[0009] Therefore, the present invention has been made to solve the above-mentioned conventional problems, and aims to provide a learning image generation device, a learning image generation method, and a program that enable the easy preparation of a large number of learning images for improving the accuracy of character recognition using AI. [Means for solving the problem]

[0010] To achieve the above objective, the invention according to claim 1 is a learning image generation device comprising: a first image generation unit that takes a known string of characters as input and generates a first image containing the string of characters; a second image generation unit that generates a second image for synthesis with the first image; an image synthesis unit that generates a learning image by synthesizing the first image and the second image; and an output unit that outputs the learning image and the correct data for the string of characters. The second image generation unit generates the second image which includes a stamp image obtained by processing the string, and the image synthesis unit generates the learning image which is obtained by overlaying the stamp image onto the string included in the first image and synthesizing them. This configuration is characterized by the following features. The invention according to claim 2 is a learning image generation device comprising: a first image generation unit that takes a known string of characters as input and generates a first image containing the string of characters; a second image generation unit that generates a second image for synthesis with the first image; an image synthesis unit that generates a learning image by synthesizing the first image and the second image; and an output unit that outputs the learning image and correct data for the string of characters, wherein the second image generation unit generates the second image containing a string of characters different from the first image. The invention according to claim 3 is a learning image generation device comprising: a first image generation unit that takes a known string of characters as input and generates a first image containing the string of characters; a second image generation unit that generates a second image for synthesis with the first image; an image synthesis unit that generates a learning image by synthesizing the first image and the second image; and an output unit that outputs the learning image and correct data for the string of characters, wherein the second image generation unit generates the second image containing an image of a string of characters different from the first image that has been horizontally flipped. The invention according to claim 4 is a learning image generation device comprising: a first image generation unit that takes a known string of characters as input and generates a first image containing the string of characters; a second image generation unit that generates a second image for synthesis with the first image; an image synthesis unit that generates a learning image by synthesizing the first image and the second image; and an output unit that outputs the learning image and the correct data of the string of characters, wherein the image synthesis unit generates a plurality of the learning images by changing the parameters when synthesizing the first image and the second image, and the output unit outputs the plurality of the learning images, the parameters including the transparency of the first image and the second image, respectively.

[0011] Claim 5 The invention relating to this is as described in claim 1 any of the following four: In the learning image generation device, the first image generation unit is characterized by generating the first image by converting the string into image data.

[0012] Claim 6 The invention relating to this claim is 5In the learning image generation device, the output unit outputs text data representing the character string as the correct answer data.

[0013] Claim 7 The invention according to claim 6 In the learning image generation device of claim

[0022] Claim 8 The invention according to claim 4 In the learning image generation device of claim

[0023] Claim 9 The invention according to claim 4 In the learning image generation device of claim

[0025] Claim 10 The invention according to claim is a learning image generation method, including a first image generation step of inputting a known character string and generating a first image including the character string, a second image generation step of generating a second image for synthesis with the first image, an image synthesis step of generating a learning image by synthesizing the first image and the second image, and an output step of outputting the learning image and the correct answer data of the character string. The second image generation step generates a second image including a stamp image obtained by processing the string, and the image synthesis step generates a training image by overlaying the stamp image onto the string included in the first image and synthesizing them. It is characterized by the above structure. The invention according to claim 11 is a method for generating a learning image, comprising: a first image generation step of inputting a known string and generating a first image containing the string; a second image generation step of generating a second image for synthesis with the first image; an image synthesis step of generating a learning image by synthesizing the first image and the second image; and an output step of outputting the learning image and the correct data for the string, wherein the second image generation step generates the second image containing a string different from the string. The invention according to claim 12 is a method for generating a learning image, comprising: a first image generation step of inputting a known string and generating a first image containing the string; a second image generation step of generating a second image for synthesis with the first image; an image synthesis step of generating a learning image by synthesizing the first image and the second image; and an output step of outputting the learning image and the correct data for the string, wherein the second image generation step generates the second image containing an image of a string different from the string that has been horizontally flipped. The invention according to claim 13 is a method for generating a learning image, comprising: a first image generation step of inputting a known string and generating a first image containing the string; a second image generation step of generating a second image for synthesis with the first image; an image synthesis step of generating a learning image by synthesizing the first image and the second image; and an output step of outputting the learning image and the correct data of the string, wherein the image synthesis step generates a plurality of the learning images by changing the parameters when synthesizing the first image and the second image, and the output step outputs the plurality of the learning images, the parameters including the transparency of the first image and the second image, respectively.

[0026] Claim 14The invention according to this is a program that causes a computer to execute a first image generation step of inputting a known character string and generating a first image including the character string, a second image generation step of generating a second image for compositing with the first image, an image compositing step of generating a learning image by compositing the first image and the second image, and an output step of outputting the learning image and the correct data of the character string. The second image generation step generates a second image including a stamp image obtained by processing the string, and the image synthesis step generates a training image by overlaying the stamp image onto the string included in the first image and synthesizing them. It is a configuration characterized by this. The invention according to claim 15 is a program that causes a computer to perform the following steps: a first image generation step of inputting a known string and generating a first image containing the string; a second image generation step of generating a second image for synthesis with the first image; an image synthesis step of generating a training image by synthesizing the first image and the second image; and an output step of outputting the training image and correct data for the string, wherein the second image generation step generates a second image containing a string different from the first string. The invention according to claim 16 is a program that causes a computer to perform the following steps: a first image generation step of inputting a known string and generating a first image containing the string; a second image generation step of generating a second image for synthesis with the first image; an image synthesis step of generating a training image by synthesizing the first image and the second image; and an output step of outputting the training image and the correct data for the string, wherein the second image generation step generates the second image containing an image of a string different from the first string that has been horizontally flipped. The invention according to claim 17 is a program that causes a computer to perform the following steps: a first image generation step of inputting a known string and generating a first image containing the string; a second image generation step of generating a second image for synthesis with the first image; an image synthesis step of generating a training image by synthesizing the first image and the second image; and an output step of outputting the training image and the correct data for the string, wherein the image synthesis step generates a plurality of the training images by changing the parameters when synthesizing the first image and the second image; the output step outputs the plurality of the training images; and the parameters include the transparency of the first image and the second image, respectively.

Effect of the Invention

[0027] According to the present invention, a large number of learning images for improving the recognition accuracy of characters in character recognition processing using AI can be easily generated.

Brief Description of the Drawings

[0028] [Figure 1] It is a diagram showing one configuration example of a character recognition system. [Figure 2] It is a diagram illustrating the hardware configuration of an information processing apparatus. [Figure 3] It is a block diagram illustrating the functional configurations of a learning image generation apparatus and an image processing apparatus. [Figure 4] It is a diagram showing an example of compositing a second image including a shadow image with a first image. [Figure 5] It is a diagram illustrating a case where the compositing position of the second image with respect to the first image is changed. [Figure 6] It is a diagram showing an example of generating M learning images from M first images. [Figure 7] It is a diagram illustrating a learning image generated by a second image including grid lines. [Figure 8] It is a diagram illustrating a learning image generated by a second image including a frame line. [Figure 9] It is a diagram illustrating a learning image generated by a second image including a frame line. [Figure 10] This figure illustrates a training image generated by a second image containing text. [Figure 11] This figure illustrates a training image generated by a second image containing a plain image. [Figure 12] This figure illustrates a second image that includes both a stamp image and a blank image. [Figure 13] This flowchart shows an example of the main processing steps performed in a learning image generation device. [Figure 14] This flowchart shows an example of a detailed processing procedure for the first image generation process. [Figure 15] This flowchart shows an example of a detailed processing procedure for the second image generation process. [Figure 16] This flowchart shows an example of a detailed processing procedure for image synthesis. [Modes for carrying out the invention]

[0029] Preferred embodiments of the present invention will be described in detail below with reference to the drawings. In the embodiments described below, elements common to all are denoted by the same reference numerals, and redundant explanations of these elements will be omitted.

[0030] Figure 1 shows an example configuration of a character recognition system according to one embodiment of the present invention. This character recognition system is a system that detects character strings contained in an image by AI-based character recognition processing, and comprises an information processing device 1 composed of a personal computer or the like, and an image processing device 2 composed of an MFP (Multifunction Peripheral) or the like, and these can communicate with each other via a network 4.

[0031] The information processing device 1 functions as a learning image generation device 5 by executing the program 17 described later. Figure 2 is a diagram illustrating the hardware configuration of the information processing device 1. As shown in Figure 2, the information processing device 1 includes a control unit 10, a display unit 11, an operation unit 12, a communication interface 13, and a storage unit 14.

[0032] The control unit 10 includes a CPU 15 and a memory 16, and controls the operation of each part. The CPU 15 is a hardware processor that executes the program 17 stored in the storage unit 14. The memory 16 is a volatile storage device that temporarily stores images, data, etc., generated as a result of the CPU 15 executing the program 17. The CPU 15 uses the memory 16 as a work area to perform various processes described later.

[0033] The display unit 11 displays various screens and images. For example, the display unit 11 is composed of a liquid crystal display or the like. The operation unit 12 accepts input operations from the user. For example, the operation unit 12 is composed of a keyboard, mouse, touch panel, handwriting input pad, or the like.

[0034] The communication interface 13 is an interface for connecting the information processing device 1 to the network 4. The information processing device 1 can communicate with the image processing device 2 via this communication interface 13.

[0035] The storage unit 14 is a non-volatile storage device composed of a hard disk drive (HDD) or a solid-state drive (SSD), etc. The storage unit 14 stores a program 17 executed by the CPU 15 and various image data 18. For example, the image data 18 includes various photographs, illustrations, graphs, tables, line drawings, plain images (monochromatic images), color images, etc.

[0036] Returning to Figure 1, the image processing device 2 includes a scanner 3 that optically reads images from original documents, such as paper documents, set by the user and generates image data. In other words, the scanner 3 is an image reading device.

[0037] Furthermore, the image processing device 2 implements an AI-based character recognition function, enabling it to recognize character strings contained in image data generated by the scanner 3 and convert the recognized strings into text data. In other words, the image processing device 2 functions as a character recognition device that performs AI-based character recognition processing. Therefore, the image processing device 2 can improve its AI-based character recognition rate by pre-training it with a variety of training images containing character strings.

[0038] The information processing device 1 functions as a learning image generation device 5, generating a large number of training images necessary to improve the character recognition rate in the image processing device 2, and providing them to the image processing device 2. Hereinafter, the information processing device 1 will be described as the learning image generation device 5.

[0039] Figure 3 is a block diagram illustrating the functional configuration of the learning image generation device 5 and the image processing device 2, respectively.

[0040] The image processing device 2 has a character recognition unit 30 that performs character recognition processing using AI. The character recognition unit 30 has an AI determination unit 31 that uses AI to determine whether or not the image components contained in the input image data are character strings. Furthermore, the AI ​​determination unit 31 has a machine learning unit 32. The machine learning unit 32 constructs a neural network model for recognizing character strings by performing machine learning such as deep learning using training images provided by the training image generation device 5.

[0041] The control unit 10 of the learning image generation device 5 functions as an input unit 20, a first image generation unit 21, a second image generation unit 22, an image synthesis unit 23, and an output unit 24, as the CPU 15 executes the program 17 described above. By activating these units, the control unit 10 generates a large number of training images for the machine learning unit 32 of the image processing device 2 to perform machine learning. For example, when a single string is input by the user, the learning image generation device 5 generates multiple training images containing that single string and provides them to the image processing device 2. Therefore, by using the learning image generation device 5, the user can input a large number of training images into the image processing device 2 at once, and efficiently perform machine learning on the machine learning unit 32.

[0042] The input unit 20 receives the input of a string Dt from the operation unit 12. That is, the input unit 20 identifies the string Dt entered by the user based on the user's operation on the operation unit 12 and accepts the identified string Dt. For example, if the user enters the string Dt on the keyboard, the input unit 20 accepts the string Dt entered by the user as text data. Once the string Dt is accepted as text data, the string Dt entered into the input unit 20 becomes a known string.

[0043] Furthermore, when a user handwrites a string Dt onto the handwriting input pad, the input unit 20 receives that string Dt as image data. In this case, since it is unknown what string is contained in the image data, the string Dt received by the input unit 20 is not known. Therefore, when a string Dt is handwritten onto the handwriting input pad, the input unit 20 further accepts text data corresponding to the handwritten string Dt from the user via a keyboard or the like, and obtains text data corresponding to the handwritten string Dt. This text data makes the handwritten string Dt a known string.

[0044] When the input unit 20 receives a known string Dt as input, it outputs the string Dt to the first image generation unit 21 and the second image generation unit 22. For example, if the input unit 20 receives a string Dt entered by the user as text data, it outputs that text data as the string Dt to the first image generation unit 21 and the second image generation unit 22, respectively. Also, if the input unit 20 receives a string Dt entered by the user as image data, it outputs that image data as the string Dt to the first image generation unit 21.

[0045] Furthermore, the input unit 20 uses text data corresponding to a known string Dt as the correct answer data Da, and outputs this correct answer data Da to the first image generation unit 21 and the second image generation unit 22. For example, if the input unit 20 receives a string Dt entered by a user as text data, the correct answer data Da will be the same data as the string Dt. In this case, the input unit 20 outputs the correct answer data Da to the output unit 24. On the other hand, if the input unit 20 receives a string Dt entered by a user as image data, the correct answer data Da will be text data different from the string Dt. In this case, the input unit 20 outputs the correct answer data Da, expressed as text data, to the second image generation unit 22 and the output unit 24, respectively.

[0046] The first image generation unit 21 obtains the string Dt output from the input unit 20 and generates a first image G1 containing the string Dt. For example, the first image generation unit 21 generates a first image G1 containing the string Dt by converting the string Dt into image data. If the string Dt is text data, the first image generation unit 21 converts the text data into image data such as JPEG or a bitmap and generates the first image G1. Also, if the string Dt is handwritten image data, the first image generation unit 21 converts the image data into image data such as JPEG or a bitmap and generates the first image G1.

[0047] When the first image generation unit 21 generates the first image G1, it generates processing parameters for processing the string Dt and processes the string Dt based on these processing parameters. The first image generation unit 21 then generates the first image G1 containing the processed string Dt. For example, if the string Dt is text data, the processing parameters may include font (typeface), size, thickness, color, density, arrangement direction, and tilt angle. If the string Dt is image data, the processing parameters may include scaling, reduction, image rotation angle, and tilt angle. The first image generation unit 21 processes the string Dt based on these processing parameters and generates the first image G1 containing the processed string Dt.

[0048] Furthermore, the first image generation unit 21 generates multiple processing parameters and processes the string Dt based on each of these processing parameters to generate multiple strings with different characteristics, thereby generating multiple first images G1, each containing a string Dt with a different characteristic. In other words, the first image generation unit 21 can generate M (where M is a natural number of 2 or more) first images G1 by applying different processing to the string Dt. The first image generation unit 21 then outputs the multiple first images G1, each containing a string Dt with a different characteristic, to the image synthesis unit 23.

[0049] The second image generation unit 22 generates a second image G2 for synthesis with the first image G1. This second image generation unit 22 generates a second image G2 that includes disturbance elements as image components that are different from the string Dt that the machine learning unit 32 will recognize when performing machine learning. This second image generation unit 22 can also generate multiple second images G2.

[0050] For example, the second image generation unit 22 generates a stamp impression image based on the text data of the string Dt, and generates a second image G2 that includes the stamp impression image. Specifically, the second image generation unit 22 converts the font of the characters included in the text data to a stamp impression-style font, and generates a stamp impression image by arranging the string in the stamp impression-style font inside a round frame or a square frame, and generates a second image G2 that includes the stamp impression image.

[0051] Furthermore, the second image generation unit 22 can also generate a second image G2 that includes a grid or border for compositing with or around a string contained in the first image G1. The second image generation unit 22 can also generate a second image G2 that includes noise components. Additionally, the second image generation unit 22 can read image data 18 stored in the storage unit 14 and generate a second image G2 that includes any image, such as a color image or a plain image, based on that image data 18. Furthermore, the second image generation unit 22 can generate a second image G2 that includes a string different from string Dt. Moreover, the second image generation unit 22 can generate a second image G2 that includes an image of a string different from string Dt that has been horizontally flipped.

[0052] When the second image generation unit 22 generates the second image G2 as described above, it outputs the second image G2 to the image synthesis unit 23.

[0053] The image synthesis unit 23 generates a training image G3 by combining the first image G1 and the second image G2. Figure 4 is a diagram illustrating the concept of the synthesis process performed by the image synthesis unit 23. Figure 4(a) shows an example of the first image G1. This first image G1 contains the string 41 "aiueokakikukeko". This string 41 is a processed string of the string Dt mentioned above. Figure 4(b) shows an example of the second image G2. This second image G2 contains a seal impression image 51. The image synthesis unit 23 generates a training image G3 as shown in Figure 4(c) by combining the first image G1 shown in Figure 4(a) and the second image G2 shown in Figure 4(b). At this time, the image synthesis unit 23 generates a training image G3 by superimposing the seal impression image 51 contained in the second image G2 onto the string 41 of the first image G1. As a result, the seal impression image 51 becomes a noise component when recognizing the string 41.

[0054] Furthermore, when the image synthesis unit 23 synthesizes the second image G2 with the first image G1, it can generate multiple training images G3 by changing synthesis parameters such as the transparency of the first image G1 and the second image G2, the synthesis position of the second image G2 relative to the first image G1, the placement angle of the second image relative to the first image, and the distortion of at least one of the first image G1 and the second image G2. Distortion refers to the distortion of a rectangular image into a parallelogram image during image synthesis.

[0055] Figure 5 illustrates the case where the composite position of the second image G2 with respect to the first image G1 is changed. Figure 5(a) shows an example where the seal impression image 51 contained in the second image G2 is composited to the beginning of the string 41 contained in the first image G1. Figure 5(b) shows an example where the seal impression image 51 contained in the second image G2 is composited to the middle of the string 41 contained in the first image G1. Furthermore, Figure 5(c) shows an example where the seal impression image 51 contained in the second image G2 is composited to the end of the string 41 contained in the first image G1. By changing the composite position of the seal impression image 51 with respect to the string 41 in this way, it is possible to change the part that reduces the readability of the string 41, and thus generate a variety of training images G3. Note that although Figure 5 illustrates the case where the second image G2 contains the seal impression image 51, the second image G2 may contain an image other than the seal impression image 51.

[0056] Furthermore, if the first image generation unit 21 generates M first images G1 containing strings 41 of different characteristics, the image synthesis unit 23 can generate at least M training images G3 by synthesizing a second image G2 with each of the multiple first images G1.

[0057] Figure 6 shows an example of generating M training images G3 from M first images G1. For example, as shown in Figure 6, if the first image generation unit 21 generates M first images G1 containing different types of strings 41, the image synthesis unit 23 synthesizes a second image G2 with these M first images G1 to generate M training images G3. In this case, M training images G3 are generated from a single string input to the input unit 20. Furthermore, by changing the synthesis parameters when synthesizing the second image G2 with the first image G1, the image synthesis unit 23 can generate multiple training images G3 from a single first image G1. In other words, the image synthesis unit 23 can generate M or more training images G3 from M first images G1. Therefore, the image synthesis unit 23 can generate many training images G3 at once for the machine learning unit 32 to perform machine learning.

[0058] Furthermore, when the second image generation unit 22 generates multiple second images G2, the image synthesis unit 23 can generate even more training images G3 by synthesizing one of the multiple second images G2 with each of the multiple first images G1. Therefore, the image synthesis unit 23 can generate a large number of training images G3 at once from a single string input by the user.

[0059] Next, we will explain the case in which the second image generation unit 22 generates a second image G2 including a grid line. Figure 7(a) shows the second image G2 generated by the second image generation unit 22. This second image G2 includes a grid line 52 that is added to the bottom of the string 41 contained in the first image G1. When the second image generation unit 22 generates this second image G2, it identifies the region (position) in the first image G1 that contains the string 41 and generates the second image G2 with the grid line 52 added to the position corresponding to the bottom of the string 41.

[0060] When a second image G2, as shown in Figure 7(a), is generated, the image synthesis unit 23 synthesizes the second image G2 with the first image G1 to generate a training image G3, as shown in Figure 7(b). As a result, the training image generation device 5 can generate a training image G3 in which a line 52, such as an underline, is added to the string 41.

[0061] Furthermore, when the image synthesis unit 23 synthesizes the second image G2 in Figure 7(a) with the first image G1, it is also possible to generate a training image G3 in which a grid line 52 is superimposed on the string 41, as shown in Figure 7(c), by changing the synthesis position or arrangement angle of the second image G2.

[0062] Next, we will describe the case in which the second image generation unit 22 generates a second image G2 including a border. Figure 8(a) shows the second image G2 generated by the second image generation unit 22. This second image G2 includes a single border 53 that surrounds the entire string 41 contained in the first image G1. When the second image generation unit 22 generates this second image G2, it identifies the region (position) in the first image G1 that contains the string 41, and generates the second image G2 with a border 53 added to the surrounding position that surrounds the entire string 41.

[0063] When a second image G2, as shown in Figure 8(a), is generated, the image synthesis unit 23 synthesizes the second image G2 with the first image G1 to generate a training image G3, as shown in Figure 8(b). As a result, the training image generation device 5 can generate a training image G3 with a frame 53 surrounding the entire string 41.

[0064] Furthermore, when the image synthesis unit 23 synthesizes the second image G2 in Figure 8(a) with the first image G1, it is also possible to generate a training image G3 in which a portion of the border line 53 overlaps the text string 41, as shown in Figure 8(c), by changing the synthesis position or arrangement angle of the second image G2.

[0065] Next, we will describe the case in which the second image generation unit 22 generates a second image G2 that includes a different border than that shown in Figure 8. Figure 9(a) shows the second image G2 generated by the second image generation unit 22. This second image G2 includes a border 54 that surrounds each character of the string 41 contained in the first image G1. When the second image generation unit 22 generates this second image G2, it identifies the region (position) in the first image G1 that contains the string 41, and generates the second image G2 with the border 54 added to the position that surrounds each character of the string 41.

[0066] When a second image G2, as shown in Figure 9(a), is generated, the image synthesis unit 23 synthesizes the second image G2 with the first image G1 to generate a training image G3, as shown in Figure 9(b). As a result, the training image generation device 5 can generate a training image G3 with a frame 54 surrounding each character of the string 41.

[0067] Furthermore, when the image synthesis unit 23 synthesizes the second image G2 in Figure 9(a) with the first image G1, it is also possible to generate a training image G3 in which each character of the string 41 is tilted relative to its individual border, or in which the border 54 overlaps at least some of the characters, as shown in Figure 9(c), by changing the synthesis position or arrangement angle of the second image G2.

[0068] Next, we will explain the case in which the second image generation unit 22 generates a second image G2 that contains a string different from the string 41. Figure 10(a) shows the second image G2 generated by the second image generation unit 22. This second image G2 contains a string 55 that is different from the string 41 contained in the first image G1. When the second image generation unit 22 generates this second image G2, it generates an arbitrary string 55 different from the string 41 based on the text data of the string 41, and generates the second image G2 to which that string 55 is added.

[0069] However, if the second image G2 contains the string 55, there is a possibility that the string 55 will be targeted for recognition by the character recognition process. To prevent this, it is preferable for the second image generation unit 22 to horizontally flip the string 55 when generating it. In this case, the second image generation unit 22 may not horizontally flip the entire string 55, but rather horizontally flip each character contained in the string 55 one by one. This makes it possible to generate an image of the string 55 in a form where it is reversed, and prevents it from being misrecognized as a string to be recognized in the character recognition process. Figures 10(a), (b), and (c) illustrate the case where the string 55 is not horizontally flipped.

[0070] When a second image G2 containing the string 55 is generated, the image synthesis unit 23 synthesizes the second image G2 with the first image G1 to generate a training image G3 as shown in Figure 10(b). As a result, the training image generation device 5 can generate a training image G3 in which a different string 55 is attached to the string 41.

[0071] Furthermore, when the image synthesis unit 23 synthesizes a second image G2, as shown in Figure 10(a), onto the first image G1, it is also possible to generate a training image G3, as shown in Figure 10(c), in which each character of the string 41 is tilted relative to its individual border, or in which the border 54 overlaps at least some of the characters, by changing the synthesis position or arrangement angle of the second image G2. In addition, when the image synthesis unit 23 synthesizes a second image G2 onto the first image G1, it is possible to display the string 55 in a translucent manner faintly by, for example, setting the transparency of the second image G2 to a value smaller than the transparency of the first image.

[0072] Next, we will describe the case in which the second image generation unit 22 generates a second image G2 that includes an image based on the image data 18. Figure 11 is a diagram illustrating a plain image as an example of an image based on the image data 18. Figure 11(a) shows the second image G2 generated by the second image generation unit 22. This second image G2 includes a plain image 56 based on the image data 18. The plain image 56 is an image in which the color and density are constant within the image plane. The second image generation unit 22 reads the image data 18 of the plain image 56 from among the multiple image data 18 stored in the storage unit 14 and generates the second image G2 based on that image data 18. At this time, the second image generation unit 22 may generate multiple second images G2 by changing the color or density of the plain image 56 to a different color or density.

[0073] When a second image G2 containing a blank image 56 is generated, the image synthesis unit 23 synthesizes the second image G2 with the first image G1 to generate a training image G3 as shown in Figure 11(b). As a result, the training image generation device 5 can generate a training image G3 in which the string 41 is synthesized on a region with the blank image 56 as the background.

[0074] Furthermore, when the image synthesis unit 23 synthesizes the second image G2, as shown in Figure 11(a), with the first image G1, it is also possible to generate a training image G3 with a different color or density of the background portion, which is a plain image 56, by changing the color or density of the second image G2, as shown in Figure 11(c). Note that although a plain image 56 is used as an example in Figure 11, the image included in the second image G2 is not limited to a plain image 56, but may be a photograph, illustration, graph, table, line drawing, or other color image.

[0075] Furthermore, the second image generation unit 22 may generate a second image G2 that includes two or more image components from the image data 18, such as the seal impression image 51, ruled lines 52, border lines 53, 54, text string 55, and blank image 56 described above. Figure 12 shows an example of a second image G2 that includes the seal impression image 51 and the blank image 56. The second image generation unit 22 can also generate a second image G2 that includes multiple image components, as shown in Figure 12. In addition, the second image generation unit 22 may randomly place dirt images 57 that mimic stains on the original document in the second image G2, as shown in Figure 12. This allows the generation of a training image G3 in which the dirt images 57 overlap the text string 41, so that the machine learning unit 32 can recognize the text string 41 even when the original document is dirty by learning from the training image G3.

[0076] The image synthesis unit 23 generates a training image G3 by combining the first image G1 and the second image G2, and then outputs the training image G3 to the output unit 24. As described above, the image synthesis unit 23 generates multiple training images G3 from a single string Dt. Therefore, the image synthesis unit 23 outputs the multiple training images G3 generated from the string Dt to the output unit 24.

[0077] When the output unit 24 acquires multiple training images G3 from the image synthesis unit 23, it outputs the multiple training images G3 to the machine learning unit 32 via the communication interface 13. At this time, the output unit 24 also outputs the correct data Da of the string Dt to the machine learning unit 32 along with the training images G3. As a result, the machine learning unit 32 recognizes that the string 41 contained in the multiple training images G3 output from the training image generator 5 is the string indicated by the correct data Da. Therefore, the machine learning unit 32 constructs a neural network model by performing machine learning so that it can recognize the string indicated by the correct data Da from each of the multiple training images G3. Consequently, the character recognition rate when character recognition processing is actually performed on the image data generated by the scanner 3 can be improved.

[0078] Next, an example of a processing procedure performed in the learning image generation device 5 described above will be explained. Figures 13 to 16 are flowcharts showing an example of a processing procedure performed in the learning image generation device 5. This process is performed in the learning image generation device 5 by the CPU 15 of the control unit 10 executing program 17.

[0079] As shown in Figure 13, when the learning image generation device 5 starts this process, the control unit 10 activates the input unit 20 and inputs the string Dt (step S1). That is, the learning image generation device 5 inputs the string Dt based on the user's operation on the operation unit 12. For example, the learning image generation device 5 inputs the string Dt as text data or image data.

[0080] The learning image generation device 5 generates correct answer data Da corresponding to the input string Dt (step S2). If the string Dt is input as text data, the correct answer data Da will be the same as that text data. If the string Dt is handwritten image data, the input unit 20 generates correct answer data Da, which is text data, based on user operations performed on a keyboard or the like.

[0081] Next, the learning image generation device 5 activates the first image generation unit 21 and executes the first image generation process (step S3). Figure 14 is a flowchart showing an example of a detailed processing procedure for the first image generation process (step S3). When the learning image generation device 5 starts the first image generation process, it obtains the string Dt (step S10) and sets the number of first images G1 to be generated M (M is a natural number greater than or equal to 2) (step S11). The number of images to be generated M can be set in advance by the user, for example. Then, the learning image generation device 5 initializes the variable i to 1 (step S12).

[0082] Next, the learning image generation device 5 generates processing parameters for processing the string Dt (step S13), and processes the string Dt based on these processing parameters (step S14). As a result, the display manner of the string Dt changes according to the processing parameters, and the string 41 included in the first image G1 is generated. Then, the learning image generation device 5 generates the first image G1 by converting the image containing the processed string 41 from the string Dt into predetermined image data (step S15).

[0083] Next, the learning image generator 5 determines whether the variable i is equal to the number of images to be generated M (step S16). If the variable i is less than the number of images to be generated M (NO in step S16), the learning image generator 5 adds 1 to the variable i (step S17) and repeats the process in steps S13 to S15. At this time, the processing parameters generated in step S13 will be different from the processing parameters generated previously. As a result, the first image G1 containing the string 41, which has been processed in a different manner than before, is repeatedly generated. By repeating the process in steps S13 to S15 M times, the learning image generator 5 can generate M first images G1, each containing the string 41 which has been processed in a different way. After that, if it is determined in step S16 that the variable i is equal to the number of images to be generated M (YES in step S16), the first image generation process ends and returns to the flowchart in Figure 13.

[0084] Next, the learning image generation device 5 activates the second image generation unit 22 and executes the second image generation process (step S4). Figure 15 is a flowchart showing an example of a detailed processing procedure for the second image generation process (step S3). When the learning image generation device 5 starts the second image generation process, it determines whether or not to generate the seal impression image 51 (step S20). For example, whether or not to generate the seal impression image 51 is set in advance by the user. If the learning image generation device 5 determines to generate the seal impression image 51 (YES in step S20), it obtains the string Dt (step S21), processes the string Dt, and generates the seal impression image 51 (step S22). However, if the string Dt is image data obtained by handwriting input, it obtains the correct answer data Da instead of the string Dt, and generates the seal impression image 51 based on the correct answer data Da. Then the learning image generation device 5 generates the second image G2 which includes the generated seal impression image 51 (step S23). At this time, the learning image generation device 5 may generate multiple second images G2 by changing the position in which the seal impression image 51 is included in the second image G2. If it is determined not to generate the seal impression image 51 (NO in step S20), the processes in steps S21 to S23 are skipped.

[0085] Next, the learning image generator 5 determines whether or not to generate a second image G2 with the ruled lines 52 (step S24). For example, whether or not to generate a second image G2 with the ruled lines 52 is predetermined by the user. If the learning image generator 5 determines to generate a second image G2 with the ruled lines 52 (YES in step S24), it analyzes the region containing the string 41 in the first image G1 (step S25) and generates an image representing the ruled lines 52 (step S26). Then the learning image generator 5 generates a second image G2 that includes the generated ruled lines 52 (step S27). At this time, the learning image generator 5 may generate multiple second images G2 by changing the position and arrangement angle in which the ruled lines 52 are included in the second image G2. If it determines not to generate a second image G2 with the ruled lines 52 (NO in step S24), the processes in steps S25 to S27 are skipped.

[0086] Next, the learning image generation device 5 determines whether or not to generate a second image G2 with borders 53 and 54 (step S28). For example, whether or not to generate a second image G2 with borders 53 and 54 is predetermined by the user. If the learning image generation device 5 determines to generate a second image G2 with borders 53 and 54 (YES in step S28), it analyzes the region containing the string 41 in the first image G1 (step S29) and generates an image representing the border 53 or border 54 (step S30). Then the learning image generation device 5 generates a second image G2 including the generated borders 53 and 54 (step S31). At this time, the learning image generation device 5 may generate multiple second images G2 by changing the position and arrangement angle in which the borders 53 and 54 are included in the second image G2. If it is determined that a second image G2 with borders 53 and 54 should not be generated (NO in step S28), then the processes in steps S29 to S31 are skipped.

[0087] Next, the learning image generator 5 determines whether or not to generate a second image G2 with a string 55 that is different from string Dt (step S32). For example, whether or not to generate a second image G2 with string 55 is predetermined by the user. If the learning image generator 5 determines to generate a second image G2 with string 55 (YES in step S32), it generates a string 55 that is different from string Dt (step S33) and generates an image containing string 55 (step S34). At this time, it is preferable for the learning image generator 5 to generate an image in which the entire string 55 is horizontally flipped. Alternatively, it may generate an image in which each character contained in string 55 is horizontally flipped. Then the learning image generator 5 generates a second image G2 containing the image of string 55 (step S35). At this time, the learning image generator 5 may generate multiple second images G2 by changing the position and arrangement angle in which the image of string 55 is included in the second image G2. If it is determined that a second image G2 with the string 55 attached is not to be generated (NO in step S32), then the processing in steps S33 to S35 is skipped.

[0088] Next, the learning image generation device 5 determines whether or not to acquire image data 18 and generate a second image G2 (step S36). For example, whether or not to acquire image data 18 and generate a second image G2 is predetermined by the user. If the learning image generation device 5 determines to acquire image data 18 and generate a second image G2 (YES in step S36), it reads and acquires the image data 18 from the storage unit 14 (step S38). At this time, the learning image generation device 5 acquires the image data 18 specified by the user. Also, the number of image data 18 read and acquired from the storage unit 14 is not limited to one, but may be multiple. Once the learning image generation device 5 acquires the image data 18, it generates processing parameters for processing and adjusting the color and density (step S38), and processes the image based on the image data 18 based on those processing parameters (step S39). Then the learning image generation device 5 generates a second image G2 that includes the processed image (step S40). As a result, for example, a second image G2 including the plain image 56 described above is generated. However, the images included in the second image G2 are not limited to the plain image 56. Furthermore, the learning image generation device 5 may generate multiple processing parameters for processing the image data 18, and generate multiple second images G2 by processing the image data 18 based on each of these processing parameters. If it is determined that the image data 18 will not be acquired and a second image G2 will not be generated (NO in step S36), the processing in steps S37 to S40 is skipped.

[0089] Next, the learning image generation device 5 determines whether or not to add noise components, such as the dirty image 57, to the second image G2 (step S41). For example, whether or not to generate noise components in the second image G2 is predetermined by the user. If the learning image generation device 5 determines to add noise components, such as the dirty image 57, to the second image G2 (YES in step S41), it reads the second image G2 generated so far (step S42), adds noise components to the second image G2 (step S43), and generates a new second image G2 (step S44). If it determines not to add noise components to the second image G2 (NO in step S41), the processes in steps S42 to S44 are skipped. This completes the second image generation process, and the process returns to the flowchart in Figure 13.

[0090] Next, the learning image generation device 5 activates the image synthesis unit 23 and performs image synthesis processing (step S5). Figure 16 is a flowchart showing an example of a detailed processing procedure for the image synthesis processing (step S5). When the learning image generation device 5 starts the image synthesis processing, it reads out the first image G1 (step S50). The learning image generation device 5 also reads out the second image G2 (step S51). Then, the learning image generation device 5 determines the synthesis parameters when synthesizing the second image G2 with the first image G1 (step S52). This determines the transparency of the first image G1 and the second image G2, the synthesis position of the second image G2 relative to the first image G1, the placement angle of the second image relative to the first image, the distortion of the first image G1 and the second image G2, and so on. The learning image generation device 5 applies the synthesis parameters determined in step S52 to synthesize the second image G2 with the first image G1 (step S53) to generate the learning image G3 (step S54).

[0091] Next, the training image generator 5 determines whether or not to generate another training image G3 without changing the first image G1 and the second image G2 (step S55). If it is to generate another training image G3 (YES in step S55), the training image generator 5 repeats the process in steps S52 to S54. At this time, the composite parameters generated in step S52 will be different from the composite parameters generated previously. This makes it possible to generate a different training image G3 from the previous one without changing the first image G1 and the second image G2. For example, if multiple variations of the composite parameters are prepared in advance, it is possible to automatically generate training images G3 to which each of these multiple variations is applied.

[0092] Furthermore, if no other training image G3 is generated (NO in step S55), the training image generator 5 determines whether or not another second image G2 has been generated (step S56). If another second image G2 has been generated (YES in step S56), the training image generator 5 repeats the processes in steps S51 to S55. This process generates the training image G3 by combining the first image G1 read in step S50 with the second image G2.

[0093] Furthermore, if it is determined that no other second image G2 exists (NO in step S56), the learning image generator 5 determines whether or not another first image G1 has been generated (step S57). If another first image G1 has been generated (YES in step S57), the learning image generator 5 repeats the processes in steps S50 to S56. As a result, another first image G1 is read out, and the second image G2 is combined with that first image G1 to generate a learning image G3. The processes in steps S51 to S56 are then performed for all first images G1, and if it is determined that no other first image G1 exists (NO in step S57), the image synthesis process ends.

[0094] Returning to the flowchart in Figure 13, the training image generation device 5 then activates the output unit 24 and outputs the training image G3 and the ground truth data Da generated in the above process as a single dataset to the machine learning unit 32 (step S6). Note that the training image generation device 5 is not limited to directly outputting the dataset containing the training image G3 and the ground truth data Da to the machine learning unit 32 of the image processing device 2; it may also output and save the dataset to a storage device such as a USB memory. This completes the processing by the training image generation device 5.

[0095] As described above, the learning image generation device 5 of this embodiment includes a first image generation unit 21 that generates a first image G1 containing a known string Dt as input, a second image generation unit 22 that generates a second image G2 for synthesis with the first image G1, an image synthesis unit 23 that generates a learning image G3 by synthesizing the first image G1 and the second image G2, and an output unit 24 that outputs the learning image G3 and the correct data Da of the string Dt. Therefore, the learning image generation device 5 can automatically generate a learning image G3 that allows the machine learning unit 32 to perform machine learning for character recognition simply by inputting a known string Dt. In particular, for the user, there is the advantage that a large number of learning images G3 can be easily prepared because a learning image G3 can be easily created simply by inputting a known string Dt into the learning image generation device 5.

[0096] Furthermore, the learning image generation device 5 of this embodiment can generate multiple first images G1 by applying various processing to a single string Dt, and then generate multiple learning images G3 by combining these multiple first images G1 with a second image G2. For example, the image data generated by the scanner 3 contains strings of various characteristics in terms of font, size, thickness, color, density, arrangement direction, and tilt angle. Also, if noise is introduced when the scanner 3 reads the document, the clarity of the strings may decrease. Therefore, by having the learning image generation device 5 generate strings 41 of various characteristics from a single string Dt, and generate multiple learning images G3 containing these strings 41, the machine learning unit 32 can learn strings of various characteristics in terms of font, size, thickness, color, density, arrangement direction, and tilt angle when performing machine learning using the learning images G3, and can also distinguish noise introduced when the scanner 3 reads the document. Thus, by using the learning images G3 generated by the learning image generation device 5, it is possible to ensure the robustness of the AI ​​and improve the character recognition rate by the AI.

[0097] Furthermore, the learning image generation device 5 of this embodiment can generate not only the first image G1, but also multiple second images G2 when generating the second image G2. For example, if M first images G1 are generated and N second images G2 are generated (where N is a natural number greater than or equal to 2), the learning image generation device 5 can generate M × N learning images G3 from a single string Dt. In this case, since the user only needs to input a single string Dt, there is the advantage that a large number of learning images G3 can be obtained at once with simple operation.

[0098] Preferred embodiments of the present invention have been described above. However, the present invention is not limited to those described in the above embodiments, and various modifications are applicable.

[0099] For example, in the above embodiment, an example was described in which an image processing device 2, composed of an MFP or the like, functions as a character recognition device that performs character recognition processing using AI. However, the character recognition device may be implemented as a separate device from the image processing device 2. Also, in the above embodiment, as an example, a case was described in which the character recognition unit 30 performs character recognition processing on image data generated by the scanner 3. However, the image data that the character recognition unit 30 processes for character recognition is not necessarily limited to image data generated by the scanner 3, but may be image data generated by another device.

[0100] Furthermore, in the above embodiment, an example was described in which an information processing device 1, which is composed of a personal computer or the like, functions as a learning image generation device 5. However, the learning image generation device 5 is not necessarily limited to being implemented by an information processing device 1 such as a personal computer. For example, the learning image generation device 5 may function in an image processing device 2 composed of an MFP or the like, or it may function in other information devices. Also, the learning image generation device 5 may be configured integrally with a character recognition device that performs character recognition processing using AI.

[0101] Furthermore, in the above embodiment, the example given was that the program 17 executed by the CPU 15 of the control unit 10 in the learning image generation device 5 is pre-stored in the storage unit 14. However, it is not limited to this, and the program 17 may be recorded on an external computer-readable recording medium. Also, the program 17 may be installed on the information processing device 1 via a network such as the Internet. [Explanation of symbols]

[0102] 1. Information Processing Device 2 Image Processing Device 5. Image generation device for learning 17 Programs 20 Input section 21 First Image Generation Unit 22 Second Image Generation Unit 23 Image Synthesis Unit 24 Output section

Claims

1. A first image generation unit that takes a known string as input and generates a first image containing the string, A second image generation unit generates a second image for synthesis with the first image, An image synthesis unit that generates a training image by combining the first image and the second image, An output unit that outputs the training image and the correct data for the string, Equipped with, The second image generation unit generates the second image which includes the stamp image obtained by processing the string, The learning image generation device is characterized in that the image synthesis unit generates the learning image by overlaying the seal impression image onto the string of characters included in the first image and synthesizing them.

2. A first image generation unit that takes a known string as input and generates a first image containing the string, A second image generation unit generates a second image for synthesis with the first image, An image synthesis unit that generates a training image by combining the first image and the second image, An output unit that outputs the training image and the correct data for the string, Equipped with, The learning image generation device is characterized in that the second image generation unit generates the second image which contains a string different from the string mentioned above.

3. A first image generation unit that takes a known string as input and generates a first image containing the string, A second image generation unit generates a second image for synthesis with the first image, An image synthesis unit that generates a training image by combining the first image and the second image, An output unit that outputs the training image and the correct data for the string, Equipped with, The learning image generation device is characterized in that the second image generation unit generates the second image which includes an image which is a horizontally flipped version of a string different from the string.

4. A first image generation unit that takes a known string as input and generates a first image containing the string, A second image generation unit generates a second image for synthesis with the first image, An image synthesis unit that generates a training image by combining the first image and the second image, An output unit that outputs the training image and the correct data for the string, Equipped with, The image synthesis unit generates a plurality of training images by changing the parameters used when synthesizing the first image and the second image. The output unit outputs a plurality of the training images, The learning image generation device is characterized in that the parameters include the transparency of the first image and the second image, respectively.

5. The learning image generation device according to any one of claims 1 to 4, characterized in that the first image generation unit generates the first image by converting the string into image data.

6. The learning image generation device according to claim 5, characterized in that the output unit outputs text data representing the string as the correct answer data.

7. The first image generation unit generates M (where M is a natural number of 2 or more) first images by applying different processing to each of the strings when converting the string into image data. The learning image generation apparatus according to claim 6, characterized in that the second image generation unit generates at least M learning images by combining the second image with each of the M first images.

8. The learning image generation apparatus according to claim 4, characterized in that the parameter includes the position of the second image being combined with the first image.

9. The learning image generation apparatus according to claim 4, characterized in that the parameter further includes the arrangement angle of the second image relative to the first image.

10. A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, It has, The second image generation step generates the second image which includes the stamp image obtained by processing the string, The method for generating a learning image is characterized in that the image synthesis step generates a learning image by overlaying the seal impression image onto the string of characters included in the first image.

11. A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, It has, The method for generating learning images is characterized in that the second image generation step generates a second image containing a string different from the string mentioned above.

12. A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, It has, The method for generating learning images is characterized in that the second image generation step generates a second image which includes an image which is a horizontally flipped version of a string different from the string.

13. A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, It has, The image synthesis step generates a plurality of training images by changing the parameters used when synthesizing the first image and the second image. The output step outputs a plurality of the training images, A method for generating learning images, characterized in that the parameters include the transparency of the first image and the second image, respectively.

14. On the computer, A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, Make it run, The second image generation step generates the second image which includes the stamp image obtained by processing the string, The program is characterized in that the image synthesis step generates a learning image by overlaying the seal impression image onto the string of characters included in the first image.

15. On the computer, A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, Make it run, The program is characterized in that the second image generation step generates a second image which contains a string different from the string mentioned above.

16. On the computer, A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, Make it run, The program is characterized in that the second image generation step generates the second image which includes an image in which a different string from the string is flipped horizontally.

17. On the computer, A first image generation step involves inputting a known string and generating a first image containing the string, A second image generation step of generating a second image to be composited with the first image, Image synthesis step of generating a training image by combining the first image and the second image, An output step which outputs the training image and the correct data for the string, Make it run, The image synthesis step generates a plurality of training images by changing the parameters used when synthesizing the first image and the second image. The output step outputs a plurality of the training images, The program is characterized in that the parameters include the transparency of the first image and the second image, respectively.