Character recognition training data generation device, trained model manufacturing device, character recognition device, character recognition training data generation method, trained model manufacturing method, character recognition method, and program

The character recognition training data generation device addresses the challenge of recognizing characters in harsh environments by generating orthogonalized images and combining them with background images, improving recognition accuracy and reducing data requirements.

JP7859730B2Active Publication Date: 2026-05-15NEC SOLUTION INNOVATORS LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC SOLUTION INNOVATORS LTD
Filing Date
2022-06-21
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing character recognition technologies struggle with identifying characters printed on products in harsh environments due to lighting conditions and angle variations, making it difficult to capture and recognize characters using conventional OCR processing.

Method used

A character recognition training data generation device that includes a parallelized image generation unit, an extraction unit, an identification unit, and an image synthesis unit to generate orthogonalized images, extract and combine background and character images, and output composite character images as training data, utilizing machine learning to improve recognition accuracy.

Benefits of technology

The solution enables easy generation of training data for character recognition, reducing the amount of data required and shortening the data collection period, thereby enhancing the recognition of characters in challenging environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859730000001
    Figure 0007859730000001
  • Figure 0007859730000002
    Figure 0007859730000002
  • Figure 0007859730000003
    Figure 0007859730000003
Patent Text Reader

Abstract

To provide a character recognition teacher data generation apparatus, a learned model production apparatus, a character recognition apparatus, a character recognition teacher data generation method, a learned model production method, a character recognition method, and a program capable of easily generating teacher data for character recognition.SOLUTION: A character recognition teacher data generation apparatus 1 comprises: a facing image generation unit 2 for generating a facing image in which a correction target image is corrected to an image viewed from a direction perpendicular to a facing reference plane using posture information of an imaging terminal at the time of acquiring the correction target image; an extraction unit 3 for extracting a facing background image and a facing character image from the facing image; an identification unit 4 for identifying characters included in the facing character image based on reference character information; an image synthesis unit 5 for generating a composite character image by combining the facing background image and the facing character image; and a teacher data output unit 6 for outputting the composite character image and a combination of characters included in the composite character image as teacher data for character recognition.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , ,

[0001] The present invention relates to a teacher data generation device for character recognition, a trained model manufacturing device, a character recognition device, a teacher data generation method for character recognition, a trained model manufacturing method, a character recognition method, and a program.

Background Art

[0002] Regarding products flowing on a production line, it is necessary to manage the products individually in order to prevent mis-shipment, mix-up, etc. In this case, it is generally carried out to attach tags such as barcodes and RFID and manage them. On the other hand, for products with a harsh environment during product processing, such as steel products, the durability of the tags is insufficient and they cannot be attached. In this case, management is performed by directly printing characters on the products (for example, engraving printing, stamp printing, stencil spraying, etc.) (for example, Patent Document 1, etc.).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In such a situation, in order to obtain identification information of a product by image recognition, an image of the product is captured. However, engraving and the like are difficult to appear in a photograph due to lighting conditions. Also, it is not always possible to capture the product from the front. And the characters printed in this way are difficult to perform character recognition by normal optical character recognition (OCR) processing. This is the same in various fields where articles are managed by engraving and the like.

[0005] For character information that is difficult to recognize using conventional OCR processing, it is conceivable to create dictionary data using machine learning to read it from images. However, character recognition from captured images requires images of characters taken from various angles, which presents the challenge of requiring a huge amount of training data.

[0006] Therefore, the present invention aims to provide a character recognition training data generation device that can easily generate training data for character recognition. [Means for solving the problem]

[0007] To achieve the above objective, the character recognition training data generation apparatus of the present invention is It includes a parallelized image generation unit, an extraction unit, an identification unit, an image synthesis unit, and a training data output unit, The orthogonalized image generation unit uses the orientation information of the imaging terminal at the time of acquiring the image to be corrected to generate an orthogonalized image in which the image to be corrected is corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane. The extraction unit extracts an orthogonal background image and an orthogonal character image from the orthogonalized image. The identification unit identifies the characters contained in the aligned character image based on the reference character information. The image synthesis unit generates a composite character image by combining the aligned background image and the aligned character image. The aforementioned training data output unit outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition.

[0008] The trained model manufacturing apparatus of the present invention is It includes a training data acquisition unit and a trained model generation unit, The aforementioned training data acquisition unit acquires the character recognition training data output by the character recognition training data generation device of the present invention as character recognition training data, The pre-trained model generation unit generates a pre-trained model that, when an image containing the character to be recognized is input, outputs the characters contained in the image to be recognized, using machine learning with the character recognition training data.

[0009] The character recognition device of the present invention is Including the character recognition target image acquisition unit and the character recognition unit, The character recognition target image acquisition unit acquires a character recognition target image that includes the character to be recognized, The character recognition unit inputs the image to be recognized into the character recognition model and recognizes the characters contained in the image. The character recognition model is either a trained model generated by machine learning using training data generated by the character recognition training data generation device of the present invention, which outputs characters included in an image of a character to be recognized when an image of a character to be recognized including the character to be recognized is input, or a trained model manufactured by the trained model manufacturing device of the present invention.

[0010] The present invention's method for generating training data for character recognition is: The process includes a parallelization image generation step, an extraction step, a classification step, an image synthesis step, and a training data output step. The aforementioned orthogonalized image generation step generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction step extracts an orthogonal background image and an orthogonal character image from the orthogonalized image. The aforementioned identification step identifies the characters contained in the aligned character image based on the reference character information, The image synthesis step generates a composite character image by combining the orthogonalized background image and the orthogonalized character image. The aforementioned training data output step outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition.

[0011] The method for manufacturing trained models according to the present invention is: This includes a training data acquisition process and a pre-trained model generation process. The aforementioned training data acquisition step involves acquiring character recognition training data output by the character recognition training data generation method of the present invention as character recognition training data, The aforementioned pre-trained model generation step generates a pre-trained character recognition model that, when an image containing the character to be recognized is input, outputs the characters contained in the image to be recognized, using machine learning with the character recognition training data.

[0012] The character recognition method of the present invention is This includes a process for acquiring an image to be recognized and a process for recognizing the character, The character recognition target image acquisition step acquires a character recognition target image that includes the character to be recognized, The character recognition step involves inputting the image to be recognized into a character recognition model and recognizing the characters contained in the image. The character recognition model is either a trained model generated by machine learning using training data generated by the character recognition training data generation method of the present invention, which outputs characters included in a character recognition target image when the character recognition target image including the character recognition target is input, or a trained model manufactured by the trained model manufacturing method of the present invention.

[0013] The first program of the present invention includes a parallelized image generation procedure, an extraction procedure, an identification procedure, an image synthesis procedure, and a training data output procedure. The above-mentioned orthogonalized image generation procedure generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction procedure extracts an orthogonalized background image and an orthogonalized character image from the orthogonalized image, The aforementioned identification procedure identifies the characters contained in the aligned character image based on the reference character information, The image synthesis procedure generates a composite character image by combining the orthogonalized background image and the orthogonalized character image, The teacher data output procedure outputs the synthesized character image and the combination of characters included in the synthesized character image as teacher data for character recognition. It is a program for causing a computer to execute each of the above procedures.

[0014] The second program of the present invention includes a teacher data acquisition procedure and a learned model generation procedure. The teacher data acquisition procedure acquires, as teacher data for character recognition, the teacher data for character recognition output by the first program. The learned model generation procedure generates, by machine learning using the teacher data for character recognition, a learned model that is a character recognition model that outputs the characters included in the character recognition target image when the character recognition target image including the character recognition target is input. It is a program for causing a computer to execute each of the above procedures.

[0015] The third program of the present invention includes a character recognition target image acquisition procedure and a character recognition procedure. The character recognition target image acquisition procedure acquires a character recognition target image including a character recognition target. The character recognition procedure inputs the character recognition target image into a character recognition model and recognizes the characters included in the character recognition target. The character recognition model is a learned model generated by machine learning using the teacher data generated by the first program, which outputs the characters included in the character recognition target image when the character recognition target image including the character recognition target is input, or a learned model manufactured by the second program, and is a program for causing a computer to execute each of the above procedures.

Advantages of the Invention

[0016] According to the present invention, teacher data for character recognition can be easily generated.

Brief Description of the Drawings

[0017] [Figure 1]Figure 1 is a block diagram showing the configuration of an example of a character recognition training data generation device according to Embodiment 1. [Figure 2] Figure 2 is a block diagram showing an example of the hardware configuration of the character recognition training data generation device of Embodiment 1. [Figure 3] Figure 3 is a flowchart showing an example of processing in the character recognition training data generation device of Embodiment 1. [Figure 4] Figure 4 is a block diagram showing an example configuration of a parallelized image generation unit (image correction device) included in the character recognition training data generation device of Embodiment 1. [Figure 5] Figure 5 is a block diagram showing an example of the hardware configuration of the aligned image generation unit (image correction device) included in the character recognition training data generation device of Embodiment 1. [Figure 6] Figure 6 is a flowchart showing an example of processing in the aligned image generation unit (image correction device) included in the character recognition training data generation device of Embodiment 1. [Figure 7] Figure 7 is an explanatory diagram illustrating an example of the use of the aligned image generation unit (image correction device) included in the character recognition training data generation device of Embodiment 1. [Figure 8] Figure 8 is a block diagram showing the configuration of an example of a character recognition training data generation device according to Embodiment 3. [Figure 9] Figure 9 is a flowchart showing an example of processing in the character recognition training data generation device of Embodiment 3. [Figure 10] Figure 10 is a schematic diagram showing an example of an image processed by the character recognition training data generation device of Embodiment 3. [Figure 11] Figure 11 is a block diagram showing the configuration of an example of a trained model manufacturing apparatus according to Embodiment 4. [Figure 12] Figure 12 is a block diagram showing an example of the hardware configuration of the trained model manufacturing apparatus of Embodiment 4. [Figure 13] Figure 13 is a flowchart showing an example of processing in the trained model manufacturing apparatus of Embodiment 4. [Figure 14]Figure 14 is a block diagram showing the configuration of an example of a character recognition device according to Embodiment 5. [Figure 15] Figure 15 is a block diagram showing an example of the hardware configuration of the character recognition device according to Embodiment 5. [Figure 16] Figure 16 is a flowchart showing an example of processing in the character recognition device of Embodiment 5. [Modes for carrying out the invention]

[0018] Next, embodiments of the present invention will be described with reference to the drawings. The present invention is not limited to the following embodiments. In the following drawings, the same parts are denoted by the same reference numerals. Furthermore, unless otherwise specified, the descriptions of each embodiment can be used interchangeably with those of the others, and unless otherwise specified, the configurations of each embodiment can be combined.

[0019] [Embodiment 1] The character recognition training data generation device of this embodiment will be described with reference to Figure 1. Figure 1 is a block diagram showing the configuration of an example of the character recognition training data generation device 1 of this embodiment. As shown in Figure 1, the character recognition training data generation device 1 (hereinafter also referred to as "this device 1") includes a parallelized image generation unit 2, an extraction unit 3, an identification unit 4, an image synthesis unit 5, and a training data output unit 6. Although not shown, this device 1 may also include, for example, a storage unit.

[0020] The device 1 may be, for example, a single device including the aforementioned parts, or it may be a device in which the aforementioned parts can be connected via a communication network. Furthermore, the device 1 can be connected to external devices described later via a communication network. The communication network is not particularly limited and can use any known network, such as wired or wireless. Examples of communication networks include the Internet, WWW (World Wide Web), telephone lines, LAN (Local Area Network), SAN (Storage Area Network), DTN (Delay Tolerant Networking), LPWA (Low Power Wide Area), L5G (Local 5G), etc. Examples of wireless communication include Wi-Fi (registered trademark), Bluetooth (registered trademark), Local 5G, LPWA, etc. The wireless communication may be in the form of direct communication between devices (Ad Hoc communication), infrastructure communication, or indirect communication via an access point. The device 1 may, for example, be incorporated into a server as part of a system. Furthermore, the device 1 may be, for example, a personal computer (PC, e.g., desktop or notebook type), smartphone, tablet terminal, etc., on which the program of the present invention is installed. Moreover, the device 1 may be in a form such as cloud computing or edge computing, where, for example, at least one of the aforementioned parts is on a server and the other parts are on a terminal. As a specific example, the device 1 may be in a form in which a device equipped with an orthogonal image generation unit 2 and a device equipped with an extraction unit 3, an identification unit 4, an image synthesis unit 5, and a training data output unit 6 are connected via a communication network. In this case, the device 10 is also called, for example, a character recognition training data generation system. In this case, the device equipped with the orthogonal image generation unit 2 is also called, for example, an orthogonal image generation device or an image correction device. The orthogonal image generation device or image correction device will be described later.

[0021] Figure 2 illustrates a block diagram of the hardware configuration of Device 1. Device 1 includes, for example, a CPU 101, memory 102, bus 103, storage device 104, input device 105, output device 106, communication device (communication unit) 107, etc. Each part of Device 1 is interconnected via the bus 103 through its respective interface (I / F).

[0022] The CPU 101 operates in conjunction with other components, such as a controller (system controller, I / O controller, etc.), and is responsible for the overall control of the device 1. In the device 1, the CPU 101 executes, for example, the program of the present invention and other programs, and also reads and writes various types of information. Specifically, for example, the CPU 101 functions as a parallelized image generation unit 2, an extraction unit 3, an identification unit 4, an image synthesis unit 5, and a training data output unit 6. The device 1 is equipped with a CPU as its computing device, but may also be equipped with other computing devices such as a GPU (Graphics Processing Unit) or an APU (Accelerated Processing Unit), or a combination of the CPU and these.

[0023] Bus 103 can also be connected to external devices, for example. Examples of such external devices include a trained model manufacturing device, a character recognition device, an external storage device (external database, etc.), a printer, an external input device, an external output device, an audio output device such as a speaker, an external imaging device such as a camera, and various sensors such as an acceleration sensor, a geomagnetic sensor, and a direction sensor. Device 1 can be connected to an external network (the aforementioned communication network) by a communication device 107 connected to bus 103, for example, and can also be connected to other devices via the external network.

[0024] Memory 102 may be, for example, main memory. When the CPU 101 performs processing, memory 102 reads various operational programs, such as the program 105 of the present invention, which is stored in the storage device 104 (described later), and the CPU 101 receives data from memory 102 and executes the program. The main memory may be, for example, RAM (random access memory). Alternatively, memory 102 may be, for example, ROM (read-only memory).

[0025] The storage device 104 is also called an auxiliary storage device, for example, in relation to the main memory (primary memory). As described above, the storage device 104 stores an operation program 105 that includes the program of the present invention. The storage device 104 may be, for example, a combination of a recording medium and a drive that reads and writes to the recording medium. The recording medium is not particularly limited and may be internal or external, for example, an HD (hard disk), CD-ROM, CD-R, CD-RW, MO, DVD, flash memory, memory card, etc. The storage device 104 may be, for example, a hard disk drive (HDD) in which the recording medium and the drive are integrated, or a solid state drive (SSD). If the device 1 includes, for example, the storage device 104 functions as the storage device. The storage device 104 may store, for example, a character recognition model and reference character information, which will be described later.

[0026] In this device 1, the memory 102 and storage device 104 can also store various types of information, such as log information, information obtained from an external database (not shown) or external devices, information generated by this device 1, and information used by this device 1 when executing processing. At least some of the information may be stored on an external server other than the memory 102 and storage device 104, or it may be stored in a distributed manner across multiple terminals using blockchain technology or the like.

[0027] The device 1 further includes, for example, an input device 105 and an output device 106. The input device 105 may include, for example, a pointing device such as a touch panel, trackpad, or mouse; a keyboard; imaging means such as a camera or scanner; a card reader such as an IC card reader or magnetic card reader; an audio input means such as a microphone; and so on. The output device 106 may include, for example, a display device such as an LED display or liquid crystal display; an audio output device such as a speaker; a printer; and so on. In this embodiment 1, the input device 105 and the output device 106 are configured separately, but the input device 105 and the output device 106 may be configured as an integrated unit, such as a touch panel display.

[0028] Next, an example of the character recognition training data generation method of this embodiment will be described based on the flowchart in Figure 3. The character recognition training data generation method of this embodiment can be implemented as follows, for example, using the character recognition training data generation device 1 shown in Figure 1 or Figure 2. Note that the character recognition training data generation method of this embodiment is not limited to the use of the character recognition training data generation device 1 shown in Figure 1 or Figure 2.

[0029] First, the orientation image generation unit 2 generates an orientation image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orientation reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected (S1, orientation image generation step). The orientation image is, for example, an orientation image of an image including an object to be recognized by this device 1. The object to be recognized is not particularly limited as long as it contains characters. The characters are not particularly limited, but are preferably used for characters that are difficult to recognize by normal OCR processing. "Characters that are difficult to recognize by normal OCR processing" are not particularly limited and include, for example, characters printed by means such as stencil spraying, stamp printing, or engraving, or handwritten characters. Specific examples of the object to be recognized include, for example, articles that are placed in harsh environments during the production process of steel products, articles that have difficult-to-read labels attached (for example, ceramics, etc.). The generation of the orientation image by the orientation image generation unit 2 will be described later in Embodiment 2.

[0030] Next, the extraction unit 3 extracts an orthogonal background image and an orthogonal character image from the orthogonalized image (S2, extraction step). The extraction unit 3 can, for example, recognize the area in the orthogonalized image where characters are written by image processing and extract the orthogonalized character image by cutting out the area where characters are written from the orthogonalized image. The extraction unit 3 can also, for example, recognize the area in the orthogonalized image where no characters are written by image processing and extract the orthogonalized background image by cutting out the area where no characters are written. The characters are not particularly limited and include, for example, English letters, numbers, symbols, hiragana, katakana, kanji, and other characters.

[0031] Next, the identification unit 4 identifies the characters contained in the aligned character image based on the reference character information (S3, identification step). The reference character information is, for example, information that associates the aligned character image with the types of characters contained in the aligned character image. The reference character information may be stored, for example, in the memory 102 or storage device 104 of the device 1, or it may be stored in an external database or server. In the latter case, the identification unit 4 obtains the reference character information from the external database or server via a communication network and performs the identification.

[0032] Next, the image synthesis unit 5 generates a composite character image by combining the aligned background image and the aligned character image (S4, image synthesis step). Specifically, the image synthesis unit 5 can generate the composite character image by, for example, randomly selecting the aligned background image and the aligned character image and combining the selected aligned background image and the aligned character image. The image synthesis unit 5 may, for example, combine one aligned character image with one aligned background image, or it may combine two or more aligned character images. In the aligned background image, for example, the position (combination position) in which the aligned character image is combined is not particularly limited and can be combined at any position. The image synthesis unit 5 may, for example, generate multiple composite character images with different combination positions for a set of aligned background image and aligned character image. The image synthesis unit 5 may also, for example, perform processing such as changing the angle or size of the aligned character image or inverting it, and then combine the processed aligned character image with the aligned background image. Furthermore, the image synthesis unit 5 may generate the synthesized character image using, for example, machine learning. The machine learning may be, for example, supervised machine learning or unsupervised machine learning, and in the latter case, a Generative Adversarial Network (GAN) may be used to generate the synthesized character image.

[0033] The training data output unit 6 then outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition (S5, training data output step). The output may be, for example, output (storage) to the memory 102 or storage device 104 of the device 1, or output to an external device via a communication network. The external device may be, for example, an external storage device, or a device that uses the training data generated by the device 1, specifically, a trained model manufacturing device or character recognition device of the present invention, which will be described later.

[0034] When performing character recognition from images of an object, the input images are captured from various angles. Therefore, generating training data for character recognition requires images of the same object from multiple angles, resulting in a massive amount of data. Furthermore, in production sites such as factories, even when attempting to capture training data, some characters, such as lot numbers, are not printed for extended periods (e.g., one year), making data collection time-consuming. In contrast, the character recognition training data generation device of this embodiment uses a composite character image, created by combining an orthogonalized background image extracted from an orthogonalized image with an orthogonalized character image, as training data. For example, even with a small number of input data types, it can generate training data for a large number of patterns. Therefore, the character recognition training data generation device of this embodiment can reduce the amount of data required for machine learning, shorten the data collection period, and easily generate training data for character recognition.

[0035] [Embodiment 2] Embodiment 2 describes the orthogonal image generation unit included in the character recognition training data generation device of Embodiment 1. In the following description, the case in which the orthogonal image generation unit is an independent image correction device capable of communicating with the character recognition training data generation device will be described as an example, but the present invention is not limited thereto, and as mentioned above, the orthogonal image generation unit may be included in the character recognition training data generation device.

[0036] The image correction device of this embodiment will be described with reference to Figure 4. Figure 4 is a block diagram showing the configuration of an example of the image correction device 2 of this embodiment. As shown in Figure 4, the image correction device 2 (hereinafter also referred to as "device 2") includes an image acquisition unit 21, a terminal information acquisition unit 22, a distance information acquisition unit 23, a reference orientation information acquisition unit 24, a reference plane setting unit 25, and an image correction unit 26. Although not shown, device 2 may also include, for example, a storage unit.

[0037] The device 2 may be, for example, a single device including the aforementioned parts, or it may be a device in which the aforementioned parts can be connected via a communication network. Furthermore, the device 2 can be connected to external devices described later via a communication network. The communication network is not particularly limited and can use any known network, such as a wired or wireless network. Examples of communication networks include the Internet, WWW (World Wide Web), telephone lines, LAN (Local Area Network), SAN (Storage Area Network), DTN (Delay Tolerant Networking), LPWA (Low Power Wide Area), L5G (Local 5G), etc. Examples of wireless communication include Wi-Fi (registered trademark), Bluetooth (registered trademark), Local 5G, LPWA, etc. The wireless communication may be in the form of direct communication between devices (Ad Hoc communication), infrastructure communication, or indirect communication via an access point. The device 2 may, for example, be incorporated into a server as part of a system. Furthermore, the device 2 may be, for example, a personal computer (PC, e.g., desktop or notebook type), smartphone, tablet terminal, etc., on which the program of the present invention is installed. The device 2 may also be an imaging terminal capable of imaging an object (e.g., a smartphone or tablet terminal with a camera), or a device capable of communicating with the imaging terminal. Moreover, the device 2 may be in the form of cloud computing or edge computing, for example, in which at least one of the aforementioned parts is on a server and the other aforementioned parts are on a terminal.

[0038] Figure 5 illustrates a block diagram of the hardware configuration of the device 2. The device 2 includes, for example, a CPU 201, memory 202, bus 203, storage device 204, input device 205, output device 206, communication device (communication unit) 207, etc. Each part of the device 2 is interconnected via the bus 203 through its respective interface (I / F).

[0039] The CPU 201 operates in conjunction with other components, such as a controller (system controller, I / O controller, etc.), and is responsible for the overall control of the device 2. In the device 2, the CPU 201 executes, for example, the program of the present invention and other programs, and also reads and writes various types of information. Specifically, for example, the CPU 201 functions as an image acquisition unit 21, a terminal information acquisition unit 22, a distance information acquisition unit 23, a reference orientation information acquisition unit 24, a reference plane setting unit 25, and an image correction unit 26. The device 2 is equipped with a CPU as its computing device, but it may also be equipped with other computing devices such as a GPU (Graphics Processing Unit) or an APU (Accelerated Processing Unit), or a combination of the CPU and these.

[0040] Bus 203 can also be connected to external devices, for example. Examples of such external devices include the character recognition training data generation device of the present invention, an external storage device (external database, etc.), a printer, an external input device, an external display device, an audio output device such as a speaker, an external imaging device such as a camera, and various sensors such as an acceleration sensor, a geomagnetic sensor, and a direction sensor. The device 2 can be connected to an external network (the communication network) by a communication device 207 connected to bus 203, for example, and can also be connected to other devices such as a user's terminal via the external network.

[0041] Memory 202 may be, for example, main memory. When the CPU 201 performs processing, memory 202 reads various operational programs, such as the program of the present invention, stored in the storage device 204 (described later), and the CPU 201 receives data from memory 202 and executes the program. The main memory may be, for example, RAM (random access memory). Alternatively, memory 202 may be, for example, ROM (read-only memory).

[0042] The storage device 204 is also called an auxiliary storage device, for example, in relation to the main memory (primary memory). As described above, the storage device 204 stores an operating program including the program of the present invention. The storage device 204 may be, for example, a combination of a recording medium and a drive for reading and writing to the recording medium. The recording medium is not particularly limited and may be internal or external, for example, an HD (hard disk), CD-ROM, CD-R, CD-RW, MO, DVD, flash memory, memory card, etc. The storage device 204 may be, for example, a hard disk drive (HDD) in which the recording medium and the drive are integrated, or a solid state drive (SSD). If the device 2 includes, for example, the storage device 204 functions as the storage unit. The storage device 204 may store, for example, at least one of the correction target image, reference pose information, object distance information, alignment reference plane, and alignment image, which will be described later.

[0043] In this device 2, the memory 202 and storage device 204 can also store various types of information, such as log information, information obtained from an external database (not shown) or external devices, information generated by this device 2, and information used by this device 2 when executing processing. At least some of this information may be stored, for example, on an external server other than the memory 202 and storage device 204, or it may be stored in a distributed manner across multiple terminals using blockchain technology or the like.

[0044] The device 2 further includes, for example, an input device 205 and an output device 206. The input device 205 may include, for example, a pointing device such as a touch panel, trackpad, or mouse; a keyboard; imaging means such as a camera or scanner; a card reader such as an IC card reader or magnetic card reader; an audio input means such as a microphone; and so on. The output device 206 may include, for example, a display device such as an LED display or liquid crystal display; an audio output device such as a speaker; a printer; and so on. In this embodiment 1, the input device 205 and the output device 206 are configured separately, but the input device 205 and the output device 206 may be configured as an integrated unit, such as a touch panel display.

[0045] Next, an example of the image correction method (aligned image generation process) of this embodiment will be described based on the flowchart in Figure 6. The image correction method of this embodiment is carried out as follows, for example, using the image correction device 2 shown in Figures 4 to 5. Note that the image correction method of this embodiment is not limited to the use of the image correction device 2 shown in Figures 4 to 5.

[0046] First, the image acquisition unit 21 of the image correction device 2 acquires the image to be corrected (S1A, image acquisition step). The image to be corrected is, for example, an image that includes an object. The image acquisition unit 21 may acquire the image to be corrected using an imaging device such as a camera provided in the device 2, or it may acquire the image to be corrected from an imaging device outside the device via a communication network. The image to be corrected may be, for example, a video or a still image, or it may be an image that has already been captured or an image preview image. If the image to be corrected is an image preview image, the image acquisition unit 21 may, for example, acquire the image preview image in real time. The image acquisition unit 21 may, for example, store the acquired image to be corrected in a storage device 204 or memory 202.

[0047] Next, the terminal information acquisition unit 22 acquires terminal orientation information (S1B, terminal information acquisition step). The terminal orientation information is information about the orientation of the imaging terminal at the time of acquiring the image to be corrected, and can be estimated from, for example, a gyro sensor, acceleration sensor, geomagnetic sensor, distance sensor (for example, optical sensors such as 3D-Lidar, millimeter-wave sensors, ultrasonic sensors, etc.) equipped in the imaging terminal. Alternatively, the terminal orientation information may also be information obtained by, for example, imaging the imaging terminal from the outside at the time of acquiring the image to be corrected, and estimating the orientation of the imaging terminal from the captured image. Preferably, the terminal orientation information includes, for example, information from the gyro sensor. The orientation information is, for example, information about the orientation coordinate system of the imaging terminal in three axes: the X axis (for example, also called the Roll axis), the Y axis (for example, also called the Pitch axis), and the Z axis (for example, also called the Yaw axis). The terminal information acquisition unit 22 may, for example, store the acquired terminal orientation information in the storage device 204 or memory 202.

[0048] The terminal orientation information may include, for example, other information. This other information may include, for example, information about the shooting location, shooting date and time, and user identification information (name, ID, terminal identification information, etc.).

[0049] Next, the distance information acquisition unit 23 acquires object distance information (S1C, distance information acquisition step). The object distance information may be, for example, a predetermined distance (for example, also called a provisional shooting distance), a distance measured from the imaging terminal to the object by a distance sensor provided by the imaging terminal (for example, an optical sensor such as a 3D-Lidar, a millimeter-wave sensor, an ultrasonic sensor, etc.), or a distance estimated from the size of the object included in the image to be corrected. The distance from the size of the object can be estimated, for example, by using distance conversion information that associates the actual distance with the number of pixels in the image, thereby calculating the distance at which the object exists. The distance conversion information may be stored in, for example, the storage unit, or it may be stored in an external database.

[0050] Next, the reference posture information acquisition unit 24 acquires reference posture information (S1D, reference posture information acquisition step). The reference posture information is information about the posture of the object at the time the image to be corrected is acquired, and may be, for example, a predetermined value set in advance, or information estimated from a gyro sensor, acceleration sensor, geomagnetic sensor, etc., provided by the object. Alternatively, the reference posture information may be, for example, information obtained by taking an image of the object from the outside at the time the image to be corrected is acquired, and estimating the posture of the object from the image taken. The image may be, for example, the image to be corrected taken by the imaging terminal, or an image taken by another device. The reference posture information is, for example, information about the posture coordinate system of the object in three axes: the X axis (for example, also called the Roll axis), the Y axis (for example, also called the Pitch axis), and the Z axis (for example, also called the Yaw axis). The reference posture information acquisition unit 24 may, for example, store the acquired reference posture information in the storage device 204 or memory 202.

[0051] Next, the reference surface setting unit 25 sets a realigned reference surface based on the terminal orientation information, the object distance information, and the reference orientation information (S1E, reference surface setting step). The realigned reference surface may be, for example, any plane on the object. If the object is, for example, a steel plate on a production line, the realigned reference surface may be, for example, the surface on the steel plate on which an identification number or the like is printed.

[0052] The image correction unit 26 then corrects the image to be corrected to an image that is aligned as viewed from a direction perpendicular to the alignment reference plane, based on the terminal orientation information and the reference orientation information (S1F, image correction step). The image correction unit 26, for example, estimates the coordinates of four points corresponding to four arbitrarily specified points in the image to be corrected, based on the terminal orientation information and the reference orientation information, when the object is viewed from a direction perpendicular to the alignment reference plane, and corrects the image to be corrected to an image that is aligned by projection transformation. The coordinates of the four points are not particularly limited, and for example, any coordinates in the image to be corrected can be specified, but it is preferable that they be coordinates of the area surrounding the feature points of the object included in the image to be corrected. The feature points are, for example, identification information of the object (for example, product management numbers that have been engraved, stenciled, stamped, etc.). The image correction unit 26 also corrects the image to the image that is aligned in real time, for example, if the image to be corrected is a shooting preview image. Furthermore, the image correction unit 26 may, for example, crop a predetermined range from the shooting preview image and correct the cropped image to an orthogonal image. The predetermined range may, for example, be the area containing characters in the shooting preview image. In this case, the image correction unit 26 can, for example, extract character candidate areas from the shooting preview image using known character recognition technology, crop a rectangular range based on the character candidate areas, and correct the cropped image to an orthogonal image. In extracting character candidate areas, for example, a trained model created by machine learning using an orthogonal image generated by the image correction device 2 (for example, a trained model generated by the trained model manufacturing device 40 of Embodiment 4 described later) may be used to extract character candidate areas from the preview image.

[0053] A specific example of image correction using the device 2 will be explained using Figure 7. In the following explanation, the image correction device 2 is a tablet terminal with a camera function, and the example will be given of capturing images of steel products on a production line using the tablet terminal, but the present invention is not limited in any way to the following example.

[0054] First, as shown in Figure 7(A), the camera function of the tablet terminal 2, which is the device 2, captures the object 30 and acquires a preview image from the camera as the image to be corrected. Next, as terminal orientation information, the orientation coordinate system of the mobile terminal (tablet terminal), indicated by the solid arrow in Figure 7(A), is acquired from the gyro sensor of the tablet terminal 2. In addition, as reference orientation information, the orientation coordinate system of the object, indicated by the dashed arrow in Figure 7(A), is acquired. Next, the device 2 detects the object included in the preview image and estimates the distance between the device 2 and the object from the size of the detected object. Next, based on the terminal orientation information, the object distance information, and the reference orientation information, the device 2 identifies the surface of the object in the preview image and sets the surface as the reference plane for orientation correction. Next, the device 2 specifies the coordinates of four arbitrarily designated points in the preview image, indicated by black circles in Figure 7(B). Then, based on the terminal orientation information and the reference orientation information, the coordinates of the four points corresponding to the object 30 when viewed from the direction perpendicular to the orienting reference plane (shown as white circles in Figure 7(B)) are estimated, and the image to be corrected is corrected to an orienting image by projection transformation.

[0055] According to the image correction device 2 of this embodiment, based on the terminal orientation information, it is possible to easily generate an orthogonalized image in which the image to be corrected is orthogonalized with respect to the orthogonalization reference plane.

[0056] [Embodiment 3] Embodiment 3 is another example of the character recognition training data generation device of the present invention.

[0057] The character recognition training data generation device of this embodiment is the same as the character recognition training data generation device 1 of Embodiment 1, except that it includes an image processing unit in addition to the configuration of the character recognition training data generation device 1 of Embodiment 1, and the description thereof can be applied accordingly. The character recognition training data generation device 1A of this embodiment includes, for example, an image processing unit that generates a processed character image by processing the composite character image, and the training data output unit further outputs the processed character image and a combination of characters contained in the processed character image as training data for character recognition.

[0058] Figure 8 is a block diagram showing an example configuration of the character recognition training data generation device 1A of this embodiment. As shown in Figure 8, the character recognition training data generation device 1A includes an image processing unit 7 in addition to the configuration of the character recognition training data generation device 1 of Embodiment 1. The hardware configuration of the character recognition training data generation device 1A is the same as that of the character recognition training data generation device 1 of Figure 2, except that the CPU 101 is configured as that of the character recognition training data generation device 1A of Figure 8 instead of that of the character recognition training data generation device 1 of Figure 1.

[0059] Next, the method for generating training data for character recognition according to this embodiment will be explained using the flowchart in Figure 9. The method for generating training data for character recognition according to this embodiment can be implemented, for example, using the training data generation device 1A for character recognition according to this embodiment shown in Figure 8. However, the method for generating training data for character recognition according to the present invention is not limited to the use of the training data generation device 1A for character recognition.

[0060] First, S1 to S4 are performed in the same manner as S1 to S4 of Embodiment 1 to generate a composite character image.

[0061] The image processing unit 7 generates a processed character image by processing the composite character image (S6, image processing step). The processing can utilize, for example, image data augmentation methods used in creating training data for general image recognition. Specific examples include changing the color, size, tilt, and perspective of the image, horizontal shifting, random shifting, horizontal inversion, vertical inversion, shear transformation, RGB channel conversion, and background removal. The image processing unit 7 may also perform processing on the composite character image, such as adding missing characters, stains, footprints, or scuffs, changing the brightness, or changing the lighting (illuminance, angle, color, etc.). Figure 10 shows an example of a oriented background image, a oriented character image, and a processed character image by the image processing unit 7.

[0062] Next, the training data output unit 6 performs S5 in the same manner as S5 in Embodiment 1, except that it further outputs the processed character image and the combination of characters contained in the processed character image as training data for character recognition, and then terminates the process (END).

[0063] The character recognition training data generation device of this embodiment can, for example, generate a processed character image by processing the composite character image using an image processing unit. Therefore, according to the character recognition training data generation device of this embodiment, for example, it is possible to further reduce the amount of training data required for character recognition and generate character recognition training data that enables highly accurate character recognition.

[0064] [Embodiment 4] Embodiment 4 is an example of a trained model manufacturing apparatus according to the present invention.

[0065] The trained model manufacturing apparatus of this embodiment will be described with reference to Figure 11. Figure 11 is a block diagram showing an example configuration of the trained model manufacturing apparatus 40 of this embodiment. As shown in Figure 11, the trained model manufacturing apparatus 40 includes a training data acquisition unit 41 and a trained model generation unit 42. Although not shown, the trained model manufacturing apparatus 40 may also include, for example, a storage unit.

[0066] The trained model manufacturing apparatus 40 may be, for example, a single device including the aforementioned parts, or it may be a device in which the aforementioned parts can be connected via a communication network. The trained model manufacturing apparatus 40 can also be connected to external devices described later via a communication network. The communication network is not particularly limited and can use a known network, for example, it may be wired or wireless. Examples of communication networks include the Internet, WWW (World Wide Web), telephone lines, LAN (Local Area Network), SAN (Storage Area Network), DTN (Delay Tolerant Networking), LPWA (Low Power Wide Area), L5G (Local 5G), etc. Examples of wireless communication include Wi-Fi (registered trademark), Bluetooth (registered trademark), Local 5G, LPWA, etc. The wireless communication may be in the form of direct communication between devices (Ad Hoc communication), infrastructure communication, indirect communication via access points, etc. The trained model manufacturing apparatus 40 may be, for example, incorporated into a server as a system. Furthermore, the trained model manufacturing apparatus 40 may be, for example, a personal computer (PC, e.g., desktop or notebook type), smartphone, tablet terminal, etc., on which the program of the present invention is installed. Moreover, the trained model manufacturing apparatus 40 may be in the form of cloud computing or edge computing, for example, in which at least one of the above-mentioned parts is on a server and the other parts are on a terminal.

[0067] Figure 12 illustrates a block diagram of the hardware configuration of the trained model manufacturing apparatus 40. As shown in Figure 12, the trained model manufacturing apparatus 40 includes, for example, a CPU 401, memory 402, bus 403, storage device 404, input device 405, output device 406, communication device 407, etc. The explanation of each component of the trained model manufacturing apparatus 40 can be made by referring to the explanation of each component of the character recognition training data generation apparatus 1. Each part of the trained model manufacturing apparatus 40 is connected via the bus 403 by its respective interface (I / F). In the trained model manufacturing apparatus 40, the CPU 401 functions as a training data acquisition unit 41 and a trained model generation unit 42.

[0068] Next, an example of a method for manufacturing the trained model of this embodiment will be described based on the flowchart in Figure 13. The method for manufacturing the trained model of this embodiment is carried out as follows, for example, using the trained model manufacturing apparatus 40 shown in Figures 11 and 12. Note that the method for manufacturing the trained model of this embodiment is not limited to the use of the trained model manufacturing apparatus 40 shown in Figures 11 and 12.

[0069] First, the training data acquisition unit 41 acquires character recognition training data output by the character recognition training data generation device of the present invention as training data for character recognition (S41, training data acquisition step). The training data acquisition unit 41 may, for example, acquire the character recognition training data from the character recognition training data generation device of the present invention via the communication network, or it may acquire the character recognition training data from an external storage device in which the character recognition training data is stored.

[0070] Next, the trained model generation unit 41 generates a trained model that outputs characters contained in an image containing characters to be recognized when an image containing characters to be recognized is input, using machine learning with the character recognition training data (S42, training process). The machine learning is not particularly limited and may include, for example, a neural network such as a Convolutional Neural Network (CNN), a Support Vector Machine (SVM), a Bayesian network, or a regression tree. The machine learning using a CNN is not particularly limited and may include, for example, Semantic Segmentation, Instance Segmentation (IS), Single Shot Detector (SSD), Weighted Single Shot Detector (WSSD), etc. The trained model generation unit 41 may also generate a retrained trained model (derived model) using, for example, the character recognition training data and an already generated trained model. Furthermore, the trained model generation unit 41 may generate a trained model obtained by transfer learning using the trained model generated using the character recognition training data, or it may generate the trained model by model compression of the trained model generated using the character recognition training data.

[0071] The trained model generated by this embodiment can be used, for example, in a character recognition device described later. This makes it possible to perform character recognition on an image of a character to be recognized using an image of the character to be recognized.

[0072] [Embodiment 5] Embodiment 5 is an example of the character recognition device of the present invention.

[0073] The character recognition device of this embodiment will be described with reference to Figure 14. Figure 14 is a block diagram showing an example configuration of the character recognition device 50 of this embodiment. As shown in Figure 14, the character recognition device 50 includes a character recognition target image acquisition unit 51 and a character recognition unit 52. Although not shown, the character recognition device 50 may also include, for example, a storage unit.

[0074] The character recognition device 50 may be, for example, a single device including the aforementioned parts, or it may be a device in which the aforementioned parts can be connected via a communication network. Furthermore, the character recognition device 50 can be connected to external devices described later via a communication network. The communication network is not particularly limited and can use any known network, such as a wired or wireless network. Examples of communication networks include the Internet, WWW (World Wide Web), telephone lines, LAN (Local Area Network), SAN (Storage Area Network), DTN (Delay Tolerant Networking), LPWA (Low Power Wide Area), L5G (Local 5G), etc. Examples of wireless communication include Wi-Fi (registered trademark), Bluetooth (registered trademark), Local 5G, LPWA, etc. The wireless communication may be in the form of direct communication between devices (Ad Hoc communication), infrastructure communication, or indirect communication via an access point. The character recognition device 50 may, for example, be incorporated into a server as part of a system. Furthermore, the character recognition device 50 may be, for example, a personal computer (PC, e.g., desktop or notebook type), smartphone, tablet terminal, etc., on which the program of the present invention is installed. Moreover, the character recognition device 50 may be configured in a form such as cloud computing or edge computing, where, for example, at least one of the aforementioned parts is on a server and the other parts are on a terminal.

[0075] Figure 15 illustrates a block diagram of the hardware configuration of the character recognition device 50. As shown in Figure 15, the character recognition device 50 includes, for example, a CPU 501, memory 502, bus 503, storage device 504, input device 505, output device 506, communication device 507, etc. The explanation of each component of the character recognition device 50 can be found by referring to the explanation of each component of the character recognition training data generation device 1. Each part of the character recognition device 50 is connected via the bus 503 by its respective interface (I / F). In the character recognition device 50, the CPU 501 functions as a character recognition target image acquisition unit 51 and a character recognition unit 52.

[0076] Next, an example of the character recognition method of this embodiment will be described based on the flowchart in Figure 16. The character recognition method of this embodiment is carried out as follows, for example, using the character recognition device 50 shown in Figures 14 and 15. Note that the method for manufacturing the trained model of this embodiment is not limited to the use of the character recognition device 50 shown in Figures 14 and 15.

[0077] First, the character recognition target image acquisition unit 51 acquires a character recognition target image by capturing the character recognition target (S51, character recognition target image acquisition step). The character recognition target image may be, for example, a still image, a video, or a still image extracted from a video. The character recognition target image acquisition unit 51 may, for example, acquire images continuously or intermittently. In the latter case, it may acquire images at predetermined time intervals or at any arbitrary timing. The character recognition target image acquisition unit 51 may, for example, acquire the character recognition target image by capturing the character recognition target with the imaging device which is the input device 506, or it may acquire the character recognition target image from an external imaging device via the communication network using the communication device 508. The character recognition target image acquisition unit 51 may, for example, store the acquired character recognition target image in the memory 502 or the storage device 504.

[0078] The character recognition unit 52 inputs the image to be recognized into the character recognition model and recognizes the characters contained in the image (S52, character recognition step). The character recognition model is, for example, a trained model generated by machine learning using training data generated by the character recognition training data generation device of the present invention, which outputs the characters contained in the image to be recognized when the image to be recognized is input. The character recognition model may also be, for example, a trained model manufactured by the trained model manufacturing device of Embodiment 4.

[0079] The character recognition model includes, for example, an input layer that inputs an image to be recognized, an output layer that outputs the character recognition result, and at least one intermediate layer provided between the input layer and the output layer. The character recognition model may also be a program module that is part of artificial intelligence software. Examples of the multilayer network include neural networks. Examples of the neural network include convolutional neural networks (CNNs), but are not limited to CNNs, and may also be neural networks other than CNNs, SVMs (Support Vector Machines), Bayesian networks, regression trees, or other pre-trained models constructed with other learning algorithms.

[0080] The character recognition model can, for example, generate training data generated by the character recognition training data generation device of the present invention through machine learning. The character recognition model may also be, for example, a pre-trained model. Furthermore, the pre-trained model may be a pre-trained model (derived model) that has been retrained using the character recognition training data and an already generated pre-trained model. Moreover, the pre-trained model may be a pre-trained model obtained by transfer learning using a pre-trained model generated using character recognition training data, or a pre-trained model generated by model compression of a pre-trained model generated using character recognition training data.

[0081] The character recognition device 50 may include, for example, an output unit. In this case, the output unit may, for example, output the character recognition result. The output unit may, for example, output the character recognition result to a terminal outside the device via the communication network, or it may output the character recognition result to an output device 507. The output character recognition result may also be stored in, for example, a memory 502 or a storage device 504.

[0082] In this embodiment of the character recognition method, the case in which steps S51 to S52 are executed sequentially has been described as an example, but the present invention is not limited thereto. Specifically, in the present invention, steps S51 and S52 may be executed simultaneously or separately, and in the latter case, the order in which they are executed is not particularly limited and is arbitrary.

[0083] According to the character recognition device of this embodiment, for example, character recognition using a character recognition model generated by machine learning becomes possible.

[0084] [Embodiment 6] The first program of this embodiment is a program that causes a computer to execute each step of the character recognition training data generation method described above. Specifically, the first program of this embodiment is a program that causes a computer to execute the orthogonal image generation procedure, the extraction procedure, the identification procedure, the image synthesis procedure, and the training data output procedure.

[0085] The above-mentioned orthogonalized image generation procedure generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction procedure extracts an orthogonalized background image and an orthogonalized character image from the orthogonalized image, The aforementioned identification procedure identifies the characters contained in the aligned character image based on the reference character information, The image synthesis procedure generates a composite character image by combining the orthogonalized background image and the orthogonalized character image, The aforementioned training data output procedure outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition.

[0086] Furthermore, the first program of this embodiment can also be described as a program that causes the computer to function as a procedure for generating aligned images, an extraction procedure, an identification procedure, an image synthesis procedure, and a training data output procedure.

[0087] The first program of this embodiment can be based on the description in the above-mentioned character recognition training data generation apparatus and character recognition training data generation method. Each of the above steps can be read as, for example, "step" or "process". The program of this embodiment may also be recorded on, for example, a computer-readable recording medium. The recording medium is, for example, a non-transitory computer-readable storage medium. The recording medium is not particularly limited and includes, for example, random access memory (RAM), read-only memory (ROM), hard disk (HD), optical disk, floppy disk (FD), and the like.

[0088] [Embodiment 7] The second program of this embodiment is a program that causes a computer to execute each step of the pre-trained model manufacturing method described above. Specifically, the second program of this embodiment is a program that causes a computer to execute the training data acquisition procedure and the pre-trained model generation procedure.

[0089] The aforementioned training data acquisition procedure acquires the character recognition training data output by the first program as character recognition training data, The aforementioned pre-trained model generation procedure generates a pre-trained character recognition model that, when an image containing the character to be recognized is input, outputs the characters contained in the image to be recognized, using machine learning with the character recognition training data.

[0090] Furthermore, the second program of this embodiment can also be described as a program that causes the computer to function as a training data acquisition procedure and a trained model generation procedure.

[0091] The second program of this embodiment can be derived from the description in the trained model manufacturing apparatus and trained model manufacturing method of the present invention. Each of the above steps can be read as, for example, "step" or "process". The program of this embodiment may also be recorded on, for example, a computer-readable storage medium. The storage medium is, for example, a non-transitory computer-readable storage medium. The storage medium is not particularly limited and includes, for example, random access memory (RAM), read-only memory (ROM), hard disk (HD), optical disk, floppy disk (FD), and the like.

[0092] [Embodiment 8] The third program of this embodiment is a program that causes a computer to execute each step of the character recognition method described above. Specifically, the third program of this embodiment is a program that causes a computer to execute the character recognition target image acquisition procedure and the character recognition procedure.

[0093] The above procedure for acquiring an image to be recognized for character recognition involves acquiring an image to be recognized for character recognition that includes the character to be recognized, The character recognition procedure involves inputting the image to be recognized into a character recognition model and recognizing the characters contained in the image. The character recognition model is either a pre-trained model generated by machine learning using training data generated by the first program, which outputs characters included in an image containing characters to be recognized when the image containing the characters to be recognized is input, or a pre-trained model manufactured by the second program.

[0094] Furthermore, the third program of this embodiment can also be described as a program that causes the computer to function as a procedure for acquiring an image to be recognized and a procedure for recognizing the character.

[0095] The third program of this embodiment can be derived from the description in the character recognition device and character recognition method of the present invention. Each of the above steps can be read as, for example, "step" instead of "process". The program of this embodiment may also be recorded on, for example, a computer-readable recording medium. The recording medium is, for example, a non-transitory computer-readable storage medium. The recording medium is not particularly limited and includes, for example, random access memory (RAM), read-only memory (ROM), hard disk (HD), optical disk, floppy disk (FD), and the like.

[0096] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments. Various modifications to the configuration and details of the present invention can be understood by those skilled in the art within the scope of the present invention.

[0097] <Note> Some or all of the above embodiments may be described as follows, but are not limited to the following: (Note 1) It includes a parallelized image generation unit, an extraction unit, an identification unit, an image synthesis unit, and a training data output unit, The orthogonalized image generation unit uses the orientation information of the imaging terminal at the time of acquiring the image to be corrected to generate an orthogonalized image in which the image to be corrected is corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane. The extraction unit extracts an orthogonal background image and an orthogonal character image from the orthogonalized image. The identification unit identifies the characters contained in the aligned character image based on the reference character information. The image synthesis unit generates a composite character image by combining the aligned background image and the aligned character image. The aforementioned training data output unit is a character recognition training data generation device that outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition. (Note 2) Including the image processing section, The image processing unit generates a processed character image by processing the composite character image, The character recognition training data generation device according to Appendix 1 further outputs the processed character image and combinations of characters contained in the processed character image as training data for character recognition. (Note 3) The aforementioned orthogonalized image generation unit includes an image acquisition unit, a terminal information acquisition unit, a distance information acquisition unit, a reference posture information acquisition unit, a reference plane setting unit, and an image correction unit. The image acquisition unit acquires the image to be corrected, The image to be corrected is an image that includes an object, The terminal information acquisition unit acquires terminal orientation information, The terminal orientation information is information about the orientation of the imaging terminal at the time the image to be corrected is acquired. The distance information acquisition unit acquires object distance information, The object distance information is information about the distance from the imaging terminal to the object, The aforementioned reference posture information acquisition unit acquires reference posture information, The aforementioned reference posture information is information about the posture of the object at the time the image to be corrected was acquired. The reference plane setting unit sets a aligned reference plane based on the terminal orientation information, the object distance information, and the reference orientation information. The image correction unit corrects the image to be corrected to an orthogonal image viewed from a direction perpendicular to the orthogonal reference plane, based on the terminal orientation information and the reference orientation information. A character recognition training data generation device as described in Appendix 1 or 2. (Note 4) The character recognition training data generation device described in Appendix 3, wherein the terminal orientation information includes information from the gyro sensor of the imaging terminal. (Note 5) The image acquisition unit acquires the shooting preview image in real time as the image to be corrected. The image correction unit is a character recognition training data generation device as described in Appendix 3 or 4, which corrects the captured preview image to the aligned image in real time. (Note 6) The image correction unit crops a predetermined range from the shooting preview image and corrects the cropped image into an orthogonal image. A character recognition training data generation device as described in any of appendices 3 to 5. (Note 7) It includes a training data acquisition unit and a trained model generation unit, The aforementioned training data acquisition unit acquires character recognition training data output by a character recognition training data generation device described in any of the appendices 1 to 6 as character recognition training data, The trained model generation unit is a trained model manufacturing device that generates a trained character recognition model, which outputs characters contained in an image containing characters to be recognized, when an image containing characters to be recognized is input, by machine learning using the character recognition training data. (Note 8) Including the character recognition target image acquisition unit and the character recognition unit, The character recognition target image acquisition unit acquires a character recognition target image that includes the character to be recognized, The character recognition unit inputs the image to be recognized into the character recognition model and recognizes the characters contained in the image. The character recognition device is a character recognition device in which the character recognition model is a trained model generated by machine learning using training data generated by a character recognition training data generation device described in any of the appendices 1 to 6, so as to output characters contained in a character recognition target image when a character recognition target image containing the character recognition target is input, or a trained model manufactured by a trained model manufacturing device described in appendice 6. (Note 9) The process includes a parallelization image generation step, an extraction step, a classification step, an image synthesis step, and a training data output step. The aforementioned orthogonalized image generation step generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction step extracts an orthogonal background image and an orthogonal character image from the orthogonalized image. The aforementioned identification step identifies the characters contained in the aligned character image based on the reference character information, The image synthesis step generates a composite character image by combining the orthogonalized background image and the orthogonalized character image. The method for generating training data for character recognition includes outputting the composite character image and combinations of characters contained in the composite character image as training data for character recognition. (Note 10) Including the image processing process, The aforementioned image processing step generates a processed character image by processing the composite character image, The method for generating training data for character recognition according to Appendix 9, wherein the training data output step further outputs the processed character image and combinations of characters contained in the processed character image as training data for character recognition. (Note 11) The orthogonalized image generation process includes an image acquisition process, a terminal information acquisition process, a distance information acquisition process, a reference orientation information acquisition process, a reference plane setting process, and an image correction process. The aforementioned image acquisition step acquires the image to be corrected, The image to be corrected is an image that includes an object, The terminal information acquisition step acquires terminal posture information, The terminal orientation information is information about the orientation of the imaging terminal at the time the image to be corrected is acquired. The distance information acquisition step acquires object distance information, The object distance information is information about the distance from the imaging terminal to the object, The above-mentioned reference posture information acquisition step acquires reference posture information, The aforementioned reference posture information is information about the posture of the object at the time the image to be corrected was acquired. The aforementioned reference plane setting step sets a aligned reference plane based on the terminal orientation information, the object distance information, and the reference orientation information. The image correction step corrects the image to be corrected to an orthogonal image viewed from a direction perpendicular to the orthogonal reference plane, based on the terminal orientation information and the reference orientation information. A method for generating training data for character recognition as described in Appendix 9 or 10. (Note 12) The method for generating training data for character recognition according to Appendix 11, wherein the terminal orientation information includes information from the gyro sensor of the imaging terminal. (Note 13) The aforementioned image acquisition step acquires a shooting preview image in real time as the image to be corrected, The image correction step is a method for generating training data for character recognition according to appendix 11 or 12, which corrects the captured preview image to the orthogonalized image in real time. (Note 14) The image correction step involves cropping a predetermined range from the captured preview image and correcting the cropped image to an orthogonal image. A method for generating training data for character recognition, as described in any of appendices 11 to 13. (Note 15) This includes a training data acquisition process and a pre-trained model generation process. The aforementioned training data acquisition step acquires character recognition training data output by the character recognition training data generation method described in any of appendices 9 to 14 as character recognition training data, The pre-trained model generation step is a method for manufacturing a pre-trained model, which generates a character recognition model that outputs characters contained in an image containing characters to be recognized when an image containing characters to be recognized is input, using machine learning with the character recognition training data. (Note 16) This includes a process for acquiring an image to be recognized and a process for recognizing the character, The character recognition target image acquisition step acquires a character recognition target image that includes the character to be recognized, The character recognition step involves inputting the image to be recognized into a character recognition model and recognizing the characters contained in the image. The character recognition method is a character recognition model which is a trained model generated by machine learning using training data generated by the character recognition training data generation method described in any of appendices 9 to 14, so as to output characters contained in the character recognition target image when the character recognition target image containing the character recognition target is input, or a trained model which is manufactured by the trained model manufacturing method described in appendice 15. (Note 17) The procedure includes a parallelized image generation procedure, an extraction procedure, an identification procedure, an image synthesis procedure, and a training data output procedure. The above-mentioned orthogonalized image generation procedure generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction procedure extracts an orthogonalized background image and an orthogonalized character image from the orthogonalized image, The aforementioned identification procedure identifies the characters contained in the aligned character image based on the reference character information, The image synthesis procedure generates a composite character image by combining the orthogonalized background image and the orthogonalized character image, The aforementioned training data output procedure outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition. A program that causes a computer to perform each of the above steps. (Note 18) Including image processing steps, The aforementioned image processing procedure generates a processed character image by processing the composite character image, The aforementioned training data output procedure is further a program as described in Appendix 17 that outputs the processed character image and the combination of characters contained in the processed character image as training data for character recognition. (Note 19) The above-mentioned orthogonalized image generation procedure includes an image acquisition procedure, a terminal information acquisition procedure, a distance information acquisition procedure, a reference orientation information acquisition procedure, a reference plane setting procedure, and an image correction procedure. The aforementioned image acquisition procedure involves acquiring the image to be corrected, The image to be corrected is an image that includes an object, The aforementioned terminal information acquisition procedure acquires terminal orientation information, The terminal orientation information is information about the orientation of the imaging terminal at the time the image to be corrected is acquired. The aforementioned distance information acquisition procedure acquires object distance information, The object distance information is information about the distance from the imaging terminal to the object, The above procedure for acquiring reference posture information involves acquiring reference posture information, The aforementioned reference posture information is information about the posture of the object at the time the image to be corrected was acquired. The above-mentioned reference plane setting procedure sets a aligned reference plane based on the terminal orientation information, the object distance information, and the reference orientation information, The image correction procedure corrects the image to be corrected to an orthogonal image viewed from a direction perpendicular to the orthogonal reference plane, based on the terminal orientation information and the reference orientation information. The program described in Appendix 17 or 18. (Note 20) The program described in Appendix 19, wherein the terminal orientation information includes information from the gyro sensor of the imaging terminal. (Note 21) The aforementioned image acquisition procedure acquires the shooting preview image in real time as the image to be corrected, The image correction procedure is a program according to Appendix 19 or 20 that corrects the captured preview image to the orthogonalized image in real time. (Note 22) The aforementioned image correction procedure involves cropping a predetermined range from the captured preview image and correcting the cropped image to an orthogonal image. The program described in any of the appendices 19 to 21. (Note 23) This includes procedures for acquiring training data and generating a trained model. The aforementioned training data acquisition procedure acquires character recognition training data output by any of the programs described in Appendix 17 to 22 as character recognition training data, The aforementioned pre-trained model generation procedure generates a pre-trained character recognition model that, when an image containing the character to be recognized is input, outputs the characters contained in the image to be recognized, using machine learning with the character recognition training data. A program that causes a computer to perform each of the above steps. (Note 24) This includes the procedure for acquiring images for character recognition and the character recognition procedure, The above procedure for acquiring an image to be recognized for character recognition involves acquiring an image to be recognized for character recognition that includes the character to be recognized, The character recognition procedure involves inputting the image to be recognized into a character recognition model and recognizing the characters contained in the image. The character recognition model is a trained model generated by machine learning using training data generated by any of the programs described in Appendix 17 to 22, which enables it to output characters contained in an image containing characters to be recognized when that image is input, or a trained model manufactured by the program described in Appendix 23, and is a program for causing a computer to execute each of the above steps. (Note 25) The procedure includes a parallelized image generation procedure, an extraction procedure, an identification procedure, an image synthesis procedure, and a training data output procedure. The above-mentioned orthogonalized image generation procedure generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction procedure extracts an orthogonalized background image and an orthogonalized character image from the orthogonalized image, The aforementioned identification procedure identifies the characters contained in the aligned character image based on the reference character information, The image synthesis procedure generates a composite character image by combining the orthogonalized background image and the orthogonalized character image, The aforementioned training data output procedure outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition. A computer-readable recording medium containing a program that causes a computer to perform each of the aforementioned steps. (Note 26) Including image processing steps, The aforementioned image processing procedure generates a processed character image by processing the composite character image, The recording medium described in Appendix 25 further outputs the processed character image and the combination of characters contained in the processed character image as training data for character recognition. (Note 27) The above-mentioned orthogonalized image generation procedure includes an image acquisition procedure, a terminal information acquisition procedure, a distance information acquisition procedure, a reference orientation information acquisition procedure, a reference plane setting procedure, and an image correction procedure. The aforementioned image acquisition procedure involves acquiring the image to be corrected, The image to be corrected is an image that includes an object, The aforementioned terminal information acquisition procedure acquires terminal orientation information, The terminal orientation information is information about the orientation of the imaging terminal at the time the image to be corrected is acquired. The aforementioned distance information acquisition procedure acquires object distance information, The object distance information is information about the distance from the imaging terminal to the object, The above procedure for acquiring reference posture information involves acquiring reference posture information, The aforementioned reference posture information is information about the posture of the object at the time the image to be corrected was acquired. The above-mentioned reference plane setting procedure sets a aligned reference plane based on the terminal orientation information, the object distance information, and the reference orientation information, The image correction procedure corrects the image to be corrected to an orthogonal image viewed from a direction perpendicular to the orthogonal reference plane, based on the terminal orientation information and the reference orientation information. Recording medium as described in Appendix 25 or 26. (Note 28) The recording medium described in Appendix 27 includes the terminal orientation information, which also includes information from the gyro sensor of the imaging terminal. (Note 29) The aforementioned image acquisition procedure acquires the shooting preview image in real time as the image to be corrected, The recording medium described in Appendix 27 or 28 corrects the captured preview image to the orthogonalized image in real time, as described in the image correction procedure. (Note 30) The aforementioned image correction procedure involves cropping a predetermined range from the captured preview image and correcting the cropped image to an orthogonal image. A recording medium as described in any of the appendices 27 to 29. (Note 31) This includes procedures for acquiring training data and generating a trained model. The aforementioned training data acquisition procedure acquires character recognition training data output by any of the programs described in Appendix 17 to 22 as character recognition training data, The aforementioned pre-trained model generation procedure generates a pre-trained character recognition model that, when an image containing the character to be recognized is input, outputs the characters contained in the image to be recognized, using machine learning with the character recognition training data. A computer-readable recording medium containing a program that causes a computer to perform each of the aforementioned steps. (Note 32) This includes the procedure for acquiring images for character recognition and the character recognition procedure, The above procedure for acquiring an image to be recognized for character recognition involves acquiring an image to be recognized for character recognition that includes the character to be recognized, The character recognition procedure involves inputting the image to be recognized into a character recognition model and recognizing the characters contained in the image. The character recognition model is a trained model generated by machine learning using training data generated by any of the programs described in Appendix 17 to 22, so as to output characters contained in an image of a character to be recognized when an image of a character to be recognized, including the character to be recognized, is input, or a trained model manufactured by the program described in Appendix 23, and the recording medium is a computer-readable recording medium that records a program for causing a computer to execute each of the above steps. [Industrial applicability]

[0098] According to the present invention, training data for character recognition can be easily generated. For this reason, the present invention is widely useful in fields that utilize image-based character recognition. [Explanation of Symbols]

[0099] 1. Character recognition training data generation device 2. Image generation unit for reversal of orientation 3 Extraction part 4. Identification Unit 5 Image Synthesis Unit 6. Training Data Output Unit 7 Image Processing Department 101 CPU 102 memory 103 Bus 104 Storage device 105 Input device 106 Output device 107 Communication devices 2. Image Correction Device (Oriented Image Generation Unit) 21 Image acquisition unit 22 Terminal Information Acquisition Unit 23 Distance information acquisition section 24 Reference attitude information acquisition unit 25 Reference plane setting section 26 Image Correction Unit 20 Character recognition device 21 Character recognition section 201 CPU 202 memory Bus 203 204 Storage device 205 Input device 206 Output device 207 Communication devices 40. Pre-trained model manufacturing equipment 41 Training Data Acquisition Unit 42 Pre-trained model generation unit 401 CPU 402 memory Bus 403 404 Storage device 405 Input device 406 Output device 407 Communication devices 50 character recognition device 51 Character Recognition Target Image Acquisition Unit 52 Character recognition section 501 CPU 502 memory Bus 503 504 Storage device 505 Input device 506 Output device 507 Communication devices

Claims

1. It includes a parallelized image generation unit, an extraction unit, an identification unit, an image synthesis unit, and a training data output unit, The orthogonalized image generation unit uses the orientation information of the imaging terminal at the time of acquiring the image to be corrected to generate an orthogonalized image in which the image to be corrected is corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane. The extraction unit extracts an orthogonal background image and an orthogonal character image from the orthogonalized image. The identification unit identifies the characters contained in the aligned character image based on the reference character information. The image synthesis unit generates a composite character image by combining the aligned background image and the aligned character image. The aforementioned training data output unit is a character recognition training data generation device that outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition.

2. Including the image processing section, The image processing unit generates a processed character image by processing the composite character image, The character recognition training data generation device according to claim 1, wherein the training data output unit further outputs the processed character image and combinations of characters contained in the processed character image as training data for character recognition.

3. Includes a character recognition training data generation unit, a training data acquisition unit, and a trained model generation unit, The character recognition training data generation unit includes the parts of the character recognition training data generation device described in claim 1 or 2, The aforementioned training data acquisition unit acquires the character recognition training data output by the character recognition training data generation device as character recognition training data, The trained model generation unit is a trained model manufacturing device that generates a trained character recognition model, which outputs characters contained in an image containing characters to be recognized, when an image containing characters to be recognized is input, by machine learning using the character recognition training data.

4. Includes a character recognition model manufacturing unit, a character recognition target image acquisition unit, and a character recognition unit, The character recognition model manufacturing unit includes the parts of the trained model manufacturing apparatus described in claim 3, and generates the character recognition model. The character recognition target image acquisition unit acquires a character recognition target image that includes the character to be recognized, The character recognition unit inputs the image to be recognized into the character recognition model and recognizes the characters contained in the image. Character recognition device.

5. The process includes a parallelization image generation step, an extraction step, a classification step, an image synthesis step, and a training data output step. The aforementioned orthogonalized image generation step generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction step extracts an orthogonal background image and an orthogonal character image from the orthogonalized image. The aforementioned identification step identifies the characters contained in the aligned character image based on the reference character information, The image synthesis step generates a composite character image by combining the orthogonalized background image and the orthogonalized character image. The method for generating training data for character recognition is characterized in that each step of the training data output step is performed by a computer, and the steps of outputting the composite character image and the combination of characters contained in the composite character image as training data for character recognition are performed by a computer.

6. A process comprising generating training data for character recognition, acquiring training data, and generating a trained model, The character recognition training data generation step includes each step of the character recognition training data generation method described in claim 5, The aforementioned training data acquisition step acquires the character recognition training data output by the character recognition training data generation step as character recognition training data, The pre-trained model generation step is a method for manufacturing a pre-trained model in which each step is performed by a computer, and the pre-trained model generates a character recognition model that outputs characters contained in a character recognition target image when an image containing the character recognition target is input, using machine learning with the character recognition training data.

7. A process comprising a character recognition model manufacturing process, a character recognition target image acquisition process, and a character recognition process, The character recognition model manufacturing step is a step of generating the character recognition model by each step of the trained model manufacturing method described in claim 6, The character recognition target image acquisition step acquires a character recognition target image that includes the character to be recognized, The character recognition step involves inputting the image to be recognized into the character recognition model and recognizing the characters contained in the image. A character recognition method in which each step is performed by a computer.

8. The procedure includes a parallelized image generation procedure, an extraction procedure, an identification procedure, an image synthesis procedure, and a training data output procedure. The above-mentioned orthogonalized image generation procedure generates an orthogonalized image by correcting the image to be corrected to an image viewed from a direction perpendicular to the orthogonalization reference plane, using the orientation information of the imaging terminal at the time of acquiring the image to be corrected. The extraction procedure extracts an orthogonalized background image and an orthogonalized character image from the orthogonalized image, The aforementioned identification procedure identifies the characters contained in the aligned character image based on the reference character information, The image synthesis procedure generates a composite character image by combining the orthogonalized background image and the orthogonalized character image, The aforementioned training data output procedure outputs the composite character image and the combination of characters contained in the composite character image as training data for character recognition. A program that instructs a computer to perform each step.

9. A procedure comprising a procedure for generating training data for character recognition, a procedure for acquiring training data, and a procedure for generating a trained model, The above-mentioned character recognition training data generation procedure includes each step of the character recognition training data generation program described in claim 8, The aforementioned training data acquisition procedure acquires the character recognition training data output by the character recognition training data generation procedure as the character recognition training data, The aforementioned pre-trained model generation procedure generates a pre-trained character recognition model that, when an image containing the character to be recognized is input, outputs the characters contained in the image to be recognized, using machine learning with the character recognition training data. A program that instructs a computer to perform each step.

10. A procedure comprising a character recognition model manufacturing procedure, a character recognition target image acquisition procedure, and a character recognition procedure, The character recognition model manufacturing procedure is a procedure for generating the character recognition model by each step of the trained model manufacturing program described in claim 9, The above procedure for acquiring an image to be recognized for character recognition involves acquiring an image to be recognized for character recognition that includes the character to be recognized, The character recognition procedure is a program that inputs the image to be recognized into the character recognition model, recognizes the characters contained in the image, and causes the computer to execute each step.