Character recognition system, character recognition method, and character recognition program
The character recognition system improves accuracy by using a recognition model that processes both handwritten characters and preprint images together, aligning and recognizing characters on preprints effectively, overcoming shape and position variations.
Patent Information
- Application Number
- JP2024508871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-03-23
AI Technical Summary
Existing character recognition systems struggle to accurately recognize handwritten characters on preprints due to variations in shape and position, and the presence of preprinted elements complicates the recognition process.
A character recognition system that utilizes a recognition model to process both the image of handwritten characters and a preprint image together, employing techniques like deep learning and spatial transformation to improve accuracy by aligning and recognizing characters on preprints.
Enhances the accuracy of character recognition on preprints by reducing the influence of preprinted elements, allowing for efficient recognition of various preprint formats without the need for preprint erasure, thus saving computational resources.
Smart Images

Figure 0007761130000001 
Figure 0007761130000002 
Figure 0007761130000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a character recognition system and the like. [Background technology]
[0002] Optical Character Recognition (OCR) is widely used. OCR scans handwritten characters on forms as images and converts them into text data by recognizing the characters in the image. Character recognition using OCR is performed, for example, by recognizing characters written on preprints of forms using a learning model generated by machine learning. However, even when the same characters are written, the shape and position of the characters written on preprints vary depending on the person who wrote them. Furthermore, images of characters written on preprints contain a mixture of preprints and characters. Therefore, a learning model for recognizing handwritten characters on forms must be able to accurately recognize characters written on preprints with a variety of shapes and positions from images that contain a mixture of preprints and characters. Therefore, a technology capable of accurately recognizing characters on preprints of forms is desirable.
[0003] The image processing system of Patent Document 1 uses a learning model to extract handwritten characters written within a preprint frame. The image processing system of Patent Document 1 extracts handwritten characters by erasing the preprint frame from an image of the handwritten characters written within the preprint frame through image processing. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-39424 Summary of the Invention [Problem to be solved by the invention]
[0005] The information processing device of Patent Document 1 may have difficulty in accurately recognizing characters written on a preprint.
[0006] In order to solve the above problems, a main object of the present invention is to provide a character recognition system and the like that can improve the accuracy of recognizing characters written on a preprint. [Means for solving the problem]
[0007] In order to solve the above problems, the character recognition system of the present invention comprises an acquisition means for acquiring an image of characters written on a preprint of a document including the preprint, a recognition means for recognizing characters written on the preprint of the acquired image from the acquired image and the preprint image using a recognition model that recognizes characters written on the preprint from the image of the characters written on the preprint and the preprint image into which the preprint is copied, and an output means for outputting the recognition results.
[0008] The character recognition method of the present invention acquires an image of characters written on a preprint of a document including the preprint, and uses a recognition model that recognizes the characters written on the preprint from the image of the characters written on the preprint and a preprint image of the preprint, recognizes the characters written on the preprint in the acquired image from the acquired image and the preprint image, and outputs the recognition results.
[0009] The recording medium of the present invention non-temporarily records a character recognition program that causes a computer to execute the following processes: acquiring an image of characters written on a preprint of a document including the preprint; recognizing characters written on the preprint of the acquired image from the acquired image and the preprint image using a recognition model that recognizes the characters written on the preprint from the image of the characters written on the preprint and the preprint image of the preprint; and outputting the recognition results. [Effects of the Invention]
[0010] According to the present invention, it is possible to improve the accuracy of recognizing characters written on a preprint. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram illustrating an example of a configuration of a first exemplary embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of a form in the first embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing an example of an image on which characters are written according to the first embodiment of the present invention. [Figure 4] FIG. 2 is a diagram showing an example of a preprint image according to the first embodiment of the present invention. [Figure 5] FIG. 2 is a diagram showing an example of an image on which characters are written according to the first embodiment of the present invention. [Figure 6] FIG. 2 is a diagram showing an example of a preprint image according to the first embodiment of the present invention. [Figure 7] 1 is a diagram illustrating an example of the configuration of a character recognition system according to a first embodiment of the present invention. [Figure 8] FIG. 2 is a diagram showing an example of an image on which characters are written according to the first embodiment of the present invention. [Figure 9] FIG. 2 is a diagram showing an example of a preprint image according to the first embodiment of the present invention. [Figure 10] FIG. 2 is a diagram showing an example of an image on which characters are written according to the first embodiment of the present invention. [Figure 11]FIG. 2 is a diagram showing an example of a preprint image according to the first embodiment of the present invention. [Figure 12] FIG. 2 is a diagram showing an example of an image on which characters are written according to the first embodiment of the present invention. [Figure 13] FIG. 2 is a diagram showing an example of a preprint image according to the first embodiment of the present invention. [Figure 14] FIG. 2 is a diagram illustrating an example of an operation flow of the character recognition system according to the first embodiment of the present invention. [Figure 15] FIG. 2 is a diagram illustrating an example of an operation flow of the character recognition system according to the first embodiment of the present invention. [Figure 16] FIG. 10 is a diagram illustrating an example of a configuration of a second exemplary embodiment of the present invention. [Figure 17] FIG. 10 is a diagram illustrating an example of the configuration of a character recognition system according to a second embodiment of the present invention. [Figure 18] FIG. 10 is a diagram schematically illustrating a flow of data processing in a second embodiment of the present invention. [Figure 19] FIG. 10 is a diagram illustrating an example of an operation flow of a character recognition system according to a second exemplary embodiment of the present invention. [Figure 20] FIG. 10 is a diagram illustrating an example of an operation flow of a character recognition system according to a second exemplary embodiment of the present invention. [Figure 21] FIG. 10 is a diagram illustrating an example of an operation flow of a character recognition system according to a second exemplary embodiment of the present invention. [Figure 22] FIG. 10 is a diagram illustrating an example of the configuration of another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0012] (First embodiment) A first embodiment of the present invention will be described in detail with reference to the drawings. FIG. 1 is a diagram showing an example of the configuration of a form processing system of this embodiment. The form processing system includes, as an example, a character recognition system 10, a scanner 20, and an information processing server 30. The character recognition system 10 is connected to the scanner 20, for example, via a network. The character recognition system 10 is also connected to the information processing server 30 via the network. There may be multiple scanners 20 and multiple information processing servers 30. The number of scanners 20 and multiple information processing servers 30 is not particularly limited.
[0013] The character recognition system 10 acquires an image of a form scanned by the scanner 20, for example. A preprint on which characters are to be written is printed on the form. The preprint is, for example, a frame or line on the form that indicates the position where characters are to be written. The character recognition system 10 acquires an image of handwritten characters written on the preprint, for example. The characters written on the preprint may be printed. The characters written on the preprint are not limited to the above example.
[0014] The character recognition system 10 uses a recognition model to recognize characters written on a preprint from an image of the characters written on the preprint acquired from the scanner 20 and a preprint image of the preprint. The recognition model is a learning model that recognizes the characters written on the preprint from an image of the characters written on the preprint and the preprint image. The character recognition system 10 outputs the recognition results of the characters written on the preprint to, for example, an information processing server 30. The information processing server 30 is a server that performs processing according to the purpose of the recognition results of the characters written on the preprint.
[0015] The character recognition system 10 uses the preprint image in addition to the image of the characters written on the preprint to be recognized to recognize the characters written on the preprint, thereby reducing the effect of the preprint on character recognition.
[0016] FIG. 2 is a diagram showing an example of a form. In the example of the form in FIG. 2, the name of the form is written at the top as "Payment Slip." The example of the form in FIG. 2 is, for example, a document that is submitted to a financial institution when depositing money into an account at the financial institution. In the example of the form in FIG. 2, entry fields for "Account Number" and "Amount" are set. In the example of the form in FIG. 2, the boxes for entering numbers in the entry fields for "Account Number" and "Amount" are preprinted.
[0017] The characters written on the preprint are, for example, characters written within the frame of the preprint. The characters written on the preprint may be written so as to overlap the frame of the preprint. An image with characters written on the preprint is an image that includes both the preprint and the characters written on the preprint. A preprint image is an image of only the preprint without any characters written on it. In the example of the form in FIG. 2, numbers are written on the preprint, but the characters written on the preprint are not limited to numbers. The characters written on the preprint may also include symbols.
[0018] FIG. 3 is a diagram showing an example of an image in which characters written on a preprint have been copied. FIG. 3 is an image in which the "account number" entry field has been extracted from the example form in FIG. 2. FIG. 4 is a preprint image of the "account number" entry field in the example form in FIG. 2. In the example of an image in which characters written on a preprint in FIG. 3 have been copied, the characters "01778543" have been handwritten on the preprint shown in FIG. 4.
[0019] An image containing only preprints may include preprinted characters. Preprinted characters are, for example, characters indicating digits of an amount, characters indicating an item, or characters indicating a unit. Preprinted characters are not limited to the above, as long as they are printed on paper as preprints.
[0020] FIG. 5 is a diagram showing an example of an image in which characters written on a preprint have been copied. FIG. 5 is an image in which the "Amount" entry field has been extracted from the example of the form in FIG. 2. FIG. 6 is a preprint image of the "Amount" entry field in the example of the form in FIG. 2. In the example of the preprint image in FIG. 6, "yen" indicating the unit of amount is printed as part of the preprint at the bottom of the right-hand frame. In the example of an image in which characters written on a preprint in FIG. 5 have been copied, the characters "40000" have been handwritten on the preprint shown in FIG. 6.
[0021] A form is a document used for procedures in, for example, a financial institution, a government agency, an educational institution, a hospital, a transportation facility, or a company. A form may also be a document attached to an item to be managed. Examples of forms are not limited to the above. A preprint indicates, for example, a position on a form where a date, name, affiliation, address, telephone number, email address, age, gender, occupation, or amount should be entered. A preprint is composed of, for example, fields to be filled in and boxes into which characters should be entered. When multiple characters are to be entered in one field, the preprint may be a series of multiple boxes. Also, preprints for multiple fields may be printed on one form. For example, when a preprint is printed on paper as a fill-in field consisting of multiple boxes, the character recognition system 10 outputs recognized characters as character string data in the order of the boxes.
[0022] The configuration of character recognition system 10 will be described. FIG. 7 is a diagram showing an example of the configuration of character recognition system 10. Character recognition system 10 basically includes an acquisition unit 11, a recognition unit 13, and an output unit 14. Character recognition system 10 also includes an image extraction unit 12, a generation unit 15, and a storage unit 16. Acquisition unit 11, image extraction unit 12, recognition unit 13, output unit 14, and storage unit 16 recognize characters written on a preprint, for example, from an image of the characters written on the preprint. Acquisition unit 11, generation unit 15, and storage unit 16 also generate a recognition model, for example.
[0023] The acquisition unit 11 acquires an image of characters written on a preprint. The acquisition unit 11 acquires an image of a form with characters written on the preprint, for example, from the scanner 20. The acquisition unit 11 may acquire an image in which a portion of the form with characters written on the preprint has been extracted. The image of the portion of the preprint with characters written is, for example, the image shown in the examples of FIGS. 3 and 5. When extracting a preprint image from a form, the acquisition unit 11 may acquire an image of the form with no characters written on the preprint. The acquisition unit 11 acquires an image of the form with no characters written on the preprint, for example, from the scanner 20.
[0024] When character recognition system 10 generates a recognition model, acquisition unit 11 may acquire training data to be used for generating the recognition model. Generation unit 15 acquires, for example, an image of characters written on a preprint and data associating the preprint image with the characters written on the preprint as the training data. The training data is input, for example, by an operator's operation, to character recognition system 10 or another terminal device connected to character recognition system 10.
[0025] The image extraction unit 12 extracts a preprint image corresponding to an image of characters written on the preprint, acquired by the acquisition unit 11. The image extraction unit 12 extracts the preprint image from, for example, form data stored in the storage unit 16. The form data includes, for example, an image of the form and definition data. The definition data includes, for example, information on the items to be written on the form and the position on the form of the preprint corresponding to the items to be written. Information on the position of the preprint, for example, information indicating the area on the form where the preprint is printed. The items to be written include, for example, one or more of name, postal code, address, telephone number, age, personal identification number, account number, amount, and date. The items to be written are not limited to the above examples.
[0026] The image extraction unit 12 identifies the position of the preprint on the form, for example, based on information about the position of the preprint included in the definition data. Then, the image extraction unit 12 extracts the preprint image by cutting out the image of the identified preprint position from the image stored in the storage unit 16. The image extraction unit 12 may extract the preprint image from the image of the form acquired by the acquisition unit 11, in which no text is written on the preprint.
[0027] The recognition unit 13 uses a recognition model to recognize characters in an image of characters written on a preprint acquired by the acquisition unit 11 from the image and the preprint image. The recognition model is a learning model that recognizes characters written on a preprint from an image of characters written on a preprint and the preprint image. For example, the recognition unit 13 inputs the image of characters written on a preprint acquired by the acquisition unit 11 and the preprint image into the recognition model. Then, the recognition unit 13 recognizes characters written on a preprint using the recognition model. The recognition unit 13 may recognize characters written on a preprint using a preprint image extracted in advance. Alternatively, the recognition unit 13 may recognize characters written on a preprint using a preprint image generated in advance as an image of the preprint portion. The recognition unit 13 recognizes characters written on a preprint using, for example, a preprint image stored in the storage unit 16.
[0028] The recognition unit 13 extracts an image showing characters written on the preprint by identifying the position of the preprint based on, for example, information on the position of the preprint included in the definition data. Then, the recognition unit 13 uses a recognition model to recognize the characters written on the preprint from the extracted image showing the characters written on the preprint and the preprint image extracted by the image extraction unit 12.
[0029] The recognition unit 13, for example, combines an image of characters written on a preprint with the preprint image into one piece of data and inputs the combined data into the recognition model. Combining the image of characters written on a preprint with the preprint image means generating image data by superimposing the two images. If the image of characters written on a preprint and the preprint image each have three RGB channels per pixel, the recognition unit 13 combines the data of the two images to generate image data with six channels per pixel. The recognition unit 13 then inputs the combined six-channel image data into the recognition model.
[0030] The recognition unit 13, for example, combines an image of characters written on a preprint with the preprint image based on preset conditions. For example, the recognition unit 13 combines two images by overlaying an image of characters written on a preprint extracted at the same size with an image of characters written on a preprint and the preprint image based on the outer periphery of the preprint image. For example, when the two images are overlaid, the recognition unit 13 combines image data of corresponding pixels. Then, the recognition unit 13 inputs the combined data into a recognition model to recognize the characters written on the preprint.
[0031] The recognition unit 13 may recognize characters other than characters written on the preprint in an image of a document. The recognition unit 13 may, for example, identify the type of document from the image of the document acquired by the acquisition unit 11. The recognition unit 13 then identifies the position of the preprint based on definition data included in the document data corresponding to the identified type of document, thereby recognizing the characters written on the preprint. The recognition unit 13, for example, identifies the type of document by recognizing the name or document number of the document printed on the document in the image of the document. The relationship between the name or document number of the document printed on the document and the type of document is preset. The recognition model used by the recognition unit 13 may also be a learning model generated outside the character recognition system 10.
[0032] FIG. 8 is a diagram showing an example of an image in which characters written on a preprint are copied. The preprint in FIG. 8 differs in appearance from the example image in FIG. 3. The preprint in the example image in FIG. 8 differs, for example, from the example image in FIG. 3 in line thickness and type. FIG. 9 is a preprint image in the example image in FIG. 8. In the example image in which characters written on a preprint in FIG. 8 are copied, the characters "13758047" are handwritten on the preprint shown in FIG. 9. When the example image in FIG. 8 and the example image in FIG. 9 are input, a recognition model outputs "13758047" as the recognition result. Even a recognition model generated using, for example, the preprint in the example image in FIG. 4 as training data can recognize the characters written on the preprint in the example image in FIG. 9. In other words, by inputting an image in which characters written on a preprint are copied and a preprint image, the recognition model can recognize characters written on an untrained preprint.
[0033] FIG. 10 is a diagram showing an example of an image in which characters describing a year in the Gregorian calendar are copied onto a preprint. In the example image of FIG. 10, the characters "Gregorian calendar" and "year" are pre-printed within a frame on the preprint. In the example image of FIG. 10, the characters "2022" are handwritten on the preprint image. FIG. 11 is a preprint image of the example image of FIG. 10. When the example image of FIG. 10 and the example image of FIG. 11 are input, the recognition model outputs "2022" as the recognition result.
[0034] FIG. 12 shows an example of an image in which the first two digits of "20," indicating the year in the Gregorian calendar, are preprinted as a preprint in the example of the image in FIG. 10 . That is, in the example of the image in FIG. 12 , "year," "20," and "year" are preprinted as a preprint. In the example of the image in FIG. 12 , "22" of "2022" is handwritten on the preprint. FIG. 13 shows a preprint image of the example of the image in FIG. 12 . When the example of the image in FIG. 12 and the example of the image in FIG. 13 are input, a recognition model outputs "22" as a recognition result. Even a learning model that does not use the preprints of the example images in FIG. 11 and FIG. 13 as training data can recognize the image written on the preprint from the input image. In this way, the recognition model can recognize characters written on various types of preprints by inputting an image containing characters written on the preprint and the preprint image. Furthermore, in the above example, examples of preprints with different line thicknesses, line types, and preprinted characters were shown, but the recognition model can also perform recognition in the case of preprints with different frame shapes and colors, even if it is not trained using all types of preprints as training data.
[0035] The output unit 14 outputs the recognition result by the recognition unit 13. The output unit 14 outputs the characters recognized by the recognition unit 13 to, for example, the information processing server 30. The output unit 14 outputs, for example, the items corresponding to the preprint and the recognized characters in association with each other. When the target of recognition is an account number as shown in the example image of FIG. 3, the output unit 14 outputs, for example, information indicating that it is an account number in association with the recognized character string. The output unit 14 may output the recognition result to a display device (not shown) connected to the character recognition system 10.
[0036] When generating a recognition model in character recognition system 10, generation unit 15 performs processing related to the generation of the recognition model. Generation unit 15 learns images of characters written on preprints and the relationship between the preprint images and the characters written on the preprints. Then, generation unit 15 generates a recognition model that recognizes characters in an image from the images of characters written on the preprints and the preprint images.
[0037] For example, the generation unit 15 generates a recognition model by learning the relationship between an image of characters written on a preprint, data obtained by combining the preprint image, and the characters written on the preprint. If the image of characters written on a preprint and the preprint image are each images with three RGB channels per pixel, the generation unit 15 combines the data of the two images to create image data with six channels per pixel. The generation unit 15 then generates a recognition model by learning the relationship between the combined six-channel image data and the characters written on the preprint.
[0038] The image of the characters written on the preprint and the preprint included in the preprint image used by the generation unit 15 as training data do not have to be the image of the preprint that is actually used. When generating a recognition model, the generation unit 15 may perform training using a randomly shaped figure as the preprint. When a randomly shaped figure is used as the preprint, the generation unit 15 generates the recognition model using, for example, an image of the characters written on the randomly shaped figure and an image of the same figure as the figure with the characters written on top as training data.
[0039] The generation unit 15 generates the recognition model by, for example, deep learning using a DNN (Deep Neural Network). The machine learning algorithm for generating the recognition model is not limited to deep learning using a DNN.
[0040] The memory unit 16 stores, for example, a recognition model used by the recognition unit 13 to recognize characters in an image. The memory unit 16 stores, for example, a preprint image. The memory unit 16 stores, for example, form data. The form data includes, for example, image data of the form and definition data. The form data may include a preprint image extracted in advance. The memory unit 16 stores, for example, an image of characters written on a preprint, the preprint image, and the characters written on the preprint as learning data. Note that the recognition model used by the recognition unit 13 may be stored in a storage means other than the memory unit 16.
[0041] The scanner 20, for example, optically reads a form and generates an image of the form. Then, the scanner 20 outputs the image of the form to the character recognition system 10. The scanner 20 may extract an image of a preprinted portion from the image of the form. When extracting an image of the preprinted portion, the scanner 20 outputs the extracted preprinted image to the character recognition system 10. Furthermore, when the form is a document attached to an item to be managed, the scanner 20 may generate an image of the form by photographing the form.
[0042] The information processing server 30 obtains the recognition results of characters written on a form from, for example, the character recognition system 10. The information processing server 30 uses the recognition results to perform processing according to the application. For example, the information processing server 30 uses the recognition results to process applications related to account management and deposits and withdrawals at a financial institution. For example, the information processing server 30 may use the recognition results to process application documents at government agencies, educational institutions, hospitals, or transportation facilities. The information processing server 30 may use the recognition results to process invoices at companies. The information processing server 30 may also use the identification results to manage goods in distribution. Examples of uses of the identification results are not limited to those mentioned above.
[0043] The following describes the operation of the character recognition system 10 when recognizing characters written on a preprint. Fig. 14 is a diagram showing an example of the operation flow when the character recognition system 10 recognizes characters written on a preprint.
[0044] The acquisition unit 11 acquires an image showing the characters written on the preprint (step S11). The acquisition unit 11 acquires, for example, from the scanner 20, an image of the form showing the characters written on the preprint.
[0045] Furthermore, the image extraction unit 12 extracts a preprint image corresponding to the image acquired by the acquisition unit 11 (step S12). The image extraction unit 12 extracts a preprint image corresponding to the image acquired by the acquisition unit 11 from, for example, data stored in the storage unit 16.
[0046] When the preprint image is extracted, the recognition unit 13 uses a recognition model to recognize characters in the image from the image acquired by the acquisition unit 11 and the preprint image (step S13). The recognition model recognizes the characters written on the preprint from the image containing the characters written on the preprint and the preprint image.
[0047] When the characters in the image are recognized, the output unit 14 outputs the recognition result (step S14). The output unit 14 outputs the recognition result to the information processing server 30, for example.
[0048] The following describes the operation of the character recognition system 10 when generating a recognition model. Fig. 15 is a diagram showing an example of the operation flow when the character recognition system 10 generates a recognition model.
[0049] The acquisition unit 11 acquires, as learning data, an image showing characters written on a preprint, the preprint image, and the characters written on the preprint (step S21).
[0050] Upon acquiring the training data, the generation unit 15 learns the image showing the characters written on the preprint and the relationship between the preprint image and the characters written on the preprint, and generates a recognition model (step S22). The generation unit 15, for example, combines the image showing the characters written on the preprint with the preprint image. The generation unit 15 then learns the relationship between the combined data and the characters written on the preprint that are included as correct answer data in the training data, and generates a recognition model.
[0051] After generating the recognition model, the generation unit 15 stores the generated recognition model (step S23). The generation unit 15 stores the generated recognition model in the storage unit 16, for example.
[0052] The character recognition system 10 of the form processing system of this embodiment uses a recognition model to recognize characters written on a preprint from an image of the characters written on the preprint and the preprint image. The character recognition system 10 recognizes characters written on the preprint by using the preprint image in addition to the image of the characters written on the preprint to be recognized, thereby reducing the influence of the preprint on character recognition. As a result, the character recognition system 10 can improve the accuracy of recognizing characters written on the preprint.
[0053] Furthermore, the recognition model used by the character recognition system 10 can recognize characters written on preprints in an untrained state by using an image of characters written on a preprint and the preprint image as inputs and recognizing the characters written on the preprint. Therefore, by using an image of characters written on a preprint and the preprint image as inputs and recognizing the characters written on the preprint, the character recognition system 10 can recognize characters written on various preprints. Furthermore, when generating a recognition model, the character recognition system 10 does not need to prepare training data for each preprint state actually used for recognition. Furthermore, when generating a recognition model, the character recognition system 10 does not need to learn training data for each preprint state actually used for recognition, which reduces the amount of training required when generating a recognition model. Therefore, the character recognition system 10 can reduce the computer resources required for generating a recognition model. Therefore, the character recognition system 10 can efficiently generate recognition models.
[0054] Furthermore, by using randomly shaped figures as preprints when generating a recognition model, character recognition system 10 can generate a recognition model that can recognize characters written on various preprint images. In other words, by using a recognition model generated using randomly shaped figures as preprints, character recognition system 10 can accurately recognize characters written on preprints even if the shapes of preprint images differ for each form.
[0055] Furthermore, if a character recognition method different from the present embodiment is used, for example, a method in which characters written on a preprint are erased from an image containing the characters before character recognition is performed, erasing the preprint may require a large amount of computer resources. Furthermore, there is a risk that some of the characters may be erased when erasing the preprint. On the other hand, the character recognition system 10 of the present embodiment recognizes characters by inputting an image containing characters written on a preprint and data combining the preprint image into a recognition model, thereby eliminating the need to erase the preprint as a preprocessing step for character recognition. Furthermore, because the preprint erasure process is not performed, the impact of the preprint erasure process on character recognition can be reduced. Therefore, the character recognition system 10 of the present embodiment can improve recognition accuracy while reducing the resources required for recognizing characters written on a preprint.
[0056] (Second embodiment) A second embodiment of the present invention will be described in detail with reference to the drawings. FIG. 16 is a diagram showing an example of the configuration of a form processing system of this embodiment. The form processing system includes, as an example, a character recognition system 40, a scanner 20, and an information processing server 30. The character recognition system 40 is connected to the scanner 20, for example, via a network. The character recognition system 40 is also connected to the information processing server 30 via the network. There may be multiple scanners 20 and multiple information processing servers 30. The number of scanners 20 and multiple information processing servers 30 is not particularly limited. The functions of the scanner 20 and the information processing server 30 of this embodiment are similar to those of the scanner 20 and the information processing server 30 of the first embodiment.
[0057] The character recognition system 10 of the first embodiment uses, for example, a recognition model to input data obtained by combining an image of characters written on a preprint with the preprint image, and recognizes the characters on the preprint. The character recognition system 10 then outputs the recognition result. In addition to this configuration, the character recognition system 40 of the present embodiment combines, for example, an image of characters written on a preprint with the preprint image, after performing a conversion process on the preprint image using a conversion model in order to improve the accuracy of overlaying the two images. The conversion model is a learning model that estimates conversion parameters used when performing the conversion process on the preprint image.
[0058] The configuration of character recognition system 40 will be described. Fig. 17 is a diagram showing an example of the configuration of character recognition system 40. Character recognition system 40 includes acquisition unit 11, image extraction unit 12, recognition unit 41, output unit 14, generation unit 42, and storage unit 16. Furthermore, recognition unit 41 includes conversion unit 51 and image recognition unit 52. The configurations and functions of acquisition unit 11, image extraction unit 12, output unit 14, and storage unit 16 of character recognition system 40 are similar to those of acquisition unit 11, image extraction unit 12, output unit 14, and storage unit 16 of character recognition system 10 of the first embodiment, respectively.
[0059] The transformation unit 51 of the recognition unit 41 transforms the preprint image using, for example, a transformation model. The transformation model performs, for example, affine transformation on the preprint image. The recognition unit 41 transforms the preprint image so that it overlaps with the image to be combined, for example, by rotating, adjusting the size, and translating the preprint image. The transformation model estimates transformation parameters used when rotating, adjusting the size, and translating the preprint image.
[0060] The conversion unit 51, for example, uses a conversion model to estimate affine transformation parameters from data obtained by combining an image showing characters written on the preprint with the preprint image under preset conditions. The conversion unit 51 then performs affine transformation on the preprint image using the estimated parameters. The conversion unit 51 combines the two images by, for example, aligning the peripheries of the image showing characters written on the preprint with the preprint image under preset conditions, thereby overlapping the two images. The conversion unit 51 then estimates transformation parameters from the data combined under preset conditions using the conversion model. The transformation parameters are parameters for transforming the preprint image so as to improve the accuracy of overlay compared to when the images are combined under preset conditions. After estimating the transformation parameters, the conversion unit 51 performs affine transformation on the preprint image using the transformation parameters, thereby improving the accuracy of overlay.
[0061] The transformation model is, for example, a learning model that uses a DNN called STN (Spatial Transformer Networks). An image transformation method using STN is described, for example, in Max Jaderberg et al. "Spatial Transformer Networks," NIPS'15: Proceedings of the 28th International Conference on Neural Information Processing Systems, Volume 2, December 2015, pp. 2017-2025.
[0062] The image recognition unit 52 of the recognition unit 41 uses a recognition model to recognize characters written on the preprint from an image showing the characters written on the preprint and the preprint image. The image recognition unit 52 combines the image showing the characters written on the preprint with the preprint image on which the transformation unit 51 has performed affine transformation. The image recognition unit 52 then uses a discriminative model to recognize the characters written on the preprint from the combined data. The transformation model and the recognition model may be learning models generated outside the character recognition system 40.
[0063] FIG. 18 is a diagram schematically illustrating a processing flow when the recognition unit 41 recognizes characters written on a preprint. In the example of FIG. 18, assume that an image showing characters written on a preprint and a preprint image are input to the recognition unit 41. The conversion unit 51 combines the image showing characters written on the preprint and the preprint image, for example, under preset conditions. The preset conditions are set, for example, so that the peripheries of the two images are aligned. After combining the images, the conversion unit 51 estimates affine transformation parameters using a transformation model. Then, the conversion unit 51 performs affine transformation on the preprint image using the estimated affine transformation parameters. The conversion unit 51 outputs the affine-transformed preprint image to the image recognition unit 52. When the affine-transformed preprint image is input, the image recognition unit 52 combines the image showing characters written on the preprint and the affine-transformed image. Once the images are combined, the image recognition unit 52 uses a recognition model to recognize the characters written on the preprint from the combined data.
[0064] The character recognition system 40 generates, for example, only the recognition model out of the conversion model and the recognition model. When generating only the recognition model out of the conversion model and the recognition model, for example, a training model generated outside the character recognition system 40 is used as the conversion model. When generating only the recognition model out of the conversion model and the recognition model, the generation unit 42 combines, for example, an image of characters written on a preprint included in the training data with the preprint image using the conversion model. Then, the generation unit 42 learns the relationship between the combined data and the characters written on the preprint included in the training data as correct data, and generates the recognition model. The generation unit 42 saves the generated conversion model and recognition model in the memory unit 16.
[0065] The character recognition system 40 may generate both a transformation model and a recognition model. When generating both a transformation model and a recognition model, the generation unit 42 uses the transformation model to estimate transformation parameters from data obtained by combining an image of characters written on a preprint with the preprint image under preset conditions. The generation unit 42 also uses the recognition model to recognize the characters written on the preprint from the combined data. The generation unit 42 updates the parameters of the transformation model so as to reduce the difference between the affine transformation parameters estimated by the transformation model and the affine transformation parameters included in the training data. The generation unit 42 also updates the parameters of the recognition model so as to reduce the difference between the identification result and the correct data.
[0066] After updating the conversion parameters of the conversion model and the parameters of the recognition model, the generation unit 42 repeats the above process using the updated models. The generation unit 42 generates the conversion model and the recognition model by repeating the above process, for example, until the estimation results of the conversion parameters of the conversion model and the accuracy of the recognition result of the recognition model satisfy a preset standard. The generation unit 42 also generates an identification model by updating the parameters of the recognition model so as to reduce the difference between the identification result and the correct data. The generation unit 42 stores the generated conversion model and recognition model in the storage unit 16, for example.
[0067] The following describes the operation of the character recognition system 40 when recognizing characters written on a preprint. Fig. 19 is a diagram showing an example of the operation flow when the character recognition system 40 recognizes characters written on a preprint.
[0068] The acquisition unit 11 acquires an image showing characters written on a preprint (step S31). The acquisition unit 11 acquires, for example, from the scanner 20, an image of the form showing characters written on a preprint.
[0069] Furthermore, the image extraction unit 12 extracts a preprint image corresponding to the image acquired by the acquisition unit 11 (step S32). The image extraction unit 12 extracts a preprint image corresponding to the image acquired by the acquisition unit 11 from, for example, data stored in the storage unit 16.
[0070] When the preprint image is acquired, the conversion unit 51 of the recognition unit 41 uses a conversion model to estimate conversion parameters to be used when converting the preprint image. Then, the conversion unit 51 converts the preprint image using the estimated conversion parameters (step S33). After the preprint image is converted, the image recognition unit 52 combines an image of the preprint with text written on it and the converted preprint image. Then, the image recognition unit 52 uses the recognition model to recognize text in the image from the combined data (step S34).
[0071] When the characters in the image are recognized, the output unit 14 outputs the recognition result (step S35). The output unit 14 outputs the recognition result to the information processing server 30, for example.
[0072] The following describes the operation of the character recognition system 40 when generating only a recognition model out of a conversion model and a recognition model. Fig. 20 is a diagram showing an example of the operation flow when the character recognition system 40 generates only a recognition model.
[0073] The acquisition unit 11 acquires, as learning data, an image showing characters written on a preprint, a preprint image, and the characters written on the preprint (step S41).
[0074] When the learning data is acquired, the generation unit 42 uses the conversion model to estimate conversion parameters to be used when converting the preprint image. Then, the generation unit 42 uses the estimated conversion parameters and the conversion model to convert the preprint image (step S42).
[0075] After converting the preprint image, the generation unit 42 combines the image of the characters written on the preprint with the converted preprint image. The generation unit 42 then learns the relationship between the combined data and the characters written on the preprint, and generates a recognition model (step S43).
[0076] After generating the recognition model, the generation unit 42 stores the generated recognition model (step S44). The generation unit 42 stores the generated recognition model in the storage unit 16, for example.
[0077] The following describes the operation of the character recognition system 40 when it generates a conversion model and a recognition model. Fig. 21 is a diagram showing an example of the operation flow when the character recognition system 40 generates a conversion model and a recognition model.
[0078] The acquisition unit 11 acquires, as learning data, data obtained by combining an image showing characters written on a preprint with the preprint image, conversion parameters, and the characters written on the preprint (step S51).
[0079] When the training data is acquired, the generation unit 42 generates a conversion model by learning the relationship between the data, which is included in the training model and is a combination of an image showing characters written on a preprint and the preprint image, and the parameters included in the training model. The generation unit 42 also generates a recognition model by learning the relationship between the data, which is included in the training model and is a combination of an image showing characters written on a preprint and the preprint image, and the characters written on the preprint (step S52).
[0080] After generating the transformation model and the recognition model, the generation unit 42 stores the generated transformation model and the recognition model (step S53). The generation unit 42 stores the generated transformation model and the recognition model in the storage unit 16, for example.
[0081] The character recognition system 40 of this embodiment uses a transformation model to combine an image showing characters written on a preprint with the preprint image. The character recognition system 40 then uses the recognition model to recognize the characters written on the preprint from the combined data. By using the preprint image converted using the transformation model, the character recognition system 40 can improve the accuracy of overlay when combining the image showing characters written on the preprint with the preprint image. By using the combined data in this manner, the character recognition system 40 can recognize the characters on the preprint using the recognition model while suppressing variations in misalignment between the image showing characters written on the preprint and the preprint image. By recognizing the characters written on the preprint using the recognition model while suppressing variations in misalignment between the two images, the character recognition system 40 can improve the accuracy of recognizing the characters written on the preprint.
[0082] Furthermore, when generating a conversion model using training data, the character recognition system 40 can generate a conversion model that suppresses misalignment between an image of characters written on a preprint and the preprint image, which may occur in actual usage situations. Therefore, the character recognition system 40 can suppress fluctuations in misalignment between an image of characters written on a preprint and the preprint image, depending on actual usage situations. Therefore, when generating a conversion model using training data, the character recognition system 40 can further improve the recognition accuracy of characters written on a preprint.
[0083] Each process in the character recognition system 10 of the first embodiment and the character recognition system 40 of the second embodiment can be realized by executing a computer program on a computer. Fig. 22 shows an example of the configuration of a computer 200 that executes a computer program that performs each process in the character recognition system 10 of the first embodiment and the character recognition system 40 of the second embodiment. The computer 200 includes a CPU (Central Processing Unit) 201, a memory 202, a storage device 203, an input / output I / F (Interface) 204, and a communication I / F 205.
[0084] The CPU 201 reads and executes computer programs for performing each process from the storage device 203. The CPU 201 may be configured with a combination of multiple CPUs. The CPU 201 may also be configured with a combination of a CPU and another type of processor. For example, the CPU 201 may be configured with a combination of a CPU and a graphics processing unit (GPU). The memory 202 is configured with a dynamic random access memory (DRAM) or the like, and temporarily stores the computer programs executed by the CPU 201 and data being processed. The storage device 203 stores the computer programs executed by the CPU 201. The storage device 203 is configured with, for example, a non-volatile semiconductor storage device. Other storage devices such as a hard disk drive may also be used for the storage device 203. The input / output I / F 204 is an interface that receives input from an operator and outputs display data, etc. The communication I / F 205 is an interface that transmits and receives data between the scanner 20 and the information processing server 30. The information processing server 30 may also have a similar configuration.
[0085] The computer program used to execute each process can also be stored and distributed on a computer-readable recording medium that non-temporarily stores data. Examples of recording media that can be used include magnetic tapes for recording data and magnetic disks such as hard disks. Optical disks such as CD-ROMs (Compact Disc Read Only Memory) can also be used as recording media. Non-volatile semiconductor storage devices can also be used as recording media.
[0086] The present invention has been described above using the above-described embodiment as an example. However, the present invention is not limited to the above-described embodiment. In other words, the present invention can be applied in various aspects that can be understood by a person skilled in the art within the scope of the present invention. [Explanation of symbols]
[0087] 10 Character Recognition System 11 Acquisition Department 12 Image extraction section 13 Recognition part 14 Output section 15 Generation part 16 Memory section 20. Scanner 30 Information Processing Server 40 Character Recognition System 41 Recognition part 42 Generation part 51 Conversion unit 52 Image Recognition Unit 100 computers 101 CPU 102 memory 103 Storage device 104 Input / Output Interface 105 Communication I / F
Claims
1. An acquisition means for acquiring an image of characters written on a preprint of a form including the preprint; a recognition means for recognizing characters written on a preprint of an acquired image from pixel data of the acquired image and data obtained by combining pixel data of the preprint image for each pixel, using a recognition model for recognizing characters written on the preprint from data obtained by combining pixel data of an image of characters written on the preprint and pixel data of a preprint image of the preprint for each pixel; an output means for outputting the result of the recognition; A character recognition system comprising:
2. further comprising a transforming means for transforming the preprint image using a transformation parameter; the recognition means recognizes characters written on the preprint of the acquired image from data obtained by combining pixel data of the acquired image and pixel data of the converted preprint image for each pixel; The character recognition system of claim 1 .
3. the conversion means converts the preprint image using a conversion model that estimates conversion parameters from data obtained by combining, for each pixel, pixel data of the image and pixel data of the converted preprint image; The character recognition system of claim 2 .
4. the recognition means identifies a type of form for which characters written on the preprint are to be recognized from the image, and recognizes the characters written on the preprint based on definition data corresponding to the identified type of form.
4. A character recognition system according to claim 1.
5. the recognition means recognizes characters written on the preprint based on definition data that defines the position of the preprint on the form; 5. A character recognition system according to claim 1.
6. The apparatus further comprises a generating means for learning the relationship between pixel data of an image of characters written on a preprint and data obtained by combining pixel data of the preprint image on a pixel-by-pixel basis, and the characters written on the preprint, and for generating a recognition model for recognizing characters written on the preprint in the image from the pixel data of the image of characters written on the preprint and data obtained by combining pixel data of the preprint image on a pixel-by-pixel basis.
6. A character recognition system according to claim 1.
7. the generating means learns the relationship between pixel data of an image of characters written on a preprint and data obtained by combining pixel data of the preprint image on a pixel-by-pixel basis, and transformation parameters, and generates a transformation model for estimating transformation parameters to be used in transforming the preprint image.
7. The character recognition system of claim 6.
8. An image of the characters written on the preprint of the document containing the preprint is acquired, using a recognition model that recognizes characters written on the preprint from pixel data of an image in which characters written on the preprint are copied and data obtained by combining pixel data of a preprint image in which the preprint is copied, the characters written on the preprint of the acquired image are recognized from pixel data of the acquired image and data obtained by combining pixel data of the preprint image, outputting the results of the recognition; Character recognition method.
9. A process of acquiring an image of characters written on a preprint of a form including the preprint; a process of recognizing characters written on a preprint of the acquired image from pixel data of the acquired image and data obtained by combining pixel data of the preprint image, using a recognition model that recognizes characters written on the preprint from data obtained by combining pixel data of an image of characters written on the preprint and pixel data of a preprint image of the preprint; a process of outputting the recognition result; A character recognition program that causes a computer to execute the following.
Citation Information
Patent Citations
Picture data processing system
JP1993266247A
OCR device, form out method, and form out program
JP2007148846A
Information processing device and information processing program
JP2020123272A
Image processing system, image processing method and program
JP2021039424A
Image processing device, image processing system, image processing method, and program
JP2021043650A