Information processing device, information processing method, and program
Patent Information
- Application Number
- JP2024513172
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2025-09-04
- Estimated Expiration
- 2044-02-27
AI Technical Summary
Conventional methods for recognizing character strings in images face challenges in improving recognition accuracy due to variations in fonts, colors, and backgrounds.
An information processing device that trains a model using first teacher data for character string recognition and second teacher data for correctness determination, enabling a single learning process to enhance both tasks, and includes mechanisms for correcting misrecognitions through additional OCR processing if needed.
Improves the accuracy of character string recognition in images by learning the tendency of misrecognition and providing alerts or alternative OCR processing when necessary, thereby enhancing overall recognition performance.
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] A method has been proposed for recognizing character strings contained in an image using a language model (see, for example, Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Darwin Bautista and Rowel Atienza. Scene text recognition with permuted autoregressive sequence models. In ECCV, pages 178-196, 2022. Summary of the Invention [Problem to be solved by the invention]
[0004] However, when recognizing a character string contained in a scene, it is necessary to recognize various fonts, colors, shapes, or backgrounds, and there are cases in which the recognition accuracy does not improve with conventional techniques.
[0005] The present invention has been made in consideration of these points, and has an object to improve the accuracy of recognizing character strings contained in images. [Means for solving the problem]
[0006] An information processing device according to a first aspect of the present invention has an acquisition unit that acquires (1) first teacher data that associates multiple image data including a character string with a correct string that is a character string included in each of the multiple image data, and (2) second teacher data that associates (a) the multiple image data, (b) a recognition target string that is a character string to be recognized, and (c) a flag indicating whether a correct string corresponding to each of the multiple image data matches the recognition target string, and a learning unit that generates a trained model that has been trained to learn (1) a string recognition task that is a task of outputting a string included in the image data, using image data associated in the first teacher data as input, based on the first teacher data, and (2) a correct / incorrect determination task that is a task of determining whether a correct string corresponding to the image data matches the recognition target string, based on the second teacher data, using image data associated in the second teacher data and the recognition target string as input, and
[0007] The character string to be recognized may include either (1) a correct character string corresponding to image data associated with the character string to be recognized, or (2) a character string that is incorrectly recognized based on the image data.
[0008] The learning unit may train the character string recognition task and the correct / incorrect determination task in a single learning process to generate the trained model.
[0009] The trained model may take target image data, which is image data including a character string to be recognized, as input, and output a predicted character string and a correct / incorrect flag indicating whether the predicted character string is correct, and the acquisition unit may further acquire target image data, which is image data including a character string to be recognized, and the information processing device may further have an output unit that inputs the target image data acquired by the acquisition unit into the trained model and outputs a predicted character string for the target image data and a correct / incorrect flag for the predicted character string.
[0010] The information processing device may further include a display control unit that, when the correct / incorrect flag output by the output unit indicates that the predicted character string is incorrect, displays a message indicating that the prediction result may be incorrect.
[0011] When the correct / incorrect flag output by the output unit indicates that the predicted character string is incorrect, the optical character recognition means may output characters contained in image data when image data is input, and may further include a recognition means control unit that inputs the target image data to the optical character recognition means that is different from the trained model, and may further include a display control unit that causes the optical character recognition means to obtain the recognized character string by optical character recognition of the target image data, and displays the obtained character string.
[0012] The system may further include a generation unit that generates the second teacher data by associating a string output by the trained model in the process of the learning unit learning the string recognition task, which string is different from a string included in input image data, with the image data input to the trained model as the recognition target string.
[0013] An information processing method according to a second aspect of the present invention includes a first acquisition step of acquiring first teacher data that associates multiple image data including a character string with a correct string that is a character string included in each of the multiple image data, executed by a computer; a second acquisition step of acquiring second teacher data that associates (a) the multiple image data, (b) a recognition target string that is a character string to be recognized, and (c) a flag indicating whether a correct string corresponding to each of the multiple image data matches the recognition target string; and a learning step of generating a trained model that has been trained to learn (1) a string recognition task that is a task of outputting a string included in the image data, using image data associated in the first teacher data as input, based on the first teacher data, and (2) a correct / incorrect determination task that is a task of determining whether a correct string corresponding to the image data matches the recognition target string, based on the second teacher data, using image data associated in the second teacher data and the recognition target string as input, the trained model outputting a predicted string that is a character string predicted to be included in the image data, the trained model receiving image data including a character string as input.
[0014] In a third aspect of the program of the present invention, a computer is caused to execute a first acquisition step of acquiring first teacher data that associates multiple image data including a character string with a correct string that is a character string included in each of the multiple image data; a second acquisition step of acquiring second teacher data that associates (a) the multiple image data, (b) a recognition target string that is a character string to be recognized, and (c) a flag indicating whether a correct string corresponding to each of the multiple image data matches the recognition target string; and a learning step of generating a trained model that has learned (1) a string recognition task that is a task of outputting a string included in the image data, using image data associated in the first teacher data as input, based on the first teacher data, and (2) a correct / incorrect determination task that is a task of determining whether a correct string corresponding to the image data matches the recognition target string, based on the second teacher data, using image data associated in the second teacher data and the recognition target string as input, the trained model outputting a predicted string that is a character string predicted to be included in the image data, the trained model being trained based on the following: Effect of the Invention
[0015] According to the present invention, it is expected that the accuracy of recognizing character strings contained in an image can be improved. [Brief description of the drawings]
[0016] [Figure 1] 1 is a diagram for explaining an overview of an information processing system S according to an embodiment. [Diagram 2] A figure showing an example of the data structure of first teacher data D1. [Diagram 3] A figure showing an example of the data structure of second teacher data D2. [Figure 4] 1 is a block diagram showing a configuration of an information processing device 1. FIG. [Diagram 5] FIG. 11 is a diagram illustrating an example of processing performed by a learning unit 132. [Figure 6] FIG. 11 is a diagram illustrating an example of processing performed by a learning unit 132. [Figure 7] FIG. 11 is a diagram illustrating an example of processing performed by a learning unit 132. [Figure 8] 13 is a diagram showing an example of a screen displayed by a display control unit 134. FIG. [Figure 9] 3 is a flowchart showing a process flow in the information processing device 1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] [Outline of Information Processing System S] 1 is a diagram for explaining an overview of an information processing system S according to an embodiment. The information processing system S is a system for providing OCR (Optical Character Recognition) that extracts characters included in an image. The information processing system S includes an information processing device 1 and an information terminal 2.
[0018] The information processing device 1 is an OCR for recognizing character strings included in an image. The information processing device 1 trains a trained model, which is a machine learning model for character recognition based on an image. The information processing device 1 also inputs image data to be inferred to the trained model and outputs characters included in the image data.
[0019] The information terminal 2 is a terminal used by a user of the information processing system S. The information terminal 2 is, for example, a smartphone, a tablet, or a personal computer. The information terminal 2 transmits an instruction to start learning to the information processing device 1 in response to a user's operation, and transmits teacher data to be used for learning. The information terminal 2 also transmits an image of a target for character string recognition, and acquires and displays the result of character string recognition by the information processing device 1.
[0020] An overview of the processing in the information processing device 1 will be described. The information processing device 1 trains a learning model on a character string recognition task and a true / false judgment task to generate a trained model. The character string recognition task is a task for predicting a character string included in image data. The true / false judgment task is a task for predicting whether or not a character string included in image data matches an input character string, based on the image data and the character string. The learning model is a pre-learned model trained to be able to execute a natural language processing task based on a large amount of data set.
[0021] The information processing device 1 acquires first teacher data D1 and second teacher data D2. The first teacher data D1 is teacher data mainly used for learning a string recognition task. FIG. 2 is a diagram showing an example of a data structure of the first teacher data D1. In the first teacher data D1, a plurality of image data D11 is associated with a correct answer string D12 corresponding to each of the plurality of image data D11. A string is captured in the image data. The correct answer string D12 indicates the correct answer of the string included in each of the plurality of image data D11.
[0022] The second teacher data D2 is teacher data used for learning the true / false judgment task. FIG. 3 is a diagram showing an example of the data structure of the second teacher data D2. The second teacher data D2 is data in which a plurality of image data D21, a judgment target string D22, and a true / false flag D23 corresponding to each of the plurality of image data are associated with each other. The judgment target string D22 of the second teacher data D2 includes either (1) a correct answer string corresponding to the image data D21 associated with the judgment target string D22, or (2) a string that is erroneously recognized based on the image data D21 in the learning process or the inference process. The true / false flag is a flag that indicates whether the correct answer string and the judgment target string match or not.
[0023] The image data (D11 and D12) in the first teacher data D1 and the second teacher data D2 are, for example, images that include a character string indicating a brand or the like attached to a product or the like. The correct character string in this example is a character string attached to the product or the like that appears in the image. The first teacher data D1 and the second teacher data D2 may be composed of a common image, or may each include different images.
[0024] By learning in this manner, the trained model takes image data containing a character string as input and outputs a predicted character string, which is a character string predicted to be contained in the image data.
[0025] By configuring the information processing system S in this way, it is expected that the accuracy of recognizing character strings included in an image can be improved. In particular, by having the information processing device 1 learn a correct / incorrect judgment task based on the second teacher data D2 including a character string (a character string different from a correct character string) that was incorrectly recognized by the information processing device 1 during the learning or inference process and image data of the target from which the character string was recognized, it is possible to learn the tendency of incorrect recognition, and it is expected that the accuracy of recognizing character strings can be improved.
[0026] [Configuration of information processing device 1] 4 is a block diagram showing the configuration of the information processing device 1. The information processing device 1 has a communication unit 11, a storage unit 12, and a control unit 13. The control unit 13 has an acquisition unit 131, a learning unit 132, an output unit 133, a display control unit 134, a recognition means control unit 135, and a generation unit 136.
[0027] The communication unit 11 is a communication interface for transmitting and receiving data to and from other devices via a network. The storage unit 12 is a storage medium including a ROM (Read Only Memory), a RAM (Random Access Memory), an SSD (Solid State Drive), a hard disk drive, etc. The storage unit 12 stores in advance a program to be executed by the control unit 13.
[0028] The control unit 13 is a processor such as a CPU (Central Processing Unit), etc. The control unit 13 executes a program stored in the storage unit 12, thereby functioning as an acquisition unit 131, a learning unit 132, an output unit 133, a display control unit 134, a recognition means control unit 135, and a generation unit 136.
[0029] The acquisition unit 131 acquires the first teacher data D1 and the second teacher data D2. As an example, the acquisition unit 131 acquires the first teacher data D1 and the second teacher data D2 from the information terminal 2. The acquisition unit 131 may acquire the first teacher data D1 and the second teacher data D2 from an external device (not shown).
[0030] The learning unit 132 learns a string recognition task based on the first teacher data D1. The learning unit 132 also learns a true / false judgment task based on the second teacher data D2. The processing of the learning unit 132 will be described with reference to FIG. 5. The learning unit 132 receives image data D11 associated in the first teacher data D1 as input, and outputs a character string included in the image data D11. As an example, in the learning model, the image data D11 is divided into a plurality of patches D31, and converted into a vector by taking a projection of each patch. As an example, the end of the input data includes data indicating that it is the end of the input data ([SEP] in FIG. 5).
[0031] The learning unit 132 inputs the prediction result output based on the input data to the learned model in an autoregressive manner. The learning unit 132 concatenates a vector corresponding to the patch D31 and a prediction result D41 output by the learning model based on the vector input immediately before, and inputs the concatenated vector to the learning model, outputting the prediction result D41. The learning model repeats the process in an autoregressive manner until a vector indicating the end of prediction ([EOS] in FIG. 5) is output.
[0032] When the prediction is completed, the learning unit 132 updates the parameters of the learning model based on the difference between the character string output by the learning model and the correct character string D12 associated in the first teacher data D1. As an example, the learning unit 132 calculates a cross-entropy error between a vector indicating the character string output by the learning model and a vector corresponding to the correct character string D12, and updates the parameters of the learning model based on the calculated cross-entropy error.
[0033] Learning of the correct / incorrect judgment task will be described with reference to FIG. 6. The learning unit 132 inputs the image data D21 and the character string D22 to be judged, which are associated in the second teacher data D2, to the learning model, and outputs a correct / incorrect flag D42. The learning unit 132 converts the character string D22 to be judged into a vector, and forms input data D32 by concatenating a vector corresponding to the patch into which the image data D21 is divided and a vector corresponding to the character string D22 to be judged. In this case, the input data D32 includes data indicating a data separator between the end of the patch and the character string to be judged ([SEP] in FIG. 6). The learning unit 132 updates the parameters of the learning model based on the difference between the correct / incorrect flag D42 output by the learning model and the correct / incorrect flag D23 associated in the second teacher data D2. The learning unit 132 updates the parameters of the learning model based on the cross entropy error between the correct / incorrect flag D42 output by the learning model and the correct / incorrect flag D23 associated in the second teacher data D2.
[0034] By configuring the information processing device 1 in this way, it is possible to improve the accuracy of recognizing character strings included in an image.
[0035] The learning unit 132 trains the character string recognition task and the correctness determination task in a single learning process to generate a trained model. The learning unit 132 trains the learning model on both the character string recognition task and the correctness determination task in the process from the start to the end of learning. The learning unit 132 may train the character string recognition task and the correctness determination task in sequence or in parallel.
[0036] By learning the prediction results of character strings output during the learning process as character strings to be judged in the second training data D2, it is possible to learn the tendency of the trained model to make erroneous recognition.
[0037] The generation unit 136 generates second teacher data D2 in which a string output by the trained model during the process in which the learning unit 132 trains the string recognition task and which differs from a string contained in the input image data is treated as a recognition target string and associated with the image data input to the trained model.
[0038] The information processing device 1 may be configured to determine whether the recognized characters are correct or not in a single prediction process.
[0039] In this case, in addition to the above learning, the learning unit 132 causes the learning model to learn a true / false judgment task based on the first teacher data D1. In other words, the learning unit 132 inputs a character string output by the learning model in response to the input image data D11 to the learning model as a character string to be judged, and causes the learning model to learn a true / false judgment task.
[0040] The process of the learning unit 132 in this case will be described with reference to FIG. 7. In this case, the learning unit 132 inputs a vector corresponding to the patch D31 obtained by dividing the image data D11 associated in the first teacher data D1 to the learning model, and causes the learning model to output a prediction result. The learning unit 132 inputs the output prediction result to the learning model in an autoregressive manner. When the prediction of the character string is completed, the learning model outputs information indicating the end of the character string ([SEP] in FIG. 7). When the information indicating the end of the character string is input, the learning model outputs a true / false flag based on the input patch and the prediction result of the character string. The learning unit 132 updates the parameters of the learning model based on the difference between the character string and the true / false flag output by the learning model, and the correct character string D12 associated with the image data D11 in the first teacher data and the true / false flag of the value of "TRUE", thereby generating a trained model. The trained model receives target image data, which is image data including a character string to be recognized, as input, and outputs a predicted character string and a true / false flag.
[0041] The acquisition unit 131 acquires target image data. The target image data is image data to be subjected to character recognition by the information processing device 1, and includes a character string to be recognized. The acquisition unit 131 may acquire the target image data from the information terminal 2. The acquisition unit 131 may acquire the target image data from an external device (not shown).
[0042] The output unit 133 inputs the target image data acquired by the acquisition unit 131 to the trained model, and outputs a predicted character string for the target image data and a true / false flag for the predicted character string. The display control unit 134 may cause the information terminal 2 to display the predicted character string output by the trained model and the true / false flag output by the trained model.
[0043] The information processing device 1 may be configured to display information to alert the user when there is a possibility that the recognition result based on the trained model is erroneous.
[0044] When the true / false flag output by the output unit 133 indicates that the predicted character string is incorrect, the display control unit 134 displays a message indicating that the prediction result may be incorrect. As an example, the display control unit 134 determines whether the true / false flag output by the trained model is "FALSE". When the true / false flag is "FALSE", the display control unit 134 displays a screen shown in FIG. 8 on the information terminal 2. The screen shown in FIG. 8 displays the target image data, the predicted character string corresponding to the target image data, and a message indicating that the prediction result may be incorrect.
[0045] The information processing device 1 is configured to learn the tendency of erroneous recognition and output information indicating that there is a possibility of erroneous recognition, so that the user can recognize that caution is required.
[0046] If there is a possibility that the recognition result of the trained model is erroneous, the information processing device 1 may be configured to input the result to another OCR.
[0047] When the correct / incorrect flag output by the output unit 133 indicates that the predicted character string is incorrect, the recognition means control unit 135 inputs the target image data to an optical character recognition means that outputs characters contained in the image data when the image data is input, and that is different from the trained model. As an example, when the correct / incorrect flag is "FALSE", the recognition means control unit 135 may input the target image data to an optical character recognition means (not shown). In addition, the storage unit 12 may store a character recognition model that is a trained model trained to perform character recognition based on a data set different from the trained model, and in this case, the recognition means control unit 135 may input the target image data to the character recognition model and output the recognition result.
[0048] The display control unit 134 acquires a character string recognized by the optical character recognition means performing optical character recognition on the target image data, and displays the acquired character string. As an example, the display control unit 134 may acquire a character string as a recognition result from the optical character recognition means to which the recognition means control unit 135 inputs the target image data. In addition, when the recognition means control unit 135 inputs the target image data to a character recognition model, the display control unit 134 acquires the recognition result output by the character recognition model. The display control unit 134 displays the acquired character string and the character string output by the trained model on the information terminal 2.
[0049] In addition, when there are multiple OCRs as input candidates, the trained model may be configured to further output OCR identification information for identifying the input destination OCR. In this case, the trained model is trained based on training data in which image data, a recognition character string, a true / false flag, and OCR identification information are associated. In this case, the OCR identification information included in the training data indicates an OCR suitable for character recognition of image data.
[0050] The information processing device 1 configured in this way has the effect of being able to perform character recognition using OCR that is more suitable for predicting character strings.
[0051] [Processing flow in information processing device 1] Fig. 9 is a flowchart showing the flow of processing in the information processing device 1. The flowchart shown in Fig. 9 starts at the point when the information processing device 1 receives an instruction to start learning from the information terminal 2.
[0052] The acquisition unit 131 acquires first teacher data (S01). As an example, the acquisition unit 131 acquires the first teacher data from the information terminal 2. The acquisition unit 131 acquires second teacher data (S02). As an example, the acquisition unit 131 acquires the second teacher data from the information terminal 2.
[0053] The learning unit 132 causes the character recognition task to be learned based on the first training data (S03). The learning unit 132 updates parameters of the learning model based on the difference between the output in the character recognition task and the correct character string D12 associated in the first training data.
[0054] The learning unit 132 causes the model to learn a true / false judgment task based on the second teacher data (S04). The learning unit 132 updates parameters of the learning model based on a difference between an output in the true / false judgment task and a true / false flag associated with the second teacher data D2.
[0055] The learning unit 132 causes the model to learn the true / false judgment task based on the first teacher data (S05). The learning unit 132 updates the parameters of the learning model based on the difference between the character string and true / false flag output in the true / false judgment task and the correct character string D12 associated with the first teacher data D1 and the true / false flag with a value of “TRUE”.
[0056] The learning unit 132 updates the parameters to generate a trained model, which is a trained model for which training has been completed, and stores the trained model in the storage unit 12 (S06). The information processing device 1 ends the process.
[0057] [Effects of the information processing device 1] By configuring the information processing device 1 in this way, it is expected that the accuracy of recognizing character strings included in an image can be improved.
[0058] Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by distributing or integrating functionally or physically in any unit. In addition, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effect of the new embodiment resulting from the combination combines the effect of the original embodiment. [Explanation of symbols]
[0059] 1. Information processing device 2. Information terminal 11 Communications Department 12 Storage section 13 Control section 131 Acquisition Department 132 Learning Department 133 Output section 134 Display control section 135 Recognition means control section 136 Generation part
Claims
1. (1) first teacher data associating a plurality of image data including character strings with correct answer character strings that are character strings included in each of the plurality of image data; (2) (a) the plurality of image data, (b) a recognition target character string which is a character string to be recognized, and (c) second teacher data in which a flag indicating whether or not a correct answer character string corresponding to each of the plurality of image data matches the recognition target character string is associated with the second teacher data; An acquisition unit for acquiring (1) a character string recognition task that is a task of inputting image data associated with the first teacher data and outputting a character string included in the image data based on the first teacher data; (2) a true / false determination task that is a task for determining whether or not a correct answer string corresponding to the image data matches the recognition target string, using image data associated in the second training data and the recognition target string as input based on the second training data; A learning unit that generates a trained model that is trained by learning the above, and that receives image data including a character string and outputs a predicted character string that is a character string predicted to be included in the image data; An information processing device having the above configuration.
2. In the character string to be recognized, (1) a correct answer character string corresponding to image data associated with the recognition target character string; and (2) a character string that is erroneously recognized based on the image data; and Including any of the following: The information processing device according to claim 1 .
3. The learning unit causes the character string recognition task and the correct / incorrect determination task to be learned in a single learning process, and generates the trained model. The information processing device according to claim 1 .
4. The trained model receives target image data, which is image data including a character string to be recognized, and outputs a predicted character string and a correct / incorrect flag indicating whether the predicted character string is correct or not; The acquisition unit further acquires target image data which is image data including a character string to be recognized; The information processing device further includes an output unit that inputs the target image data acquired by the acquisition unit into the trained model and outputs a predicted character string for the target image data and a true / false flag for the predicted character string. The information processing device according to claim 1 .
5. The present invention further includes a display control unit that, when the correct / incorrect flag output by the output unit indicates that the predicted character string is incorrect, displays a message indicating that the prediction result may be incorrect. The information processing device according to claim 4.
6. When the correct / incorrect flag output by the output unit indicates that the predicted character string is incorrect, an optical character recognition unit outputs characters contained in the image data when the image data is input, and a recognition unit control unit inputs the target image data to the optical character recognition unit different from the trained model, The optical character recognition means acquires a character string recognized by performing optical character recognition on the target image data, and a display control unit displays the acquired character string. The information processing device according to claim 4.
7. The method further includes a generation unit that generates the second teacher data by associating a character string output by the trained model in a process in which the learning unit learns the character string recognition task, the character string being different from a character string included in the input image data, with the image data input to the trained model as the recognition target character string. The information processing device according to claim 2 .
8. The computer executes a first acquisition step of acquiring first teacher data in which a plurality of image data including character strings and a correct answer character string that is a character string included in each of the plurality of image data are associated with each other; a second acquisition step of acquiring second teacher data in which (a) the plurality of image data, (b) a recognition target character string which is a character string to be recognized, and (c) a flag indicating whether or not a correct answer character string corresponding to each of the plurality of image data matches the recognition target character string; (1) a character string recognition task that is a task of inputting image data associated with the first teacher data and outputting a character string included in the image data based on the first teacher data; (2) a true / false determination task that is a task for determining whether or not a correct answer string corresponding to the image data matches the recognition target string, using image data associated in the second training data and the recognition target string as input based on the second training data; A learning step of generating a trained model that is trained by learning the above, the trained model inputting image data including a character string and outputting a predicted character string that is a character string predicted to be included in the image data; An information processing method comprising the steps of:
9. On the computer, a first acquisition step of acquiring first teacher data in which a plurality of image data including character strings and a correct answer character string that is a character string included in each of the plurality of image data are associated with each other; a second acquisition step of acquiring second teacher data in which (a) the plurality of image data, (b) a recognition target character string which is a character string to be recognized, and (c) a flag indicating whether or not a correct answer character string corresponding to each of the plurality of image data matches the recognition target character string; (1) a character string recognition task that is a task of inputting image data associated with the first teacher data and outputting a character string included in the image data based on the first teacher data; (2) a true / false determination task that is a task for determining whether or not a correct answer string corresponding to the image data matches the recognition target string, using image data associated in the second training data and the recognition target string as input based on the second training data; A learning step of generating a trained model that is trained by learning the above, the trained model inputting image data including a character string and outputting a predicted character string that is a character string predicted to be included in the image data; A program that executes the following.