Information processing apparatus, control method for information processing apparatus, and control program for information processing apparatus
The information processing apparatus addresses the challenge of accurately reading non-standard identification documents by using a combination of printed character recognition, keyword specification, and handwritten character analysis to achieve high accuracy in character recognition.
Patent Information
- Application Number
- JP2024044992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-03-21
AI Technical Summary
Existing OCR technologies struggle with high accuracy in reading non-standard identification documents, such as maternal and child health handbooks, which have varying formats and a mix of handwritten and printed items.
An information processing apparatus that acquires images of documents, recognizes printed characters, specifies keyword positions, sets target areas for handwritten strings, extracts and recognizes handwritten characters, calculates recognition reliability, and determines the most reliable handwritten string for detection.
Achieves highly accurate recognition of characters in documents with non-standard formats, regardless of whether they are handwritten or printed, by utilizing a combination of printed character recognition, keyword specification, and handwritten character analysis.
Smart Images

Figure 0007698089000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, a control method for the information processing apparatus, and a control program for the information processing apparatus.
Background Art
[0002] Conventionally, by OCR (Optical Character Recogination / Reader), characters described in each item of identification documents such as driver's licenses and insurance certificates have been read from an image of the identification document (for example, Patent Document 1). In recent years, with the development of artificial intelligence technology, the use of AI-OCR (Artificial Intelligence - Optical Character Recognition) that utilizes AI technologies such as deep learning to improve character recognition accuracy and layout analysis accuracy has also been spreading.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, among identification documents, there are non-standard documents such as maternal and child health handbooks whose formats vary depending on the municipality, and some documents have a mixture of handwritten items and printed items. In such cases, it has been difficult to perform highly accurate reading even using AI-OCR.
[0005] There is a demand for highly accurate recognition of characters described in documents regardless of whether they are standard or non-standard, and whether they are handwritten or printed.
Means for Solving the Problems
[0006] An information processing apparatus according to an embodiment of the present invention includes an acquisition unit that acquires an image of a document including printed characters and handwritten characters, a printed character recognition unit that recognizes printed characters in the image, a specifying unit that specifies the positions of keywords in the image based on the recognition results of the printed character recognition unit, a setting unit that sets a target area, which is an area including a handwritten string to be detected, based on the positions of the keywords and the features of the image, an extraction unit that extracts a plurality of strings included in the target area, a handwritten character recognition unit that recognizes the handwritten characters of the plurality of strings and calculates the reliability of the recognition results, a determination unit that determines the handwritten string to be detected among the plurality of strings based on the reliability, and an output unit that outputs the recognition results of the handwritten character recognition unit for the handwritten string to be detected determined by the determination unit.
[0007] In the information processing apparatus according to an embodiment of the present invention, the image includes a plurality of handwritten strings to be detected, and the setting unit may use the position of one target area including one handwritten string to be detected, which is specified by the position of the keyword, as the position of the keyword when setting another target area including another handwritten string to be detected.
[0008] In the information processing apparatus according to an embodiment of the present invention, the setting unit may set the target area based on grid lines included in the image as a feature of the image.
[0009] In the information processing apparatus according to an embodiment of the present invention, the acquisition unit may further acquire information regarding the type of the document, and the specifying unit may specify the positions of keywords according to the type of the document.
[0010] In the information processing apparatus according to an embodiment of the present invention, the extraction unit may combine bounding boxes surrounding characters or strings included in the target area based on predetermined conditions and extract them as a plurality of strings.
[0011] In an information processing apparatus according to an embodiment of the present invention, a setting unit sets, as a target area, an area including a printed character string to be detected. A printed character recognition unit recognizes printed characters of a plurality of character strings included in the target area extracted by an extraction unit, calculates a confidence level of the recognition result of the printed characters, and a determination unit determines, based on the confidence level of the recognition result of the printed characters, the printed character string to be detected among the plurality of character strings. An output unit may output the recognition result by the printed character recognition unit for the printed character string to be detected determined by the determination unit.
[0012] In an information processing apparatus according to an embodiment of the present invention, the document may be a Maternal and Child Health Handbook.
[0013] A control method of an information processing apparatus according to an embodiment of the present invention causes the information processing apparatus to execute an acquisition step of acquiring an image of a document including printed characters and handwritten characters, a printed character recognition step of recognizing the printed characters in the image, a specifying step of specifying the position of a keyword in the image based on the recognition result in the printed character recognition step, a setting step of setting, as a target area, an area including a handwritten character string to be detected based on the position of the keyword and the features of the image, an extraction step of extracting a plurality of character strings included in the target area, a handwritten character recognition step of recognizing handwritten characters of the plurality of character strings and calculating a confidence level of the recognition result, a determination step of determining, based on the confidence level, the handwritten character string to be detected among the plurality of character strings, and an output step of outputting the recognition result by the handwritten character recognition step for the handwritten character string to be detected determined by the determination step.
[0014] A control program for an information processing apparatus according to an embodiment of the present invention causes the information processing apparatus to implement an acquisition function for acquiring an image of a document including printed characters and handwritten characters, a printed character recognition function for recognizing printed characters in the image, a specification function for specifying the positions of keywords in the image based on the recognition results of the printed character recognition function, a setting function for setting a target area, which is an area including a handwritten string to be detected, based on the positions of the keywords and the features of the image, an extraction function for extracting a plurality of strings included in the target area, a handwritten character recognition function for recognizing the handwritten characters of the plurality of strings and calculating the reliability of the recognition results, a determination function for determining, based on the reliability, the handwritten string to be detected among the plurality of strings, and an output function for outputting the recognition results of the handwritten character recognition function for the handwritten string to be detected determined by the determination function.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an embodiment of the invention according to the present disclosure (also referred to as the present invention) will be described with reference to the drawings. Note that the drawings are merely examples, and the present invention is not limited to those shown in the drawings. For example, the number of servers, user terminals, etc., data sets (tables), flowcharts, and documents shown are merely examples, and the present invention is not limited thereto.
[0017] <System Configuration> FIG. 1 is a diagram showing a configuration example of an information processing system according to an embodiment of the present invention. The information processing system 600 is a system that extracts information for determining whether a user satisfies the conditions for using a predetermined service from an image 10 of the user's certificate documents. That is, in the information processing system 600, necessary character information (text information) may be output from the image 10 by performing character recognition processing on the image 10. The certificate documents may vary depending on the predetermined service provided. For example, in addition to public personal identification documents such as driver's licenses, health insurance cards, passports, my number cards, resident registers, mother and child health handbooks, national pension handbooks, student IDs, and school notebooks, employee IDs, business cards, postal items, deposit passbooks, various forms, etc. may also be used.
[0018] The information processing system 600 includes a server (character recognition device) 100, a service server 101, and one or more user terminals 200. The server 100 is an information processing device (character recognition device) that performs character recognition processing. The service server 101 is an information processing device that performs processing related to a predetermined service provided to the user. The service server 101 is connected to the server 100 and passes the image 10 transmitted from the user terminal 200 to the server 100. The server 100 performs character recognition processing on the image of the document 10 transmitted from the user terminal 200 and outputs the result (character recognition result) of extracting the characters included in the document 10.
[0019] In FIG. 1, the server 100 and the service server 101 are shown as separate entities, but this is not limiting. That is, each function described as being provided by the server 100 may be realized by a plurality of servers or by a single server. Further, the server 100 may be, for example, a distributed server system that cooperates by communicating via a network, or a so-called cloud server. That is, the server 100 is not limited to a physical server and may include a virtual server implemented by software. Further, the server 100 may be any device that can realize the functions described in each embodiment, and may include, for example, a server device, a computer (by way of non-limiting example, a desktop, a laptop, a tablet, etc.), a communication platform, etc.
[0020] The user terminal 200 and the service server 101 are connected via the network 500. The network 500 includes a wireless network and a wired network. Specifically, for example, the network 500 may be a wireless LAN (WLAN), a wide area network (WAN), CDMA (code division multiple access), LTE (long term evolution), LTE-Advanced, 4th generation communication (4G), 5th generation communication (5G), and a mobile communication system such as 6th generation communication (6G) and later. Note that the network 500 is not limited to these examples, and may normally be, for example, Bluetooth (registered trademark), an optical line, etc. Further, the network 500 may be a combination of these.
[0021] In FIG. 1, a smartphone is shown as the user terminal 200, but the user terminal 200 may be any terminal that can transmit the image 10 to the service server 101.
[0022] Next, with reference to FIG. 2, the hardware configuration and functional configuration of the server 100 will be described. (1) Hardware Configuration of the Server Server 100 includes a control unit 110, a communication unit 120, an input / output unit 130, and a storage unit 170.
[0023] The control unit 110 is typically a processor, including a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), a microprocessor, etc., and is realized by a logic circuit (hardware) formed on an integrated circuit (IC (Integrated Circuit) chip, LSI (Large Scale Integration)) or a dedicated circuit. The control unit 110 reads the program (software) stored in the storage unit 170 and executes the code and instructions included in the program to execute the functions and methods shown in each embodiment.
[0024] The storage unit 170 stores various programs and various data required for the operation of the server 100. The storage unit 170 may include, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. Further, the storage unit 170 may include a memory (RAM (Random Access Memory), ROM (Read Only Memory), etc.) that provides a working area for the control unit 110.
[0025] The communication unit 120 is implemented as hardware such as a network adapter, communication software, and combinations thereof. The communication unit 120 transmits and receives various data to and from the service server 101 via the network 500.
[0026] The input / output unit 130 includes an input device that receives various operations from an administrator or the like for the server 100, and an output unit that outputs the processing result processed by the server 100. The input device may include, for example, a keyboard and a microphone, and the output device may include, for example, a display and a speaker.
[0027] (2) Functional Configuration of the Server As functions realized by the control unit 110, the server 100 may include an acquisition unit 111, a character detection unit 112, a printed character recognition unit 113, a handwritten character recognition unit 114, a specifying unit 115, a setting unit 116, an extraction unit 117, and a determination unit 118. Among the functional units described in FIG. 2, functional units that are not essential in each of the embodiments described hereinafter may not be provided. Also, the functions or processes of each functional unit may be realized by machine learning or AI within a realizable range.
[0028] <Server Control Flow> Together with the description of each functional unit of the server 100, the control flow of the server 100 according to an embodiment of the present invention will be described with reference to FIGS. 3 to 5 as well.
[0029] The acquisition unit 111 acquires an image 10 of a document including printed characters and handwritten characters (step S11). Hereinafter, the service server 101 provides a service that requires proof of parent-child relationship, and the image 10 will be described as being related to the proof of birth notification in the Maternal and Child Health Handbook. Specifically, the character recognition process according to an embodiment of the present invention may aim to detect the name of the parent, the name of the child, and the child's date of birth from the image of the proof of birth notification in order to determine whether the information acquired at the time of application for the service is correct. However, the documents targeted by the present invention are not limited to the Maternal and Child Health Handbook.
[0030] Fig. 4(a) shows an example of Image 10, which is an image of a page (hereinafter also simply referred to as the "birth notification payment certificate") including the already filled-in birth notification payment certificate in the maternal and child health handbook. As described above, the format of the birth notification certificate is not unified by the local government, and the locations for writing the names of the parents and the child are different. Also, in some local governments, a seal with the filled-in birth notification certificate is pasted at the original location of the birth notification certificate. At this time, since the paper of the seal is thin, the characters under the seal may be seen through. For these reasons, it is difficult to recognize the names of the parents and the child and the date of birth that need to be detected. Furthermore, factors that make the character recognition process difficult include the mixture of printed characters and handwritten characters, the furigana written above the name, and the presence of circled characters in the same line as the character string, etc.
[0031] The printed character recognition unit 113 recognizes the printed characters in the image 10 (step S12). Note that prior to the recognition of the printed characters by the printed character recognition unit 113, the character detection unit 112 detects the characters in the image 10. Note that a known character detection method may be used for the character detection process. For example, it may be by deep learning (deep neural network) (R-CNN: Recrrent CNN (Convolutional Neural Networks), YOLO, SSD) or by the sliding window method, etc. Also, for the character recognition process, a known character recognition method may be used. A character recognition model set by prior learning by deep learning may be used, or a pattern matching method or the like may be used.
[0032] Based on the recognition result by the printed character recognition unit 113, the specific part 115 specifies the position of the keyword in the image 10 (step S13). A keyword is a character or object for estimating the position where the target detection object (here, the name of the parent, the name of the child, and the date of birth) is described, and serves as an anchor. As shown in FIG. 4(a), in the birth registration certificate, characters such as "mother", "father", and "name" exist adjacent to the column for entering the name of the parent. Also, characters such as "child's name" and "birth registration certificate" exist adjacent to the column for entering the name of the child. These keywords may be stored in the storage unit 170 for each type of document.
[0033] FIG. 5 is an example of a data table that stores keywords when the document is a birth registration certificate and is stored in the storage unit 170. Note that the figure is an example, and there may be more keywords. The target area will be described later. The specific part 115 may specify the characters stored in the data table TB10 among the characters recognized by the printed character recognition unit 113 as the keyword 31.
[0034] The setting unit 116 sets a target area, which is an area including the handwritten string of the detection target, based on the position of the keyword 31 and the features of the image 10 (step S14). Here, the data table TB10 in FIG. 5 also stores information regarding the relative positional relationship between the position of the keyword and the target area. For example, the data table TB10 indicates that there is a target area including the "name of the parent", which is the detection target, on the right side of the keyword "mother" (here, the right side as viewed from above the image 10). Note that the right side is used for the purpose of explanation, and actually, the relative position of the target area with respect to the keyword is stored in the data table TB10 in a format that can be understood by the information processing device.
[0035] Fig. 4(b) shows an example of the target area 33 set according to the position of the keyword. Note that the setting unit 116 may use the ruled lines that divide each item and entry field of the birth notification certificate, which are detected by image processing of the image 10, for setting the target area. That is, the target area 33 is detected on the right side of the keyword "mother" and is set as a rectangular area (cell) divided by ruled lines.
[0036] The extraction unit 117 extracts a plurality of character strings included in the target area (step S15). This will be described with reference to Fig. 4(c). As a first step, the extraction unit 117 may extract the character strings included in the target area 33 using bounding boxes. Fig. 4(c) shows an example in which bounding boxes b1 to b4 are extracted. Note that since the bounding boxes are known, the description thereof is omitted. As a second step, the extraction unit 117 combines at least one or more bounding boxes to extract a plurality of character strings. Note that the extraction unit 117 may combine the bounding boxes surrounding the characters included in the target area based on a predetermined condition and extract them as a plurality of character strings.
[0037] Here, a predetermined condition regarding the combination of bounding boxes will be described. First, candidate bounding boxes may be tentatively determined from their heights and position information and grouped by row. As a result, rows of typeset characters and rows of handwritten characters can be distinguished, and furigana and Chinese characters can be distinguished. Also, based on the height of the row and the relative position from the anchor, etc., rows may be combined or split and selected as the final rows. Further, the predetermined condition may be, for example, a condition for predicting the date of birth, separating furigana from the name, removing circled characters, and combining the surname and given name. Specifically, when predicting the date of birth, the year, month, and day are typeset characters, and the numbers are handwritten characters. At this time, it is considered that handwritten characters tend to be larger than typeset characters, and after appropriately selecting the combination of recognized characters and their heights, they may be combined. Also, when separating furigana from the name, it may be split by deleting the central part among the overlapping parts of the bounding boxes. Also, when removing circled characters, for example, in the case of a mother and child handbook, since there are circles for male and female, it may be deleted when "male and female" is within the same bounding box or when only "female" is recognized and the male cannot be read. Furthermore, when combining the surname and given name, it may be determined whether to combine the bounding boxes according to the height and position of the bounding boxes.
[0038] For example, in FIG. 4(c), it is assumed that the bounding boxes located in the same row are combined, and the character strings of "SoftBank" consisting of only bounding box b3, "SoftBank Hanako" consisting of bounding boxes b3 and b4, and "Hanako" consisting of only bounding box b4 may be extracted. Also, the character strings of "Nanyin" consisting of only bounding box b1, "Nanyin Hanako" consisting of bounding boxes b1 and b2, and "Hanako" consisting of only bounding box b2 may be extracted. Note that the process of the extraction unit 117 conforms to the process of extracting the region of interest (ROI) in object detection.
[0039] The handwritten character recognition unit 114 recognizes the handwritten characters of the plurality of character strings extracted by the extraction unit 117 and calculates the confidence of the recognition result (step S16). As the recognition process of handwritten characters, existing methods may be used, such as using a model trained by machine learning with a dataset of handwritten characters, or a pattern matching method using a dictionary of handwritten characters. The confidence of the recognition result is a numerical value indicating the certainty of the recognized handwritten characters, and may be, for example, a confidence score used in YOLO. According to an embodiment of the present invention, a score of the certainty (how reliable it is) for the prediction result may be used, and when the score is lower than the threshold value, the recognition result may be returned as "no character string". For example, in an area where blanks are allowed, such as when only one of two parent name entry fields is filled, the above logic may be used.
[0040] The determination unit 118 determines the handwritten character string to be detected among the plurality of character strings based on the confidence (step S17). That is, the character string with the highest confidence may be used as the detection target. The output unit 119 outputs the recognition result by the handwritten character recognition unit 114 for the handwritten character string to be detected determined by the determination unit 118 (step S18).
[0041] As described above, according to an embodiment of the present invention, a target area including the character string to be detected is set from the keyword specified by the recognition process of printed characters with a lower processing load compared to the recognition of handwritten characters. Therefore, it is possible to highly accurately estimate the position where the object to be detected is described for documents with different formats.
[0042] Note that the setting unit 116 may use the position of a target area including a handwritten string of one detection target, which is specified by the position of the keyword, as the position of the keyword when setting another target area including a handwritten string of another detection target. This will be described with reference to FIG. 4(b). For example, assume that the keyword "mother" is not recognized by the printed character recognition unit 113 due to light reflection or the like during the shooting of image 10. At this time, the setting unit 116 may set a target area 33 including the mother's name using the target area 34 set using "father" as the keyword. In this case, prior information that the mother's name is written above the father's name may be used.
[0043] Thereby, even when the estimation accuracy of the position of the keyword by printed characters is low, the target area including the string to be detected can be estimated.
[0044] Note that in one embodiment of the present invention, the acquisition unit 111 may further acquire information regarding the type of document, and the specifying unit 115 may specify the position of the keyword according to the type of document. That is, information on keywords serving as anchors may be stored for each type of document in the storage unit 170 and may be used for specifying the position of the keyword by the specifying unit 115. Thereby, it becomes possible to detect the string to be detected for various types of documents.
[0045] Note that the above-described processing according to one embodiment of the present invention may be realized by an End-to-End model based on deep learning. The End-to-End model is a method of training an input and an output in one pipeline. FIG. 6 shows an example of the pipeline processing for character recognition according to one embodiment of the present invention.
[0046] When pipeline 20 receives the input of image 10, the detection process (S1) of printed characters and the detection process (S4) of handwritten characters are performed. After the recognition process (S3) is performed on the detected printed characters, they are used for the extraction of the target area (S3). For the extracted target area, a process (S4) for removing noise other than the character string and a refinement process (S5) of the target area are performed. Note that the result of the handwritten character detection (S4) may also be used for the refinement process of the target area. Here, the refinement process is an improvement and division process of the target area, and may be a process of bringing the target area closer to the detection target or a process of dividing overlapping target areas. Thereafter, a recognition process (S6) is performed on the handwritten characters included in the target area, and the recognition result is output through a filter process (S7). Note that the filter process may include removal of noise such as furigana and circled characters, correction of the date format, and the like.
[0047] Although the present invention has been described based on the drawings and embodiments, it should be noted that those skilled in the art can easily make various modifications and corrections based on the present disclosure. Therefore, it should be noted that these modifications and corrections are included in the scope of the present invention. For example, the functions included in each component, each step, etc. can be rearranged so as not to be logically inconsistent, and a plurality of components and steps can be combined into one or divided. Also, it is also possible to appropriately combine the configurations shown in the above embodiments. For example, each component described as being provided in server 100 may be realized by being distributed among a plurality of servers.
[0048] The program of each embodiment of the present disclosure may be provided in a state stored in a storage medium readable by an information processing apparatus. The storage medium can store the program in a "non-transitory tangible medium". The program includes, for example, software programs and control programs. When each functional unit of the server 100 is realized by software, the server 100 functions as an acquisition unit 111, a character detection unit 112, a printed character recognition unit 113, a handwritten character recognition unit 114, a specifying unit 115, a setting unit 116, an extraction unit 117, a determination unit 118, and an output unit 119 by executing the program loaded on the memory by the processor.
[0049] When appropriate, the storage medium can include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs), application specific ICs (ASICs), etc.), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid state drives (SSDs), RAM drives, secure digital cards or drives, any other appropriate storage medium, or any suitable combination of two or more of these. The storage medium can be volatile, non-volatile, or a combination of volatile and non-volatile when appropriate.
[0050] Also, the program of the present disclosure may be provided to the server 100 via any transmission medium (such as a communication network or a broadcast wave) capable of transmitting the program.
[0051] Also, each embodiment of the present disclosure can also be realized in the form of a data signal embedded in a carrier wave in which the program is embodied by electronic transmission. Note that the program of the present disclosure may be implemented using, for example, script languages such as JavaScript (registered trademark), Python, C language, Go language, Swift, Kotlin, Java (registered trademark), etc.
[0052] According to each aspect of the present disclosure described above, by providing a highly convenient personal authentication method, it is possible to contribute to the achievement of Sustainable Development Goal (SDG) 11, "Sustainable cities and communities".
Explanation of symbols
[0053] 100 Server (information processing device) 111 Acquisition unit 112 Character detection unit 113 Typeface character recognition unit 114 Handwritten character recognition unit 115 Identification unit 116 Setting unit 117 Extraction unit 118 Decision unit 119 Output unit 120 Communication unit 130 Input / output unit 170 Storage device 101 Service server 200 User terminal (terminal device) 500 Network 600 Information processing system
Claims
1. An acquisition unit for acquiring an image of a document including printed characters and handwritten characters; a printed character recognition unit for recognizing printed characters in the image; an identification unit that identifies a position of a keyword within the image based on a recognition result by the printed character recognition unit; a setting unit that sets a target area that is an area including a handwritten character string to be detected based on a position of the keyword and a ruled line included in the image; An extraction unit that extracts a plurality of character strings included in the target area; a handwritten character recognition unit that recognizes handwritten characters of the plurality of character strings and calculates a reliability of a recognition result; a determination unit that determines a handwritten character string to be detected from among the plurality of character strings based on the reliability; an output unit that outputs a recognition result by the handwritten character recognition unit for the handwritten character string to be detected that has been determined by the determination unit; An information processing device, wherein the document is a maternal and child health handbook.
2. An acquisition unit for acquiring an image of a document including printed characters and handwritten characters; a printed character recognition unit for recognizing printed characters in the image; an identification unit that identifies a position of a keyword within the image based on a recognition result by the printed character recognition unit; a setting unit that sets a target area that is an area including a handwritten character string to be detected based on a position of the keyword and a ruled line included in the image; An extraction unit that extracts a plurality of character strings included in the target area; a handwritten character recognition unit that recognizes handwritten characters of the plurality of character strings and calculates a reliability of a recognition result; a determination unit that determines a handwritten character string to be detected from among the plurality of character strings based on the reliability; an output unit that outputs a recognition result by the handwritten character recognition unit for the handwritten character string to be detected that has been determined by the determination unit; The acquisition unit further acquires information regarding the type of the document, The identification unit is an information processing device that identifies a position of a keyword according to a type of the document.
3. An acquisition unit for acquiring an image of a document including printed characters and handwritten characters; a printed character recognition unit for recognizing printed characters in the image; an identification unit that identifies a position of a keyword within the image based on a recognition result by the printed character recognition unit; a setting unit that sets a target area that is an area including a handwritten character string to be detected based on a position of the keyword and a ruled line included in the image; An extraction unit that extracts a plurality of character strings included in the target area; a handwritten character recognition unit that recognizes handwritten characters of the plurality of character strings and calculates a reliability of a recognition result; a determination unit that determines a handwritten character string to be detected from among the plurality of character strings based on the reliability; an output unit that outputs a recognition result by the handwritten character recognition unit for the handwritten character string to be detected that has been determined by the determination unit; The extraction unit combines bounding boxes surrounding characters or character strings included in the target area based on predetermined conditions to extract the multiple character strings.
4. the image includes a plurality of handwritten character strings to be detected, The information processing device according to any one of claims 1 to 3, wherein the setting unit uses a position of one target area including a handwritten character string of a detection target identified by the position of the keyword as a position of a keyword when setting another target area including a handwritten character string of another detection target.
5. The setting unit sets an area including a character string to be detected as the target area, the printed character recognition unit recognizes printed characters of a plurality of character strings included in the target area extracted by the extraction unit, and calculates a reliability of a recognition result of the printed characters; the determination unit determines a print character string to be detected from among a plurality of character strings based on a reliability of a recognition result of the print character; The information processing apparatus according to claim 1 , wherein the output unit outputs a recognition result by a print character recognition unit for the print character string to be detected determined by the determination unit.
6. An information processing device comprising: acquiring an image of a document containing printed and handwritten characters; a typographical character recognition step of recognizing typographical characters within the image; a step of identifying a position of a keyword in the image based on a recognition result of the print character recognition step; a setting step of setting a target area which is an area including a handwritten character string to be detected based on a position of the keyword and a ruled line included in the image; An extraction step of extracting a plurality of character strings included in the target region; a handwritten character recognition step of recognizing handwritten characters of the plurality of character strings and calculating a reliability of a recognition result; determining a handwritten character string to be detected from among the plurality of character strings based on the reliability; an output step of outputting a recognition result of the handwritten character string to be detected determined in the determination step, The method for controlling an information processing device, wherein the document is a maternal and child health handbook.
7. An information processing device comprising: an acquisition function for acquiring images of documents containing printed and handwritten characters; a typographical character recognition function for recognizing typographical characters in the image; an identifying function for identifying a position of a keyword within the image based on a recognition result by the printed character recognition function; A setting function for setting a target area, which is an area including a handwritten character string to be detected, based on the position of the keyword and the ruled lines included in the image; An extraction function for extracting a plurality of character strings included in the target region; a handwritten character recognition function for recognizing handwritten characters of the plurality of character strings and calculating a reliability of a recognition result; a determination function for determining the handwritten character string to be detected from among the plurality of character strings based on the reliability; an output function for outputting a recognition result by the handwritten character recognition function for the handwritten character string to be detected that is determined by the determination function; A control program for an information processing device, wherein the document is a maternal and child health handbook.
8. An information processing device comprising: acquiring an image of a document containing printed and handwritten characters; a typographical character recognition step of recognizing typographical characters within the image; a step of identifying a position of a keyword within the image based on a recognition result from the print character recognition step; a setting step of setting a target area which is an area including a handwritten character string to be detected based on a position of the keyword and a ruled line included in the image; An extraction step of extracting a plurality of character strings included in the target region; a handwritten character recognition step of recognizing handwritten characters of the plurality of character strings and calculating a reliability of a recognition result; determining a handwritten character string to be detected from among the plurality of character strings based on the reliability; an output step of outputting a recognition result of the handwritten character string to be detected determined in the determination step, The obtaining step further includes obtaining information regarding the type of the document; The control method for an information processing apparatus, wherein the specifying step specifies a position of a keyword according to a type of the document.
9. An information processing device comprising: an acquisition function for acquiring images of documents containing printed and handwritten characters; a typographical character recognition function for recognizing typographical characters in the image; an identifying function for identifying a position of a keyword within the image based on a recognition result by the printed character recognition function; A setting function for setting a target area, which is an area including a handwritten character string to be detected, based on the position of the keyword and the ruled lines included in the image; An extraction function for extracting a plurality of character strings included in the target region; a handwritten character recognition function for recognizing handwritten characters of the plurality of character strings and calculating a reliability of a recognition result; a determination function for determining the handwritten character string to be detected from among the plurality of character strings based on the reliability; an output function for outputting a recognition result by the handwritten character recognition function for the handwritten character string to be detected that has been determined by the determination function; The acquiring function further acquires information regarding the type of the document; The identifying function identifies a position of a keyword according to the type of the document. A control program for an information processing device.
10. An information processing device comprising: acquiring an image of a document containing printed and handwritten characters; a typographical character recognition step of recognizing typographical characters within the image; a step of identifying a position of a keyword within the image based on a recognition result from the print character recognition step; a setting step of setting a target area which is an area including a handwritten character string to be detected based on a position of the keyword and a ruled line included in the image; An extraction step of extracting a plurality of character strings included in the target region; a handwritten character recognition step of recognizing handwritten characters of the plurality of character strings and calculating a reliability of a recognition result; determining a handwritten character string to be detected from among the plurality of character strings based on the reliability; a step of outputting a recognition result of the handwritten character string to be detected determined in the determination step; A control method for an information processing device, wherein the extraction step combines bounding boxes surrounding characters or character strings included in the target area based on a predetermined condition to extract the multiple character strings.
11. An information processing device comprising: an acquisition function for acquiring images of documents containing printed and handwritten characters; a typographical character recognition function for recognizing typographical characters in the image; an identifying function for identifying a position of a keyword within the image based on a recognition result by the printed character recognition function; A setting function for setting a target area, which is an area including a handwritten character string to be detected, based on the position of the keyword and the ruled lines included in the image; An extraction function for extracting a plurality of character strings included in the target region; a handwritten character recognition function for recognizing handwritten characters of the plurality of character strings and calculating a reliability of a recognition result; a determination function for determining the handwritten character string to be detected from among the plurality of character strings based on the reliability; a function of outputting a recognition result by the handwritten character recognition function for the handwritten character string to be detected that is determined by the determination function; The extraction function is a control program for an information processing device that combines bounding boxes surrounding characters or character strings included in the target area based on predetermined conditions to extract the multiple character strings.
Citation Information
Patent Citations
Method and device for positioning answering area in test question image and electronic equipment
CN111507251A
Method for reading insurance policy, system thereof, and insurance policy recognition system
JP2007140703A
Document processing device and document processing method
JP2012155662A
Image processing system, image processing method, and program
JP2021033577A
Layout analysis device, analysis program thereof and analysis method thereof
JP2021167990A