Information processing device, method, recording medium and computer program product
By preprocessing the acquired image data in the information processing device in units of local areas, the problem of difficulty in improving the local area recognition accuracy of OCR processing in the prior art is solved, and a higher processing result accuracy is achieved.
Patent Information
- Application Number
- CN202010756418.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-12
- Filing Date
- 2020-07-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-07-31
AI Technical Summary
When processing images, it is difficult to effectively improve the recognition accuracy of local areas, especially in areas with low confidence, and preprocessing does not pay enough attention to these areas.
An information processing device is designed to preprocess the acquired image data and receive information of the local area in the later-stage processing, and perform specific preprocessing for local areas with low confidence to improve the accuracy of the processing results.
By pre-treatment in units of local areas, the accuracy of the results in the post-stage side processing is significantly improved, ensuring that high-quality processing results can be obtained in areas with low confidence.
Smart Images

Figure CN113255673B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing device, a recording medium and an information processing method. Background Art
[0002] There is a known technique for converting characters printed on a printed matter or handwritten character images into character codes that can be used in a computer. This technique is called OCR (=Optical Character Recognition) processing. In acquiring an image corresponding to a document including characters, a so-called scanner or digital camera is used.
[0003] Although it is also possible to directly output the image data captured by a scanner or a digital camera to the OCR process, in order to increase the value representing the accuracy of the result of character recognition performed by the OCR process (hereinafter referred to as "reliability"), additional processing is sometimes performed in advance. For example, a cleaning process of removing noise or color blocks included in the image is sometimes performed before the OCR process. In addition, the resolution of the image at the time of acquisition is sometimes set to be higher. Hereinafter, the processing performed before the OCR process is referred to as preprocessing.
[0004] Patent Document 1: Japanese Patent Application Publication No. 2015-146075
[0005] Currently, the reliability of OCR processing is calculated using the entire image or data file as a unit. Therefore, even if an area with reduced reliability of OCR processing is included, if the overall reliability is high, preprocessing will not focus on the area including the low reliability. Summary of the invention
[0006] An object of the present invention is to improve the accuracy of a result obtained in a subsequent-stage process in units of local regions, compared to a case where information on the local regions obtained in the subsequent-stage process is not notified to the previous-stage process.
[0007] The invention described in Scheme 1 is an information processing device, which has a processor, and the processor performs the following processing: performs preprocessing on the acquired image data; when receiving information on at least one local area determined in an image corresponding to the image data from the later-stage processing of the preprocessed image data, the processor performs specific preprocessing on the determined local area as an object.
[0008] The invention described in claim 2 is the information processing device described in claim 1, wherein the information for specifying the local area includes information on accuracy, and the accuracy is the accuracy of a result of processing the image data corresponding to the local area.
[0009] The invention described in claim 3 is the information processing device described in claim 2, wherein the information related to the accuracy is given in units of local areas.
[0010] The invention described in claim 4 is the information processing device described in claim 2 or 3, wherein the information related to the accuracy indicates that the accuracy of the processing result is lower than a predetermined threshold value.
[0011] The invention described in claim 5 is the information processing device described in claim 1, wherein the information for specifying the local area includes information on the cause of low accuracy of the processing result.
[0012] The invention described in claim 6 is the information processing device described in claim 5, wherein the information related to the cause is information related to a character or a background included in the partial area.
[0013] The invention described in claim 7 is the information processing device described in claim 1, wherein the information for specifying the local area includes information indicating the content of preprocessing.
[0014] The invention described in claim 8 is the information processing device described in claim 7, wherein the information indicating the content of the preprocessing includes information indicating parameter values used in the preprocessing.
[0015] The invention described in claim 9 is the information processing device described in claim 1, wherein the processor performs a pre-processing different in content from that performed previously on the local area specified based on the information for specifying the local area.
[0016] The invention described in Scheme 10 is, in the information processing device described in Scheme 9, when the information for determining the local area includes information indicating that the accuracy of the processing result taking the local area as the object is lower than a predetermined threshold, the processor performs preprocessing on the corresponding local area that is different from that performed last time.
[0017] The invention described in claim 11 is the information processing device described in claim 10, wherein the processor estimates the cause based on the information indicating that the accuracy of the processing result is lower than a predetermined threshold value.
[0018] The invention described in claim 12 is the information processing device described in claim 11, wherein the processor includes estimating the cause based on the accuracy of a result of a post-processing that targets the entire image data.
[0019] The invention described in claim 13 is the information processing device described in claim 10, wherein the processor infers the cause based on information obtained in the preprocessing process.
[0020] The invention described in claim 14 is the information processing device described in claim 10, wherein the processor estimates the cause based on a history of pre-processing of other image data similar to the image data.
[0021] The invention described in claim 15 is the information processing device described in claim 10, wherein the processor estimates the cause based on the difference in accuracy between local areas of the same type.
[0022] The invention described in claim 16 is the information processing device described in claim 9, wherein the processor notifies a subsequent processing step of information identifying a local area on which a specific pre-processing has been performed.
[0023] The invention described in claim 17 is the information processing device described in claim 9, wherein the processor shares information for identifying the pre-processed image data with a subsequent processing.
[0024] The invention described in claim 18 is the information processing device described in claim 9, wherein when the information determining the local area includes information related to the reason why the accuracy of the local area is low, the processor performs preprocessing corresponding to the reason.
[0025] The invention described in claim 19 is that in the information processing device described in claim 9, when the information determining the local area includes information indicating the content of preprocessing, the processor performs preprocessing of the indicated content.
[0026] The invention described in Scheme 20 is a recording medium, which records a program for enabling a computer to implement the following functions: a function of performing preprocessing on acquired image data; and a function of performing specific preprocessing on the determined local area as an object when receiving information on at least one local area determined in an image corresponding to the image data from a post-processing of the preprocessed image data.
[0027] The invention described in Scheme 21 is an information processing method, which includes the following steps: a step of performing preprocessing on the acquired image data; and a step of performing specific preprocessing on the determined local area as an object when receiving information on at least one local area determined in an image corresponding to the image data from a later-stage processing of the preprocessed image data.
[0028] Effects of the Invention
[0029] According to the first aspect of the present invention, the accuracy of the result obtained in the subsequent-stage processing can be improved in units of local regions, compared with a case where information on the local region obtained in the subsequent-stage processing is not notified to the previous-stage processing.
[0030] According to the second aspect of the present invention, the influence of the content of the preprocessing executed in the specified local area on the result of the post-processing can be confirmed on the preprocessing side.
[0031] According to the third aspect of the present invention, it is possible to confirm the influence of the content executed in the pre-processing on the result of the post-processing in units of local areas.
[0032] According to the fourth aspect of the present invention, it is possible to confirm the presence of a local area where the accuracy of the post-processing result is low.
[0033] According to the fifth aspect of the present invention, it is possible to effectively determine the content of preprocessing that improves the accuracy of the post-processing result.
[0034] According to the sixth aspect of the present invention, it is possible to effectively determine the content of preprocessing that improves the accuracy of the post-processing result.
[0035] According to the seventh aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0036] According to the eighth aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0037] According to the ninth aspect of the present invention, it is possible to change the accuracy of the post-processing result corresponding to the identified local area.
[0038] According to the tenth aspect of the present invention, when the accuracy of the processing result is notified regardless of whether the accuracy is high or low, the content of the preprocessing can be changed only in the local area where the accuracy is low.
[0039] According to the eleventh aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0040] According to the twelfth aspect of the present invention, it is possible to identify the cause caused by the preprocessing that targets the entire image.
[0041] According to the thirteenth aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0042] According to the fourteenth aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0043] According to the fifteenth aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0044] According to the sixteenth aspect of the present invention, unnecessary post-processing can be omitted.
[0045] According to the seventeenth aspect of the present invention, it is possible to omit unnecessary processing in both pre-processing and post-processing.
[0046] According to the eighteenth aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0047] According to the nineteenth aspect of the present invention, the accuracy of the post-processing result can be effectively improved.
[0048] According to the twentieth aspect of the present invention, compared with a case where information on a local area obtained in a downstream process is not notified to a upstream process, the accuracy of a result obtained in a downstream process can be improved in units of local areas.
[0049] According to the 21st aspect of the present invention, compared with a case where information on a local area obtained in a downstream process is not notified to a upstream process, the accuracy of a result obtained in a downstream process can be improved in units of local areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Embodiments of the present invention will be described in detail with reference to the following drawings.
[0051] Figure 1 is a diagram showing a configuration example of an information processing system used in the embodiment;
[0052] Figure 2 A diagram for explaining an example of the hardware configuration of the image processing device used in Embodiment 1;
[0053] Figure 3 A diagram for explaining an overview of processing performed in Implementation Example 1;
[0054] Figure 4 is a flowchart for explaining an example of processing performed by the image processing apparatus in Embodiment 1;
[0055] Figure 5 A diagram for explaining an example of a document to be read;
[0056] Figure 6 A diagram illustrating an example of an object separated from image data;
[0057] Figure 7 This is a flowchart for explaining an example of the processing performed in step S11 of the first embodiment;
[0058] Figure 8 A diagram for explaining an overview of processing performed in Implementation Example 2;
[0059] Fig. 9 is a flowchart for explaining an example of processing performed by the image processing apparatus in Embodiment 2;
[0060] Fig.10This is a flowchart for explaining an example of the processing performed in step S21 of the second embodiment;
[0061] Fig.11 A diagram for explaining an overview of processing performed in Implementation Example 3;
[0062] Fig.12 is a flowchart for explaining an example of processing performed by the image processing device in Embodiment 3;
[0063] Fig.13 This is a flowchart for explaining an example of the processing performed in step S32 of the third embodiment;
[0064] Fig.14 This is a flowchart for explaining another example of the process performed in step S32 of the third embodiment;
[0065] Fig.15 This is a diagram for explaining an example of a cause estimated when the target data is a table area. Fig.15 In the figure, (A) represents the combination of the credibility reported about a table area, and (B) represents the inferred cause;
[0066] Fig.16 A diagram for explaining an overview of processing performed in Implementation Example 4;
[0067] Fig.17 is a flowchart for explaining an example of processing performed by the image processing apparatus in Embodiment 4;
[0068] Fig.18 This is a flowchart for explaining an example of the processing performed in step S42 of the fourth embodiment;
[0069] Fig.19 This is a diagram for explaining an example of a reason for being notified;
[0070] Fig. 20 A diagram for explaining an overview of processing performed in Implementation Example 5;
[0071] Fig.21 is a flowchart for explaining an example of processing performed by the image processing apparatus in Embodiment 5;
[0072] Fig. 22 A diagram for explaining an overview of processing performed in Implementation Example 6;
[0073] Fig.23 is a flowchart for explaining an example of processing performed by the image processing apparatus in Embodiment 6;
[0074] Fig.24is a flowchart for explaining an example of processing performed by the image processing apparatus in Embodiment 7;
[0075] Fig.25 This is a diagram for explaining the outline of the processing performed in Implementation Example 8.
[0076] Explanation of symbols
[0077] 1-information processing system, 10-image processing device, 11-control unit, 11A-processor, 12-storage device, 13-image reading unit, 14-image processing unit, 15-image forming unit, 16-operation receiving unit, 17-display unit, 18-communication device, 20-OCR processing server, 30-cloud network. DETAILED DESCRIPTION
[0078] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0079] <Implementation Method>
[0080] <System Structure>
[0081] Figure 1 It is a diagram showing a configuration example of an information processing system 1 used in the embodiment.
[0082] Figure 1 The information processing system 1 shown includes an image processing apparatus 10 , an OCR processing server 20 that recognizes characters included in image data supplied from the image processing apparatus 10 , and a cloud network 30 as a network environment.
[0083] The image processing apparatus 10 in this embodiment has functions such as generating a copy of a manuscript, printing a document or an image on paper, optically reading a manuscript and generating image data, and transmitting and receiving faxes.
[0084] exist Figure 1 The upper part of the main body of the image processing device 10 shown is equipped with a mechanism for feeding documents one by one to a position where information is optically read. This mechanism is called ADF (=Auto Document Feeder). ADF is used to read copied documents or faxed documents.
[0085] The function of printing a document or image on paper is also used in generating a copy. The data of the document or image is not only optically read in the image processing device 10, but also supplied from a recording medium connected to the main body or an external information terminal.
[0086] The functions provided in the image processing device 10 are not limited to the above functions. However, in the case of the present embodiment, as long as the image processing device 10 is provided with a function of optically reading information of a document and generating image data, other functions are optional.
[0087] The manuscript in this embodiment may be a document or image with handwritten characters in addition to paper with characters or images printed on it. The handwritten characters may be part of the document or image. That is, not all characters in the document need to be handwritten.
[0088] In this embodiment, the handwritten document is assumed to be a handwritten form such as an application, a bill, a delivery note, an invoice, etc. In the handwritten form, characters are written in a pre-printed frame. The handwritten document is not limited to the form type, and may also be a communication memo, a circulation document, a postcard, a sealed letter, etc.
[0089] The image processing device 10 in this embodiment also has a function of removing noise, color patches, etc. from an image read from a document. In this embodiment, image data after the noise and the like have been removed is sent to the OCR processing server 20 .
[0090] Figure 1 Although only one image processing apparatus 10 is shown in the figure, a plurality of image processing apparatuses 10 may constitute the information processing system 1. The image processing apparatus 10 in the present embodiment is an example of an information processing apparatus.
[0091] The OCR processing server 20 in this embodiment is designed to perform OCR processing on the image data given from the image processing device 10, and transfer the text data as the processing result to the image processing device 10. The image processing device 10 to which the text data is transferred performs post-processing on the received text data. The post-processing includes, for example, language processing, processing for associating the text data with the correct location in management, searching for documents related to the text data, and searching for the path used in delivering the item. In addition, the post-processing content is set according to the content of the original to be read or the required processing content.
[0092] In addition, the OCR processing server 20 of the present embodiment is also provided with a function of feeding back information related to the credibility of the OCR processing result to the pre-processing for each local area. This function is provided to improve the accuracy of the character recognition result obtained by the OCR processing, or to improve the quality or precision of the post-processing result.
[0093] However, in the present embodiment, the operator of the image processing apparatus 10 and the operator of the OCR processing server 20 may be the same as or different from each other.
[0094] In the present embodiment, an OCR processing server 20 dedicated to OCR processing is used, but a general-purpose server corresponding to multiple functions may be used. In addition, the computer that performs OCR processing is not limited to the server. The computer that performs OCR processing may be, for example, a desktop computer or a notebook computer, or may be a smart phone or a tablet terminal.
[0095] exist Figure 1 In the case of , there is only one OCR processing server 20, but there may be plural OCR processing servers 20 constituting the information processing system 1. The plural OCR processing servers 20 may perform distributed processing on one image data. The OCR processing server 20 in this embodiment is an example of a device that performs a subsequent-stage process.
[0096] In the present embodiment, the cloud network 30 is used for communication between the image processing device 10 and the OCR processing server 20, but the communication is not limited to the communication via the cloud network 30. For example, a mobile communication system referred to as LAN (Local Area Network), 4G or 5G may be used for communication between the image processing device 10 and the OCR processing server 20.
[0097] <Structure of image processing device>
[0098] Figure 2 This is a diagram for explaining an example of the hardware configuration of the image processing device 10 used in the first embodiment. Figure 2 The image processing device 10 shown has: a control unit 11 that controls the entire device; a storage device 12 that stores image data, etc.; an image reading unit 13 that optically reads a manuscript and generates image data; an image processing unit 14 that adds grayscale conversion processing or color correction processing to the image data; an image forming unit 15 that forms an image corresponding to the image data on a sheet of paper; an operation receiving unit 16 that receives user operations; a display unit 17 that displays a user interface screen, etc.; and a communication device 18 that is used for communication with the outside. The control unit 11 and each unit are connected via a bus 19 or a signal line not shown.
[0099] The control unit 11 in this embodiment includes a processor 11A, a ROM (=Read Only Memory) (not shown) storing firmware or BIOS (=Basic Input Output System), etc., and a RAM (=Random Access Memory) (not shown) used as a work area. The control unit 11 functions as a so-called computer. The above-mentioned pre-processing or post-processing is realized by the processor 11A executing a program.
[0100] The storage device 12 is constituted by a hard disk device or a nonvolatile rewritable semiconductor memory, etc. The storage device 12 stores, for example, image data read by the image reading unit 13. The storage device 12 can store an application program.
[0101] The image reading unit 13 includes, for example, a CIS (Contact Image Sensor) which includes an LED (Light Emitting Diode) that emits illumination light, a photosensor that receives light reflected from a document, and an optical system that collects the light reflected from the document on the photosensor.
[0102] When the image is read while the document is conveyed to the reading position by the ADF, the CIS is used in a state fixed at the reading position. In a mode of reading an image with the document placed on a light-transmitting glass surface, the CIS is controlled to move relative to the document.
[0103] The image processing unit 14 is composed of a GPU (Graphics Processing Unit) or an FPGA (Field Programmable Gate Array) that executes a process of converting grayscale or a process of correcting color.
[0104] The image forming unit 15 has mechanisms corresponding to the following methods: an electronic photography method, in which the colorant transferred to the paper is fixed by heating, thereby forming an image corresponding to the image data on the paper surface; and an inkjet method, in which droplets are ejected onto the paper, thereby forming an image corresponding to the image data on the paper surface.
[0105] The operation receiving unit 16 is constituted by a touch sensor, a physical switch, a button, or the like arranged on the display surface of the display unit 17 .
[0106] The display unit 17 is composed of, for example, a liquid crystal display or an organic EL display. A device in which the operation receiving unit 16 and the display unit 17 are integrated is also called a touch panel. The touch panel is used to receive user operations on keys displayed in a software manner (hereinafter also referred to as "soft keys").
[0107] The communication device 18 is composed of a module conforming to a communication standard based on a wired or wireless system. For example, an EtherNet (registered trademark) module, USB (=Universal Serial Bus), wireless LAN, a fax modem, etc. are used as the communication device 18.
[0108] <Processing Content>
[0109] Hereinafter, a process executed by cooperation between the image processing apparatus 10 and the OCR processing server 20 will be described.
[0110] <Overview of the Process>
[0111] Figure 3 The diagram is a diagram for explaining the outline of the processing performed in the first embodiment. The processing in the present embodiment is composed of five types of processing. The five types of processing are processing of acquiring image data of the original document, preprocessing of the acquired image data, OCR processing of the preprocessed image data, postprocessing of the result of the OCR processing, i.e., text data, and storing the result of the postprocessing in the storage device 12 (refer to Figure 2 ) in the processing.
[0112] In the case of this embodiment, the OCR processing server 20 (refer to Figure 1 ) performs OCR processing, and the other four types of processing are performed by the image processing device 10 (reference Figure 1 )implement.
[0113] In the case of this embodiment, as pre-processing, a cleaning process of removing noise or color blocks or a process of separating objects is performed. On the other hand, as post-processing, a dictionary in which a combination of values (=value) corresponding to keys (=skey) is registered is referred to to perform a process of extracting values corresponding to keys or keys corresponding to values. The keys and values in this embodiment correspond to characters or images. For example, in the case where the key is a name, Fujitaro corresponds to the value. In other words, the key is a character or graphic representing an item, and the value is a character or graphic representing the specific content corresponding to the item.
[0114] In the case of this embodiment, information that specifies the processing target between the pre-processing and the OCR processing is notified from the pre-processing to the OCR processing. Figure 3 In the case of the preprocessed image data, the file name as information for determining the processing object is notified as data attached to the preprocessed image data. In the case of the present embodiment, the file name is composed of the reading date and time or the user name who has performed the reading operation, information for distinguishing the image processing device 10 used for reading, etc. However, the information for determining the file name is not limited to this.
[0115] On the other hand, when the OCR process gives feedback to the preprocessing, it clearly indicates the processing object using, for example, a file name. By notifying the file name, the preprocessing and the OCR process can cooperate. For example, when the preprocessing and the OCR process process a plurality of image data in parallel, the processing object can be distinguished by using the file name. In addition, as long as the processing object can be determined, the information from the preprocessing notification to the OCR process is not limited to the file name.
[0116] In the case of this embodiment, a process of separating image data into a plurality of objects is also performed in the pre-processing.
[0117] In the case of this embodiment, four areas are used as objects: an area for characters (hereinafter referred to as the "character area"), an area for tables (hereinafter referred to as the "table area"), an area for graphics (hereinafter referred to as the "graphic area"), and an area for figures (hereinafter referred to as the "figure area").
[0118] For example, the area including the title, characters, and numerical values of the manuscript is cut as the character area. The table itself or the title attached to the table is cut as the table area. The area with patterned company names, etc. is cut as the graphic area or the figure area. The area other than this is the background. Each object is an example of a local area.
[0119] In addition, the background, graphic area and figure area are excluded from the OCR processing object. Therefore, the image data corresponding to the character area and table area is sent as object data from the image processing device 10 to the OCR processing server 20. In addition, each object data is given information for identifying each local area.
[0120] In the case of this embodiment, the pre-processed image data is sent from the pre-processing to the OCR processing in units of local areas.
[0121] Furthermore, information identifying the content of the preprocessing to be executed may be notified from the preprocessing to the OCR processing. The preprocessing content can be used when the OCR processing side estimates the cause of low reliability.
[0122] In the case of this embodiment, the content of the preprocessing requested in local area units is fed back from the OCR processing to the preprocessing, or information indicating that the reliability of the result after the OCR processing is low is fed back. The reliability is an example of information related to accuracy.
[0123] In the case of this embodiment, the local area is used in the sense of each area cut as an object. In other words, when there are multiple character areas, different information may be fed back for each character area. The same is true for the table area. In addition, for the table area, different information may also be fed back in rows or columns.
[0124] In the case of this embodiment, information identifying each local area is notified from the preprocessing to the OCR processing. Therefore, the feedback from the OCR processing to the preprocessing includes the information identifying each local area. However, in the OCR processing, it is also possible to regard multiple local areas of the same object type as one local area and calculate the reliability, and feedback the information indicating that the reliability is low together with the information identifying the multiple local areas for which the reliability is calculated.
[0125] In the case of this embodiment, no feedback is performed on the local area whose credibility exceeds the predetermined threshold. Therefore, for all local areas, when the credibility exceeds the predetermined threshold, no feedback from OCR processing to preprocessing is performed. This is because text data with high credibility is obtained.
[0126] In this embodiment, the average value of the credibility calculated for each character extracted from the corresponding local area is obtained as the credibility of the local area. The characters here also include numbers or symbols. Different weights can be used to calculate the average value for each local area. For example, different weights can be used in the case of a character area and in the case of a table area. Moreover, in the local area of the same type, different weights can also be used in the title part and the text.
[0127] For each type of object corresponding to the local region, the threshold used in the credibility evaluation of the local region can be different or the same. For example, different weights can be used to calculate the credibility in the character region and the table region.
[0128] In addition, the feedback from the OCR process to the pre-processing may be feedback that does not specify a local area. In this case, the OCR process also identifies which local area has a high reliability and which local area has a low reliability.
[0129] Thus, in the OCR process, only the local area with low credibility in the previous OCR process can be selected from the image data to which the new pre-processing is added, and the change in credibility can be confirmed. In addition, in the OCR process, only the text data with a credibility higher than a threshold can be selectively output to post-processing.
[0130] <Processing performed by the image processing device>
[0131] Figure 4 This is a flowchart for explaining an example of processing executed by the image processing device 10 in Embodiment 1. The symbol S shown in the figure indicates a step. Figure 4 The processing shown is performed by the processor 11A (refer to Figure 2 )implement.
[0132] Figure 4 The processing shown is started by receiving an instruction to read a document accompanied by an OCR process. The reading instruction to the image processing apparatus 10 uses, for example, an operation of a start button.
[0133] In the reading instruction, the conditions or prerequisites for reading can be set. For example, the type of the document to be read can be specified. When the type of the document is specified, the content of the preprocessing prepared corresponding to the type of the document is selected by the processor 11A. In addition, when the relationship between the type of the document and the content of the preprocessing that can obtain high reliability is learned by machine learning, the processor 11A selects the content of the preprocessing corresponding to the specified type of the document.
[0134] However, it is also possible to issue a reading instruction without setting a reading condition or premise. In this case, the processor 11A selects a preprocessing of the content corresponding to the type of the manuscript estimated by reading or the characteristics of the manuscript. Furthermore, when the image processing device 10 can read the title of the manuscript, the processor 11A selects a preprocessing of the content corresponding to the read title.
[0135] Then, upon receiving a document reading instruction involving OCR processing, the processor 11A acquires image data of the document (step S1). The image data is output in a predetermined format such as PDF (=Portable Document Format).
[0136] Figure 5 This is a diagram for explaining an example of a document that is a reading target. Figure 5 The manuscript shown is titled Quotation, with color blocks appended throughout the paper. Figure 5 The quotation shown includes two tables. The upper part is Table A and the lower part is Table B. Figure 5 Table A and Table B shown are each composed of 3 lines. The item names of the titles of Table A and Table B are printed with blank characters on a black background. The second and third lines of Table A are printed with black characters on a white background. The second and third lines of Table B are printed with characters on a colored background. The characters may be any of black characters, blank characters, and colored characters. In addition, it is also possible to assume that the background is a shadow.
[0137] Return to Figure 4 Description.
[0138] Next, the processor 11A performs preprocessing on the acquired image data (step S2). In the case of the present embodiment, object separation is performed in the preprocessing. A known technique is used in the object separation. In addition, a cleaning process selected in advance or determined in the initial setting is also performed.
[0139] Figure 6 This is a diagram for explaining an example of an object separated from image data. Figure 6In the case of , the region including the character strings "quotation", "ABC Industry", "XYZ Chamber of Commerce", and "total amount 16,000 yen" is separated from the image data as a character region. Furthermore, the region including the table corresponding to the characters of table A and the table corresponding to the characters of table B is separated from the image data as a table region. Furthermore, the logo arranged at the lower right of the image data is separated from the image data as a graphic region or a figure region.
[0140] Preprocessing can be performed before or after object separation. In this embodiment, object separation is performed after preprocessing.
[0141] Return to Figure 4 Description.
[0142] Next, the processor 11A sends the target data to the OCR processing server 20 (step S3 ). In this embodiment, the target data is image data corresponding to the character area and the table area. That is, the image data of the portion determined to be the graphic area and the diagram area is not sent to the OCR processing server 20 .
[0143] Then, the processor 11A determines whether information is fed back from the OCR processing server 20 (step S4).
[0144] For example, if there is no information feedback within a predetermined time, the processor 11A obtains a negative result in step S4. In the case of this embodiment, the absence of information feedback means that the reliability of the results of the OCR processing in all local areas is high. As described above, in this embodiment, the reliability is calculated in units of local areas. In addition, the reliability can be calculated in units of the entire table, or in units of rows or columns constituting the table.
[0145] If a negative result is obtained in step S4, the processor 11A retrieves the data from the storage device 12 (reference Figure 2 ) in step S1 (step S5). This is because it is not necessary to perform preprocessing again on the target image data. Here, the deletion is the deletion of the image data used for preprocessing. Thus, the image data can be saved for other purposes.
[0146] Next, the processor 11A performs post-processing on the text data acquired from the OCR processing server 20 (step S6). Then, the processor 11A stores the processing result in the storage device 12 (step S7). In addition, step S5 may be performed after step S6 or step S7.
[0147] In the case of this embodiment, during or after executing steps S5 to S7, the processor 11A learns that a high degree of reliability is obtained in the contents of the preprocessing performed last time on the entire image data or a specific local area. The learning unit is the same as the case of low degree of reliability described later.
[0148] In the case of a positive result in step S4, the processor 11A determines the object data of the feedback (step S8). In the case of the present embodiment, when there is a local area whose credibility does not exceed a predetermined threshold, there is feedback from the OCR processing server 20 to the image processing device 10. In the case of the present embodiment, the information fed back from the OCR processing server 20 includes information for identifying the target local area. In the information for identifying the local area, for example, coordinates or serial numbers representing the position in the image data of the original document are used. The coordinates are assigned, for example, in the form of one or more coordinate points of the outer edge of the specified area. In the case where the local area is rectangular in shape, the coordinate point of the local area, for example, the upper left corner, is used. The information for identifying the local area is included in the object data sent to the OCR processing server 20 in step S3. If the object data is determined, the type of the target image data or object is also determined.
[0149] Next, the processor 11A determines whether the fed-back information includes an indication of the pre-processing content (step S9). Figure 3 As described in the above, the OCR processing server 20 of the present embodiment is provided with a pre-processing function of feeding back information indicating that the reliability is lower than a threshold value or the processing content of the pre-processing request to the image processing device 10. The requested processing content includes, for example, the type of cleaning processing, the intensity of the cleaning processing, and the parameter value used in the cleaning processing. The types of cleaning processing include, for example, a process of removing color blocks or shadows, a process of removing stains, and a process of removing background colors and converting blank characters or color characters to black characters.
[0150] In addition, in removing color blocks or shadows, for example, a method called a generative adversarial network (GAN) is applied. Since the technology of removing noise using GAN has been put into practical use, detailed description is omitted.
[0151] If a positive result is obtained in step S9, processor 11A executes the instructed preprocessing content (step S10). In the case of this embodiment, new preprocessing is performed only on the target data determined in step S8. However, the preprocessing target of the new content may be the entire image data of the original.
[0152] After executing step S10, the processor 11A returns to step S3. In the case of this embodiment, the processor 11A sends the pre-processed image data as object data to the OCR processing server 20 only for the local area determined in step S8. At this time, the processor 11A notifies the OCR processing server 20 of the information that the local area for re-performing pre-processing is determined.
[0153] The object data sent to the OCR processing server 20 in step S3 may include other object data than the object data determined in step S8. Even if other object data is included, the OCR processing server 20 can selectively extract object data corresponding to a local area with low reliability.
[0154] When a negative result is obtained in step S9, processor 11A determines and executes the content of preprocessing to be executed (step S11).
[0155] Figure 7 This is a flowchart for explaining an example of the process executed in step S11 in Embodiment 1. The symbol S shown in the figure indicates a step.
[0156] The processor 11A that started step S11 determines the content of the preprocessing that has been executed on the determined target data (step S111). When the preprocessing has been executed multiple times on the same target data of the same document, the processor 11A determines the content of the multiple preprocessing.
[0157] Next, the processor 11A selects pre-processing of different contents from the previous one for the target data (step S112 ). This is because, in the case of the present embodiment, only information indicating that the reliability is lower than the threshold is fed back from the OCR processing server 20 .
[0158] Next, the processor 11A performs preprocessing of the selected content on the object data (step S113). By performing preprocessing of different content, the reliability in the OCR processing server 20 may exceed the threshold. However, since the reason for the low reliability is not clear, the reliability may decrease.
[0159] In addition, in step S113, the pre-processing selected in step S112 can also be performed on the entire image data.
[0160] Next, the processor 11A uses the content of the preprocessing performed on the object data last time and the information on the reliability to learn the relationship between the local area and the content of the preprocessing (step S114). Here, the information on the reliability is information indicating that the reliability is low.
[0161] In the case of this embodiment, the content of the preprocessing performed on the local area last time is mechanically learned as the teacher data. The learning unit is not limited to the local area, but can be the unit of the type of object, the unit of the type of original, or the unit of similar images. In addition, even if the object type is the same, such as Figure 6 In Table A and Table B, if the background or character combination is different, the content of the preprocessing that helps to improve the credibility is different. Therefore, in this embodiment, machine learning is performed in units of local areas.
[0162] In this embodiment, reinforcement learning is used in machine learning. In reinforcement learning, learning is performed in a way that rewards increase. Therefore, no reward or only a low reward is given to the preprocessed content that only obtains a credibility below the threshold. On the other hand, if a negative result is obtained in step S4, a high reward is given to the preprocessed content when the credibility is higher than the threshold.
[0163] The result of machine learning is used to determine the content of preprocessing to be performed next and thereafter. The preprocessing to be performed next and thereafter includes re-execution of preprocessing with feedback and preprocessing of image data of a newly read document.
[0164] If the image data corresponding to the local area is given to the completed learning model after reinforcement learning, the preprocessing content used in step S2 is output. By improving the accuracy of reinforcement learning, the number of re-performations is also reduced. In addition, the completed learning model after machine learning can also be applied to the selection of preprocessing content in step S112. Compared with the case of randomly selecting preprocessing content, the possibility of reliability exceeding the threshold can also be increased.
[0165] After step S114, the processor 11A determines whether the processing is completed with respect to all the notified object data (step S115).
[0166] When a negative result is obtained in step S115 , the processor 11A returns to step S111 and repeats a series of processes on object data corresponding to another local area.
[0167] On the other hand, in the case of an affirmative result in step S115, the processor 11A returns to step S3.
[0168] The above process is repeated until a negative result is obtained in step S4 executed after step S10 or step S11. As a result, text data with a credibility higher than the threshold value is given in the post-processing performed by the image processing device 10, thereby improving the accuracy or reliability of the post-processing result. In addition, the labor and time of manually confirming or manually correcting the recognized text data are reduced.
[0169] <Implementation method 2>
[0170] In the above-mentioned embodiment, the case where the requested pre-processing content is specifically fed back from the OCR processing server 20 to the image processing device 10 is described. However, in order to specifically determine the pre-processing content, it is necessary to prepare a function corresponding to the following processing on the OCR processing server 20 side: a process of estimating the cause of low reliability in the OCR processing server 20; and a process of determining the content of the process of eliminating the estimated cause, etc. However, the OCR processing server 20 may not always have the same function.
[0171] Figure 8 This is a diagram for explaining the outline of the processing executed in the second embodiment. Figure 8 The mark in the middle corresponds to Figure 3 The corresponding parts are shown by symbols.
[0172] exist Figure 8 In the case of the processing shown, the information fed back from the OCR processing to the pre-processing is different from that in the first embodiment.
[0173] In the case of this embodiment, information requesting a change in the preprocessing content is fed back, that is, the specific processing content is not fed back in response to the preprocessing request.
[0174] The request to change the pre-processing content can be outputted when the content of the local area with low reliability exists, and the above-mentioned estimation and other processing are not necessary.
[0175] Fig. 9 This is a flowchart for explaining an example of processing executed by the image processing device 10 in the second embodiment. Fig. 9 The mark in the middle corresponds to Figure 4 The corresponding parts are shown by symbols.
[0176] In the case of the present embodiment, only information requesting a change in the content of the pre-processing is fed back from the OCR processing server 20 to the image processing apparatus 10 .
[0177] Therefore, after executing step S8, the processor 11A determines and executes the content of preprocessing to be performed on the determined target data (step S21).
[0178] Fig.10 This is a flowchart for explaining an example of the process executed in step S21 in the second embodiment. Fig.10 The mark in the middle corresponds to Figure 7 The symbol S shown in the figure refers to a step.
[0179] The processor 11A which has started step S21 determines the content of the pre-processing to be performed on the determined object data (step S111 ).
[0180] Next, the processor 11A estimates the reason for the reduction in reliability (step S211). The processor 11A estimates the reason by referring to the history of the content of the pre-processing completed and acquired in step S111, for example.
[0181] In addition, the image processing device 10 stores original image data as a processing target. Therefore, the processor 11A reads information such as the presence of color patches, the presence of background, font size, the presence of stains, the presence of folds, the relationship between background and character colors, and the type of manuscript from the image data and uses it to estimate the cause.
[0182] Furthermore, the reliability of the entire image data can also be referred to in estimating the cause. The reliability of the entire image data can be calculated in the image processing device 10 .
[0183] The credibility of the entire image data can be calculated, for example, based on the area ratio of each local area sent as object data on the original and the level of credibility relative to each local area in step S3. For example, the sum of the following values is calculated: the area of the local area that obtains a credibility higher than the threshold multiplied by the value of "1" as a weight; and the area of the local area that obtains a credibility lower than the threshold multiplied by the value of "0" as a weight. Then, the total area of the local area that will be the object of OCR processing is divided by the calculated value for standardization, and the credibility is calculated by comparing the standardized value with the threshold. For example, if the standardized value is higher than the threshold, it is determined that the credibility of the entire image data is high, and if the standardized value is lower than the threshold, it is determined that the credibility of the entire image data is low.
[0184] In addition, it is considered that the reliability of the local area that does not require the change of the preprocessing content is high, and the reliability of the local area that requires the change of the preprocessing content is low.
[0185] For example, when the credibility of the entire image data is high and only the credibility of a specific local area is low, a cause inherent in the specific local area can be considered. On the other hand, when the credibility of not only the specific local area but also the entire image data is low, a common cause that does not depend on the difference in the type of object can be inferred. For example, it can be inferred that stains or creases may be the cause.
[0186] Furthermore, in the case where there is a history of preprocessing contents used for local areas of similar or same type and information related to their credibility, the cause can also be inferred based on the preprocessing contents when high credibility is obtained. Here, local areas of similar or same type means that the contents of the image data corresponding to the local areas are similar or of the same type. However, in this case, it is not necessary to infer the cause, and the preprocessing contents when high credibility is obtained can be assigned to step S212.
[0187] In the estimation in step S211 , for example, a correspondence table prepared in advance, a completed learning model updated by machine learning, or a determination program is used.
[0188] The correspondence table stores combinations of features extracted from the entire image data or features of a local area and, when the reliability is low, the reasons assumed. However, the content of preprocessing recommended for each combination may be stored.
[0189] Furthermore, when the completed learning model is used, if the image data corresponding to the local area is input to the completed learning model, the cause is output. However, it is also possible to output the recommended preprocessing content if the image data corresponding to the local area is input to the completed learning model.
[0190] Furthermore, when using a determination program, the reason for the reduction in reliability is outputted by repeating the divergence caused by a single determination once or a plurality of times. In this case, not only the reason but also the recommended preprocessing contents may be outputted.
[0191] If the cause is estimated in step S211, the processor 11A performs preprocessing on the object data to eliminate the estimated cause (step S212). The relationship between the estimated cause and the preprocessing content that has an effect on the elimination is stored, for example, in the storage device 12. In addition, as described above, when the estimation of the cause is skipped and the preprocessing content that eliminates the cause of lowering the credibility is determined, the determined preprocessing content is executed.
[0192] Next, the processor 11A uses the content of the preprocessing performed last time on the object data and the information on the reliability to learn the relationship between the local area and the content of the preprocessing (step S213).
[0193] In the case of this embodiment, the information related to the credibility is not directly notified from the OCR processing server 20. Therefore, the processor 11A determines the information related to the credibility for each local area. As described above, the credibility of the local area determined not to require the change of the preprocessing content is high, while the credibility of the local area determined to require the change of the preprocessing content is low.
[0194] After step S213, the processor 11A determines whether the processing is completed with respect to all the notified object data (step S115).
[0195] When a negative result is obtained in step S115 , the processor 11A returns to step S111 and repeats a series of processes on object data corresponding to another local area.
[0196] On the other hand, in the case of an affirmative result in step S115, the processor 11A returns to step S3.
[0197] The above process is repeated until a negative result is obtained in step S4. As a result, text data with a credibility higher than the threshold is given in the post-processing performed by the image processing device 10, and the accuracy or reliability of the post-processing result is improved. In addition, the labor and time of manually confirming the recognized text data or manually correcting it are reduced.
[0198] In addition, Fig.10 In the flowchart shown, preprocessing of the content that eliminates the cause of lowering the reliability is performed on the target data, but it is also possible to select and execute one of the preprocessing contents that have not been performed on the same local area in the same image data without estimating the cause.
[0199] In this case, the cause of the decrease in credibility may not be eliminated, but it is expected that the decrease in credibility will be eliminated in the process of repeatedly changing the preprocessing content. Fig.10 Even if the processing shown is increased, the load on computing resources can be reduced by an amount equivalent to the amount by which processing such as estimation is not required.
[0200] <Implementation method 3>
[0201] Fig.11 This is a diagram for explaining the outline of the processing executed in the third embodiment. Fig.11 The mark in the middle corresponds to Figure 8 The corresponding parts are shown by symbols.
[0202] In the present embodiment, part of the information fed back from the OCR process to the pre-processing is different from that in Embodiment 1. Specifically, the reliability of the result after the OCR process is fed back. The reliability is sent in two cases: when the reliability is low and when the reliability is high.
[0203] Fig.12 This is a flowchart for explaining an example of processing executed by the image processing device 10 in the third embodiment. Fig.12 The mark in the middle corresponds to Fig. 9 The symbol S shown in the figure refers to a step.
[0204] In the case of the present embodiment, the reliability of the OCR processing result is fed back to the image processing device 10 from the OCR processing server 20 each time.
[0205] Therefore, after executing step S3, the processor 11A determines whether the reliability is greater than a threshold value (step S31). Here, the threshold value may be the same value regardless of the difference between the objects, or may be a different value for each object.
[0206] If a positive result is obtained in step S31, the processor 11A proceeds to step S5. This is because, if the reliability is higher than the threshold value for all the object data sent to the OCR processing server 20, it is not necessary to perform the pre-processing again.
[0207] When a negative result is obtained in step S31, processor 11A determines and executes the content of preprocessing to be executed (step S32).
[0208] Fig.13 This is a flowchart for explaining an example of the process executed in step S32 in the third embodiment.
[0209] First, the processor 11A determines the content of pre-processing to be performed with respect to object data of low reliability (step S321 ).
[0210] Next, the processor 11A selects preprocessing different from the previous one for the object data with low credibility (step S322). However, as in the case of Embodiment 2, the cause of the reduced credibility may be estimated, and preprocessing that eliminates the estimated cause may be selected. This example will be described later.
[0211] Next, the processor 11A performs pre-processing of the selected content on the object data (step S323).
[0212] Next, the processor 11A uses the content of the preprocessing performed on the object data last time and the information related to the credibility to learn the relationship between the local area and the preprocessing content (step S324). In the case of this embodiment, the relationship between the local area and the preprocessing content is learned not only in the case of low credibility, but also for the object data determined to be highly credible. However, it is also possible to learn only one of them.
[0213] If the processing is completed, the processor 11A returns to step S3. The above processing is repeated until a positive result is obtained in step S31.
[0214] As a result, text data with a reliability higher than a threshold value is given in the post-processing performed by the image processing device 10, and the accuracy or reliability of the post-processing result is improved. In addition, the labor and time of manually confirming the recognized text data or manually correcting it is reduced.
[0215] Fig.14 This is a flowchart for explaining another example of the process executed in step S32 in the third embodiment. Fig.14 The mark in the middle corresponds to Fig.13 The corresponding parts are shown by symbols.
[0216] exist Fig.14 A combination of multiple confidence levels is used in the illustrated process.
[0217] First, the processor 11A uses a combination of credibility and estimates the reason for lowering the credibility (step S325). The combination of credibility can be a combination of a plurality of credibility acquired for a plurality of local areas common to the object type, or a combination of credibility of a row or column unit constituting the local area. Furthermore, it can also be a combination of a plurality of credibility integrated in units of processing objects, i.e., image data.
[0218] Fig.15 This is a diagram for explaining an example of the cause estimated when the object data is a table area. In the figure, (A) shows a combination of the reliability reported back about a table area, and (B) shows the estimated cause. In addition, Fig.15 The example is an example of a case where the reliability is calculated and fed back for each row. In the case where the reliability is fed back in units of the entire table area, it is difficult to infer that Fig.15 The detailed reasons are shown.
[0219] Fig.15 The example assumes Figure 5 Table A or Table B in the table. Therefore, the number of rows is 3. Fig.15 In the case of , there are 8 combinations of credibility for one table area. Fig.15 In , these are represented by combinations 1 to 8. The number of combinations depends on the number of rows constituting the table area or the number of confidences notified about a table area.
[0220] In addition, if Fig.15 The background color of each row corresponding to the combinations 1 to 8 is different in the case where the background color of table A or table B is different in the even-numbered rows and the odd-numbered rows, even if the number of rows increases, the reliability can be calculated in units of odd-numbered rows and even-numbered rows.
[0221] Combination 1 is a case where each reliability corresponding to rows 1 to 3 is higher than the threshold value. In this case, since there is no problem with the content of the preprocessing performed last time, there is no need to estimate the cause of the low reliability.
[0222] Combination 2 is a case where each reliability corresponding to the first and second rows is higher than the threshold, but the reliability corresponding to the third row is lower than the threshold. In this case, it can be estimated that the value cell is a color background as a cause of lowering the reliability.
[0223] Combination 3 is a case where the reliability corresponding to the first and third rows is higher than the threshold, and the reliability corresponding to the second row is lower than the threshold. In this case, the cause of the reduction in reliability may be that the value cell is estimated to be a color background.
[0224] Combination 4 is a case where the reliability corresponding to the first row is higher than the threshold, while the reliability corresponding to the second row and the third row is lower than the threshold. In this case, it can be inferred that only the value unit is a shadow as the cause of the lowered reliability. Here, the shadow also includes a color block.
[0225] Combination 5 is a case where the reliability corresponding to the second and third rows is higher than the threshold, and the reliability corresponding to the first row is lower than the threshold. In this case, it can be estimated that the item name unit is a blank character as a cause of the reduction in reliability.
[0226] Combination 6 is a case where the reliability corresponding to row 2 is higher than the threshold, but the reliability corresponding to row 1 and row 3 is lower than the threshold. In this case, it can be estimated that the item name unit is a blank character and the value unit is a colored background as the cause of the lowered reliability.
[0227] Combination 7 is a case where the reliability corresponding to row 3 is higher than the threshold, but the reliability corresponding to row 1 and row 2 is lower than the threshold. In this case, the cause of the lowered reliability is estimated to be that the item name unit is a blank character and the value unit is a colored background.
[0228] Combination 8 is a case where the reliability corresponding to the first to third lines is lower than the threshold. In this case, the reason for the reduction in reliability can be estimated that the entire screen is a color block or the entire screen is a color background, and each character is a color character.
[0229] In addition, the above estimation focuses on the features of the original manuscript image. Therefore, in order to estimate the reduction in reliability due to the influence of stains or folds, other information is also required. For example, the reliability of the manuscript unit or the information of the original image data is required.
[0230] Return to Fig.14 Description.
[0231] If a cause is estimated in step S325, the processor 11A executes pre-processing for eliminating the content of the estimated cause on the target data (step S326).
[0232] Next, the processor 11A uses the content of the preprocessing performed on the object data last time and the information related to the credibility to learn the relationship between the local area and the preprocessing content (step S324). In the case of this embodiment, both the relationship between the object data with low credibility and the content of the preprocessing performed and the relationship between the object data with high credibility and the content of the preprocessing performed are learned. However, it is also possible to learn only one of them.
[0233] After the processing is completed, the processing content is as follows: Fig.12 Explained.
[0234] <Implementation method 4>
[0235] Fig.16 This is a diagram for explaining the outline of the processing executed in the fourth embodiment. Fig.16 The mark in the middle corresponds to Fig.11 The corresponding parts are shown by symbols.
[0236] In the case of this embodiment, not the reliability but the estimated cause is fed back from the OCR process to the pre-processing. In this case, the estimation performed in the third embodiment is performed on the OCR processing server 20 side.
[0237] Fig.17 This is a flowchart for explaining an example of processing executed by the image processing device 10 in the fourth embodiment. Fig.17 The mark in the middle corresponds to Fig.12 The symbol S shown in the figure refers to a step.
[0238] In the case of the present embodiment, after the processor 11A transmits the object data to the OCR processing server 20 (step S3 ), it determines whether the reason has been fed back (step S41 ).
[0239] The feedback reason is not limited to the case where there is a local area with low reliability in the object data. Therefore, if a negative result is obtained in step S41, the processor 11A transfers to step S5 and then executes the same Fig.12 The same treatment is applied to the situation.
[0240] In contrast, in the case of an affirmative result in step S41, processor 11A determines and executes the content of preprocessing to be performed (step S42).
[0241] Fig.18 This is a flowchart for explaining an example of the processing executed in step S42 in the fourth embodiment.
[0242] First, the processor 11A performs pre-processing of eliminating the content of the notified cause on the target data (step S421).
[0243] Fig.19 This is a diagram for explaining an example of the reason for being notified.
[0244] Fig.19 Causes 1 to 5 shown are Fig.15The reasons are as shown. Reason 1 indicates that the value cell has a colored background. Reason 2 indicates that the value cell has a colored background. Reason 3 indicates that the item name cell has a blank character. Reason 4 indicates that the item name cell has a blank character and the value cell has a colored background. Reason 5 indicates that the entire surface is a color block or the entire surface has a colored background and each character is a colored character, etc. In addition, stains or folds, etc. may also be notified as a cause.
[0245] Return to Fig.18 Description.
[0246] If the preprocessing corresponding to the cause is performed on the object data, the processor 11A learns the relationship between the local area and the preprocessing content using the content of the preprocessing performed last on the object data and information on the reliability (step S422).
[0247] In addition, the reason notified means that the reliability of the corresponding local area is lower than the threshold, and the reason not notified means that the reliability of the corresponding local area is higher than the threshold. Therefore, the processor 11A determines the reliability level according to whether the reason has been notified.
[0248] After the processing is completed, the processing content is as follows: Fig.12 Explained.
[0249] <Implementation method 5>
[0250] Fig. 20 This is a diagram for explaining the outline of the processing executed in the fifth embodiment. Fig. 20 The mark in the middle corresponds to Figure 3 The corresponding parts are shown by symbols.
[0251] The present embodiment is different from the first embodiment in that feedback after post-processing is added to pre-processing.
[0252] Fig.21 This is a flowchart for explaining an example of processing executed by the image processing device 10 in the fifth embodiment. Fig.21 The mark in the middle corresponds to Figure 4 The corresponding parts are shown by symbols.
[0253] In the present embodiment, when a negative result is obtained in step S4, the processor 11A does not execute step S5, but sequentially executes steps S6 and 7. That is, when a negative result is obtained in step S4, the processor 11A executes post-processing based on the text data obtained from the OCR processing server 20, and stores the processing result in the storage device 12.
[0254] In this embodiment, after executing step S7, a notification of completion of the post-processing is received (step S51). After receiving the notification, the processor 11A deletes the image data (step S5). Since the image data is deleted after confirming the notification of completion of the post-processing, the image data will not be requested after the image data is deleted.
[0255] In this embodiment, feedback of post-processing is added to the first embodiment, but it may be added to any of the second to fourth embodiments.
[0256] <Implementation method 6>
[0257] Fig. 22 This is a diagram for explaining the outline of the processing executed in the sixth embodiment. Fig. 22 The mark in the middle corresponds to Figure 3 The corresponding parts are shown by symbols.
[0258] The present embodiment is different from the first embodiment in that a function of completing storage and feeding back to preprocessing is added.
[0259] Fig.23 This is a flowchart for explaining an example of processing executed by the image processing device 10 in the sixth embodiment. Fig.23 The mark in the middle corresponds to Figure 4 The corresponding parts are shown by symbols.
[0260] In the present embodiment, when a negative result is obtained in step S4, the processor 11A does not execute step S5, but sequentially executes steps S6 and 7. That is, when a negative result is obtained in step S4, the processor 11A executes post-processing based on the text data obtained from the OCR processing server 20, and stores the processing result in the storage device 12.
[0261] In this embodiment, after executing step S7, a notification of completion of storing the processing result is received (step S61). After receiving the notification, the processor 11A deletes the image data (step S5). Since the image data is deleted after the processing result is stored, the image data will not be requested after the image data is deleted.
[0262] In the present embodiment, feedback of the completion of the storage processing result is added to the first embodiment, but it may be added to any of the second to fourth embodiments.
[0263] <Implementation method 7>
[0264] Fig.24 This is a flowchart for explaining an example of processing executed by the image processing device 10 in the seventh embodiment. Fig.24 The mark in the middle corresponds to Figure 4 The corresponding parts are shown by symbols.
[0265] exist Figure 4 In the case of the flowchart shown, the case where only the pre-processing content is re-performed in step S10 or step S11 is described.
[0266] However, if Fig.24 As shown, when performing preprocessing again, you can start again from object separation. Fig.24 In FIG. 1 , steps S10 and 11 including re-separation of the objects are shown as steps S10A and 11A.
[0267] In addition, in the aforementioned other embodiments, when the preprocessing is performed again, the object separation may be performed again.
[0268] <Implementation method 8>
[0269] Fig.25 This is a diagram for explaining the outline of the processing performed in Implementation Example 8. Fig.25 The mark in the middle corresponds to Figure 3 The corresponding parts are shown by symbols.
[0270] In the case of the aforementioned embodiment, the OCR process performs feedback to the pre-processing, but in the case of the present embodiment, the OCR process performs feedback to the process of acquiring the image data of the original.
[0271] For example, if the resolution used to obtain the image data is smaller than the size of the characters printed or written on the original, the reliability of the OCR processing result may be reduced. If the resolution inconsistency is the cause of the reduced reliability, the reliability will not be improved even if the content of the pre-processing is changed.
[0272] Therefore, in the present embodiment, when the size of the font included in the image data to be processed by OCR is considered to be the cause of low reliability, the change in the resolution of the image data is fed back to the process of acquiring the image data of the original. Fig.25 In the example of , the instruction is to change from 200 dpi to 600 dpi. In addition, a technique for detecting the size of a font size is known.
[0273] The feedback described in this embodiment can also be combined with any of the above-mentioned embodiments.
[0274] <Other embodiments>
[0275] The embodiments of the present invention are described above, but the technical scope of the present invention is not limited to the scope described in the aforementioned embodiments. According to the description of the technical scope of the present invention, it is clear that various changes or improvements added to the aforementioned embodiments are also included in the technical scope of the present invention.
[0276] (1) For example, in the above-mentioned embodiment, OCR processing is assumed as an example of the post-processing of the pre-processed image data, but the post-processing is not limited to OCR processing. For example, the post-processing described in the above-mentioned embodiment 7 or the storage processing described in the embodiment 8 is also included in the post-processing.
[0277] Furthermore, the combination of the preprocessing and the post-processing is not limited to the combination of the cleaning process and the OCR process. For example, the preprocessing may extract feature quantities for face recognition, and the post-processing may be face recognition using the extracted feature quantities. In this case, the credibility is information indicating the accuracy of the result of face recognition. Thus, in the aforementioned embodiment, the preprocessing content is described on the premise of performing the OCR process, but the combination of the preprocessing and the post-processing may be arbitrary.
[0278] (2) In the above-mentioned embodiment, Figure 5 The illustrated manuscript is used as a premise, and the reliability is calculated in units of rows constituting the table, but it is also applicable when the reliability is calculated in units of columns.
[0279] (3) In the above-mentioned embodiment, as an example of a device for performing additional preprocessing on the image data given to the OCR processing server 20, an image processing device 10 having a function of optically reading an original document and generating image data is illustrated, but an image scanner dedicated to reading image data corresponding to an original document may be used as the image processing device 10. An ADF (=Auto Document Feeder) may be provided in the image scanner.
[0280] Furthermore, as a device for performing additional pre-processing on the image data given to the OCR processing server 20, in addition to a smartphone or a digital camera for photographing the document, a computer that obtains the image data of the document photographed from the outside can also be used. Here, the computer is used for pre-processing and post-processing of the data after OCR processing, and does not need to have a function of photographing the document image or a function of optically reading the document information.
[0281] (4) In the above embodiment, the image processing device 10 and the OCR processing server 20 are described as being independent devices, but the OCR processing function may be built into the image processing device 10. In this case, all the processes including pre-processing, OCR processing, and post-processing are executed inside the image processing device 10.
[0282] (5) In the above embodiment, the image processing device 10 is described as executing the process of separating the image area corresponding to the image data for each object. However, the OCR processing server 20 may execute the process.
[0283] (6) In the above-mentioned embodiment, the case where the post-processing of the text data obtained by the OCR processing is transferred to the image processing device 10 that has performed the pre-processing is described, but the text data obtained by the OCR processing can also be output to a processing device different from the image processing device 10 that has performed the pre-processing.
[0284] (7) The processor in the aforementioned embodiments refers to a processor in a broad sense, and includes not only general-purpose processors (such as CPU (=Central Processing Unit)), but also special-purpose processors (such as GPU, ASIC (=Application Specific Integrated Circuit), FPGA, programmable logic devices, etc.).
[0285] Furthermore, the actions of the processors in the aforementioned embodiments may be performed by a single processor alone, or may be performed collaboratively by a plurality of processors located in physically separate locations. Furthermore, the order in which the actions of the processors are performed is not limited to the order described in the aforementioned embodiments, but may be changed independently.
[0286] The above-mentioned embodiments of the present invention are provided for the purpose of illustration and description. In addition, the embodiments of the present invention do not include the present invention in its entirety and in detail, and do not limit the present invention to the disclosed methods. Obviously, various modifications and changes are self-evident to those skilled in the art to which the present invention belongs. The present embodiment is selected and described in order to most easily understand the principles of the present invention and its application. Thus, other technicians in the art can understand the present invention through various modified examples optimized for specific uses assumed to be various embodiments. The scope of the present invention is defined by the above claims and their equivalents.
Claims
1. An information processing device comprising a processor, The processor performs the following processing: performing preprocessing on the acquired image data; When receiving information identifying at least one local area in an image corresponding to the image data from a subsequent-stage process for processing the pre-processed image data, performing pre-processing on the identified local area as an object, The post-stage processing is optical character recognition processing, and the information of the local area is determined to include information related to a reason why the accuracy of the processing result is low and information indicating the content of pre-processing corresponding to the reason, The local area does not include a graphic area and a graph area in the image data. When the information for determining the local area includes information indicating that the accuracy of the processing result with the local area as the object is lower than a predetermined threshold, the processor performs a preprocessing different from that performed previously on the local area determined based on the information for determining the local area, The processor estimates a cause of low accuracy of a processing result based on a history of preprocessing of other image data similar to the image data.
2. The information processing device according to claim 1, wherein: The information for determining the local area includes information related to accuracy, where the accuracy is the accuracy of a result of processing the image data corresponding to the local area.
3. The information processing device according to claim 2, wherein: The information related to the accuracy is given in units of local areas.
4. The information processing device according to claim 2 or 3, wherein: The information related to the accuracy indicates that the accuracy of the processing result is lower than a predetermined threshold value.
5. The information processing device according to claim 1, wherein: The information related to the cause is information related to the characters or background included in the partial area.
6. The information processing device according to claim 1, wherein: The information indicating the content of the preprocessing corresponding to the cause includes information indicating a parameter value used in the preprocessing.
7. The information processing device according to claim 1, wherein: The processor estimates a cause of low accuracy of the processing result based on the information indicating that the accuracy of the processing result is lower than a predetermined threshold value.
8. The information processing device according to claim 7, wherein: The processor includes a step of estimating a cause of low accuracy of a result of a subsequent process based on the accuracy of the entire image data.
9. The information processing device according to claim 1, wherein: The processor estimates a cause of low accuracy of a processing result based on information obtained in a preprocessing process.
10. The information processing device according to claim 1, wherein: The processor notifies a subsequent stage of processing of information identifying the local area on which the preprocessing has been performed.
11. The information processing device according to claim 1, wherein: The processor shares information identifying the pre-processed image data with a subsequent processing step.
12. The information processing device according to claim 1, wherein: In a case where it is determined that the information on the local area includes information on a reason why the accuracy of the local area is low, the processor performs preprocessing corresponding to the reason.
13. The information processing device according to claim 1, wherein: In a case where it is determined that the information of the local area includes information indicating pre-processing content, the processor performs pre-processing of the indicated content.
14. An information processing device comprising a processor, The processor performs the following processing: performing preprocessing on the acquired image data; When receiving information identifying at least one local area in an image corresponding to the image data from a subsequent-stage process for processing the pre-processed image data, performing pre-processing on the identified local area as an object, The post-stage processing is optical character recognition processing, and the information of the local area is determined to include information related to a reason why the accuracy of the processing result is low and information indicating the content of pre-processing corresponding to the reason, The local area does not include a graphic area and a graph area in the image data. When the information for determining the local area includes information indicating that the accuracy of the processing result with the local area as the object is lower than a predetermined threshold, the processor performs a preprocessing different from that performed previously on the local area determined based on the information for determining the local area, The processor estimates a cause of low accuracy of a processing result based on a difference in accuracy between local areas of the same type.
15. A recording medium having recorded thereon a program for causing a computer to implement the following functions: A function for performing preprocessing on the acquired image data; In the case where information identifying at least one local area in an image corresponding to the image data is received from a post-processing process for processing the pre-processed image data, a function of performing pre-processing on the local area identified as an object, the post-processing being an optical character recognition process, the information identifying the local area includes information related to a reason for low precision of a processing result and information indicating a pre-processing content corresponding to the reason, and the local area does not include a graphic area and a graph area in the image data; A function of performing a preprocessing different from that performed previously on the local area determined based on the information for determining the local area, when the information for determining the local area includes information indicating that the accuracy of the processing result with the local area as the object is lower than a predetermined threshold value; and A function of estimating the cause of low accuracy of a processing result based on a history of preprocessing of other image data similar to the image data.
16. An information processing method comprising the following steps: A step of performing pre-processing on the acquired image data; and A step of performing preprocessing on the determined local area as an object, in the case where information for determining at least one local area in an image corresponding to the image data is received from a subsequent side process for processing the preprocessed image data, the subsequent side process being an optical character recognition process, the information for determining the local area includes information related to a reason for low precision of a processing result and information indicating a preprocessing content corresponding to the reason, and the local area does not include a graphic area and a graph area in the image data; When the information for determining the local area includes information indicating that the accuracy of the processing result with the local area as the object is lower than a predetermined threshold, a step of performing a preprocessing different from that performed previously on the local area determined based on the information for determining the local area; and A step of estimating a cause of low accuracy of a processing result based on a history of preprocessing of other image data similar to the image data.
17. A computer program product that enables a computer to implement the following functions: A function for performing preprocessing on the acquired image data; In the case where information identifying at least one local area in an image corresponding to the image data is received from a post-processing process for processing the pre-processed image data, a function of performing pre-processing on the local area identified as an object, the post-processing being an optical character recognition process, the information identifying the local area includes information related to a reason for low precision of a processing result and information indicating a pre-processing content corresponding to the reason, and the local area does not include a graphic area and a graph area in the image data; A function of performing a preprocessing different from that performed previously on the local area determined based on the information for determining the local area, when the information for determining the local area includes information indicating that the accuracy of the processing result with the local area as the object is lower than a predetermined threshold value; and A function of estimating the cause of low accuracy of a processing result based on a history of preprocessing of other image data similar to the image data.
Citation Information
Patent Citations
Accounting data input support system, method, and program
JP2015146075A
Character recognition processing device, character recognition processing method, and computer program
JP2007086954A