Information processing system, method, and program

JP2024107598A5Pending Publication Date: 2025-11-13PFU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023011609
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing optical character recognition (OCR) technologies face challenges in achieving high accuracy due to factors such as copy-forgery-inhibited patterns, noise, ruled lines, and faded characters, making it difficult to configure optimal image processing settings for character recognition.

Method used

An information processing system that analyzes captured images to determine recommended image processing settings by selecting candidate values for various setting items, repeatedly trying image processing, and determining settings that yield the highest character recognition accuracy.

Benefits of technology

Enables easy specification of image processing settings that enhance character recognition accuracy, allowing users to obtain high-quality images suitable for OCR without requiring expert knowledge, and reduces processing time to determine optimal settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To easily determine an image processing setting capable of acquiring an image suitable for character recognition.SOLUTION: An information processing system is provided with: an image acquisition unit which acquires a photographed image obtained by photographing a manuscript; and an analysis unit which determines, by using the photographed image, a recommended setting for a plurality of setting items in image processing means for converting an image obtained in a result of image processing performed on the read image of the manuscript read by the reading means to an image suitable to character recognition. The analysis unit selects at least one setting value as a candidate of the recommended setting for each of at least one setting item among the plurality of setting items by executing the analysis process using the photographed image, and determines the recommended setting by repeatedly making trial of image processing on the photographed image while changing the setting value of each of the plurality of setting items after limiting the setting value for the at least one setting item to the setting value selected as the candidate of the recommended setting.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a technique for configuring a device. [Background technology]

[0002] Conventionally, a character recognition parameter optimization method (see Patent Document 1) has been proposed, which includes a first means for retaining image information relating to characters in the same form obtained by scanning the same form only once so that the image information can be read out repeatedly; a second means for repeating a character recognition process a predetermined number of times by reading the image information relating to the characters in the form and automatically set parameters relating to character recognition accuracy; a third means for outputting the image information each time the second means repeats the character recognition process, as if it had been obtained by actually scanning the same form; and a fourth means for measuring the accuracy of character recognition based on the result of the character recognition process and correct answer information relating to the characters in the form each time the result is output from the second means. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2000-322515 A Summary of the Invention [Problem to be solved by the invention]

[0004] Conventionally, when converting documents (paper documents) such as papers or slips into digital data, character recognition processing (OCR (Optical Character Recognition) processing) is performed on the image obtained by reading the document using an image reading device such as a scanner. However, character recognition accuracy can be reduced due to various factors such as background patterns and noise on the document, ruled lines, overlapping seal impressions, and blurred characters. There are various image processing settings (settings related to image reading devices) to eliminate (improve) these factors that reduce character recognition accuracy, but it is difficult for users to combine these settings effectively to obtain an image suitable for character recognition.

[0005] In view of the above problems, an object of the present disclosure is to easily identify image processing settings that can obtain an image suitable for character recognition. [Means for solving the problem]

[0006] An example of the present disclosure is an information processing system including: an image acquisition means for acquiring an image of a document; and an analysis means for determining, using the captured image, recommended settings for a plurality of setting items in the image processing means, so that an image obtained as a result of image processing by an image processing means on a read image of the document read by an image reading means becomes an image suitable for character recognition, wherein the analysis means includes: a candidate selection means for selecting at least one setting value that is a candidate for the recommended setting from settable setting values ​​for at least one of the plurality of setting items by performing analysis processing using the captured image; and a recommended setting determination means for determining the recommended setting for the plurality of setting items by repeatedly attempting image processing on the captured image while changing the setting values ​​of each of the plurality of setting items and limiting the setting value for the at least one setting item to the at least one setting value selected as the candidate for the recommended setting.

[0007] The present disclosure can be understood as an information processing device, a system, a method executed by a computer, or a program executed by a computer. The present disclosure can also be understood as a program recorded on a recording medium readable by a computer or other device, machine, etc. Here, a recording medium readable by a computer, etc. refers to a recording medium that stores information such as data and programs by electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer, etc. Effect of the Invention

[0008] According to the present disclosure, it is possible to easily identify image processing settings that can obtain an image suitable for character recognition. [Brief description of the drawings]

[0009] [Figure 1] 1 is a schematic diagram showing a configuration of a system according to a first embodiment. [Diagram 2] 1 is a diagram illustrating an outline of a functional configuration of an information processing device according to a first embodiment. [Diagram 3] 5A to 5C are diagrams illustrating an example of image processing setting items related to OCR according to the embodiment and options thereof. [Figure 4] FIG. 11 is a diagram showing an example of a captured image that has been converted into a grayscale image according to the embodiment; [Diagram 5] FIG. 13 is a diagram illustrating an example of a histogram of an edge image according to the embodiment. [Figure 6] 11 is a diagram showing an example of a line segment extracted from a captured image according to the embodiment; FIG. [Figure 7] FIG. 2 is a diagram showing an example of a binarized image (an image obtained by cutting out an OCR area) of a captured image according to the embodiment. [Figure 8] FIG. 13 is a diagram illustrating an example of setting values ​​according to an estimated noise amount according to the embodiment. [Figure 9] 11 is a diagram showing an example of a plurality of setting items and options after narrowing down according to the embodiment; FIG. [Figure 10]1 is a diagram for explaining a method for evaluating a character recognition result according to an embodiment; [Figure 11] 11 is a diagram for explaining a method for calculating an evaluation value based on a reliability according to an embodiment. FIG. [Figure 12] FIG. 13 is a diagram illustrating an example of a presetting screen according to the embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of a recommended setting decision screen according to the embodiment. [Figure 14] FIG. 13 is a diagram showing an example of a progress display screen according to the embodiment. [Figure 15] FIG. 13 is a diagram illustrating an example of a recommended setting save screen according to the embodiment. [Figure 16] 10 is a flowchart showing an overview of the flow of a recommended setting determination process according to the first embodiment. [Figure 17] FIG. 11 is a schematic diagram showing a configuration of a system according to a second embodiment. [Figure 18] FIG. 11 is a diagram illustrating an outline of a functional configuration of a server according to a second embodiment. [Figure 19] FIG. 13 is a schematic diagram showing a configuration of a system according to a third embodiment. [Figure 20] FIG. 11 is a diagram illustrating an outline of the functional configuration of a scanner according to a third embodiment. [Figure 21] FIG. 13 is a diagram illustrating an outline of a functional configuration of an information processing device according to a fourth embodiment. [Figure 22] FIG. 4 is a diagram illustrating an example of a document scan screen according to the embodiment. [Diagram 23] FIG. 11 is a diagram showing an example of a pre-setting screen (before setting) according to the embodiment. [Figure 24] FIG. 13 is a diagram showing an example of a pre-setting screen (after setting) according to the embodiment. [Diagram 25] FIG. 13 is a diagram showing an example of an evaluation result display screen according to the embodiment. [Figure 26] FIG. 13 is a diagram showing an example of a display screen of an evaluation result according to the embodiment (when a correct text is acquired). [Figure 27] FIG. 13 is a diagram showing an example of a display screen of an evaluation result according to the embodiment (when the correct text is not acquired). [Figure 28] 13 is a flowchart showing an overview of the flow of an evaluation result display process according to the fourth embodiment. [Figure 29] 13 is a flowchart showing an outline of the flow of a pop-up display process according to the fourth embodiment. [Diagram 30] FIG. 13 is a diagram illustrating an outline of a functional configuration of an information processing device according to a fifth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of an information processing system, an information processing device, a method, and a program according to the present disclosure will be described with reference to the drawings. However, the embodiment described below is an example of an embodiment, and the information processing system, the information processing device, the method, and the program according to the present disclosure are not limited to the specific configuration described below. In carrying out the present disclosure, a specific configuration according to the embodiment may be appropriately adopted, and various improvements and modifications may be made.

[0011] [First embodiment] In the first to third embodiments, an information processing system, an information processing device, a method, and a program according to the present disclosure will be described as being implemented in a system for estimating (determining) image processing settings for a scanner to make an image obtained by scanning a document suitable for character recognition (OCR). However, the information processing system, the information processing device, the method, and the program according to the present disclosure can be widely used in a technology for estimating image processing settings for obtaining an image suitable for character recognition, and the application of the present disclosure is not limited to the examples shown in the embodiments.

[0012] Conventionally, there is an automatic binarization (binary image processing technology) as a technology for outputting an image suitable for OCR. This technology is a technology (function) that automatically determines binarization parameters (parameter values) for outputting an appropriate binary black-and-white image according to the document (document to be read) by analyzing some of the features of the document during scanning. However, this feature analysis alone may not provide sufficient recognition accuracy when the output image is OCR processed. For example, this technology does not distinguish between the background and the text (determines using a grayed histogram), so in the case of a document with a complex background pattern or watermark, the background may remain in the output image or the text may partially disappear, and in such cases, the recognition accuracy of the OCR is not sufficient. In addition, this technology analyzes the document during scanning and determines the binarization parameters, so there is an issue with the processing time when high-speed and large-volume scanning is required. Therefore, the issue is how to generate a profile with higher accuracy (higher recognition accuracy) in advance according to the document, rather than determining the parameters during scanning.

[0013] As a solution to this problem, it is possible to carry out image processing and OCR processing for all combinations of multiple image processing (image processing settings) related to OCR, and select the combination with the highest OCR recognition accuracy from all combinations as the image processing settings suitable for OCR. However, there are many settings related to OCR (affecting OCR accuracy), and simply combining (multiplying) multiple image processing related to OCR results in a huge number of combinations, so it is not realistic to carry out the above-mentioned processing for all these combinations. Therefore, it is possible to reduce the number of combinations, but there is a possibility that a setting suitable for OCR cannot be obtained by randomly thinning out the combinations. For example, if even a small amount of noise remains in the output image, it affects the recognition accuracy of OCR, so it is desirable to fine-tune the parameter values ​​(image processing settings) so that as little noise remains as possible, but if the settings suitable for OCR are thinned out by randomly thinning out the combinations, it becomes difficult to perform the fine adjustment.

[0014] In view of such circumstances, the information processing system, information processing device, method, and program according to the present embodiment selects candidates (setting values) for recommended settings by performing analysis processing using captured images, and then repeatedly tries image processing on the captured images while changing the respective setting values ​​of multiple setting items, limiting the setting values ​​selected as candidates for recommended settings (image processing settings for making an acquired image suitable for character recognition), thereby determining recommended settings for multiple setting items, thereby making it possible to easily identify image processing settings capable of acquiring an image suitable for character recognition processing. This makes it possible to determine in advance (generate in advance a profile) image processing settings with higher accuracy (higher recognition accuracy) according to the original.

[0015] <System configuration> 1 is a schematic diagram showing the configuration of a system 9 according to this embodiment. The system 9 according to this embodiment includes a scanner 8 and an information processing device 1 that are communicatively connected to each other via a network or other communication means.

[0016] The information processing device 1 is a computer including a central processing unit (CPU) 11, a read only memory (ROM) 12, a random access memory (RAM) 13, a storage device 14 such as an electrically erasable and programmable read only memory (EEPROM) or a hard disk drive (HDD), an input device 15 such as a keyboard, a mouse, or a touch panel, an output device 16 such as a display, and a communication unit 17 such as a network interface card (NIC). However, the specific hardware configuration of the information processing device 1 can be omitted, replaced, or added as appropriate depending on the embodiment. In addition, the information processing device 1 is not limited to a device consisting of a single housing. The information processing device 1 may be realized by multiple devices using so-called cloud or distributed computing technology.

[0017] The scanner 8 is a device (image reading device) that captures an image (image data) by capturing an image of a document, business card, receipt, or photo / illustration, etc., set by the user. In this embodiment, a scanner is exemplified as an image reading device, but the image reading device is not limited to a scanner and may be a multifunction device or the like. The scanner 8 according to this embodiment has a function of transmitting image data obtained by capturing an image to the information processing device 1 via a network. The scanner 8 may further have a user interface, such as a touch panel display or a keyboard, for enabling character input / output and item selection, as well as a web browsing function and a server function. The communication means and hardware configuration of the scanner that can employ the method according to this embodiment are not limited to those exemplified in this embodiment.

[0018] 2 is a diagram showing an outline of the functional configuration of the information processing device 1 according to the present embodiment. The information processing device 1 functions as a device including an image acquisition unit 31, a reception unit 32, an analysis unit 33, a storage unit 34, and a presentation unit 35, by a program recorded in the storage device 14 being read into the RAM 13 and executed by the CPU 11, and each piece of hardware included in the information processing device 1 is controlled. The image acquisition unit 31 includes a read image acquisition unit 41 and a read image processing unit 42. The reception unit 32 includes a character region acquisition unit 43 and a correct answer information acquisition unit 44. The analysis unit 33 includes a candidate selection unit 45 and a recommended setting determination unit 46. Note that in this embodiment and other embodiments described later, each function included in the information processing device 1 is executed by the CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.

[0019] The image acquisition unit 31 acquires a captured image (original image) of an original. In this embodiment, the image acquisition unit 31 corresponds to a driver (scanner driver) of the scanner 8 (the "reading means" in this embodiment), and acquires a captured image of an original by controlling the scanner 8 to capture an image of the original set by a user. Specifically, the image acquisition unit 31 includes a read image acquisition unit 41 and a read image processing unit 42 (the "image processing means" in this embodiment), and the read image acquisition unit 41 acquires a read image generated by the scanner 8 reading an original, and the read image processing unit 42 acquires an image that has been subjected to image processing (processed image) by performing image processing on the read image. Note that in this embodiment, the read image refers to an image that has not been subjected to image processing (raw image).

[0020] The document read by the scanner 8 (a document used for the analysis process described later, hereinafter referred to as the "read document") may be any document, for example, a document used when operating the scanner 8 (a customer-operated document). The read document may be one or more documents. When multiple documents are read (imaged) by the scanner, the read image acquisition unit 41 acquires an image of each of the multiple documents. The image processing performed on the processed image may be any image processing. When the scanner 8 is equipped with an image processing means (read image processing unit 42), the image acquisition unit 31 acquires not only the read image but also the processed image from the scanner 8.

[0021] The reception unit 32 receives input of the specification of the OCR area and the correct string for the read document by the user selecting a field in the read document (captured image) for character recognition (a character area (OCR area) that is an area including a character string that the user wishes to character recognize) and inputting the correct string written in that area. In other words, when the user specifies an OCR area and inputs a correct string (correct text read by the user from the OCR area), the character area acquisition unit 43 acquires the OCR area, and the correct information acquisition unit 44 acquires the correct string (correct information) for each OCR area. Note that either one or multiple OCR areas may be selected.

[0022] The analysis unit 33 uses the captured image (the read image or the processed image) to determine (estimate) the recommended image processing settings (recommended settings suitable for the read document) for obtaining an image (binarized image) suitable for character recognition. Specifically, the analysis unit 33 uses the captured image to determine the recommended settings (recommended values) for a plurality of setting items in the image processing means so that the image obtained as a result of image processing performed by the image processing means (the read image processing unit 42) on the read image of the read document read by the scanner 8 becomes an image suitable for character recognition. Here, the setting items for which the recommended settings are determined are image processing setting items related to character recognition (OCR), and more specifically, image processing setting items that may affect character recognition (character recognition results) (setting items for which the character recognition results for the image obtained as a result of image processing may differ depending on the setting contents). In this embodiment, the setting items for which the recommended settings are determined are exemplified by image processing setting items related to character thickness, background pattern removal, extraction of specific characters (special characters), dropout color, binarization sensitivity, and noise removal. However, the setting items for which the recommended settings are determined are not limited to this example, and may be any setting items and any number of setting items. Furthermore, the setting items for which the recommended settings are determined may include setting items other than image processing setting items related to character recognition.

[0023] Here, the image processing setting items include setting items that have a large effect on the entire document (captured image of the document). For such setting items, it is possible to (roughly) grasp (understand) the characteristics of the document by performing document analysis (captured image analysis) or by trying image processing while changing the setting values, and it is possible to narrow down recommended setting candidates (setting value candidates that can become recommended settings (suitable as recommended settings)) from multiple configurable setting values.

[0024] FIG. 3 is a diagram showing an example of image processing setting items related to OCR and their options according to the present embodiment. FIG. 3 shows setting items (image processing setting items) related to OCR, multiple setting values ​​(options) that can be set for each setting item, and the number of the multiple setting values ​​(total number of options). As shown in FIG. 3, since there are multiple setting values ​​(options) that can be set for the setting items related to OCR, the number of combinations by simply multiplying the options of these multiple setting items is huge, and if all of these combinations are verified (recommended setting determination process described later), a huge amount of time is required (it does not fit within a finite time). Therefore, in this embodiment, an analysis process using a captured image is performed to narrow down the setting values ​​from multiple settable setting values ​​(narrow down to setting values ​​of recommended setting candidates). In other words, an analysis process using a captured image (analysis process capable of capturing the characteristics of the read document) is performed, and a setting value according to the result of the analysis process (a setting value according to the characteristics of the read document) is determined as a recommended setting candidate. This narrows down the number of combinations to be verified, and in the recommended settings determination process described below, it becomes unnecessary to verify (obtain image processing and character recognition results) all combinations by simple cross-combinations of the options shown in Fig. 3. In other words, the number of times (repetitions) image processing and obtaining character recognition results can be reduced, making it possible to quickly determine recommended settings.

[0025] As described above, in this embodiment, the analysis unit 33 performs an analysis process using the captured image to select candidates (candidate values) for recommended settings, and determines the recommended settings using the selected candidates for recommended settings. Below, we will explain each of the candidate selection unit 45 that selects candidates for recommended settings and the recommended setting determination unit 46 that determines the recommended settings.

[0026] (Candidate Selection) The candidate selection unit 45 performs an analysis process using the captured image to select a setting value that is a candidate for the recommended setting from a plurality of settable setting values ​​for at least one setting item among a plurality of setting items for which the recommended setting is determined. The setting value selected as the candidate for the recommended setting may be one or more. In this embodiment, the analysis process using the captured image is performed to capture the amount of background pattern, the presence or absence of specific characters (white-out characters, shaded characters, characters with overlapping imprints, etc.), the presence or absence of ruled lines (color of ruled lines), the amount of noise (presence or absence of noise), etc. as features of the read document (captured image). In this embodiment, as an example, the candidate for the recommended setting for the image processing setting items related to background pattern removal, extraction of specific characters, dropout color, binarization sensitivity, and noise removal is selected. When the candidate selection unit 45 selects a candidate value for a certain setting item, it performs an analysis process that can capture features related to the setting item (features of the read document).

[0027] There are mainly two methods for selecting candidates (candidate values) of recommended settings. The first method (method 1) is a method for selecting candidate values ​​by performing image analysis on a captured image. The second method (method 2) is a method for performing image processing using configurable setting values ​​on a captured image, and selecting candidate values ​​based on a character recognition result for an image obtained as a result of the trial (hereinafter referred to as "single verification"). In this embodiment, in order to select candidates of recommended settings for a plurality of setting items by method 1 and / or method 2, the candidate selection unit 45 includes an image analysis unit 51, a first image processing unit 52, a first recognition result acquisition unit 53, and a selection unit 54. In method 1, the image analysis unit 51 performs image analysis on the captured image. In method 2, the first image processing unit 52 performs (trials) image processing on the captured image, and the first recognition result acquisition unit 53 acquires an image obtained as a result of the trial (a captured image on which image processing has been performed) and a character recognition result (OCR result) for the captured image. The selection unit 54 selects candidate values ​​based on the result of the image analysis by the image analysis unit 51 or the character recognition result acquired by the first recognition result acquisition unit 53.

[0028] The first recognition result acquisition unit 53 may acquire the character recognition result by performing character recognition processing (OCR processing), or may acquire the character recognition result from another device that performs character recognition processing (device equipped with an OCR engine). A method for selecting recommended setting candidates for each setting item will be described below.

[0029] (Remove background pattern) The image processing setting item related to background pattern removal (hereinafter referred to as the "background pattern removal item") is a setting item for image processing aimed at removing background patterns (including watermarks) contained in a document (scanned image). In the case of a document containing a background pattern, the influence of the background pattern may deteriorate the character recognition accuracy of an image captured from the document. Therefore, in order to obtain an image suitable for character recognition, it is desirable to set the background pattern removal item in a way suitable for the document. Candidates for recommended settings for the background pattern removal item can be selected by method 1 or method 2.

[0030] (Background Pattern Removal: Method 1) In the case of method 1, first, the candidate selection unit 45 (image analysis unit 51) performs image analysis to determine the amount of the background pattern in the captured image. Based on the result of this image analysis, the candidate selection unit 45 (selection unit 54) can estimate the amount of the background pattern contained in the read document as a feature of the read document. The candidate selection unit 45 (selection unit 54) selects a setting value for the background pattern removal item according to the result of the image analysis (estimated result of the document's features) as a candidate for the recommended setting. In this embodiment, the amount of the background pattern in the read document is determined (estimated) by performing edge analysis (histogram analysis of the edge image) on the captured image. Here, a typical background pattern has a gradation value (pixel value) when the captured image is grayscaled that is lighter than the character part (black color), and often looks like countless thin lines.

[0031] FIG. 4 is a diagram showing an example of a grayscaled captured image (part) according to the present embodiment. As shown in FIG. 4, in the grayscaled captured image, the background pattern in the background of the characters is thinner than the characters and looks like countless thin lines. Therefore, in this embodiment, an edge analysis is performed on the captured image (grayscaled image), and the amount of the background pattern included in the read document is estimated based on the detected edge amount. Specifically, the candidate selection unit 45 first uses an edge filter (such as a Laplacian filter) to extract edge parts (amount of change in pixel value from surrounding pixels (edge ​​amount)) from the grayscaled captured image (generation of edge image). Then, a histogram (edge ​​amount histogram) of the edge image is generated, and the appearance of peaks in the generated histogram is analyzed to estimate the amount of the background pattern.

[0032] FIG. 5 is a diagram showing an example of a histogram of an edge image according to this embodiment. In the histogram shown in FIG. 5, the horizontal axis indicates the edge amount (gradation value (pixel value) in the edge image), and the vertical axis indicates the number of pixels. In the histogram shown in FIG. 5, three peaks appear in ascending order of gradation value. Since not many edges are detected in the background portion (solid portion) of the captured image, it can be assumed that the peaks in the areas with low gradation values ​​(areas with few edge amounts) correspond to the background portion. Furthermore, since many edges are detected in the character portion (around the characters) of the captured image, it can be assumed that the peaks in the areas with high gradation values ​​(areas with many edge amounts) correspond to the character portion. And, as described above, the background pattern is thinner than the character portion and looks like countless thin lines, so it can be assumed that the edge amount detected in the area where the background pattern is present is greater than the edge amount detected in the background portion and less than the edge amount detected in the character portion. Therefore, when a peak exists (at an intermediate position) between a peak corresponding to a background portion (solid portion) and a peak corresponding to a character portion in the edge amount histogram as shown in Fig. 5, the candidate selection unit 45 infers (determines) that the peak is a peak corresponding to a background pattern and that a background pattern exists in the read document. Then, when it is inferred that a background pattern exists, the candidate selection unit 45 estimates (guesses) the amount of the background pattern included in the read document based on the amount of the portion corresponding to the background pattern (the number of pixels (frequency) around the peak corresponding to the background pattern).

[0033] Then, the candidate selection unit 45 (selection unit 54) selects a setting value according to the image analysis result (prediction result) from among the settable setting values ​​as a candidate for the recommended setting. For example, if the result of the above prediction is that the document does not contain a background pattern, the candidate selection unit 45 selects, for example, "no background pattern removal (background pattern removal process disabled)" as a candidate for the recommended setting for the background pattern removal item. If the result of the above prediction is that the document contains a small amount of background pattern, the candidate selection unit 45 selects, for example, two setting values ​​("background pattern removal level 1 (Lv1)" and "background pattern removal level 2 (Lv2)") in order of the degree to which background pattern removal is performed as a candidate for the recommended setting for the background pattern removal item. If it is determined (judged) as a result of the above estimation that the document contains a large amount of background patterns, the candidate selection unit 45 selects, as candidates (candidate values) for the recommended setting for the background pattern removal item, for example, two setting values ​​("background pattern removal level 2 (Lv2)" and "background pattern removal level 3 (Lv3)") in order of the degree to which background pattern removal is performed.

[0034] In addition, any method such as peak search may be used to detect the peak in the histogram. The above-mentioned method is an example of image analysis for determining the amount of background pattern, and various methods (any method) may be used for image analysis for determining the amount of background pattern. In addition, the filter used to generate the edge image is not limited to the Laplacian filter, and any filter may be used.

[0035] (Texture removal: Method 2) In the case of method 2, the candidate selection unit 45 tries image processing by possible setting values ​​(for example, "without background pattern removal", "background pattern removal level 1", "background pattern removal level 2", and "background pattern removal level 3") for the background pattern removal item on the captured image, and selects a candidate for the recommended setting based on the character recognition result for the image obtained as a result of the trial. For example, if the captured image used is an image acquired with "without background pattern removal (setting that does not perform background pattern removal processing)", the candidate selection unit 45 tries image processing (background pattern removal processing) by each of "background pattern removal level 1", "background pattern removal level 2", and "background pattern removal level 3" on the captured image, and selects a candidate value for the background pattern removal item based on the character recognition result for each image (three images) obtained as a result of the trial and the captured image (image corresponding to "without background pattern removal"). In other words, the candidate selection unit 45 selects candidate values ​​for the background pattern removal item by comparing character recognition results for images corresponding to each setting value (a captured image for "no background pattern removal" and images obtained as a result of image processing using each setting value for "background pattern removal level 1", "background pattern removal level 2", and "background pattern removal level 3"). For example, when the character recognition results are compared and the character recognition result for an image obtained by trialing image processing using "background pattern removal level 3" is the best, it is possible to determine (infer) that a large amount of background patterns is included in the scanned document. In this case, the candidate selection unit 45 selects "background pattern removal level 2" and "background pattern removal level 3", which are setting values ​​that perform a high degree of background pattern removal, as candidates for the recommended settings.

[0036] That is, the candidate selection unit 45 selects a predetermined number (one or more) of setting values ​​(e.g., two setting values) selected in order of best character recognition results (recognition rate) for an image obtained as a result of trial image processing from among the settable setting values ​​as candidates for the recommended setting for the background pattern removal item. Note that, as a method for evaluating the character recognition results, an evaluation method performed when determining the recommended setting, which will be described later, may be used. That is, character recognition results may be compared by using evaluation method 1 or evaluation method 2, which will be described later. Also, candidates may be selected by comparing the CC number in addition to the character recognition result (OCR recognition rate). For example, a predetermined number (e.g., two) of setting values ​​that are best in the order of character recognition result and CC number are selected as candidate values.

[0037] (Extracting specific characters) Image processing setting items related to character extraction (function) (hereinafter referred to as "character extraction items") are setting items for image processing that aim to obtain an image with high character visibility even when specific characters that are difficult to recognize as they are exist. When specific characters such as white-out characters, characters with a shaded background, and characters with overlapping seal imprints exist in a document, the accuracy of character recognition for an image obtained by capturing the document may deteriorate due to the influence of the specific characters. Therefore, in order to obtain an image suitable for character recognition, it is desirable to set the character extraction items suitable for the document. In this embodiment, as examples of character extraction items, image processing setting items related to white-out character extraction, image processing setting items related to shaded character extraction, and image processing setting items related to overlapping seal imprint character extraction are shown. Candidates for recommended settings for the character extraction items can be selected by the method 2.

[0038] The candidate selection unit 45 tries image processing with a settable setting value (for example, "ON (enabled)" and "OFF (disabled)") for the character extraction item on the captured image, and selects a candidate for the recommended setting based on the character recognition result for the image obtained as a result of the trial. For example, when the captured image used is an image acquired with "OFF (setting not to perform character extraction processing)", the candidate selection unit 45 tries image processing (character extraction processing) with the setting value "ON (enabled)" on the captured image, and selects a candidate value for the character extraction item based on the character recognition result for the image obtained as a result of the trial (one image) and the captured image (image corresponding to "OFF (disabled)"). In other words, the candidate selection unit 45 selects a candidate value for the character extraction item by comparing the character recognition results for the images corresponding to each setting value (the captured image for "OFF" and the image obtained as a result of image processing with the setting value "ON" for "ON"). For example, when comparing the character recognition results when the image processing setting item for white-out character extraction is "ON" with the character recognition results when it is "OFF," if the character recognition results for the image obtained by trialing image processing with the setting value "ON" are better, it is possible to determine (infer) that white-out characters are included in the scanned document. In this case, the candidate selection unit 45 selects the setting value for white-out character extraction ("ON") as a candidate for the recommended setting.

[0039] That is, the candidate selection unit 45 selects, from among possible setting values ​​(e.g., ON, OFF), a setting value (e.g., ON) that provides the best character recognition result (character recognition rate) for an image obtained as a result of trial image processing, as a candidate for a recommended setting for a character extraction item. Note that, as a method for evaluating the character recognition result, an evaluation method that is used when determining a recommended setting, which will be described later, may be used. That is, the character recognition results may be compared with each other by using evaluation method 1 or evaluation method 2, which will be described later.

[0040] (Dropout Color) Image processing setting items related to dropout color (hereinafter referred to as "dropout color items") are setting items for image processing aimed at preventing a specified color from appearing in an image (making it difficult to appear). For example, in the case of a document containing ruled lines, the influence of the ruled lines may deteriorate the character recognition accuracy of an image obtained by capturing the document. Therefore, in order to obtain an image suitable for character recognition, it is desirable to set the dropout color item in a way suitable for the document, for example, by setting the color of the ruled lines as the dropout color to erase the ruled lines. Candidates for recommended settings for the dropout color item can be selected by the method 1.

[0041] First, the candidate selection unit 45 (image analysis unit 51) performs image analysis to determine the presence or absence of ruled lines in the captured image. Based on the result of this image analysis, the candidate selection unit 45 (selection unit 54) can estimate the presence or absence of ruled lines (whether or not ruled lines exist) in the read document as a feature of the read document. The candidate selection unit 45 (selection unit 54) selects a setting value for the dropout color item according to the result of the image analysis (estimated result of the document's features) as a candidate for the recommended setting. In this embodiment, the presence or absence of ruled lines in the read document is determined (estimated) by performing a line segment extraction process on the captured image. Any method may be used for the line segment extraction process (process of extracting lines in an image). For example, edge extraction and Hough transform are performed on the captured image to extract lines (line segment list).

[0042] FIG. 6 is a diagram showing an example of line segments extracted from a captured image according to the present embodiment. As a result of the line segment extraction process performed by the candidate selection unit 45 on the captured image, line segments are extracted as shown by thick lines in FIG. 6. The candidate selection unit 45 performs line segment extraction process (analysis for determining the presence or absence of line segments) on the captured image, and estimates the presence or absence of ruled lines in the read document based on the result. Specifically, when a line segment is extracted as shown in FIG. 6 as a result of the line segment extraction process, the candidate selection unit 45 estimates (determines) that a ruled line exists in the read document. Then, when it is estimated that a ruled line exists, the candidate selection unit 45 estimates the color of the ruled line included in the read document by judging the color of the line segment corresponding to the ruled line (performing a ruled line color analysis). Note that the color of one of the extracted line segments may be estimated as the color of the ruled line, or the color of the ruled line may be estimated based on the colors of the extracted multiple line segments. For example, the color information constituting the line segments is converted into a histogram, and the color that appears most frequently is estimated as the color of the ruled line.

[0043] Then, the candidate selection unit 45 (selection unit 54) selects a setting value according to the image analysis result (estimated result) from among possible setting values ​​(setting values ​​for each of RGB (values ​​from 0 to 255)) as a candidate for the recommended setting. For example, if the result of the estimation indicates that a ruled line exists in the document, the candidate selection unit 45 selects a setting value corresponding to the color of the ruled line estimated from the color of the extracted line segment as a candidate (candidate value) for the recommended setting for the dropout color item.

[0044] Note that some OCR systems use ruled lines for form recognition, and in such cases it is not appropriate to erase the ruled lines. Therefore, the user may be allowed to select in advance whether or not to remove the ruled lines (whether or not to set the ruled line color as the dropout color).

[0045] (binary sensitivity, noise reduction) Auto binary is an image processing that automatically adjusts a threshold value suitable for binarizing an image to binarize it, and is a function that separates characters from the background and creates an image with good contrast. The image processing setting item related to binarization sensitivity (hereinafter referred to as the "binarization sensitivity item") is an item for setting the sensitivity (effect) of this auto binary, and is an item for the purpose of removing background noise and clarifying characters. For example, if the effect (sensitivity) of auto binary is too large, noise is likely to occur. If a lot of noise occurs (if the document is one that is prone to generating noise in the captured image), the noise may affect the character recognition results for the image of the document. Therefore, in order to obtain an image suitable for character recognition (an image with less noise), it is desirable to set the binarization sensitivity item to be suitable for the document, such as lowering the auto binary sensitivity (binarization sensitivity) when a lot of noise occurs. In addition, the image processing setting item related to noise removal (specifying dust removal) (hereinafter referred to as the "noise removal item") is an image processing setting item for the purpose of removing isolated points after binarization (automatic binarization) (performing fine adjustments when noise remains). For the noise reduction item, it is desirable to set the item appropriately for the manuscript for the same reason as for the binarization sensitivity item. Candidates for the recommended settings for the binarization sensitivity item and noise reduction item can be selected using Method 1.

[0046] The candidate selection unit 45 (image analysis unit 51) performs image analysis (noise analysis) to determine the amount of noise in the captured image. Based on the result of this image analysis, the candidate selection unit 45 (selection unit 54) can estimate the amount of noise generated when the scanned document is captured as a feature of the scanned document. The candidate selection unit 45 (selection unit 54) selects setting values ​​for the binarization sensitivity item and the noise elimination item according to the result of the image analysis (estimated result of the document's features) as candidates for the recommended settings of the binarization sensitivity item and the noise elimination item, respectively. In this embodiment, noise analysis is performed on the binarized image of the captured image by the following method. Note that in this embodiment, a noise analysis is performed on the captured image (binarized image) that has been subjected to image processing using a candidate value for the background pattern elimination item, thereby determining candidate values ​​for the binarization sensitivity item and the noise elimination item that correspond to (are combined with) the candidate value for the background pattern elimination item. However, without being limited to this example, the noise amount may be estimated by performing the following noise analysis on the binarized image of the captured image, and candidate values ​​for the binarization sensitivity item and the noise removal item may be determined.

[0047] First, the user inputs a field (OCR area) to be character-recognized in the read document and a correct character string written in the area, and the reception unit 32 acquires the OCR area and the correct character string for the read document in advance. Then, the candidate selection unit 45 calculates the number of black blocks (black connected pixel blocks) that are connected components (CCs) (hereinafter referred to as the "CC number") in each OCR area in an image obtained by performing image processing (background pattern removal processing) based on a candidate value for the background pattern removal item on a captured image (binarized image). That is, the candidate selection unit 45 calculates the CC number for each image (partial image) obtained by cutting out the OCR area of ​​the image. Note that, when the candidate value for the background pattern removal item is "no background pattern removal," the candidate selection unit 45 calculates the CC number in each OCR area in the binarized image of the captured image that has not been subjected to background pattern removal processing. In addition, the candidate selection unit 45 calculates the expected value of the CC number (hereinafter referred to as the "CC number expected value") in each OCR area based on the correct character string for the OCR area. Then, the candidate selection unit 45 estimates the amount of noise in the read document (the amount of noise generated when the read document is imaged) by comparing the calculated CC number with the CC number expected value.

[0048] The CC number expected value may be calculated using one of the following two methods. In the first method, the CC number expected value is calculated using data (CC number dictionary data) that summarizes the CC number expected value for each character. The candidate selection unit 45 searches for the CC number expected value for each character included in the correct character string from the CC number dictionary data, and calculates the CC number expected value for the OCR area by adding up the CC number expected values ​​searched for for each character. In the second method, the CC number expected value is calculated from the language of the characters to be recognized (the language of the characters in the OCR area) and the number of characters in the correct character string. It can be seen that the number of CCs per character is somewhat related to the language. For example, the number of CCs is often large for Chinese, and small for English. Therefore, the candidate selection unit 45 sets a coefficient (weighting coefficient) per character for each language, and calculates the CC number expected value based on the coefficient and the correct character string. For example, if the coefficient per character for English is set to 1.2, the CC number expected value of the correct character string "abcde" is calculated to be 6 (=1.2×5 (characters)). Also, for example, the coefficient for English is 1.2, whereas for Chinese, a higher coefficient is set, such as 2.5, for each character.

[0049] FIG. 7 is a diagram showing an example of a binarized image (an image obtained by cutting out an OCR region) of a captured image according to the present embodiment. In the case of the image of the OCR region shown in FIG. 7, for example, the expected CC number is calculated to be 14, and the actual CC number is calculated to be 1260. In this case, when the actual CC number is compared with the expected CC number, the candidate selection unit 45 estimates (determines) that the amount of noise is large because the CC number is much larger than the expected CC number. For example, the candidate selection unit 45 may estimate the amount of noise by comparing (actual CC number) / (expected CC number) with a predetermined threshold value (one or more threshold values). For example, when (actual CC number) / (expected CC number) is less than 1, it is estimated that there is no noise, when (actual CC number) / (expected CC number) is 1 or more and less than 5, it is estimated that the amount of noise is small, when (actual CC number) / (expected CC number) is 5 or more and less than 10, it is estimated that the amount of noise is medium, and when (actual CC number) / (expected CC number) is 10 or more, it is estimated that the amount of noise is large. In addition, when multiple OCR areas are set, the noise amount is estimated by, for example, comparing (the total value of the actual CC counts in the multiple OCR areas) / (the total value of the expected CC counts in the multiple OCR areas) with a predetermined threshold value (one or multiple threshold values). In this case, other representative values, etc. may be used instead of the total value.

[0050] Then, the candidate selection unit 45 (selection unit 54) selects a setting value according to the image analysis result (prediction result) from among possible setting values ​​(for example, binarization sensitivity -50 to 50) as a candidate for the recommended setting. For example, if the result of the above prediction is that it is predicted (determined) that there is no noise generated when the scanned document is imaged, a setting value of 0 or a positive direction (a direction that makes characters stand out) is selected as a candidate for the recommended setting (candidate value) for the binarization sensitivity item. Also, for example, if the result of the above prediction is that it is predicted (determined) that noise will be generated when the scanned document is imaged, a setting value in the negative direction (a direction that eliminates noise) according to the predicted amount of noise is selected as a candidate value.

[0051] Fig. 8 is a diagram showing an example of setting values ​​according to the estimated noise amount according to this embodiment. Fig. 8 illustrates the value (range) of (actual CC number) / (expected CC number) and setting values ​​(candidates for recommended settings) according to the value. As shown in Fig. 8, the value of (actual CC number) / (expected CC number), i.e., the setting values ​​of the binarization sensitivity item and the noise removal item according to the estimated noise amount, are selected as candidates for recommended settings for each item.

[0052] The above-mentioned method of noise analysis is an example, and any method may be used for noise analysis. In the present embodiment, the recommended setting candidates for the binarization sensitivity item and the noise reduction item are selected based on the noise analysis, but the recommended setting candidates for only one of the binarization sensitivity item and the noise reduction item may be selected.

[0053] In this embodiment, among the items shown in Fig. 3, the image processing setting item related to character thickness (a setting item intended for fine adjustment when characters are blurred) is not included in the selection of candidate values ​​for recommended settings (a target for narrowing down options), but the item may also be included in the selection of candidate values. Furthermore, when the candidate selection unit 45 selects candidates only by method 1, it does not necessarily have to include the first image processing unit 52 and the first recognition result acquisition unit 53. On the other hand, when the candidate selection unit 45 selects candidates only by method 2, it does not necessarily have to include the image analysis unit 51.

[0054] (Recommended settings decision) The candidate selection unit 45 narrows down the setting values ​​(candidates) that can be the recommended settings (recommended values). In response to this, the recommended settings determination unit 46 determines the recommended settings by making detailed adjustments (fine adjustments such as noise removal settings and character thickness for the purpose of completely removing noise, leaving characters, and matching the characteristics of the OCR engine). Specifically, the recommended settings determination unit 46 limits the setting values ​​for the setting items (at least one setting item among the multiple setting items) for which the setting values ​​have been narrowed down to the setting values ​​selected by the candidate selection unit 45 as candidates for the recommended settings, and then repeatedly tries image processing on the captured image (read image or processed image) while changing each of the multiple setting items, thereby determining the recommended settings for the multiple setting items. Specifically, the recommended settings for the multiple setting items are determined based on the character recognition results for the multiple images obtained by repeatedly trying image processing on the captured image while changing each of the multiple setting items. In this embodiment, in order to determine the recommended settings, the recommended settings determination unit 46 includes a second image processing unit 55, a second recognition result acquisition unit 56, and a determination unit 57.

[0055] The second image processing unit 55 performs (trials) image processing on the captured image, and the second recognition result acquisition unit 56 acquires an image obtained as a result of the trial (captured image subjected to image processing) and a character recognition result (OCR result) for the captured image. The determination unit 57 determines the recommended settings based on the character recognition result acquired by the second recognition result acquisition unit 56. The second recognition result acquisition unit 56 may acquire the character recognition result by performing character recognition processing (OCR processing), or may acquire the character recognition result from another device that performs character recognition processing. In this embodiment, the recommended setting determination unit 46 first uses the candidate values ​​of the recommended settings selected by the candidate selection unit 45 to create a combination table by simple multiplication of candidate values ​​of multiple setting items (parameters). However, the candidate values ​​for the setting item related to the character size are left as the setting values ​​that can be set for the setting item related to the character size. In this embodiment, the candidate values ​​for the binarization sensitivity item and the noise removal item are determined for each candidate value for the background pattern removal item. Therefore, when creating the above combinations (combination table), no combinations of setting values ​​for the background pattern removal item, binarization sensitivity item, and noise removal item will be created other than combinations of candidate values ​​for the background pattern removal item and candidate values ​​for the binarization sensitivity item and noise removal item corresponding to those candidate values.

[0056] FIG. 9 is a diagram showing an example of multiple setting items according to this embodiment and options after narrowing down. As shown in FIG. 9, it can be seen that options for some of the multiple setting items (binarization sensitivity, background pattern removal, noise removal, white-out character extraction, shaded character extraction, seal imprint overlap character extraction, and dropout color) are reduced by the candidate selection process by the candidate selection unit 45. The candidate selection unit 45 creates (generates) all combinations (combination table) of the setting values ​​of the multiple setting items by simply multiplying all options (candidate values) after narrowing down for the multiple setting items. The generated combinations are combinations for performing the above-mentioned detailed adjustment. At this time, if the number of combinations becomes enormous, the setting values ​​(candidate values) may be further thinned out. For example, the candidate values ​​0 to 50 of the binarization sensitivity may be thinned out so that the setting values ​​are in increments of 5. Although a combination table is created in this embodiment, the creation of the combination table is optional since image processing and character recognition may be performed using all combinations.

[0057] Next, the recommended setting determination unit 46 (second image processing unit 55) performs (trials) image processing for each combination included in all combinations after narrowing down (all combinations shown in the combination table) on the captured image. Then, the recommended setting determination unit 46 (second recognition result acquisition unit 56) acquires character recognition results for images corresponding to each combination (images obtained as a result of image processing by each combination). Then, the recommended setting determination unit 46 (determination unit 57) determines the combination (combination of setting values ​​of multiple setting items) when an image with the best character recognition result (character recognition rate) is obtained as the recommended setting for the multiple setting items. In this embodiment, an evaluation value (evaluation index) based on the character recognition result is calculated for each character recognition result, and the combination with the highest evaluation value is determined as the recommended setting. Below, two methods for evaluating the character recognition result (method of calculating the evaluation value) are illustrated.

[0058] (Evaluation method 1) In the first method, the user inputs a field (OCR area) in the scanned document to be character-recognized and a correct character string written in the field, and the reception unit 32 acquires the OCR area and the correct character string for the scanned document in advance. In the above-mentioned candidate selection process, if the OCR area and the correct character string have already been acquired, they may be used. Then, the recommended setting determination unit 46 determines for each OCR area whether or not the recognized character string, which is the character recognition result acquired for the OCR area, completely matches the correct character string for the OCR area, and calculates the number of OCR areas (number of fields) in which the recognized character string and the correct character string completely match. Hereinafter, the number of OCR areas in which the character strings completely match is referred to as the "field recognition rate." In addition, the recommended setting determination unit 46 calculates the number of characters that match between the recognized character string for all OCR areas and the correct character string for all OCR areas (the number of characters that match between the recognized character and the correct character). Hereinafter, the number of characters that match between the recognized character and the correct character (the recognition rate for each character) is referred to as the "character recognition rate."

[0059] For example, assume that the correct answer character strings for three OCR areas (OCR areas 1 to 3) in a scanned document (captured image) are "PFU Inc." for OCR area 1, "invoice" for OCR area 2, and "¥10,000" for OCR area 3. Two results are shown as examples of character recognition results (recognized character strings) obtained when character recognition is performed on these OCR areas. In the first character recognition result, the recognized character strings for each OCR area are "PFU Inc.", "invoice," and "¥IO,OOO." In this case, the recognized character strings and the correct answer character strings completely match for OCR area 1 and OCR area 2, so the field recognition rate is calculated as 2 / 3. For OCR area 3, in the recognized character string, "1 (number one)" is erroneously recognized as "I (letter I)," and "0 (number zero)" is erroneously recognized as "O (letter O)." Note that the recognized character and the correct character match for the other characters. Therefore, the character recognition rate is calculated as 11 / 16.

[0060] In the second character recognition result, assume that the recognized character strings for each OCR region are "PF Corporation", "Invoice 1", and "¥I0,000". In this case, since the recognized character strings do not exactly match the correct character strings in any of the OCR regions, the field recognition rate is calculated as 0 / 3. Also, for OCR region 1, the character "U" is not recognized in the recognized character string. For OCR region 2, the character "書" is misrecognized as "書1". For OCR region 3, the character "1 (the digit one)" is misrecognized as "I (the letter eye)". Therefore, the character recognition rate is calculated as 13 / 16.

[0061] Based on the calculated field recognition rate and character recognition rate, which are the evaluation values, the recommendation setting determination unit 46 determines (evaluates) the quality of the character recognition result. For example, a method of selecting the better ones in the order of the field recognition rate and the character recognition rate may be adopted. In this method, first, the field recognition rates for all character recognition results are compared, and the character recognition result with the best field recognition rate is determined as the best character recognition result. However, if there are multiple OCR regions with the same field recognition rate, the character recognition rates are compared among those OCR regions, and the character recognition result with the best character recognition rate is determined as the best character recognition result. When this method is used, for the first and second character recognition results described above, the first character recognition result with a better field recognition rate is determined as the better character recognition result. Note that the method of determination based on the field recognition rate and the character recognition rate is not limited to the above method, and any other arbitrary method may be used. For example, another evaluation value (evaluation index) may be obtained based on the field recognition rate and the character recognition rate, and a method of determination based on that evaluation value may be used.

[0062] FIG. 10 is a diagram for explaining a method for evaluating a character recognition result according to the present embodiment. FIG. 10 shows image processing results (images for selected OCR areas, which have been subjected to image processing using combinations of candidate values), OCR results, and character recognition rates for each combination of setting values ​​(candidate values) for multiple setting items. As shown in FIG. 10, the recommended setting determination unit 46 calculates the character recognition rate for each OCR area for each combination, and determines the combination that provides the best character recognition result as the recommended setting. Note that FIG. 10 illustrates only the character recognition rate for one OCR area, but as described above, when multiple OCR areas are set, the character recognition rate and field recognition rate for the multiple OCR areas may be calculated to determine the setting (combination of candidate values) that provides the best character recognition result.

[0063] (Evaluation method 2) In the second method, the evaluation value is calculated based on the reliability of each character acquired from the OCR engine. FIG. 11 is a diagram for explaining a method of calculating an evaluation value based on reliability according to the present embodiment. FIG. 11(a) is a diagram illustrating a method of calculating an evaluation value for a character recognition result (case 1), and FIG. 11(b) is a diagram illustrating a method of calculating an evaluation value for a character recognition result (case 2). As shown in FIG. 11, an evaluation value is calculated based on the reliability obtained from the OCR engine for the character recognition result (recognition value) for each character (correct value) of a correct character string. In the example of FIG. 11, the average value of the reliability of each character is calculated as the evaluation value. The recommended setting determination unit 46 judges (evaluates) the quality of the character recognition result based on the calculated evaluation value. For example, in the case of FIG. 11(a), 77, which is the average value of the reliability of each character, is calculated as the evaluation value, and in the case of FIG. 11(b), 91, which is the average value of the reliability of each character, is calculated as the evaluation value. Therefore, the character recognition result with a higher evaluation value (Case 2) is determined to be a better character recognition result. Note that the evaluation value is not limited to the average value of the reliability of each character, and other representative values ​​may be used.

[0064] In addition, when the captured image used in the analysis process in the candidate selection process is a read image (raw image), the captured image to be subjected to image processing in the recommended setting determination process may be the read image used in the candidate selection process, or may be an image obtained by performing image processing on the captured image (read image) used in the candidate selection process. Similarly, when the captured image to be subjected to image processing in the recommended setting determination process is a read image (raw image), the captured image used in the analysis process in the candidate selection process may be the read image used in the recommended setting determination process, or may be an image obtained by performing image processing on the captured image (read image) used in the recommended setting determination process.

[0065] The storage unit 34 stores recommended settings (recommended values) for the multiple setting items determined by the analysis unit 33. For example, the storage unit 34 stores the recommended settings for the multiple setting items determined using a scanned document as a profile suitable for the scanned document. Thereby, when scanning the scanned document and documents of the same type as the scanned document, it becomes possible to perform scanning using the stored profile (to perform image processing settings suitable for the document) thereafter.

[0066] The presentation unit 35 presents (suggests) the user with recommended settings (setting items and recommended values ​​determined for the setting items) for the multiple setting items determined by the analysis unit 33. Any method may be used to present the recommended settings. For example, a method of presenting the recommended settings by displaying the recommended settings in a list format on a setting screen or the like via the output device 16, a method of providing information on the recommended settings to the user via the communication unit 17, a method of displaying a screen that prompts (suggests) the user to register (save) the recommended settings as a profile (set of settings) to be used in the future, and the like may be used. When presenting the recommended settings to the user, the presentation unit 35 may present (display) to the user an image in which the recommended settings are reflected, or a character recognition result (OCR result) for the image in which the recommended settings are reflected. Below, various screens that are user interfaces (UIs) for presenting the recommended settings to the user by the presentation unit 35 will be exemplified. Below, a screen in which the user is asked to input an OCR area to be recognized and a correct answer character string in advance, and an analysis process is performed using the OCR area and the correct answer character string will be exemplified.

[0067] FIG. 12 is a diagram showing an example of a pre-setting screen according to the present embodiment. In the pre-setting screen displayed when starting the recommended setting determination process, the user performs pre-setting for performing the recommended setting determination process. In the example screen shown in FIG. 12, the user can select (set) the language (Japanese, English, Chinese, etc.) for which character recognition (OCR) is desired, the reading resolution (240 dpi, 300 dpi, 400 dpi, 600 dpi, etc.), and whether or not to output an image excluding ruled lines in a document (whether or not to erase ruled lines), as pre-settings. Note that the screen shown in FIG. 12 is displayed, for example, when the presenting unit 35 receives an instruction from the user to create a profile suitable for character recognition (to determine the recommended setting). After the pre-setting is performed on the screen shown in FIG. 12, when the user presses a button ("scan" button) for acquiring a scanned image (captured image) by the scanner 8, the scanner 8 reads the scanned document to acquire the captured image. On the other hand, when the user presses the "cancel" button, the recommended setting determination process is terminated, and the pre-setting screen is closed (hidden).

[0068] FIG. 13 is a diagram showing an example of a recommended setting determination screen according to the present embodiment. The screen shown in FIG. 13 is a screen that is displayed when a captured image is acquired as a result of the user pressing the "Scan" button on the screen of FIG. 12. As shown in FIG. 13, the recommended setting determination screen displays the captured image that has been acquired. Note that FIG. 13 illustrates an example in which a scanned image (captured image) of one document is acquired, but when multiple documents are read, multiple scanned images (captured images) of multiple documents are displayed on the recommended setting determination screen. With the captured image acquired and displayed as shown in FIG. 13, the user specifies an area in which character recognition is desired (see the four bold frames in the figure). Furthermore, on the screen of FIG. 13, the user inputs a correct answer character string for the specified area (a correct character string written in the area) (see fields [1] to [4] in the figure). After completing the specification of the area to be character recognized and the input of the correct answer character string for the area, the user presses a button ("Create profile" button) for performing the recommended setting determination process, and the recommended setting determination process is started. Also, as shown in FIG. 13, the recommended settings determination screen displays a button ("Register Profile" button) for registering (storing) the recommended settings (profile), and the recommended settings (profile) are registered when the user presses this button. If the user presses the "Cancel" button on the screen shown in FIG. 13, the recommended settings are not registered. In this case, the recommended settings determination process (profile creation process) may be performed again by the user changing (e.g., adding) the OCR area and then pressing the "Create Profile" button again. Also, if the user presses the "Back" button on the screen shown in FIG. 13, it is possible to display the pre-setting screen shown in FIG. 12 and perform the process of acquiring a captured image again.

[0069] FIG. 14 is a diagram showing an example of a progress display screen according to the present embodiment. The screen shown in FIG. 14 is a screen that is displayed when the user presses the "Create profile" button on the screen of FIG. 13. As shown in FIG. 14, when the analysis process is being performed by the analysis unit 33, the progress display screen displays progress information that is information indicating that the analysis process is being performed and / or information indicating the progress of the analysis process. In the example screen of FIG. 14, the characters "Progress: 36%" and a progress bar are displayed as information indicating the progress (for example, 36% of 100% has been completed). In addition, in the example screen of FIG. 14, the characters "Predicted remaining time: 2 minutes" are displayed as a time that is predicted to be the remaining time until the analysis process is completed. Note that the screen shown in FIG. 14 may be closed (hidden) when the profile creation (determination of the recommended value) is completed. In addition, when the profile creation (determination of the recommended value) is completed, the screen shown in FIG. 13 or FIG. 14 may display a character recognition result for an image to which the recommended settings are reflected (such as a captured image to which image processing with the recommended settings is applied).

[0070] FIG. 15 is a diagram showing an example of a recommended setting save screen according to the present embodiment. The screen shown in FIG. 15 is a screen that is displayed when the user presses the "Register profile" button on the screen shown in FIG. 13. As shown in FIG. 15, the recommended setting save screen displays, for example, a display (button, etc.) for newly saving the recommended setting (profile) and a display (button, etc.) for overwriting and saving the recommended setting (profile). The user can select whether to newly save or overwrite the recommended setting (profile) determined by the analysis process on the screen shown in FIG. 15 and then save it. When the user presses the "OK" button on the screen shown in FIG. 15, a profile (driver profile) consisting of the recommended setting determined by the analysis process is registered (stored) in the storage device 14. This allows the user to perform a scan process (image processing) using the registered profile (set of settings), so that an image suitable for character recognition can be obtained. Note that this registered profile can be used not only for the scanned document read by the operation on the screen of FIG. 12, but also for scanning documents of the same type as the scanned document (such as the same fixed-type form). It should be noted that if the user presses the "Cancel" button on the screen shown in FIG. 15, the process of registering the recommended settings is terminated, and the recommended settings save screen is closed (displayed).

[0071] In this embodiment, the presentation unit 35 generates and displays the recommended settings generation screen, but this is not limited to this example, and a display control unit (not shown) separate from the presentation unit 35 that presents the recommended settings may generate and display the recommended settings generation screen.

[0072] <Processing flow> Next, a flow of processing executed by the information processing system according to the present embodiment will be described. Note that the specific contents and processing order of the processing described below are an example for implementing the present disclosure. The specific contents and processing order may be appropriately selected according to the embodiment of the present disclosure.

[0073] Fig. 16 is a flowchart showing an outline of the flow of the recommended setting determination process according to this embodiment. The process shown in this flowchart is started when the information processing device 1 receives an instruction to determine the recommended settings from the user. When the instruction from the user is received, the presentation unit 35 displays a pre-setting screen (see Fig. 12) on the output device 16 (display means). Note that the process shown in this flowchart illustrates an example in which the candidate values ​​for the background pattern removal are determined by the above-mentioned method 2 (single verification).

[0074] In step S101, an image is acquired. For example, when a user presses the "scan" button on the screen shown in Fig. 12, the image acquisition unit 21 performs a reading process on the read document, thereby acquiring a captured image of the read document. Also, in step S101, it is assumed that an area in which the user wishes to perform character recognition (OCR area), a correct answer character string for the OCR area, and an OCR language are input by the user and acquired by the reception unit 32. After that, the process proceeds to step S102.

[0075] In step S102, a ruled line color analysis is performed. The analysis unit 33 performs an analysis to determine the presence or absence of ruled lines in the captured image acquired in step S101, and if it is determined that ruled lines exist, the analysis unit 33 estimates the color of the ruled lines included in the read document (captured image) by performing a ruled line color analysis. As a result, the analysis unit 33 determines candidate values ​​(candidate parameter values) for the dropout color item. After that, the process proceeds to step S103.

[0076] In step S103, the expected CC number for each OCR region is calculated. The analysis unit 33 calculates an appropriate CC number (expected CC number) for each OCR region, for example, from the number of characters in the correct string and the OCR language. After that, the process proceeds to step S104.

[0077] In step S104, it is determined whether the processing for all patterns of background pattern removal and character extraction has been completed (executed). For example, the analysis unit 33 determines whether the processing (image processing (step S105), OCR recognition rate calculation (step S106), and CC number calculation (step S107)) for each of the seven patterns of all patterns of background pattern removal (none, four patterns of Lv1 to 3) and all patterns of character extraction (three patterns of outline character extraction ON, shaded character extraction ON, and seal imprint overlap character extraction ON) has been completed. If the processing for all patterns has been completed (YES in step S104), the process proceeds to step S108. On the other hand, if the processing for all patterns has not been completed (NO in step S104), the process proceeds to step S105.

[0078] In step S105, image processing related to background pattern removal or character extraction is performed. The analysis unit 33 performs image processing for the pattern determined in step S104 as not having been completed on the captured image acquired in step S101. For example, if the processing for "background pattern removal Level 4" is not completed, image processing for background pattern removal Level 4 (background pattern removal processing) is performed. Also, for example, if the processing for "seal imprint overlapping character extraction ON" is not completed, image processing for seal imprint overlapping character extraction ON (seal imprint overlapping character extraction processing) is performed. Note that image processing does not need to be performed for "no background pattern removal". After that, the process proceeds to step S106.

[0079] In step S106, the OCR recognition rate is calculated. The analysis unit 33 acquires the character recognition result for the captured image (each OCR area) after the image processing is performed in step S105. Note that the image corresponding to "without background pattern removal" is the captured image acquired in step S101, and in the case of "without background pattern removal", the character recognition result for the captured image (each OCR area) acquired in step S101 is acquired. Then, the analysis unit 33 calculates the OCR recognition rate (for example, field recognition rate or character recognition rate) based on the character recognition result (recognized character string) for each OCR area. Note that various methods may be used to calculate the OCR recognition rate. Thereafter, the process proceeds to step S107.

[0080] In step S107, the number of CCs is calculated. The analysis unit 33 calculates the number of CCs for the captured image (each OCR area) after the image processing is performed in step S105. Note that the image corresponding to "without background pattern removal" is the captured image acquired in step S101, and in the case of "without background pattern removal", the number of CCs for the captured image (each OCR area) acquired in step S101 is calculated. Note that in step S107, the number of CCs for each pattern of background pattern removal (the number of CCs for the image corresponding to each setting of background pattern removal) is calculated, but the number of CCs for each pattern of character extraction (the image corresponding to each setting of character extraction) is not calculated. In other words, if the image processing performed in step S105 is image processing related to character extraction, the calculation process of CCs is not performed in step S107. After that, the process returns to step S104.

[0081] In steps S108 to S110, candidate values ​​of some parameters are determined (narrowing down of parameter value candidates). In step S108, a candidate value (parameter value candidate) for the background pattern removal item is determined based on the OCR recognition rate and the number of CCs. In this embodiment, the analysis unit 33 compares the OCR recognition rate calculated in step S106 and the number of CCs calculated in step S107 between all patterns (setting values) of the background pattern removal item to select a predetermined number (for example, two) of setting values ​​(patterns) having the best OCR recognition rate and number of CCs in that order as candidate values. Note that the candidate values ​​may be selected by comparing only the OCR recognition rate between all patterns. Note that when comparing the OCR recognition rate and the number of CCs between patterns, the OCR recognition rate and the number of CCs in all OCR areas are taken into consideration. For example, the representative value (average value, etc.) or the total value of the number of CCs calculated in each OCR area is compared between patterns. Thereafter, the process proceeds to step S109.

[0082] In step S109, a candidate value (candidate parameter value) for the character extraction item is determined based on the OCR recognition rate. In this embodiment, the analysis unit 33 compares the OCR recognition rates when character extraction is ON and OFF, and determines whether the recognition rate is improved when character extraction is ON, thereby determining the candidate value (ON or OFF) for character extraction. For example, the OCR recognition rate calculated in step S106 when "white-outline character extraction ON" is compared with the OCR recognition rate calculated in step S106 when "white-outline character extraction OFF" is compared, and if the recognition rate is higher (improved) in "white-outline character extraction ON", the candidate value (setting value) for "white-outline character extraction" is determined to be "ON". Note that the "OCR recognition rate calculated in step S106 when "character extraction OFF"" is the OCR recognition rate calculated for the image acquired in step S101, and the OCR recognition rate calculated in step S106 may be used in the case of the pattern "without background pattern removal" (when all character extractions are OFF). Furthermore, when comparing the OCR recognition rates, the OCR recognition rates in all OCR regions are taken into consideration. After that, the process proceeds to step S110.

[0083] In step S110, candidate values ​​(candidate parameter values) for each of the binarization sensitivity item and the noise elimination item are determined based on the number of CCs and the expected number of CCs. The analysis unit 33 determines candidate values ​​for each of the binarization sensitivity item and the noise elimination item corresponding to the candidate values ​​for the background pattern elimination item determined in step S108. For example, assume that the candidate values ​​for the background pattern elimination item are determined to be "Lv1" and "Lv2" in step S108. In this case, the analysis unit 33 determines candidate values ​​for the binarization sensitivity item and the noise elimination item corresponding to "Lv1" (for example, the candidate value for the binarization sensitivity item is "-10 to 10" and the candidate value for the noise elimination item is "0 to 10") by comparing the number of CCs calculated in step S107 and the expected value of the number of CCs calculated in step S103 when the image processing (background pattern removal processing) according to "Lv1" is performed in step S105. Similarly, when image processing (background pattern removal processing) according to "Lv2" is performed in step S105, the analysis unit 33 compares the CC number calculated in step S107 with the CC number expected value calculated in step S103 to determine candidate values ​​(for example, the candidate value of the binarization sensitivity item is "-30 to -10" and the candidate value of the noise removal item is "0 to 20") for the binarization sensitivity item and the noise removal item corresponding to "Lv2". In this way, the analysis unit 33 compares the calculated CC number with the CC number expected value for each of the candidate values ​​for the background removal item to determine candidate values ​​for each of the binarization sensitivity items and the noise removal items corresponding to each candidate value.

[0084] When determining candidate values ​​for the binarization sensitivity item and noise removal item corresponding to "without background pattern removal", the number of CCs calculated in step S107 in the case of the pattern "without background pattern removal" (when all character extraction is OFF) is compared with the expected number of CCs. When comparing the number of CCs with the expected number of CCs, the number of CCs and the expected number of CCs in all OCR areas are taken into consideration. For example, the total number of CCs calculated in each OCR area is compared with the total expected number of CCs calculated in each OCR area. Then, the process proceeds to step S111.

[0085] In step S111, combinations (combination table) are generated. The analysis unit 33 generates combinations (combination table) of setting values ​​(candidate values) of multiple parameters by simple multiplication of the candidate values ​​of multiple parameters (all parameters) using the candidate values ​​determined in steps S102 and S108 to S110. After that, the process proceeds to step S112.

[0086] In step S112, the recommended settings are determined. For each combination generated in step S111, the analysis unit 33 performs image processing on the captured image acquired in step S101 using that combination, thereby determining the recommended settings for a plurality of setting items. After that, the process shown in this flowchart ends.

[0087] Note that image processing of all patterns of background pattern removal and image processing of all patterns of character extraction may be performed at different times. For example, image processing of all patterns of background pattern removal may be performed, candidate values ​​for background pattern removal may be determined, and then image processing of all patterns of character extraction may be performed to determine candidate values ​​for character extraction. Steps S106 and S107 may be performed in any order, and steps S108 and S109 may be performed in any order.

[0088] Furthermore, in this embodiment, if the user is not satisfied (dissatisfied) with the image processing setting proposal (presentation of recommended settings) by the presenting unit 35, the OCR area, etc. may be changed and the above-mentioned analysis process may be performed again to determine image processing settings (recommended settings) suitable for OCR again and present them to the user. This process may be repeated until a result (character recognition result) that satisfies the user is obtained. This makes it possible to perform image processing settings with higher accuracy. Note that, if a newly set OCR area is included in the changed OCR area by changing the OCR area, input of the correct character string for this newly set OCR area is accepted from the user in advance when the above-mentioned analysis process is performed.

[0089] As described above, according to this embodiment, by performing an analysis process using a captured image, candidates (setting values) of recommended settings are selected, and image processing is repeatedly performed on the captured image while changing the respective setting values ​​of multiple setting items, with the setting values ​​being limited to those selected as candidates of recommended settings (image processing settings for making an acquired image suitable for character recognition). This allows the recommended settings for multiple setting items to be determined in advance (a profile to be generated in advance) that are more accurate (higher recognition accuracy) according to the original. Furthermore, according to this embodiment, even a non-expert user (a user who does not understand image processing parameters) can perform settings (scan settings) that are suitable for the original and are optimal for character recognition (OCR) simply by scanning the original. Furthermore, in this embodiment, since the setting values ​​(parameter values) optimal for character recognition are determined using the actual character recognition results, it is possible to reliably obtain setting values ​​(image processing parameter values) that are suitable for character recognition. In addition, since the possible setting values ​​are narrowed down (candidate values ​​are selected) before combinations are generated and image processing is attempted, the processing time required to determine the recommended settings is realistic, and it is possible to obtain desirable results within this time.

[0090] In this embodiment, a method for determining recommended settings suitable for one document (one type of document) by reading the document has been exemplified. As a result, even when scanning a large amount of documents of the same type as the document (for example, standard forms), image processing can be performed using the determined recommended settings for each scan. However, in a site where large-volume scanning is performed, it is assumed that not only one type of document but also multiple types of documents (forms) are scanned at one time (mixed cases). Even in such cases, it is desirable to perform image processing using recommended settings suitable for the document for each scan. Two methods for dealing with such cases will be described below.

[0091] The first method is a method of combining with an automatic profile selection function (existing function) using ruled line information. In this method, first, a plurality of captured images (captured images corresponding to each document) obtained by capturing images of a plurality of types of documents (a plurality of sheets of documents) are captured by the image capture unit 21. Then, by the above-mentioned method, a recommended setting (optimum profile) is determined for each document (each type of document), and the recommended setting (profile) is registered for each document (type of document). At this time, the storage unit 34 may store the recommended setting and document identification information in association with each document. Then, in the scan settings, the automatic profile selection function (a function of performing document identification and selecting (using) a profile (setting information) registered for the identified document) is enabled, and information for identifying the document (for example, ruled line information) is registered. During operation, the captured document is identified based on the captured image and the registered document identification information, and a profile registered for the identified document is selected based on the document identification information, and scanning (image processing) according to the profile is performed. This means that even in cases of mixed loads, scanning (image processing) can be performed using recommended settings appropriate for each document (document type), making it possible to obtain images suitable for character recognition.

[0092] The second method is to determine (suggest) one recommended setting (profile) applicable to any document. In this method, first, a plurality of captured images (captured images corresponding to each document) obtained by capturing images of a plurality of types of documents (a plurality of sheets of documents) are acquired by the image acquisition unit 21. Then, by the above-mentioned method, for each document (captured image), a candidate selection process (narrowing down of setting values), creation of a combination of setting values ​​(combination table) based on the selected candidate values, and calculation of an evaluation value for each combination (evaluation value for the character recognition result corresponding to each combination) are performed. Then, for all documents, the combination with the highest evaluation value is determined as a recommended setting (profile) applicable to these multiple types of documents. As a result, even in the case of mixed loading, scanning (image processing) can be performed using the recommended setting applicable to these multiple types of documents, and an image suitable for character recognition can be obtained.

[0093] In addition, in the present embodiment, a method for determining recommended settings suitable for one document by reading the document has been exemplified, but a method for determining recommended settings suitable for a document having a predetermined format by using multiple documents (multiple documents of the same type) having the predetermined format may be used. In this case, multiple captured images of multiple documents are acquired by the image acquisition unit 31, but the captured images used to select candidate values ​​in the candidate selection process may be different from the captured images that are the targets of image processing trials in the recommended setting determination process. For example, the candidate selection process may be performed using the captured image of the first document, and the recommended setting determination process may be performed using the captured image of the second document.

[0094] [Second embodiment] In the first embodiment, the analysis process is performed in the information processing device 1 having a driver for the scanner 8 (read image processing unit 42), but the configuration of the system 9 is not limited to this configuration, and the analysis process may be performed in an information processing device that is communicably connected to the information processing device 1 and does not have a driver for the scanner 8. In the present embodiment, an example will be given of the case where the analysis process is performed in an information processing device (for example, a server) that does not have a driver for the scanner 8.

[0095] <System configuration> FIG. 17 is a schematic diagram showing the configuration of a system 9 according to this embodiment. The system 9 according to this embodiment includes a scanner 8, an information processing device 1, and a server 2, which are communicatively connected to each other via a network or other communication means. In FIG. 17, the information processing device 1 is connected to the scanner 8 via a router (or gateway) 7. Note that, although FIG. 17 illustrates an example in which one scanner 8 and one information processing device 1 are connected to the server 2, a plurality of scanners 8 and a plurality of information processing devices 1 may be connected to the server 2. Note that the configurations of the scanner 8 and the information processing device 1 are roughly similar to those described in the above-described embodiment, and therefore description thereof will be omitted.

[0096] The server 2 acquires the captured image acquired by the information processing device 1 and performs an analysis process using the captured image to determine the above-mentioned recommended settings. The server 2 is a computer equipped with a CPU 21, a ROM 22, a RAM 23, a storage device 24, an input device 25, an output device 26, a communication unit 27, and the like. However, the specific hardware configuration of the server 2 can be omitted, replaced, or added as appropriate depending on the embodiment. In addition, the server 2 is not limited to a device consisting of a single housing. The server 2 may be realized by a plurality of devices using so-called cloud or distributed computing technology, etc.

[0097] FIG. 18 is a diagram showing an outline of the functional configuration of the server 2 according to the present embodiment. The server 2 functions as a device including an image acquisition unit 31, a reception unit 32, an analysis unit 33, a storage unit 34, and a presentation unit 35, by a program recorded in the storage device 24 being read into the RAM 23 and executed by the CPU 21, and each hardware included in the server 2 is controlled. The analysis unit 33 includes a candidate selection unit 45 and a recommended setting determination unit 46. In this embodiment and other embodiments described later, each function included in the server 2 is executed by the CPU 21, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors. The functional configuration (each functional unit) of the server 2 is generally similar to the functional configuration (each functional unit) of the information processing device 1 in the first embodiment, and therefore description thereof will be omitted. However, in this embodiment, the image acquisition unit 31 acquires a captured image from the information processing device 1 via a network. However, in this embodiment, the image acquisition unit 31 may acquire the captured image by reading out the captured image stored in the storage device 24. Moreover, in this embodiment, the reception unit 32 acquires the OCR area specified by the user in the information processing device 1 and the input correct answer character string from the information processing device 1. Moreover, in this embodiment, the presentation unit 35 may present the recommended settings and the captured image reflecting the recommended settings to the user by transmitting them to the information processing device 1.

[0098] [Third embodiment] In the first embodiment, the analysis process is performed in the information processing device 1 having a driver for the scanner 8, but the configuration of the system 9 is not limited to this configuration, and the analysis process may be performed in the scanner 8. In the present embodiment, the case where the analysis process is performed in the scanner 8 will be illustrated as an example.

[0099] <System configuration> FIG. 19 is a schematic diagram showing the configuration of a system 9 according to this embodiment. The system according to this embodiment includes a scanner 8b. The configuration of the scanner 8b is roughly the same as that of the first embodiment described above, and therefore will not be described. However, the scanner 8b is a computer (information processing device) including a CPU 81, a ROM 82, a RAM 83, a storage device 84, an input device 85, an output device 86, a communication unit 87, a reading unit (a unit that reads an original (an image of an original) by an imaging element) 88 (image reading means), and the like. However, the specific hardware configuration of the scanner 8 can be omitted, replaced, or added as appropriate depending on the embodiment.

[0100] 20 is a diagram showing an outline of the functional configuration of the scanner according to this embodiment. The scanner 8b functions as a device including an image acquisition unit 31, a reception unit 32, an analysis unit 33, a storage unit 34, and a presentation unit 35, by a program recorded in a storage unit 84 being read into a RAM 83 and executed by a CPU 81, which controls each piece of hardware included in the scanner 8b. The analysis unit 33 includes a candidate selection unit 45 and a recommended setting determination unit 46. In this embodiment, each function included in the scanner 8b is executed by the CPU 81, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.

[0101] The functional configuration (each functional unit) of the scanner 8b is generally similar to the functional configuration (each functional unit) of the information processing device 1 in the first embodiment, and therefore a description thereof will be omitted. However, in this embodiment, the image acquisition unit 31 includes an image reading unit (image reading means) 47 and a read image processing unit (image reading means) 42. The image reading unit 47 reads an original (image of the original) using an imaging element, and the read image processing unit 42 performs image processing on the read image generated by reading the original by the image reading unit 47. In this way, the image acquisition unit 31 acquires a captured image. Also, in this embodiment, the presentation unit 35 may present the recommended settings or a captured image reflecting the recommended settings to the user by displaying them on, for example, a touch panel provided in the scanner 8b.

[0102] [Fourth embodiment] In the fourth embodiment, an embodiment will be described in which the information processing system, information processing device, method, and program according to the present disclosure are implemented in a system for evaluating whether the image processing to be evaluated is suitable for character recognition (is the image processing suitable for acquiring an image suitable for character recognition) or not. However, the information processing system, information processing device, method, and program according to the present disclosure can be widely used for techniques for evaluating character recognition results (character recognition accuracy), and the application of the present disclosure is not limited to the examples shown in the embodiments.

[0103] Conventionally, an OCR engine performs character recognition processing on an image obtained by reading a document with an image reading device, but the character recognition rate of the OCR engine is not 100% because the OCR engine may misread the text. Therefore, a user checks whether the OCR result is correct by comparing the OCR result (recognized character string) with the correct text (correct character string). However, even if there is a difference between the recognized character string and the correct character string, if the different characters are similar characters, the user may erroneously determine that they are the same character. If the user makes an erroneous determination in this way, the OCR result cannot be evaluated correctly.

[0104] In view of such circumstances, the information processing system, information processing device, method, and program according to the present embodiment control the display of a screen (a screen showing a collation result between a correct answer character string and a recognized character string) for confirming the character recognition result for an image subjected to the image processing to be evaluated so that the user can evaluate whether the image processing to be evaluated is suitable for character recognition or not, so that the accuracy of the evaluation of the character recognition result by the user can be improved. This makes it possible to assist the user in determining (evaluating) the OCR result (OCR accuracy). Note that the configuration of the system 9 according to the present embodiment is roughly the same as the configuration of the system 9 according to the first embodiment described above with reference to FIG. 1, and therefore a description thereof will be omitted.

[0105] 21 is a diagram showing an outline of the functional configuration of the information processing device 1 according to the present embodiment. The information processing device 1 functions as a device including an image acquisition unit 61, a reception unit 62, a recognition result acquisition unit 63, a matching unit 64, and a display control unit 65, by a program recorded in the storage device 14 being read into the RAM 13 and executed by the CPU 11, and each piece of hardware included in the information processing device 1 is controlled. The image acquisition unit 61 includes a read image acquisition unit 71 and an image processing unit 72. The reception unit 62 includes a character region acquisition unit 73 and a correct answer information acquisition unit 74. Note that in this embodiment and other embodiments described later, each function included in the information processing device 1 is executed by the CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.

[0106] The image acquisition unit 61 acquires a captured image of an original. The image acquisition unit 61 is generally similar to the image acquisition unit 31 in the first embodiment, and therefore a description thereof will be omitted. However, in this embodiment, the read image acquired by the read image acquisition unit 71 is subjected to image processing (image processing of an evaluation target, which is an evaluation target for whether the image processing is suitable for character recognition) by the image processing unit 72 (corresponding to the "image processing means" in this embodiment), and the image that has been subjected to image processing (processed image) is acquired as a captured image.

[0107] The reception unit 62 receives the specification of the OCR area for the read document and the input of the correct character string by the user selecting a field (a character area (OCR area) that is an area containing characters) in which character recognition is desired in the read document (captured image) and inputting the correct character string written in that area. Note that the reception unit 62 is generally similar to the reception unit 32 in the first embodiment, and therefore a description thereof will be omitted.

[0108] The recognition result acquisition unit 63 acquires a character recognition result for the captured image (processed image). Specifically, the recognition result acquisition unit 63 acquires a character recognition result (recognized character string) for a character area (OCR area) in the captured image (processed image). The recognition result acquisition unit 63 may acquire the character recognition result by performing character recognition processing (OCR processing), or may acquire the character recognition result from another device that performs character recognition processing (device equipped with an OCR engine).

[0109] The collation unit 64 collates the correct character string with the recognized character string. The collation unit 64 collates (compares) the correct character string with the recognized character string for the same OCR region, and determines whether the correct character string and the recognized character string completely match, and if the correct character string and the recognized character string do not completely match, identifies characters that do not match (characters that differ) between the two character strings.

[0110] The display control unit 65 displays on the display means (corresponding to the output device 16 in FIG. 1) one or more screens showing the matching results between the correct answer character string and the recognized character string (the evaluation results for the character recognition results) so that the user can evaluate whether the image processing to be evaluated is suitable for character recognition. In this embodiment, a screen (first screen) showing the matching results for all OCR areas specified by the user in the captured image and a screen (second screen (pop-up screen)) showing the matching results for each OCR area are displayed. In this embodiment, when the screens showing the matching results are displayed, the display of at least one screen is controlled to be different depending on the matching results between the correct answer character string and the recognized character string. In this embodiment, as a method for controlling the display of the screen to be different depending on the matching result, a method (method 1) of changing the display mode of a predetermined screen component related to the screen (first screen and / or second screen) depending on the matching result and a method (method 2) of changing the display content of a predetermined screen component related to the screen (first screen and / or second screen) depending on the matching result will be described. In this embodiment, the second screen is displayed by hovering the mouse over the OCR area on the first screen, but the second screen may be displayed by processing the OCR area other than hovering the mouse. For example, the second screen may be displayed by processing (clicking) to select the OCR area on the first screen.

[0111] (Method 1: Display of OCR area frame) In this embodiment, as described later, a captured image (processed image) is displayed on the first screen, and a frame (border line) (hereinafter referred to as "OCR area frame") indicating an OCR area (character area) designated by a user is superimposed on the captured image. The display control unit 65 controls the display mode of the OCR area frame superimposed on the captured image to be different depending on the collation result. Specifically, at least one of the line color of the OCR area frame, the line thickness of the OCR area frame, the line type of the OCR area frame (dotted line, solid line, etc.), and the background color (overlay) within the OCR area frame is controlled to be different depending on the collation result.

[0112] (Method 1: Pop-up screen frame display mode) In this embodiment, as described later, when the user hovers the mouse over an OCR area (OCR area frame) on the first screen, a screen (second screen) showing the collation result for the OCR area is popped up (display of a pop-up screen). The display control unit 65 controls the display mode of the screen frame (frame of the pop-up screen) surrounding this second screen to be different depending on the collation result for the OCR area. Specifically, at least one of the line color of the screen frame, the line thickness of the screen frame, the line type of the screen frame (dotted line, solid line, etc.), and the background color (overlay) within the screen frame is controlled to be different depending on the collation result. Note that, in this embodiment, the display mode of the frame of the second screen is changed, but the display mode of the frame of the first screen may be changed depending on the collation result for all OCR areas specified by the user.

[0113] (Method 1: Display of characters that do not match between strings) In this embodiment, the second screen (pop-up screen) displays (arranges) an icon, text indicating a matching result, a recognized character string (OCR text), and a correct character string (correct text) related to the OCR area of ​​the second screen. The display control unit 65 controls the display mode of a character in the recognized character string that is determined not to match (different from) a character in the correct character string (hereinafter referred to as a "mismatched character") to be different depending on the matching result for the OCR area. Specifically, the display control unit 65 controls at least one of the decoration of the mismatched character (color, size, thickness, italics, underline, etc.), the background color of the mismatched character, and the font of the mismatched character to be different depending on the matching result for the OCR area. In this embodiment, the display mode of the mismatched character displayed on the second screen is different, but in the case of an embodiment in which the recognized character string is displayed on the first screen, the display mode of the mismatched character in the recognized character string displayed on the first screen may be controlled to be different depending on the matching result for the OCR area.

[0114] (Method 2: Icon type) As described above, in this embodiment, an icon for indicating the matching result is displayed on the second screen. The display control unit 65 controls the type of icon (circle, triangle, square, etc.) to be different depending on the matching result for the OCR area. For example, when the correct answer character string does not match the recognized character string in the OCR area, an icon (a mark other than a circle, etc.) that can call the user's attention more than when the correct answer character string matches the recognized character string is used. Note that, in this embodiment, the display mode of the icon displayed on the second screen is made different, but in the case of an embodiment in which the icon is displayed on the first screen, the display mode of the icon displayed on the first screen may be controlled to be different depending on the matching result for the OCR area.

[0115] (Method 2: Text content showing match results) As described above, in this embodiment, text showing the matching result (text for notifying the user of the matching result) is displayed on the second screen. The display control unit 65 controls the content of this text (content of the sentence) to be different depending on the matching result for the OCR area. For example, when the correct answer character string and the recognized character string do not match in the OCR area, text showing the matching result "Correct text cannot be obtained" is displayed, and when the correct answer character string and the recognized character string match, text showing the matching result "Correct text has been obtained" is displayed. Note that, in this embodiment, the display mode of the text displayed on the second screen is made different, but in the case of an embodiment in which the text is displayed on the first screen, the display mode of the text displayed on the first screen may be controlled to be different depending on the matching result for the OCR area.

[0116] As described above, by making the display of the screen showing the matching result different (changing) depending on the matching result between the correct character string and the recognized character string, it is possible to alert the user to the OCR area where the correct character string does not match the recognized character string among the multiple OCR areas. Note that, although the above describes a case where the display mode and display contents of multiple screen components are changed depending on the matching result, it is sufficient that the display mode or display contents of at least any of the above-mentioned multiple screen components are changed depending on the matching result. Below, various screens (user interfaces (UIs)) displayed on the display means by the display control unit 65 are illustrated.

[0117] Fig. 22 is a diagram showing an example of a document scan screen according to the present embodiment. As shown in Fig. 22, a button for scanning a document (a "scan" button) is displayed (arranged) on the scan screen, and when a user presses the "scan" button, the document is scanned and a captured image (document image) is generated. As a result, the image acquisition unit 61 acquires the captured image.

[0118] FIG. 23 is a diagram showing an example of a pre-setting screen (pre-setting state) according to this embodiment. The screen shown in Fig. 23 is a screen (initial screen) for setting an OCR area and inputting a correct answer character string in advance. As shown in Fig. 23, a button ("Add" button) for setting (adding) a captured image and an OCR area is displayed (arranged) on the advance setting screen (pre-setting state), and the user can set the OCR area and input a correct answer character string by pressing the "Add" button.

[0119] FIG. 24 is a diagram showing an example of a pre-setting screen (after setting) according to the present embodiment. The screen shown in FIG. 24 is a pre-setting screen in a state where an OCR area on a captured image has been set and a correct answer character string has been input. FIG. 24 illustrates an example where five locations (circled characters 1 to 5) in the figure are specified as OCR areas. As shown in FIG. 24, the pre-setting screen (after setting) displays (arranges) a captured image, an OCR area specification frame, an input form for a correct answer character string for each OCR area (input frame and correct answer character string), and a button for performing character recognition and evaluating the character recognition result ("Start evaluation" button). As shown in FIG. 24, the user can set (specify) an OCR area on a captured image by pressing the "Add" button in FIG. 23 and surrounding (specifying) an area to be OCRed with a rectangular frame. Also, as shown on the right side of the screen in FIG. 24, the user can input a character string (correct answer character string) that can be read from each OCR area into the input form for the correct answer character string. In the example of Fig. 24, the correct answer character strings "01234567", "001234", "4-4-5 Minatomirai Nishi-ku Yokohama-shi Kanagawa-ken", "TO123456789012", and "172,769" are input for each of the five OCR areas (circled characters 1 to 5) in the figure. When the user presses the "Start evaluation" button on this screen, the screen transitions to a screen displaying the evaluation results, and evaluation of the character recognition results begins.

[0120] FIG. 25 is a diagram showing an example of an evaluation result display screen according to the present embodiment. On the evaluation result display screen (the above-described first screen), display is made according to the collation result between the correct character string and the recognized character string for each OCR region. As shown in FIG. 25, on the evaluation result display screen, an imaging image, an OCR region frame for each OCR region superimposed on the imaging image, and a character recognition result are displayed (arranged). In the example of FIG. 25, for each of the five OCR regions (circled characters 1 to 5) in the figure, texts of "01234567", "001234", "Minato Mirai 4-4-5, Yurigaoka Ward, Yokohama City, Kanagawa Prefecture", "TO123456789012", and "172,769" are extracted (recognized character strings are acquired). As a result of collating the correct character string and the recognized character string in each OCR region, it is determined that the correct character string and the recognized character string do not match in the OCR region indicated by the circled character 3. Specifically, the character "West" described in the OCR region indicated by the circled character 3 is misread and read as "Rooster". As a result, the display control unit 65 causes the display mode of the OCR region frame of the OCR region indicated by the circled character 3 to be displayed in a display mode corresponding to the fact that the correct character string and the recognized character string do not match. On the other hand, for each of the OCR regions indicated by the circled characters 1, 2, 4, and 5, since the same text (OCR text) as the correct character string can be obtained from the image, the display control unit 65 causes the display mode of the OCR region frame of each of the OCR regions indicated by the circled characters 1, 2, 4, and 5 to be displayed in a display mode corresponding to the fact that the correct character string and the recognized character string match.

[0121] For example, the OCR region frame of the OCR region indicated by the circled character 3 is displayed in red, thick line, and with a background color (overlay), and the OCR region frames of each of the OCR regions indicated by the circled characters 1, 2, 4, and 5 are displayed in green, thin line, and without a background color. In this way, the display control unit 65 may compare the display mode of the OCR region frame when the correct character string and the recognized character string do not match with the display mode when they match, and make it a mode that can call the user's attention more.

[0122] FIG. 26 is a diagram showing an example of a display screen of the evaluation result according to this embodiment (when the correct text is acquired). In addition to the first screen shown in FIG. 25, FIG. 26 shows a screen (a pop-up screen) that is displayed when the mouse is placed over the OCR area (character area) shown by the circled character 5 on the screen of FIG. 25, and shows a screen (the above-mentioned second screen) showing the collation result for the OCR area shown by the circled character 5. As described above, in the OCR area shown by the circled character 5, the same text (OCR text) as the correct character string can be acquired from the image. In this case, the display of the second screen is a display corresponding to the match between the correct character string and the recognized character string.

[0123] For example, the display mode of the screen frame of the second screen, the text indicating the matching result displayed on the second screen, and the type of icon displayed on the second screen are displayed (display mode, display contents) according to the match between the correct character string and the recognized character string. For example, the screen frame of the second screen is displayed in green with thin lines and a white background. In addition, the text indicating the matching result, "The correct text has been obtained," is displayed. In addition, a green round icon is displayed.

[0124] FIG. 27 is a diagram showing an example of a display screen of the evaluation result according to the present embodiment (when the correct text is not obtained). In addition to the first screen shown in FIG. 25, FIG. 27 shows a screen (a pop-up screen (pop-up screen)) that is displayed when the mouse is placed over the OCR area (character area) shown by the circled character 3 on the screen of FIG. 25, and shows a screen (the above-mentioned second screen) showing the collation result for the OCR area shown by the circled character 3. As described above, in the OCR area shown by the circled character 3, the same text (OCR text) as the correct character string cannot be obtained from the image. In this case, the display of the second screen is a display corresponding to the fact that the correct character string and the recognized character string do not match.

[0125] For example, the display mode of the screen frame of the second screen, the display mode of the mismatched characters displayed on the second screen, the text showing the collation result displayed on the second screen, and the type of icon displayed on the second screen are displayed (display mode, display content) according to the mismatch between the correct character string and the recognized character string. For example, the screen frame of the second screen is displayed in red, with a bold line, and with a red background. In addition, the mismatched characters are displayed in italics, bold, and red, and the background color of the mismatched characters is displayed in a darker red than the background color of the screen. In addition, the text showing the collation result, "Correct text cannot be obtained," is displayed. In addition, a red triangular icon is displayed. As can be seen by comparing the pop-up screens of FIG. 26 and FIG. 27, the display control unit 65 compares the display mode and display content of the pop-up screen when the correct character string and the recognized character string do not match with the display mode and display content when the correct character string and the recognized character string match, and sets the display mode and display content to be more alert to the user.

[0126] In addition, in Figures 25 to 27, the character recognition result (recognized character string) is displayed on the right side of the screen (circled characters 1 to 5), but the correct character string and / or the recognized character string may be displayed in this area, or the correct character string and the recognized character string may not be displayed.

[0127] Fig. 28 is a flowchart showing an outline of the flow of evaluation result display processing according to this embodiment. The processing shown in this flowchart is started when, in a state where a user scans a document for which OCR is to be performed on the information processing device 1 and acquires a captured image (image data), an OCR area is specified and a correct answer character string is input. For example, the processing is started when the user presses the "Start evaluation" button on the screen shown in Fig. 24.

[0128] In step S201, it is determined whether all OCR areas have been determined. The collation unit 64 determines whether a determination has been made for all OCR areas specified by the user as to whether the recognized character string matches the correct character string. If a determination has been made for all OCR areas as to whether the recognized character string matches the correct character string (YES in step S201), the process shown in this flowchart ends. On the other hand, if a determination has not been made for all OCR areas as to whether the recognized character string matches the correct character string (NO in step S201), the process proceeds to step S202.

[0129] In step S202, an undetermined OCR area is acquired. The recognition result acquisition unit 63 acquires one OCR area (an image related to the OCR area) from the OCR area determined in step S201 as an area for which it has not been determined whether the recognized character string matches the correct character string. Then, the process proceeds to step S203.

[0130] In step S203, the recognized character string for the undetermined OCR area is acquired. The recognition result acquisition unit 63 acquires the recognized character string for the OCR area acquired in step S202. After that, the process proceeds to step S204.

[0131] In step S204, it is determined whether the recognized character string matches the correct character string. The collation unit 64 collates (compares) the recognized character string acquired in step S203 with the correct character string for the OCR area acquired in step S202, which has been input in advance by the user, and determines whether these character strings match. If the recognized character string matches the correct character string (YES in step S204), the process proceeds to step S205. On the other hand, if the recognized character string does not match the correct character string (NO in step S204), the process proceeds to step S206.

[0132] In step S205, the OCR area (OCR area frame) is displayed in a display corresponding to the match (display mode indicating the match). The display control unit 65 displays the OCR area frame for the OCR area acquired in step S202 in a display (display mode) corresponding to the match between the recognized character string and the correct character string (see FIG. 25). Then, the process returns to step S201.

[0133] In step S206, the OCR area (OCR area frame) is displayed in a display corresponding to the lack of match (display mode indicating the lack of match). The display control unit 65 displays the OCR area frame for the OCR area acquired in step S202 in a display (display mode) corresponding to the lack of match between the recognized character string and the correct answer character string (see FIG. 25). After that, the process returns to step S201.

[0134] Fig. 29 is a flowchart showing an outline of the flow of the pop-up display process according to this embodiment. The process shown in this flowchart is started when the user places the mouse over the OCR area, for example. For example, the process is started when the user places the mouse over the OCR area on the screen shown in Fig. 25.

[0135] In step S301, it is determined whether the recognized character string matches the correct character string. The collation unit 64 determines whether the recognized character string for the mouse-over OCR area matches the correct character string. If the recognized character string matches the correct character string (YES in step S301), the process proceeds to step S302. On the other hand, if the recognized character string does not match the correct character string (NO in step S301), the process proceeds to step S303.

[0136] In step S302, a pop-up screen is displayed in accordance with the match (display mode and / or display contents indicating the match). The display control unit 65 displays the screen components (screen frame of the pop-up screen, icons, text indicating the match result, and mismatched characters) of the pop-up screen indicating the result (matching result) determined in step S301 in a display (display mode and / or display contents) corresponding to the match between the recognized character string and the correct character string (see FIG. 26). After that, the process shown in this flowchart ends.

[0137] In step S303, differences in the character strings are extracted. The collation unit 64 extracts differences (mismatched characters) between the recognized character strings that are determined not to match in step S301 and the correct character string. After that, the process proceeds to step S304.

[0138] In step S304, a pop-up screen is displayed in accordance with the fact that there is no match (a display format and / or display contents indicating that there is no match). The display control unit 65 displays the screen components (screen frame of the pop-up screen, icons, text indicating the matching result, and mismatched characters) of the pop-up screen indicating the result (matching result) determined in step S301 in a display (a display format and / or display contents) in accordance with the fact that the recognized character string and the correct character string do not match (see FIG. 27). After that, the process shown in this flowchart ends.

[0139] In addition, the user who has confirmed the matching result (screens in Figs. 25 to 27) may change the image processing settings and repeat the process of confirming the matching result until a satisfactory result (character recognition result) is obtained. Specifically, a case is assumed in which the user who confirmed the matching result determines that the result is not satisfactory. In this case, the user obtains a captured image (processed image) different from the captured image used for the matching result by performing image processing (image processing based on different image processing settings) different from the image processing (image processing settings by the image processing unit 72) performed on the captured image used for the matching result on the read image. Then, the above-mentioned processing is performed on the newly obtained captured image to obtain a matching result for this newly obtained captured image, and a screen (screens in Figs. 25 to 27) showing the matching result is displayed by the above-mentioned display control processing. Then, the user checks the matching result again to confirm whether or not a satisfactory result has been obtained. If a result satisfactory to the user is obtained by repeating these processes, the image processing settings at the time when the satisfactory result was obtained may be saved and used for subsequent operations. For example, the information processing device 1 according to the present embodiment has a functional unit (e.g., an evaluation acquisition unit (not shown)) that acquires an evaluation result from a user that the character recognition result is a satisfactory result, that is, that the performed image processing is image processing suitable for character recognition. When the user presses a button (e.g., an "OK" button) on a screen showing the collation result, which is pressed when the character recognition result is a satisfactory result (when the performed image processing is image processing suitable for character recognition), the evaluation acquisition unit may acquire the evaluation result that the performed image processing is image processing suitable for character recognition. When the user presses this "OK" button, the storage unit (not shown) may store the image processing settings when a satisfactory result is obtained. In the above, the image processing (image processing settings) may be changed manually by the user or automatically by a function on the program.

[0140] As described above, according to the present embodiment, the display of a screen (a screen showing a collation result between a correct answer character string and a recognized character string) for confirming the character recognition result for an image subjected to the image processing to be evaluated in order for the user to evaluate whether the image processing to be evaluated is suitable for character recognition is controlled to be different depending on the collation result, thereby making it possible to improve the accuracy of the user's evaluation of the character recognition result. In other words, it is possible to prevent the user from making an erroneous judgment when comparing the correct answer character string with the recognized character string. This makes it possible to assist the user in making a judgment (evaluation) of the character recognition result. In addition, according to the present embodiment, the correctness of the OCR text (recognized character string) is judged not by the reliability of the recognized character string but by a comparison with the correct answer text input in advance by the user, so that it is possible to judge the correctness of the OCR text (whether it matches or not) with high accuracy (100% accuracy). In addition, according to the present embodiment, the display of the screen showing the collation result is made different (changed) depending on the collation result between the correct answer character string and the recognized character string, so it is possible to call the user's attention to an OCR area where the correct answer character string does not match the recognized character string among multiple OCR areas.

[0141] [Fifth embodiment] In this embodiment, an embodiment combining the first embodiment and the fourth embodiment (a system for evaluating whether the determined recommended settings are suitable for character recognition (image processing based on the recommended settings (image processing by the recommended settings) is suitable for character recognition)) will be described. In this embodiment, first, the recommended settings are determined by the method according to the first embodiment. Then, a character recognition result for an image in which the recommended settings are reflected is obtained by the method according to the fourth embodiment, and a screen showing an evaluation result of the obtained character recognition result (a result of matching between the recognized character string and the correct character string) is displayed. Note that, by the method according to the fourth embodiment, the display of this screen is controlled to be different depending on the result of matching between the recognized character string and the correct character string. Note that the configuration of the system 9 according to this embodiment is roughly similar to the configuration of the system 9 according to the first embodiment described above with reference to FIG. 1, and therefore a description thereof will be omitted.

[0142] FIG. 30 is a diagram showing an outline of the functional configuration of the information processing device according to the present embodiment. The information processing device 1 functions as a device including an image acquisition unit 31, a reception unit 32, an analysis unit 33, a storage unit 34, a presentation unit 35, and a display control unit 65, by a program recorded in the storage device 14 being read into the RAM 13 and executed by the CPU 11, and each hardware included in the information processing device 1 is controlled. The image acquisition unit 31 includes a read image acquisition unit 41 and a read image processing unit 42. The reception unit 32 includes a character region acquisition unit 43 and a correct answer information acquisition unit 44. The analysis unit 33 includes a candidate selection unit 45 and a recommended setting determination unit 46. The candidate selection unit 45 includes an image analysis unit 51, a first image processing unit 52, a first recognition result acquisition unit 53, and a selection unit 54. The recommended setting determination unit 46 includes a second image processing unit 55, a second recognition result acquisition unit 56, and a determination unit 57. In this embodiment and other embodiments described later, each function of the information processing device 1 is executed by a CPU 11 which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.

[0143] The image acquisition unit 31, the reception unit 32, the analysis unit 33, the storage unit 34, and the presentation unit 35 in this embodiment are generally similar to the image acquisition unit 31, the reception unit 32, the analysis unit 33, the storage unit 34, and the presentation unit 35 in the first embodiment, and therefore their explanations will be omitted. Also, the display control unit 65 in this embodiment is generally similar to the display control unit 65 in the fourth embodiment, and therefore their explanations will be omitted. Note that the second image processing unit 55 corresponds to the "image processing means" in the fourth embodiment, the correct answer information acquisition unit 44 corresponds to the "correct answer information acquisition means" in the fourth embodiment, the second recognition result acquisition unit 56 corresponds to the "recognition result acquisition means" in the fourth embodiment, and the "collation means" in the fourth embodiment corresponds to the means (functional unit) included in the determination unit 57 in this embodiment.

[0144] In this embodiment, when the analysis unit 33 (recommended setting determination unit 46) determines the recommended settings (image processing settings suitable for character recognition), the display control unit 65 causes the display means to display a screen for the user to evaluate whether or not the image processing based on the recommended settings is suitable for character recognition. For example, the display control unit 65 causes the evaluation result display screen as shown in Fig. 25 to be displayed in accordance with the result of matching between the correct answer character string and the recognized character string for each OCR area. In this embodiment, the evaluation result display screen displays an image reflecting the recommended settings (an image that has been subjected to image processing based on the recommended settings), an OCR area frame, and character recognition results.

[0145] In this embodiment, the character recognition result for the image that has been subjected to image processing using the recommended settings (image processing settings that will be determined as the recommended settings later) that have already been acquired by the second recognition result acquisition unit 56 in the recommended settings determination process is displayed on the screen. However, after the recommended settings are determined, the second image processing unit 55 may again perform image processing based on the recommended settings on the captured image, and the character recognition result for the resulting processed image may be acquired by the second recognition result acquisition unit 56, and the acquired character recognition result may be displayed on the screen.

[0146] In addition, in this embodiment, it is assumed that the above-mentioned evaluation method 1 is used in the recommended setting determination process. In this case, in the recommended setting determination process, a comparison (determination of whether the character strings match each other) between the correct answer character string and the recognized character string for each OCR area in the image in which the recommended setting (the image processing setting that will be determined as the recommended setting later) is reflected has already been performed. Therefore, the display control unit 65 can control the display of the evaluation result display screen to be a display according to the result of the matching process that has already been performed, without performing a matching process after the recommended setting is determined. In other words, when the recommended setting is determined by the process shown in the flowchart shown in FIG. 16, the process of step 205 or step S206 shown in the flowchart shown in FIG. 28 is executed according to the result of the matching process that has already been performed on the recommended setting.

[0147] In addition, when the above-mentioned evaluation method 2 is used in the recommended setting determination process, the correct answer character string and the recognized character string for each OCR area in the image in which the recommended setting (the image processing setting that will be later determined as the recommended setting) are not compared (determined whether the character strings match each other) in the recommended setting determination process. In this case, the comparison unit 64 described in the fourth embodiment compares the correct answer character string and the recognized character string for each OCR area in the image in which the recommended setting is reflected, and the display control unit 65 controls so that the display is performed according to the result. That is, when the recommended setting is determined by the process shown in the flowchart shown in FIG. 16, the process shown in the flowchart shown in FIG. 28 is executed. In this case, in step S203 in FIG. 28, after the recommended setting is determined, the recognized character string for the image in which the recommended setting is reflected may be newly acquired, or the recognized character string for the recommended setting (the image processing setting that will be later determined as the recommended setting) that has already been acquired in the recommended setting determination process may be acquired from the storage device 14 or the like.

[0148] In addition, the user who has confirmed the collation result (screens in Figs. 25 to 27) may change the image processing settings and repeat the process of confirming the collation result until a satisfactory result (character recognition result) is obtained, thereby determining image processing settings more suitable for character recognition. Specifically, a case is assumed in which the user who has confirmed the collation result regarding the recommended settings determines that the result is not satisfactory. In this case, the user corrects (changes) the recommended settings, and, for example, the second image processing unit 55 performs image processing based on the corrected recommended settings on the scanned image, thereby obtaining an image different from the image in which the recommended settings are reflected. Then, for example, the second recognition result obtaining unit 56 obtains a character recognition result (recognized character string) for the newly obtained image, and, for example, the determination unit 57 (collation means) compares the recognized character string with the correct answer character string (obtains the collation result), and the display control unit 65 displays a screen (screens in Figs. 25 to 27) showing the collation result. Then, the user checks the collation result again to confirm whether or not a satisfactory result has been obtained. If the user is satisfied with the result obtained by repeating these processes, the image processing settings at the time when the satisfactory result was obtained may be saved and used for subsequent operations. For example, as described above, the information processing device 1 according to this embodiment may include an evaluation acquisition unit, so that when a user presses a button (e.g., an "OK" button) on a screen showing the collation result, which is to be pressed when the performed image processing is suitable for character recognition, the evaluation acquisition unit may acquire an evaluation result that the performed image processing is suitable for character recognition. In addition, when the user presses this "OK" button, the storage unit 34 may store the image processing settings when a satisfactory result is obtained.

[0149] In the above, the recommended settings may be changed manually by the user, or automatically by a program function. For example, as described in the first embodiment, if the user is not satisfied with the image processing settings proposed by the presenting unit 35 (presentation of recommended settings), it is possible to determine image processing settings (recommended settings) suitable for OCR again by changing the OCR area, etc., and performing the above-mentioned analysis process again. The recommended settings may be automatically changed by using the re-determined recommended settings.

[0150] The display control method in this embodiment (method of controlling the screen display to be different depending on the collation result) is generally similar to the method described in the fourth embodiment, and therefore the description will be omitted. Also, the flow of the pop-up display process in this embodiment is generally similar to the flow of the pop-up display process in the fourth embodiment described with reference to Fig. 29, and therefore the description will be omitted.

[0151] According to the present embodiment, the display of a screen (a screen showing the result of matching between a correct answer character string and a recognized character string) for the user to evaluate whether the image processing by the recommended settings is suitable for character recognition is controlled to be different depending on the result of the matching, so that the user can easily evaluate whether the image processing by the recommended settings determined to obtain an image suitable for character recognition is suitable for character recognition. More specifically, even if image processing suitable for character recognition is performed, text that cannot be read by OCR or misreading may occur, so the user may actually check the character recognition accuracy of the image that has been subjected to image processing suitable for character recognition to check whether there is any misreading. Even in this case, according to the present embodiment, it becomes easier for the user to determine whether the "text read by the user" and the "text read by OCR" match, so that it is possible to assist the user in checking and to make the image processing settings for OCO more efficient. Furthermore, according to the present embodiment, when the user determines whether to change the recommended settings (whether to perform the process of determining the recommended settings again) based on the character recognition result, it is possible to prevent misreading, so that it is possible to appropriately determine whether to change the recommended settings.

[0152] 1. Information processing device 2 Server 8. Scanner

Claims

1. Computer, an image acquisition means for acquiring a captured image of a document; image processing means for performing image processing for character recognition on the read image of the document read by the image reading means; an analysis means for determining recommended settings for a plurality of setting items in the image processing means using the captured image; A program that functions as The analysis means a recommended setting determination means for determining the recommended settings for the plurality of setting items by repeatedly performing image processing on the captured image while changing the setting values ​​of the plurality of setting items, after limiting the setting values ​​for the setting items included in the plurality of setting items to at least one setting value selected as a candidate for the recommended setting; program.

2. The analysis means further comprises a candidate selection means for selecting at least one setting value that is a candidate for the recommended setting from settable setting values ​​for the setting items included in the plurality of setting items by performing an analysis process using the captured image. The program according to claim 1.

3. the candidate selection means selects setting values ​​that are candidates for the recommended setting by performing image analysis on the captured image. The program according to claim 2.

4. the candidate selection means performs image analysis to determine the amount of background pattern in the captured image, and selects at least one setting value for a setting item related to background pattern removal according to a result of the image analysis as a candidate for the recommended setting for the setting item related to background pattern removal. The program according to claim 3.

5. the candidate selection means performs an edge analysis on the captured image in the image analysis for determining the amount of the background pattern; The program according to claim 4.

6. the candidate selection means performs image analysis on the captured image to determine whether or not a ruled line is present, and selects at least one setting value for a setting item related to a dropout color according to a result of the image analysis as a candidate for the recommended setting for the setting item related to the dropout color. The program according to claim 3.

7. When it is determined that the ruled line exists as a result of performing image analysis to determine the presence or absence of the ruled line, the candidate selection means determines the color of the ruled line and selects a setting value for the setting item related to the dropout color corresponding to the determined color as the candidate for the recommended setting for the setting item related to the dropout color. The program according to claim 6.

8. the candidate selection means performs image analysis to determine the amount of noise using the imaging analysis, and selects at least one setting value for each setting item related to binarization sensitivity and / or noise reduction according to the image analysis result as a candidate for the recommended setting for each setting item related to the binarization sensitivity and / or noise reduction; The program according to claim 3.

9. the candidate selection means, in the image analysis for determining the amount of noise, calculates the number of black connected pixel blocks in a character region in the binarized image of the captured image, and selects at least one setting value for a setting item related to binarization sensitivity and / or noise reduction according to a comparison result between the number of black connected pixel blocks and an expected number of black connected pixel blocks derived from a correct character string for the character region, as a candidate for the recommended setting for the setting item related to binarization sensitivity and / or noise reduction; The program according to claim 8.

10. the candidate selection means calculates the expected number of black connected pixel blocks based on the language of characters in the character area and the number of characters in the correct character string for the character area; The program according to claim 9.

11. the candidate selection means performs image processing using the settable setting values ​​on the captured image, and selects setting values ​​that are candidates for the recommended setting based on a character recognition result for the image obtained as a result of the image processing. The program according to claim 2.

12. the candidate selection means selects, as the recommended setting candidates, a predetermined number of setting values ​​selected from the settable setting values ​​in descending order of the character recognition results. The program according to claim 11.

13. the recommended setting determination means determines the recommended settings for the plurality of setting items based on character recognition results for a plurality of images obtained by repeatedly performing the image processing on the captured image while changing the setting values ​​of the plurality of setting items. The program according to any one of claims 1 to 12.

14. the recommended setting determination means determines, as the recommended settings for the plurality of setting items, a combination of setting values ​​for the plurality of setting items when the image with the best character recognition result is obtained. The program according to claim 13.

15. In addition to the image acquisition means, the image processing means, and the analysis means, the computer is caused to function as a storage means for storing the recommended settings for the determined plurality of setting items in association with identification information related to the document. The program according to any one of claims 1 to 12.

16. The plurality of setting items include image processing setting items related to at least one of pattern removal, specific character extraction, dropout color, binarization sensitivity, and noise removal. The program according to any one of claims 1 to 12.

17. In addition to the image acquisition means, the image processing means, and the analysis means, the computer is caused to function as an output means for outputting the recommended settings for the determined plurality of setting items. The program according to any one of claims 1 to 12.

18. In addition to the image acquisition means, the image processing means, and the analysis means, the computer functions as a display control means for displaying on a display means one or more screens that show the results of matching a recognized character string, which is the result of character recognition of a character area in an image in which the recommended settings are reflected, with a correct character string for the character area, on which a screen for a user to evaluate whether or not the image processing according to the determined recommended settings is suitable for character recognition, the display control means controls the display of at least one of the one or more screens to be different depending on the comparison result; The program according to any one of claims 1 to 12.

19. A computer, an image acquisition means for acquiring a plurality of captured images obtained by capturing images of a plurality of documents; an image processing means for performing image processing for character recognition on the plurality of read images of the plurality of documents read by the image reading means; an analysis means for determining recommended settings for a plurality of setting items in the image processing means using the plurality of captured images; A program that functions as The analysis means a recommended setting determination means for determining the recommended settings for the plurality of setting items by repeatedly performing image processing on any of the plurality of captured images while changing the setting values ​​of the plurality of setting items, after limiting the setting values ​​for the setting items included in the plurality of setting items to at least one setting value selected as a candidate for the recommended setting; program.

20. an image acquisition means for acquiring a captured image of a document; image processing means for performing image processing for character recognition on the read image of the document read by the image reading means; an analysis means for determining recommended settings for a plurality of setting items in the image processing means using the captured image; Equipped with The analysis means a recommended setting determination means for determining the recommended settings for the plurality of setting items by repeatedly performing image processing on the captured image while changing the setting values ​​of the plurality of setting items, after limiting the setting values ​​for the setting items included in the plurality of setting items to at least one setting value selected as a candidate for the recommended setting; Information processing system.