Information processing system, its control method, and program

The information processing system improves OCR accuracy in non-standard forms by specifying and correcting document ranges, addressing incomplete character detection and reducing user burden.

JP7824544B1Active Publication Date: 2026-03-05CANON MARKETING JAPAN INC +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing optical character recognition (OCR) systems struggle to comprehensively detect all character data in image data, particularly in non-standard forms like invoices and delivery notes, leading to incomplete character detection.

Method used

An information processing system that includes a receiving means for specifying a document range, a determining means to ensure the range is suitable for OCR, and a notification means to confirm the suitability, allowing for user correction and storage of correction ranges for future reference.

Benefits of technology

Enhances OCR accuracy by ensuring appropriate range selection for character detection, reducing missed detections and user effort through automated and manual correction processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007824544000001_ABST
    Figure 0007824544000001_ABST
Patent Text Reader

Abstract

The present invention provides an information processing system relating to form analysis that realizes appropriate OCR processing, and a control method and program for the same. [Solution] In an OCR support system in which a client terminal 101, a scanner 102, and a file server 103 are connected to each other so that they can communicate with each other via a network 100, the client terminal, which is an information processing device, is equipped with a reception means for receiving a range specification for a document, a determination means for determining whether the range specified by the reception means meets predetermined criteria indicating a size suitable for OCR processing, and a notification means for notifying the determination result by the determination means.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system for document analysis, a control method thereof, and a program. [Background technology]

[0002] Conventionally, it has been known to perform optical character recognition (hereinafter referred to as OCR) to obtain character string information written in form image data.

[0003] Patent Document 1 describes a method of searching for item names related to the item value to be extracted from within a form image, identifying information having a character string type corresponding to the item value to be extracted from nearby character strings as the item value, and assisting in the task of confirming whether the acquired item value is correct. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-11019 DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]

[0005] However, in character detection, there are cases where it is not possible to comprehensively detect all character data in image data. Therefore, an object of the present invention is to realize appropriate OCR processing. [Means for solving the problem]

[0006] a receiving means for receiving a range specification for a document; a determining means for determining whether the range specified by the receiving means satisfies a predetermined standard indicating a size suitable for OCR processing; a notification means for notifying a result of the determination by the determination means; The present invention is characterized by comprising: [Effects of the Invention]

[0007] According to the present invention, appropriate OCR processing can be achieved. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an example of the configuration of an OCR support system. [Figure 2] FIG. 2 is a block diagram showing an example of the hardware configuration of a client terminal 101. [Figure 3] 10 is a flowchart illustrating an example of an outline of a process. [Figure 4] 10 is a flowchart illustrating an example of a process for correcting an OCR result using a recorded correction range A. [Figure 5] 10 is a flowchart illustrating an example of a process for receiving a correction range B from a user and correcting an OCR result using the correction range B. [Figure 6] 10 is a screen showing an example of an OCR result before correction by a user input and before correction by a recorded correction range A. [Figure 7] FIG. 10 is a diagram illustrating an example of a process for determining the type of a document to be subjected to OCR. [Figure 8] FIG. 10 illustrates an example of a list of missing items. [Figure 9] 10 is a screen showing an example of an operation performed by a user when inputting a correction range B for an OCR result. [Figure 10] FIG. 10 is a diagram illustrating an example of data in a correction range A. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0010] In this embodiment, the invention will be explained using OCR processing for non-standard forms, which are forms whose layout cannot be determined in advance (examples of forms include invoices, purchase orders, and delivery notes.) Note that the present invention can also be applied to standard forms, which are forms whose layout can be determined in advance.

[0011] It should be noted that a form is an example of a document to be subjected to OCR, and the present invention is also applicable to documents other than forms.

[0012] FIG. 1 is a diagram showing an example of the configuration of an OCR support system according to an embodiment of the present invention.

[0013] In the information processing system of the present invention, a client terminal 101, a scanner 102, and a file server 103 are communicably connected via a network 100.

[0014] The client terminal 101 may be any device that incorporates the functions shown in Figure 2, such as a personal computer (hereinafter referred to as PC), a mobile terminal such as a smartphone, a tablet terminal, a wearable terminal such as a smart watch, a PDA, or a game console.

[0015] The scanner 102 scans the form, converts it into an image file, and transmits it to the client terminal 101 via the network.

[0016] The network 100 can take the form of a wired LAN, a wireless LAN, a USB, or the like, depending on the physical interface that the scanner 102 has.

[0017] The file server 103 can store image files captured by a scanner and can store the range of OCR corrections for forms.

[0018] As a method for importing an image scanned by the scanner 102 into the client terminal 101, either the image can be sent directly from the scanner 102 to the client terminal 101, or the image file imported by the scanner 102 can be temporarily stored in the file server 103, and the client terminal 101 can retrieve the image file from the file server 103.

[0019] 2 is a block diagram showing an example of the hardware configuration of a client terminal 101 (a client terminal is an example of an information processing device) according to an embodiment of the present invention. The scanner 102 and file server 103 also have the same configuration.

[0020] As shown in FIG. 2, each information processing device is connected to a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 via a system bus 204.

[0021] The CPU 201 comprehensively controls each device and controller connected to the system bus 204 .

[0022] ROM202 or external memory 211 stores the BIOS (Basic Input / Output System) and OS (Operating System), which are control programs executed by CPU201, computer-readable and executable programs for realizing this information processing method, and various necessary data (including data tables).

[0023] The RAM 203 functions as a main memory, a work area, etc. for the CPU 201. The CPU 201 loads programs and the like required for executing processing from the ROM 202 or the external memory 211 into the RAM 203, and executes the loaded programs to realize various operations.

[0024] The input controller 205 controls input from input devices such as a keyboard 209 and pointing devices such as a mouse, touchpad, etc. (not shown). If the input device is a touch panel, the user can issue various instructions by pressing (touching with a finger, etc.) icons, cursors, or buttons displayed on the touch panel.

[0025] The touch panel may also be a touch panel capable of detecting positions touched by multiple fingers, such as a multi-touch screen.

[0026] The video controller 206 controls the display on an external output device such as a display 210. The display also includes the display of a notebook PC integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. In addition, for the device capable of receiving the above-mentioned touch operation, an input device is also provided.

[0027] In the explanation using the flowcharts below, the display destination when displaying is assumed to be the display 210 unless otherwise specified.

[0028] The video controller 206 can control a video memory (VRAM) for display control, and can use part of the RAM 203 as a video memory area, or can provide a separate dedicated video memory.

[0029] The memory controller 207 controls access to the external memory 211. The external memory may be an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a CompactFlash (registered trademark) memory connected to a PCMCIA card slot via an adapter.

[0030] The communication I / F controller 208 connects to and communicates with external devices via a network 214 (for example, the network 101 shown in FIG. 1), and executes communication control processing on the network 214. For example, communication using TCP / IP, telephone lines such as ISDN, and 3G, 4G, and 5G lines of mobile phones are possible.

[0031] The CPU 201 enables display on the display 210 by, for example, executing a process of expanding (rasterizing) an outline font into a display information area in the RAM 203. The CPU 201 also enables user instructions using a mouse cursor (not shown) or the like on the display 210.

[0032] Next, the processing executed by the client terminal 101 in the embodiment of the present invention will be described with reference to FIGS.

[0033] 3 to 5 are flowcharts showing the flow of processing executed by the client terminal 101 in the embodiment of the present invention.

[0034] The flowchart in Fig. 3 shows the overall processing when the CPU 201 of the client terminal 101 reads and executes a predetermined control program. The processing of each step is executed by the CPU 201 of each device. When the client terminal 101 acquires an image to be OCR-targeted from the scanner 102 or the file server 103, the processing in Fig. 3 starts.

[0035] In S301, the CPU 201 acquires an image related to a form (hereinafter, referred to as an OCR target image). The OCR target image may be acquired directly from the scanner 102, or may be acquired from an image file stored in the file server 103.

[0036] In S302, the CPU 201 performs preprocessing for the subsequent processing on the OCR target image acquired in S301. The preprocessing includes processing necessary for the subsequent processing, such as binarization and noise removal.

[0037] In S303, the CPU 201 determines the type of the form related to the OCR target image. For example, when determining the type of form using the layout, this can be realized by using a method (so-called template matching) in which a template image is prepared for each type of form, the similarity between the layout of each template image and the OCR target image is calculated, and the type of form is determined using the similarity.

[0038] The method for classifying the type of form is not limited to template matching, and any method that can determine the type of form may be used. For example, the type may be determined based on the title of the form, such as invoice, receipt, or order form. In this case, the type of form is determined based on the result of OCR of the acquired form title. For this reason, the processing of this step may be performed after character recognition in S305. Also, a combination of multiple classification methods may be used.

[0039] FIG. 7 is a diagram illustrating an example of a process for determining the type of a form related to an OCR target image.

[0040] The template image 701 is data stored in the file server 103 .

[0041] OCR target images 702 and 703 are OCR target images to be processed.

[0042] Comparing template image 701 and OCR target image 702 reveals three slightly larger characters in the upper left corner of the document, with eight similarly sized characters below them. Furthermore, a table is included in the lower half of the document, with no text or other content below the table. This reveals that the overall layout, including the size and placement of the characters and the placement of the table, is similar. Therefore, template matching can be used to determine that the two images are the same type. Comparing template image 701 and OCR target image 703, OCR target image 703 reveals that three slightly larger characters are located in the center of the document, and the table is also located approximately in the center of the document. There is also a rectangular area below the table, with smaller characters below that. Thus, because the layouts of template image 701 and OCR target image 703 are not similar, template matching does not determine that they are the same type. The type may be determined by calculating the similarity between each stored template image 701 and the OCR target images 702 and 703 and determining that the image is of the same type as the template image with the highest similarity, or by presenting types with similarities exceeding a threshold value and accepting the user's selection of an appropriate type.

[0043] Furthermore, when identifying the type by subject, for example, the largest character string in the upper part of the document (presumably set in advance, such as the upper half or the upper one-third) is acquired, and a determination can be made based on whether the character strings are identical (or the similarity between the character strings). The character string "invoice" is acquired in template image 701, and the character string "invoice" is also acquired from OCR target image 702. Since both are the same character string, template image 701 and OCR target image 702 are determined to be of the same type. Since the character string "receipt" is acquired in OCR target image 703, template image 701 and OCR target image 703 are determined to be of different types.

[0044] It is also possible to combine template matching and title-based determination (combining multiple methods to determine type). In this case, an evaluation value may be calculated by multiplying the similarity calculated by template matching by the result of the title-based determination, and the calculated evaluation value may be used to determine whether the types are the same.

[0045] The information on the type of form obtained here is used to obtain the history of correction range A in S402 and to save correction range B in S508.

[0046] The correction range in this embodiment refers to an area (a part of the OCR target image) designated by the user as an area where OCR processing should be performed again, for an area in which no characters were detected in the OCR processing in S304 (OCR processing on the entire OCR target image). Of the correction ranges, "correction range A" refers to the correction range designated by the user in a form that was previously subjected to OCR processing. Correction range A is stored in the file server 103 in association with the form. "Correction range B" refers to the correction range designated by the user in the OCR target image acquired in S301 (i.e., the OCR target image being processed).

[0047] The CPU 201 performs the first character detection process in step S304, and performs the second and subsequent character detection processes in steps S307 and thereafter.

[0048] In S304, the CPU 201 performs character detection from the entire image to be OCR processed. Specifically, the CPU 201 finds the outlines of characters using methods such as template matching and edge detection, and detects areas containing characters. The CPU 201 further divides the detected areas into individual areas, taking into account the distance between characters, the shape of the characters, and other factors. The CPU 201 extracts features from the detected individual areas to identify characters and then performs character detection.

[0049] In the character detection in steps S307 and thereafter, the same method as in S304 may be used, or a different detection method may be used.

[0050] In S305, the CPU 201 performs character recognition on the characters detected in S304. Specifically, the CPU 201 uses template matching, pattern recognition techniques, etc. to match the shapes and patterns of the characters, identify the correct characters, and recognize them as text. The CPU 201 associates the character recognition results with each individual region.

[0051] In the character recognition in steps S307 and thereafter, the same method as in S305 may be used, or a different recognition method may be used.

[0052] In this embodiment, the series of steps from character detection in S304 to character recognition in S305 is referred to as OCR. OCR in this embodiment includes OCR using AI (Artificial Intelligence) (hereinafter referred to as AI-OCR) and OCR without AI.

[0053] FIG. 6 shows an example of a screen displayed by the CPU 201 showing the OCR result of the OCR target image.

[0054] A screen 600 displays the OCR results of a form that has undergone OCR.

[0055] A form 601 is the OCR target image acquired in step S301. The CPU 201 displays the OCR results on the display unit. This screen can also accept input of the correction range B from the user.

[0056] 6, character strings that have undergone OCR processing are displayed surrounded by a dashed rectangle, as shown in character string 602. For example, the character string "Bill amount: XXX yen" shown in character string 603 is not surrounded by a dashed rectangle, which indicates that OCR has not been performed (characters could not be detected).

[0057] The method of identifying and displaying the character string that has been OCRed does not have to be a method of enclosing it in a dashed rectangle; for example, it can be a rectangle with solid lines, wavy lines, or dotted lines, or it can be a method of enclosing it in a circle (including an ellipse) instead of a rectangle, or it can be a method of coloring the area related to the character string by shading it, etc.

[0058] The end button 604 is a button that accepts an instruction to end the input of the correction range B. The display of the end button 604 may be any display that indicates that the end of the process is accepted, and may be a button that displays "Done," "Save," "Confirm," etc.

[0059] 6, there are still character strings that should be OCR'd but have not been OCR'd, such as character string 603. A method for efficiently performing OCR processing on character strings that were not OCR'd in the OCR processing of S304 and S305 will be described below.

[0060] One of the reasons for missed detections of character strings that should be OCRed is that the character size is too large or too small relative to the range of the OCR processing. Specifically, if the range of the OCR processing is too large, the relative character size (character size relative to the OCR range) will be small, resulting in characters that cannot be recognized and are likely to be missed. On the other hand, if the range of the OCR processing is too small, the relative character size will be large, resulting in characters that cannot be recognized and are likely to be missed. This decrease in character detection and recognition accuracy due to the inappropriate range specification is caused by the characteristics of OCR engines, particularly machine learning models (trained models, AI) in general. Machine learning models typically resize input images to a specified size before processing them, such as object detection. Therefore, it is necessary to ensure that the relative size of the characters contained in the input image is appropriate, rather than the absolute size. Because character size increases or decreases with resizing, it is necessary to select a range of the OCR processing that is appropriate for character detection (neither too large nor too small). Specifying an appropriate size facilitates character recognition and reduces missed detections. As described above, the success or failure of OCR depends on the size of the range to be OCR. Therefore, in this embodiment, a process that allows the user to select a range of a size suitable for OCR (which reduces the number of missed characters) will be described below.

[0061] In S306, the CPU 201 determines whether a correction range has been specified in the past for a form of the same type as the OCR target form. Specifically, this determination is made based on whether a correction range A has been registered for a form of the same type as the OCR target form registered in the file server 103. If a correction range has been specified in the past for a form of the same type as the OCR target form, the process proceeds to S307; if not, the process proceeds to S308.

[0062] In S307, the CPU 201 performs re-OCR processing on the form 601 based on the correction range A stored in the file server 103. Details of the processing in S307 will be described with reference to the flowchart in FIG.

[0063] In S308, the CPU 201 accepts the correction range specification manually input by the user. Details of the processing in S308 will be described with reference to the flowchart in FIG.

[0064] In S309, the file server 103 stores the character strings that have been OCR processed up to this point.

[0065] This concludes the explanation of Figure 3.

[0066] FIG. 4 is a flowchart showing the flow of OCR processing using the recorded correction range A.

[0067] In S401, the CPU 201 obtains the missing item list 800. Items to be OCRed are defined in advance, and if OCR has not been performed on character strings related to the defined items, the CPU 201 extracts the items for which OCR has not been performed and creates the missing item list 800.

[0068] FIG. 8 is a diagram showing an example of a missing item list.

[0069] The missing item list 800 in FIG. 8 is a list of items that should be OCR processed but for which OCR processing has not been performed at this stage, and is made up of an item name 801 and a processing execution flag 802.

[0070] Item name 801 is the name of an item that should be OCR processed but has not yet been OCR processed. Missing item list 800 in Figure 8 lists the customer name, invoice amount, subtotal, total amount, etc.

[0071] The processing execution flag 802 registers a flag indicating whether OCR processing using correction range A in the flowchart of FIG. 4 has been executed. In FIG. 8, when the flag is set (OCR processing using correction range A has been executed), it is shown as "processed." If it has not been executed, it is left blank. In FIG. 8, OCR processing using correction range A has been executed for the "Customer Name" and "Billing Amount" items, so it is registered that these items have been processed (the processing execution flag is set). OCR processing using correction range A has not been executed for the "Subtotal" and "Total Amount" items, so it is not registered that they have been processed.

[0072] In the missing item list 800, all the fields in the processing execution flag 802 column are blank at the time of S401.

[0073] In S402, the CPU 201 reads the correction range A for the same type of form as the type of form identified in S303 from the file server 103. Thereafter, the following steps S403 to S409 are repeated for the area in the OCR target image that corresponds to the acquired correction range A. In S409, if the CPU 201 has performed OCR processing for all items registered in the missing item list 800, the loop processing ends.

[0074] In S404, the CPU 201 extracts one item whose processing execution flag 802 is blank from the missing item list 800 acquired in S401.

[0075] In S405, the CPU 201 performs OCR processing on the area corresponding to the item acquired in S404. As for the specific method of OCR, the same method as in S304 to S305 may be used, or a different method may be used.

[0076] As described above, by performing OCR processing using correction range A, which is a portion of the OCR target image rather than the entire image, OCR processing is performed on relatively large characters, which is expected to improve OCR processing accuracy. As a result, OCR processing can be performed on character strings that would have been missed if OCR processing had been performed on the entire OCR target image. Furthermore, since correction range A is an area that has been specified by the user in the past on the same type of form, it can be said that it is an area that is prone to being missed on that type of form. Since OCR processing can be performed on such areas without the user having to specify it each time, it is possible to reduce user effort and achieve efficient OCR processing.

[0077] In S406, the CPU 201 sets a flag in the processing execution flag 802 field corresponding to the item for which re-OCR has been executed.

[0078] In S407, the CPU 201 determines whether the re-OCR in S405 was successful. If the process was successful, the process proceeds to S408. If not, the process in S408 is skipped and the process proceeds to S409.

[0079] In S408, the CPU 201 deletes the successfully processed item from the missing item list 800.

[0080] When re-OCR has been performed on all items in the list, the process of this flowchart ends. If there are items that have not yet been re-OCRed, the process from S404 onwards is performed on the next item. When all items in the missing item list 800 have been deleted by repeating the processes of S404 to S408, all items that were missing from OCR in the OCR target image have been OCR'd. On the other hand, even if missing items remain in the missing item list 800, if the "processed" flag is set for all items, it means that OCR processing has been attempted for all missing items.

[0081] Note that the number of times the OCR process in step S405 can fail (hereinafter referred to as the maximum failure count) may be set in advance, and if the number of times the OCR process fails exceeds the maximum failure count, the loop may be terminated even if OCR process has not been performed on the areas corresponding to all missing items. That is, if the number of times the CPU 201 performs OCR process on the areas corresponding to the items acquired in step S404 but fails to detect characters in the areas exceeds the maximum failure count, the process proceeds to a step in which the user inputs correction range B and performs OCR process using that correction range B ( FIG. 5 ). For example, if the maximum failure count is set to two, OCR process is performed three times on the areas corresponding to the three items in missing item list 800—"Customer Name," "Bill Amount," and "Subtotal"—but no characters are detected. In this case, OCR process has not yet been performed on the areas corresponding to the missing items after "Total Amount," but because the OCR process has failed twice, the maximum failure count is exceeded. Therefore, the CPU 201 terminates the loop when the third OCR process on the area corresponding to the "Subtotal" item fails. The loop is terminated midway because if OCR processing fails multiple times, it is unlikely that OCR processing will be successful in areas corresponding to other missing items. By setting the number of attempts, the user can move on to the process in Figure 5 without waiting until OCR processing is successful (even though the chances of success are low), which reduces the work time.

[0082] In the process shown in Figure 4, if corrections have previously been made to a document of the same type as the document to be OCRed, the corrections can be made using the recorded correction range A, which reduces the number of areas that the user needs to correct manually, thereby reducing the burden on the user.

[0083] This concludes the explanation of Figure 4.

[0084] FIG. 5 is a flowchart showing an example of the flow of OCR processing for the correction range B received from the user.

[0085] In S501, the CPU 201 displays the results of the OCR processing performed in S304 to S307 on the image to be subjected to OCR on the screen. An example of the screen displaying the results of the OCR processing is shown in FIG.

[0086] In S502, the CPU 201 accepts an operation by the user to specify the correction range B or an operation to end the specification of the correction range B.

[0087] In S503, the CPU 201 determines whether the user operation received in S502 is an operation to end the designation of the correction range B. If it is an operation to end, the process proceeds to S508. If it is not an operation to designate the correction range, the process proceeds to S504.

[0088] The operation to end the specification of the correction range is specifically an operation to press, for example, by clicking the end button 604 on the screen shown in FIG.

[0089] The input operation for correction range B can be accepted, for example, by a drag operation using a mouse. Specifically, an instruction for a diagonal line of a rectangle related to correction range B is accepted by the drag operation, and the rectangle with the diagonal line from the start point to the end point of the drag operation is set as correction range B. Alternatively, a range specified based on a user operation may be set as correction range B, such as an area related to a figure that has been freely drawn by a drag operation. In this way, by allowing the user to specify / select the range that requires OCR processing, it is possible to perform OCR processing without omission for the range that requires OCR processing.

[0090] In S504, it is determined whether the correction range B accepted in S502 is a size suitable for OCR processing. If it is not a suitable size (S504: NO), the process proceeds to S505, and if it is a suitable size (S504: YES), the process proceeds to S506.

[0091] One method for determining whether correction range B is a suitable size for OCR processing is to set a reference value in advance, and determine that the size is suitable if it meets the reference value, and not suitable if it does not. As mentioned above, since OCR may fail if the size of the characters to be OCRed is relatively small, for example, a reference value A as an upper limit of the size may be set, and if the size of correction range B is less than reference value A, it may be determined to be a suitable size, and if it is equal to or greater than reference value A, it may be determined to be unsuitable. Alternatively, a reference value B as a lower limit of the size may be set, and if the size of correction range B is equal to or greater than reference value B, it may be determined to be suitable, and if it is less than reference value B, it may be determined to be unsuitable. These two reference values ​​A and B may be set, and if the size of correction range B is less than reference value A and equal to or greater than reference value B, it may be determined to be suitable, and if it is equal to or greater than reference value A or less than reference value B, it may be determined to be unsuitable.

[0092] In S505, the CPU 201 displays correction range B in a form (a form according to the determination result) that allows the user to recognize that it is not suitable for OCR processing. Specifically, for example, the frame line representing correction range B or the area of ​​correction range B may be displayed in a predetermined color (for example, red), or a notification in the form of text such as "inappropriate" or "NG" may be displayed near correction range B. Also, a display urging the user to make the size appropriate (for example, displaying text such as "bigger" or "smaller") may be displayed.

[0093] In S506, the CPU 201 displays the correction range B in a form that allows the user to recognize that the correction range B is suitable for OCR processing. Specifically, for example, the frame line representing the correction range B or the area of ​​the correction range B may be displayed in a predetermined color (e.g., blue, a color different from that displayed when the correction range B is not suitable), or text such as "suitable" or "OK" may be displayed near the correction range B.

[0094] In this embodiment, the configuration is such that the judgment is made as either suitable or unsuitable (two options), but the judgment may be made in three stages such as "optimal," "suitable," and "unsuitable," or the degree of suitability may be expressed numerically (for example, 1 to 100). When the above-mentioned color-based identification display is used, three stages may be used, such as three colors "blue," "yellow," and "red," or when expressed numerically, the appropriateness may be expressed in a recognizable manner by a method such as displaying a gradation from blue to red.

[0095] In this way, by displaying a display that allows the user to recognize whether the size of correction range B is appropriate, the user will try to adjust the size of correction range B so that it is appropriate, thereby improving the accuracy of OCR processing for correction range B.

[0096] The method of notifying the user whether the size of the correction range B is appropriate or inappropriate may be a method other than displaying using color or text, and may be, for example, a method of notifying the user using sound. Specifically, if the size of the correction range B is inappropriate, a method of notifying the user by voice such as "The selected range is too large. Please reduce it," and if the size of the correction range B is appropriate, a method of notifying the user by voice such as "The selected range is appropriate" may be considered. Also, if the size of the correction range B is inappropriate, an alert sound may be sounded. Any method may be used as long as it notifies the user that the correction range B selected by the user is appropriate or inappropriate.

[0097] In S507, CPU 201 performs OCR on correction range B. The specific OCR method may be the same as in S304 to S305, or a different method. If correction range B is determined to be an appropriate size in S504, OCR is performed. However, if correction range B is determined to be an inappropriate size, OCR may be performed as is, even though it is not appropriate, or a notification may be given to reset correction range B without performing OCR. Furthermore, if OCR is performed, a notification may be given that accuracy may be reduced, and the user may be asked to choose whether or not to perform OCR anyway.

[0098] In S508, correction range B on which OCR has been performed is saved in the file server 103. The items saved at this time are the coordinate information of correction range B, the type of image to be OCRed, and the type of item that was recognized by OCR in correction range B. Correction range B saved by the file server 103 will be used the next time the same type of form is OCRed. In other words, it will be used as correction range A in the process shown in FIG. 4.

[0099] As described above, the flowchart in FIG. 5 shows a process flow in which input of correction range B is accepted (S502), followed by determining whether correction range B is an appropriate size (S504) and displaying the determination result in a recognizable manner (S505, S506). However, it is also possible to determine whether the size of correction range B is appropriate (performing the process of S504) while accepting input of correction range B (e.g., while a drag operation is being performed, i.e., from the start to the end of the accepting operation) and display the determination result while accepting input of correction range B (performing the process of S505, S506). By notifying the user of whether the size of correction range B is appropriate while accepting input in this way, the user can recognize the appropriate size while entering the input. This prevents the user from having to re-enter the input because the size was not appropriate after completing the input, thereby reducing the hassle of having to re-enter the input.

[0100] Furthermore, OCR processing for correction range B (S507) may also be performed while the input of correction range B is being accepted. Specifically, when the input of correction range B is accepted by a drag operation, once the drag operation is started, acquisition of the area specified by the operation is repeated at predetermined intervals (for example, every 0.1 seconds) until the operation is completed. Then, for each acquired area, OCR processing (which may be just character detection processing, as described below) is performed. Then, the area related to the character string detected by the OCR processing is displayed identifiably, for example, by enclosing it in a rectangle. Furthermore, the result of the OCR processing (character strings resulting from character recognition) may be displayed near the rectangular area.

[0101] Even when performing OCR in real time like this, OCR may be performed if the size of correction range B is suitable for OCR processing, and not performed if the size is not suitable. For example, OCR may not be performed from the time input acceptance for correction range B begins until correction range B becomes a size suitable for OCR processing, but OCR may be performed once the size becomes suitable. If the size of correction range B again becomes a size that is not suitable for OCR processing, OCR may not be performed.

[0102] By recognizably displaying the character string that has been OCR processed in real time in this way, the user can understand whether OCR processing is being performed on the character string that they want to OCR (whether the characters are detected and recognized appropriately). As a result, the user can set correction range B so that the character string that they want to OCR is detected and recognized, reducing the effort of having to reset correction range B.

[0103] Since OCR processing is a high-load process, it is also possible to control the system so that only character detection is performed and character recognition is not performed while the input of correction range B is being accepted, and character recognition is performed after the acceptance of the input has finished.Also, for the OCR processing performed while the input of correction range B is being accepted, an OCR engine with low detection and recognition accuracy but high processing speed and low processing load may be used, and after correction range B is determined, an OCR engine with high detection and recognition accuracy (but slow processing speed and heavy processing load) may be used.

[0104] FIG. 9 shows how the input operation for correction range B is being accepted.

[0105] Fig. 9 is an example of an input screen for correction range B using the screen of Fig. 6. Fig. 9 is based on Fig. 6, and some screens are omitted for the sake of explanation.

[0106] Figure 9-1 is a diagram showing a screen displaying the OCR target image 601 before the user specifies the correction range B. This screen displays the results of the processing in steps S304 and S305 (as explained in Figure 6, character strings for which OCR has been performed are indicated by dashed rectangles), and the user can check on this screen character strings on the form that have not been OCRed.

[0107] FIG. 9-2 illustrates a situation in which a user specifies a correction range B for a character string for which OCR has not been performed. A rectangle 901 indicating the specified correction range B is displayed on the screen. In FIG. 9, a border 902 of the rectangle 901 is displayed in a different shape from the border surrounding the character string 602 for which OCR has been performed (although a dashed line is used in FIG. 9, the rectangle 901 may be distinguished from the rectangle representing the character string 602 by, for example, a different border thickness, a different spacing between dots in the dotted line, or a solid or dotted line). By distinguishing the border shape, the user can recognize which part of the OCR target image 601 is being specified as correction range B without confusing it with the border surrounding the character string 602 for which OCR has been performed or the border 903. The interior of the specified rectangle 901 may also be displayed in any color. This allows the user to recognize which part of the form 601 is being specified as correction range B. A frame 903 is a dotted line that displays the results of the OCR process (S507) performed on the correction range B while the input for the correction range B is being accepted. The character string enclosed by the frame 903 is the character string for which the OCR process was successful while the input for the correction range B was being accepted. The shape of the frame 903 may be any shape that can be distinguished from the frame 902 and the dotted line surrounding the character string 603, and the frame may be made to stand out by blinking or adding a color to the frame. By distinguishing the shape of the frame, the user can recognize which character string in the OCR target image 601 has been successfully OCR processed, without confusing it with the frame 902 or the frame surrounding the character string 602 for which OCR has already been performed.

[0108] FIG. 9-3 is an example of a screen displaying the results of OCR performed in S507 on correction range B. Character strings 904 to 906 have their dotted frame lines 903 changed to dashed lines, indicating that the OCR was successful.

[0109] FIG. 10 is a diagram illustrating an example of a data table for the correction range A. As shown in FIG.

[0110] Data table 1000 in Figure 10 is a data table in which correction range B, for which OCR was successful, is saved as correction range A to be used when OCRing the same type of document from the next time onwards (S508), and consists of document type 1001, item name 1002, and coordinate information 1003.

[0111] The form type 1001 indicates the type of form from which the items and coordinate information for which OCR processing was successful in S507 were OCRed, and is composed of a combination of the form name and alphabet. Since the coordinate information for items varies depending on the type of invoice, even for the same invoice, the type is categorized by combining the form name and alphabet, such as invoice AAA or invoice AAB. Note that the code used to categorize the form does not have to be alphabetic; any other character string, number, symbol, or other character can be used as long as it can identify the type of form. In the data table 1000 in Figure 10, "Customer Name," "Message from Person in Charge," and "Invoice Amount" are OCRed from Invoice AAA, while "Total Amount" and "Person in Charge Name" are OCRed from Receipt AAC.

[0112] Item name 1002 is the name of the item for which OCR was successful in S507. In the data table 1000 in Fig. 10, OCR was successful for the business partner name, message from the person in charge, and billing amount from invoice AAA, and for the person in charge contact information and person in charge name from receipt AAC.

[0113] Coordinate information 1003 is coordinate information indicating where on the OCR target image the character string corresponding to the item for which OCR processing was successful in S507 was located. Specifically, since this is coordinate information indicating the rectangular area in which the character string existed, it is coordinate information for at least two points, such as the top left and bottom right vertices of the rectangular area. Furthermore, the coordinate information may be expressed as a ratio (e.g., (0.8, 0.15)) when the values ​​of the vertical and horizontal widths of the form are 1. Because the size of the image acquired by the CPU 201 in S301 is not constant, expressing the coordinate information as a ratio can accommodate different sizes.

[0114] Depending on the correction range B entered by the user, it is possible that some characters within the correction range B will be missed by OCR. In this case, the user may be able to perform OCR by entering an additional correction range B that surrounds the character string that has been missed. By repeating this process until there are no more characters that have been missed by OCR, the character strings on the form that have been missed by OCR may be corrected.

[0115] As described above, according to this embodiment, it is possible to reduce the burden on the user in correcting the OCR result.

[0116] The present invention can be embodied as, for example, a system, an apparatus, a method, a program, a recording medium, etc. Specifically, the present invention may be applied to a system consisting of multiple devices, or may be applied to an apparatus consisting of a single device.

[0117] The various controls described above as being performed by CPU 201 may be performed by a single piece of hardware, or the entire device may be controlled by multiple pieces of hardware (e.g., multiple processors or circuits) sharing the processing.

[0118] Furthermore, although the present invention has been described in detail based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0119] In the above-described embodiment, the present invention has been described as being applied to a PC, but the present invention is not limited to this example and can be applied to any device or software capable of executing OCR processing. For example, the present invention can be applied to PDAs, mobile phone terminals (smartphones), wearable terminals, tablet terminals, game consoles, etc.

[0120] (Other embodiments) The present invention can also be realized by executing the following process. That is, software (programs) that realize the functions of the above-described embodiments are supplied to a system or device via a network or various storage media, and the computer (or CPU, MPU, etc.) of the system or device reads and executes the program code. In this case, the program and the storage media storing the program constitute the present invention.

Claims

1. an accepting means for accepting a specification of a range from which characters are to be acquired from a document; a determining means for determining whether the size of the range specified by the receiving means satisfies a predetermined standard set based on the size of characters for the range from which characters are to be acquired; a notification means for notifying a result of the determination by the determination means; An information processing device comprising:

2. A display control means for controlling the display of the range accepted by the accepting means together with the result of character acquisition within the range, in order to correct the range accepted by the accepting means to a range in which character acquisition is successful; The information processing apparatus according to claim 1 , further comprising:

3. The receiving means receives the designation of the range on a screen that displays the result of character acquisition.

2. The information processing device according to claim 1,

4. A storage means for storing the range accepted by the accepting means; an acquisition means for acquiring an acquisition result of characters in the range before the specification by the acceptance means is accepted; Furthermore, 2. The information processing apparatus according to claim 1, wherein the storage means stores the range accepted by the accepting means as the range of character acquisition in the document, depending on the character acquisition result acquired by the acquiring means.

5. The storage means stores the range in association with a document type, 5. The information processing apparatus according to claim 4, wherein the storage means stores the range accepted by the accepting means as the range of characters to be acquired from the document corresponding to the type.

6. An item specifying means for specifying an item from which characters are to be acquired; Furthermore, 2. The information processing apparatus according to claim 1, wherein the accepting unit accepts the specification of the range relating to the item for which character acquisition is not performed.

7. The range to be acquired is a target range for identifying an image for acquiring characters by OCR, The receiving means receives a designation of a target range for OCR; 7. The information processing apparatus according to claim 1, wherein the determining means determines whether the size of the OCR target range specified by the receiving means satisfies a predetermined standard set based on the size of characters relative to the OCR range.

8. further comprising a detection means for detecting characters from the document; The detecting means detects characters from the range determined by the determining means to satisfy the predetermined criterion.

2. The information processing device according to claim 1,

9. The detecting means does not detect characters from the range determined by the determining means not to satisfy the predetermined criterion.

9. The information processing device according to claim 8,

10. the predetermined standard is a first standard value indicating an upper limit of the magnitude of the range; The determining means determines whether the size of the range specified by the accepting means is equal to or smaller than the first reference value.

2. The information processing device according to claim 1,

11. the predetermined standard is a second standard value indicating a lower limit of the size of the range, The determining means determines whether the size of the range specified by the accepting means is equal to or larger than the second reference value.

2. The information processing device according to claim 1,

12. the predetermined criteria are a first reference value indicating an upper limit value of the size of the range and a second reference value indicating a lower limit value of the size of the range, The determining means determines whether the size of the range specified by the accepting means is equal to or smaller than the first reference value and equal to or larger than the second reference value.

2. The information processing device according to claim 1,

13. The notification means notifies the determination result by displaying the range, the designation of which has been accepted by the acceptance means, in a manner corresponding to the determination result by the determination means.

2. The information processing device according to claim 1,

14. The notification means notifies the determination result by changing the color of the range depending on whether the predetermined standard is met or not.

2. The information processing device according to claim 1,

15. The notification means notifies the determination result by changing the color of the frame line indicating the range depending on whether the predetermined criterion is satisfied or not.

2. The information processing device according to claim 1,

16. The notification means issues a notification based on the satisfaction of the predetermined criterion when the size of the range satisfies the predetermined criterion.

2. The information processing device according to claim 1,

17. When the size of the range does not satisfy the predetermined standard, the notification means issues a notification based on the fact that the predetermined standard is not satisfied.

2. The information processing device according to claim 1,

18. The notification means issues a notification urging the user to adjust the range to a size that satisfies the predetermined standard as a notification based on the fact that the predetermined standard is not satisfied. The information processing device according to claim 17,

19. The notification means displays the range specified by the reception means in a manner according to the size of the range, thereby notifying the user so that the user can recognize whether the size of the range is suitable for OCR processing.

2. The information processing device according to claim 1,

20. a receiving step in which a receiving means of the information processing device receives a specification of a range from which characters are to be acquired from the document; a determination step in which a determination means of the information processing device determines whether the size of the range specified in the receiving step satisfies a predetermined standard set based on the size of characters for the range from which characters are to be acquired; a notification step in which a notification means of the information processing device notifies a result of the determination step; An information processing method comprising:

21. A program for causing a computer to function as each of the means according to any one of claims 1 to 6.

22. A program for causing a computer to function as each of the means described in claim 7.

23. A program for causing a computer to function as each of the means described in any one of claims 8 to 19.

Citation Information

Patent Citations

  • Format parameter generating method for ocr

    JP1997034989A

  • Information processor, character recognizing program, and recording medium

    JP2005043995A

  • Character recognition device and program

    JP2017138703A

  • Data input assistance device, data input assistance method and program

    JP2022011019A