Image forming apparatus
The image forming apparatus addresses the issue of missed item value detections in unstructured documents by displaying confidence scores and enabling separate corrections, improving extraction accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-10-03
- Publication Date
- 2026-04-15
AI Technical Summary
Existing image recognition systems fail to directly prompt users to confirm the extraction position of item values, leading to detection failures, especially in unstructured documents where item names are absent.
An image forming apparatus with display control means to show position and OCR confidence scores, enabling separate correction of string positions and characters based on confidence levels, using a position correction and string correction mechanism.
Enhances the accuracy of string extraction by allowing direct correction of positions and characters, reducing the likelihood of missed detections and shortening correction times.
Smart Images

Figure 2026065448000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image forming apparatus, an information processing system, a control method of an image forming apparatus, and a program.
Background Art
[0002] There is a system that performs information extraction from a document image using an optical character recognition technique and displays the extraction result together with the certainty of the extraction result, so that only suspicious extraction results are confirmed and corrected by the user, reducing the burden on the user for confirmation and correction.
[0003] In an unformatted form such as an invoice or estimate in which the position of the total amount changes between multiple companies or the detail lines change every time, such as an invoice or estimate, since the description position of the item to be extracted is not fixed, two patterns of assumed errors can be considered. One is a pattern in which the character recognition of the character string to be extracted is incorrect, and the other is a pattern in which the position of the item to be extracted is incorrect. In order to reduce the burden on the user for confirmation and correction, since these two error patterns have different appropriate correction methods, they need to be presented to the user separately.
[0004] Patent Document 1 discloses a technique for information extraction based on the positional relationship between an item value and an item name. By separately highlighting a suspicious item value and an item name, it shows the user whether confirmation should be centered on the confirmation of the character string of the item value, or whether it is necessary to confirm the extraction position of the item name.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] The technology described in Patent Document 1 indirectly prompts the user to confirm the extraction position of item values by prompting them to confirm the extraction position of item names, but it does not directly prompt the user to confirm the extraction position of item values, which is the information that is actually desired. Therefore, even if the extraction position of the item name is correct, detection failures occur when the extraction position of the item value is incorrect. Furthermore, it is difficult to prompt the user to confirm the extraction position of item values that do not have item names, such as titles.
[0007] The purpose of this disclosure is to enable the display of information for correcting the position of strings extracted from image data. [Means for solving the problem]
[0008] The image forming apparatus includes a display control means that controls the display of information for correcting the position of a string extracted from image data, or information for correcting a string extracted from image data, based on a position confidence score indicating the certainty of the position of a string extracted from image data and an OCR confidence score indicating the certainty of the extracted string; a position correction means that corrects the position of a string extracted from image data based on an instruction to correct the position of a string extracted from image data; and a string correction means that corrects a string extracted from image data based on an instruction to correct a string extracted from image data. [Effects of the Invention]
[0009] According to this disclosure, information for correcting the position of strings extracted from image data can be displayed. [Brief explanation of the drawing]
[0010] [Figure 1] This is a diagram showing an example of the configuration of an information processing system. [Figure 2] This figure shows an example of the hardware configuration of an information processing system. [Figure 3] This figure shows an example of the functional configuration of an information processing system. [Figure 4]It is a flowchart for explaining the entire processing of an information processing system. [Figure 5] It is a diagram showing an example of a confirmation / correction screen. [Figure 6] It is a diagram showing an example of a correction icon and the screen display of the correction icon. [Figure 7] It is a diagram showing an example of a character correction screen. [Figure 8] It is a diagram showing an example of a screen where OCR errors have been corrected. [Figure 9] It is a diagram showing an example of a position correction screen. [Figure 10] It is a diagram showing an example of a character correction screen. [Figure 11] It is a diagram showing an example of a screen where extraction position errors have been corrected. [Figure 12] It is a diagram showing an example of a character correction screen. [Figure 13] It is a diagram showing an example of a position correction screen.
Mode for Carrying Out the Invention
[0011] Hereinafter, this embodiment will be described with reference to the drawings. Note that the following embodiments do not limit the scope of the claims, and not all combinations of the features described in the embodiments are essential for the solution means.
[0012] (First Embodiment) <Overall Configuration> FIG. 1 is a diagram showing an example of the configuration of an information processing system 100 in the first embodiment. The information processing system 100 includes an image forming apparatus 110, an image processing server 120, and a storage server 130. These devices and servers are interconnected by a network 140 and can communicate with each other.
[0013] The image forming apparatus 110 of this embodiment can request to transmit the image data of the scanned document to the storage server 130 via the image processing server 120, etc.
[0014] In addition, in this embodiment, the image forming apparatus 110 will be described by taking a multi-functional device having a scanning function, a printing function, a copying function, a FAX function, etc. as an example, but the present invention is not limited to the multi-functional device. Any device having the above scanning function can execute the processes of this embodiment described later.
[0015] Here, the scanning function is a function of transmitting, to the outside, image data generated by reading a document with a scanner included in the image forming apparatus 110. The printing function is a function of printing image data received from a user terminal or the like. The copying function is a function of obtaining a copy of a document by printing the image data of the document read by the scanner. The FAX function is a function of transmitting and receiving image data using a telephone line.
[0016] The image processing server 120 is a server that performs information extraction processing from image data.
[0017] The storage server 130 is a server that stores image data and the processing results of the image processing server 120.
[0018] Note that the information processing system 100 of this embodiment is configured to include the image forming apparatus 110, the image processing server 120, and the storage server 130, but is not limited thereto. For example, the image forming apparatus 110 may also serve as the image processing server 120. Further, the image processing server 120 may also serve as the storage server 1) and may be arranged in a connection form with a server on a LAN instead of on the Internet. Further, the storage server 130 may be replaced with a mail server or the like, and the scanned image may be attached to an email and transmitted. Further, a configuration in which there are a plurality of image processing servers 120 and a plurality of storage servers 130 with respect to the image forming apparatus 110 may also be possible.
[0019] <Hardware Configuration> FIG. 2 is a diagram showing an example of the hardware configuration of the information processing system 100 in this embodiment.
[0020] The image forming apparatus 110 includes a printer 201, a scanner 202, an operating unit 203, a CPU 211, RAM 212, an HDD 213, a network interface 214, a printer interface 215, a scanner interface 216, an operating unit interface 217, and an expansion interface 218.
[0021] The CPU 211 can exchange data with the RAM 212, HDD 213, network I / F 214, printer I / F 215, scanner I / F 216, control unit I / F 217, and expansion I / F 218. Furthermore, the CPU 211 reads instructions (computer programs) from the HDD 213, expands them into the RAM 212, and controls the execution of each process described later by executing the instructions in the RAM 212.
[0022] In this embodiment, one CPU 211 uses one memory (RAM 212 or HDD 213) to execute each process shown in the flowchart described later, but this is not limited to this. For example, multiple CPUs or multiple RAMs or HDDs may work together to execute each process.
[0023] The HDD213 can store instructions that can be executed by the CPU211, setting values used in the image forming apparatus 110, and data related to processing requested by the user.
[0024] RAM212 is an area for temporarily storing instructions read by the CPU211 from the HDD213. RAM212 can also store various data necessary for executing instructions. For example, in image processing, processing can be performed by loading the submitted data into RAM212.
[0025] Network I / F 214 is an interface for communication with network 140. Network I / F 214 can notify CPU 211 that data has been received and can transmit data from RAM 212 to network 140 according to instructions from CPU 211.
[0026] The printer interface 215 can send image data to be printed to the printer 201 according to instructions from the CPU 211, and can also transmit the status of the printer 201 received from the printer 201 to the CPU 211.
[0027] The scanner interface 216 transmits image reading instructions received from the CPU 211 to the scanner 202 and transmits the image data received from the scanner 202 to the CPU 211. The scanner interface 216 can also transmit information about the status of the scanner 202 received from the scanner 202 to the CPU 211.
[0028] The control unit I / F 217 can transmit user instructions made via the control unit 203 to the CPU 211, and can also display screen information for user operation on the control unit 203.
[0029] The expansion I / F 218 is an interface that enables the connection of external devices to the image forming apparatus 110. The expansion I / F 218 is equipped with, for example, a USB (Universal Serial Bus) interface. When an external storage device such as a USB memory is connected to the expansion I / F 218, the image forming apparatus 110 can read data stored in the external storage device and write data to the external storage device.
[0030] Printer 201 can print image data received via printer I / F 215 onto paper, and can also transmit its status to printer I / F 215.
[0031] Scanner 202 can transmit image data obtained by scanning a document (paper) placed on the scanner according to image reading instructions received via scanner I / F 216 to scanner I / F 216. Scanner 202 can also transmit its status to scanner I / F 216.
[0032] The control unit 203 is an interface for issuing various instructions to the image forming apparatus 110 based on user operations. For example, the control unit 203 is equipped with a touch panel liquid crystal screen, which displays an operation screen and accepts user operations.
[0033] The image processing server 120 has a CPU 221, RAM 222, HDD 223, and network I / F 224.
[0034] The CPU 221 controls the entire image processing server 120 and can control the exchange of data between the RAM 222, HDD 223, and network I / F 224. The CPU 221 also reads control programs (instructions) from the HDD 223, loads them into the RAM 222, and executes them.
[0035] The HDD223 of the image processing server 120 is a large-capacity storage unit that stores image data, various programs, and data generated during image processing.
[0036] Network I / F224 is an interface for communication on network 140.
[0037] In this embodiment, one CPU 221 uses one memory (RAM 222 or HDD 223) to execute each process shown in the flowchart described later, but this is not limited to this configuration.
[0038] The storage server 130 has a CPU 231, RAM 232, HDD 233, and network I / F 234.
[0039] The CPU 231 controls the entire storage server 130 and can control the exchange of data between the RAM 232, HDD 233, and network interface 234. The CPU 231 also reads control programs (instructions) from the HDD 233, loads them into the RAM 232, and executes them.
[0040] The HDD 233 of the storage server 130 is capable of storing image data and extracted information received from the image forming apparatus 110.
[0041] Network I / F234 is an interface for communication on network 140.
[0042] In this embodiment, one CPU 231 uses one memory (RAM 232 or HDD 233) to execute each process shown in the flowchart described later, but this is not limited to this configuration.
[0043] <Functional Configuration> Figure 3(a) shows an example of the functional configuration of the image forming apparatus 110. The software (program) of the image forming apparatus 110 is stored in the HDD 213, transferred to the RAM 212, and executed by the CPU 211. This realizes each of the functional components shown in Figure 3(a).
[0044] The image forming apparatus 110 includes an image reading means 311, a position correction determination means 312, a character correction determination means 313, a screen generation means 314, a UI display means 315, an input receiving means 316, and a transmission means 317.
[0045] The image reading means 311 reads the document placed on the scanner 202 using the scanner 202, converts it into image data, and stores it in the HDD 213.
[0046] The position correction determination means 312 determines whether to prompt the user to correct their position based on the position confidence level obtained by the position confidence level acquisition means 322, which will be described later.
[0047] The character correction judgment means 313 determines whether to prompt the user to correct the characters based on the OCR confidence level obtained by the OCR confidence level acquisition means 323, which will be described later.
[0048] The screen generation means 314 generates a screen to be displayed on the operation unit 203, which includes extracted results, image data, user-operated buttons and other operational components, and UI components that display information such as progress status, according to the judgment results of the position correction judgment means 312 and the character correction judgment means 313. The screen generation means 314 also generates a screen corresponding to the input received by the input reception means 316, which will be described later.
[0049] The UI display means 315 displays the screen generated by the screen generation means 314.
[0050] The input receiving means 316 receives input to the UI components displayed on the operation unit 203 by the UI display means 315 and executes processing corresponding to the input.
[0051] The transmission means 317 transmits image data, print attributes, and extracted information stored in the HDD 213 to other devices on the network 140, such as the image processing server 120 and the storage server 130, via the network I / F 214.
[0052] Figure 3(b) shows an example of the functional configuration of the image processing server 120. The software (program) of the image processing server 120 is stored in the HDD 223, transferred to the RAM 222, and executed by the CPU 221. This realizes each of the functional components shown in Figure 3(b).
[0053] The image processing server 120 includes a character extraction result acquisition means 321, a position confidence acquisition means 322, and an OCR confidence acquisition means 323.
[0054] The character extraction result acquisition means 321 performs character extraction on the received image data and acquires the extraction results. In this embodiment, the character extraction result acquisition means 321 is realized by performing entity extraction from the results of OCR (Optical Character Recognition) on the entire image data. Entity extraction, as referred to here, is a technique that automatically extracts strings of specific items, such as company names and dates, and their attributes, using a model trained with Deep Learning using the full-surface OCR results of the image data as training data. Internally, the confidence level of each attribute is calculated as a continuous value for each string, and the string with the highest confidence level for each attribute from the entire string is output as the extraction result.
[0055] However, the method of character extraction is not particularly limited, and known techniques can be used. For example, key-value extraction, which finds a string that corresponds to the key of the information to be extracted from the strings that exist in the data and detects strings adjacent to that key string as values, or string search using pattern matching can be used. The acquired extraction results and full OCR results are stored on HDD223. Here, the extraction results include the extracted string, the position information of the extracted string, the attributes of the character, language information, and the confidence level of the character output by the OCR engine and the confidence level of each item output by entity extraction.
[0056] The position confidence acquisition means 322 calculates the position confidence of the extracted results extracted by the character extraction result acquisition means 321 and acquires that position confidence. Here, position confidence represents the certainty of the position of the extracted item. In this embodiment, the position confidence takes a continuous value from 0 to 1. Position confidence below an arbitrary threshold is classified as "low confidence," and position confidence above an arbitrary threshold is classified as "high confidence." There may be multiple thresholds, in which case three or more types of labels will be output.
[0057] Furthermore, if the position confidence cannot be calculated, it may be recorded as "confidence unknown." Also, the position confidence is not a continuous value, and labels such as "low confidence" or "high confidence" may be directly output as a result of classifying whether or not a condition is met. For example, a position confidence of "low confidence" indicates that there is a high possibility that the position of the extracted string is incorrect. Note that whether the extracted string matches the string written in the actual image data is irrelevant to the position confidence. In this embodiment, the reciprocal of the number of items with a confidence level exceeding an arbitrary threshold among the item candidates output by entity extraction is used as the position confidence, with the more items exceeding the threshold, the less confident the result is.
[0058] However, other methods for calculating position confidence are also acceptable, and the method is not limited. For example, the position confidence can be calculated based on the degree of match with the position information of items previously approved by the user in image data with the same layout. Alternatively, the position confidence can be calculated based on the positional relationship between the extracted string and the surrounding strings. The position confidence can also be calculated by combining the above methods.
[0059] The OCR confidence level acquisition means 323 calculates the OCR confidence level of the extraction results extracted by the character extraction result acquisition means 321 and acquires that OCR confidence level. Here, the OCR confidence level represents the certainty of the string of the extracted item. In this embodiment, the OCR confidence level takes a continuous value from 0 to 1. When the OCR confidence level falls below an arbitrary threshold, it is classified as "low confidence," and when the OCR confidence level exceeds an arbitrary threshold, it is classified as "high confidence." There may be multiple thresholds, in which case three or more types of labels will be output.
[0060] Furthermore, if the OCR confidence level cannot be calculated, it may be labeled as "confidence level unknown." Also, the OCR confidence level is not a continuous value, and labels such as "low confidence level" or "high confidence level" may be directly output as a result of classifying whether or not the conditions were met. For example, an OCR confidence level of "low confidence level" indicates that there is a high possibility that the extracted string is incorrect. Note that whether the extracted item was actually the item that should have been extracted is irrelevant to the OCR confidence level. In this embodiment, a second extracted string is obtained using a second OCR engine different from the first OCR engine used in the character extraction result acquisition means 321, and the OCR confidence level is calculated from the degree of matching by comparing it with the first extracted string.
[0061] However, other methods for calculating OCR confidence are also acceptable, and the method is not limited. For example, the OCR confidence of the entire string can be calculated based on the confidence of each character output by the first OCR engine. Alternatively, the OCR confidence can be calculated using strings of items previously approved by the user in image data with the same layout, based on the degree of match with those strings, whether the type of the item's attributes (such as company name or date) matches, or whether it contains a specific string. The OCR confidence can also be calculated by combining the above methods.
[0062] <Overall processing flow> This embodiment describes a system in which scanned image data is automatically filenamed using strings within the image data when it is converted into a file and sent, and the user confirms and corrects the filename.
[0063] Furthermore, the image data to be processed is assumed to be unstructured forms with no fixed layout. Since the position of the items to be extracted is not fixed in unstructured forms, two types of errors are possible. One is a pattern where the character recognition of the string to be extracted is incorrect (OCR error), and the other is a pattern where the position of the items to be extracted is incorrect (extraction position error).
[0064] To verify and correct OCR errors, one only needs to examine an image of a portion of the text. However, to verify and correct errors in extraction position, one needs to examine the entire image to determine the original target item. To reduce the burden on users in verifying and correcting errors, these two error patterns should be presented to the user separately, and the optimal correction method should also be suggested.
[0065] Figure 4 is a flowchart showing the processing flow from the scanning image data in the image forming apparatus 110 to the file and transmission to the storage server 130. This section will focus on the communication between each device. The control methods for the image forming apparatus 110 and the image processing server 120 will be described below.
[0066] In its normal state, the image forming apparatus 110 displays a main screen on the touch panel of the operation unit 203, which contains buttons for performing each of the functions it provides.
[0067] By installing an additional application on the image forming apparatus 110 for sending image data to the image processing server 120, a button to use the application's functions appears on the main screen of the image forming apparatus 110. When this button is pressed, a screen for sending the scanned form to the image processing server 120 is displayed, and when the "Start Scan" button is pressed, the process shown in Figure 4 is performed.
[0068] In step S401, when the "Start Scan" button is pressed, the image reading means 311 in the image forming apparatus 110 reads the document placed on the scanner 202 and generates image data according to the various scan settings set on the scan settings screen.
[0069] In step S402, the image forming apparatus 110 transmits the image data generated in step S401 to the image processing server 120 via the transmission means 317.
[0070] In step S403, the character extraction result acquisition means 321 in the image processing server 120 performs character extraction on the image data transmitted in step S402 and acquires the extraction result.
[0071] In step S404, the position confidence acquisition means 322 in the image processing server 120 calculates the position confidence for the extraction results extracted in step S403 and acquires the position confidence. The position confidence indicates the certainty of the position of the string extracted from the image data.
[0072] In step S405, the OCR confidence level acquisition means 323 in the image processing server 120 calculates the OCR confidence level for the extraction results extracted in step S403 and acquires the OCR confidence level. The OCR confidence level indicates the likelihood of the string extracted from the image data.
[0073] In step S406, the position correction determination means 312 in the image forming apparatus 110 determines whether to prompt the user to correct the position for each extraction result, according to the position confidence obtained in step S404.
[0074] In step S407, the character correction determination means 313 in the image forming apparatus 110 determines whether to prompt the user to correct the characters for each extraction result, according to the OCR confidence level obtained in step S405.
[0075] In step S408, in the image forming apparatus 110, the screen generation means 314 generates a screen that includes the image data generated in step S401, the extraction results extracted in step S403, operation components such as icons prompting the use of the correction means and buttons for user operation, and UI components that display information such as progress status, according to the determination result determined by the position correction determination means 312 in step S406 and the determination result determined by the character correction determination means 313 in step S407. The screen generation means 314 also generates a screen corresponding to the input received by the input receiving means 316.
[0076] In step S409, the UI display means 315 in the image forming apparatus 110 functions as a display control means and controls the display of the screen generated by the screen generation means 314. The file name, which was automatically generated from the extraction results extracted in step S403, is displayed on the initial screen of the confirmation / correction screen. Details of the file name confirmation and correction operations performed by the user will be described later.
[0077] In step S410, the input receiving means 316 in the image forming apparatus 110 receives input to the operating component displayed in step S409 and executes processing corresponding to the input. Examples include operations to correct the file name, instructions to transition to a confirmation / correction screen, and instructions to send the generated file.
[0078] In step S411, the input receiving means 316 in the image forming apparatus 110 determines whether the input received in step S410 is a correction instruction operation or not.
[0079] If a correction instruction is received in step S411, the process proceeds to step S412.
[0080] In step S412, the input receiving means 316 determines whether the correction instruction is for position correction or character correction.
[0081] If step S412 is a position correction, proceed to step S413.
[0082] In step S413, the input receiving means 316 displays a position correction screen to the user and accepts corrections to the position of the extraction result extracted in step S403.
[0083] If the correction in step S412 is text, proceed to step S414.
[0084] In step S414, the input receiving means 316 displays a character correction screen to the user and accepts corrections to the characters extracted in step S403.
[0085] In both steps S413 and S414, once the acceptance of the correction is complete, the process returns to step S408, and the screen generation means 314 generates a screen that reflects the correction instructions received from the user.
[0086] Returning to step S411, we will continue the explanation. If an operation other than a correction instruction is received in step S411, proceed to step S415.
[0087] In step S415, the input receiving means 316 determines whether the received input is an operation other than pressing the "Send" button.
[0088] If, in step S415, an input other than pressing the "Send" button is received, the process returns to step S408, and the screen generation means 314 generates a screen corresponding to the input.
[0089] If the "Send" button is pressed in step S415, the process proceeds to step S416.
[0090] In step S416, the transmission means 317 sends the image data file generated in step S401, the extraction results extracted in step S403, or the extraction results modified by the user in step S413 or S414 to the storage server 130.
[0091] The CPU 231 of the storage server 130 receives the files and other data sent in step S416 and saves them to the HDD 233.
[0092] This concludes the series of steps shown in Figure 4.
[0093] <Details of the confirmation / correction screen> Here, we will explain in detail the display screen and its transitions in the process performed by the UI display means 315 in step S409 of Figure 4, the process performed by the input receiving means 316 in steps S401 to S415, and the process performed by the transmission means 317 in step S416.
[0094] In this embodiment, the confirmation and correction of the file names of scanned image data are performed on the touch panel. In the following processes, the UI display means 315 handles the display on the touch panel, the input receiving means 316 handles the reception of user input to the control components and processing of that input, and the transmission means 317 handles the transmission process.
[0095] Figure 5 shows an example of the confirmation / correction screen 501 presented to the user. The confirmation / correction screen 501 is displayed as the initial screen in step S409. In the file name field 502, the string formed by concatenating the extraction results extracted in step S403 with underscores is displayed as the file name.
[0096] The file name in file name field 502 is an example of a name that is a combination of multiple strings extracted from the image data, and is the file name of the image data.
[0097] The scanned image display screen 503 is a screen that can enlarge and reduce the image data scanned by the image forming apparatus 110. In Figure 5, the top part of the quotation is shown enlarged. Extracted items 504 to 506 are items on the scanned image display screen 503 that correspond to the string displayed in the file name field 502.
[0098] The extracted items 504, which have a low OCR confidence level, and 505, which have a low positional confidence level, are highlighted in different colors compared to extracted item 506, which has a high confidence level. Items with an unknown OCR confidence level or positional confidence level may also be highlighted in a similar manner.
[0099] Extracted string 507 is the string that is highlighted by a change in color in the file name field 502. Extracted string 507, like extracted items 504 and 505, is highlighted because its confidence level is "low".
[0100] The back button 508 is a button that transitions to the previous screen, and in this embodiment, it transitions to the scan start screen.
[0101] The send button 509, when pressed, completes the confirmation and correction process and sends the data to the storage server 130. In this embodiment, it confirms the file name, sends the image data to the storage server 130, and transitions to the transmission completion screen. Pressing the "Next" button on the transmission completion screen transitions to the scan start screen.
[0102] Figure 6(a) is a magnified view of the icons displayed on the screen. Icon group 601 is displayed on the screen when an extracted string displayed in the file name field 502 is selected. Icon group 601 may also be displayed when an item is selected on the scanned image display screen 503.
[0103] The position correction icon 602, when pressed, transitions to the extraction position correction screen. If the confidence level of the selected extraction string's position is "low confidence," it is highlighted to prompt the user to make corrections. For example, the icon's line becomes thicker and its color changes. Note that either one of these actions is sufficient.
[0104] The text correction icon 603, when pressed, transitions to the text correction screen. If the OCR confidence level of the selected extracted text is "low confidence," it is highlighted to prompt the user to make corrections. For example, the line becomes thicker and the color changes. Note that either one of these actions is sufficient.
[0105] The deselection icon 604 is an icon used to deselect an item.
[0106] Simply displaying icons like those shown in icon group 601 to the user does not allow the user to immediately understand which correction method to choose. Therefore, by highlighting the icons according to the position confidence level and OCR confidence level to prompt the user to choose a correction method, the user can immediately understand which correction method to select and shorten the time it takes to make corrections. Alternatively, explanations of the icons and text prompting corrections can be displayed simultaneously with the icons.
[0107] Here, let's consider the case where both the position confidence level and the OCR confidence level are "low confidence." While it would be possible to highlight both the position correction icon 602 and the text correction icon 603, there is a high possibility of an extraction position error, so correcting the extracted string on the text correction screen first would likely be pointless. Therefore, when both confidence levels are "low confidence," the position confidence level takes precedence over the OCR confidence level. As a result, both icons will not be highlighted simultaneously; only the position correction icon 602 will be highlighted, and the text correction icon 603 will not. Similarly, items with "uncertain confidence" for both OCR confidence and position confidence levels are also highlighted to encourage the user to make corrections.
[0108] Figure 6(b) shows the screen where the extracted string 605, which has an OCR confidence level of "low confidence," has been selected. Because the OCR confidence level is "low confidence," the character correction icon 603 is highlighted. Also, the extracted string 605 is displayed in bold to indicate that it is selected. The extracted string 605 actually has an OCR error, mistaking "quotation" for "estimate." Because the extracted string 605 is selected, the area around the extracted item 606 is highlighted in bold. The highlighting of the character correction icon 603 is an example of information for correcting the string extracted from the image data, and is displayed in step S409.
[0109] Figure 6(c) shows the screen where the extracted string 607, which has a position confidence level of "low confidence," has been selected. Because the position confidence level is "low confidence," the position correction icon 602 is highlighted. Also, the extracted string 607 is displayed in bold to indicate that it is selected. The extracted string 607 has an extraction position error, and the correct extraction position is "×× Corporation" instead of "△△ Corporation." The extracted item 608 is highlighted in bold around the area because the extracted string 607 is selected. The highlighting of the position correction icon 602 is an example of information for correcting the position of the string to be extracted from the image data, and is displayed in step S409.
[0110] Here, we will explain the process of correcting OCR errors using Figures 7 and 8.
[0111] Figure 7 shows an example of a text correction screen. The text correction screen 701 is displayed by selecting the highlighted text correction icon 603 from the icon group 601 in Figure 6(b). The text input field 702 displays the user input received by the input receiving means 316 via the soft keyboard 703. Due to the user's operation, the correctly corrected "Quotation" is displayed instead of "Estimate". The cancel button 704, when pressed, resets the corrected content and returns to the screen in Figure 6(b). The confirm button 705, when pressed, confirms the corrected content and transitions to Figure 8.
[0112] Figure 8 shows the screen after the OCR error has been corrected. The extracted string 801 has changed to the corrected string in Figure 7, and the color and bold highlighting have been removed. Similarly, the extracted item 802 has also lost its highlighting. This is because the user's confirmation and correction process was completed when the confirmation button 705 was pressed in Figure 7, and the display returned to normal. In this way, the CPU 211 functions as a string correction means, correcting the string extracted from the image data based on instructions to correct the string extracted from the image data.
[0113] Next, Figures 9 to 11 illustrate the process of correcting errors in extraction location.
[0114] Figure 9 shows an example of the position correction screen. The position correction screen 901 is displayed by selecting the highlighted position correction icon 602 from the icon group 601 in Figure 6(c). Extracted items 902 and 903 are already selected as file names and therefore cannot be selected during position correction. Selected extracted item 904 is currently selected, so the area around the text is highlighted in bold. The candidate group of position correction destinations 905 has its text area enclosed in a rectangle and can be selected. Selecting it changes the selected extracted item 904. The correct answer item 906 is the correct answer in this embodiment. When the cancel button 907 is pressed, the corrected content is reset and the screen returns to Figure 6(c). When the confirm button 908 is pressed, the correction position is confirmed with the selected extracted item 904 and the screen transitions to Figure 10.
[0115] Figure 10 is the same screen 1001 as the text correction screen shown in Figure 7, so a detailed explanation is omitted. During position correction, the extraction position of the string to be used as the file name changes, so the user needs to confirm and correct the string. In the text input field 1002, the corrected "×× Corporation" is displayed instead of "△△ Corporation" due to the user's operation. The cancel button 1004, when pressed, resets the corrected content and returns to the screen in Figure 9. The confirm button 1005, when pressed, confirms the corrected content and transitions to Figure 11.
[0116] Figure 11 shows the screen after the extraction position error has been corrected. The extracted string 1101 has changed to the string selected in Figure 9 and confirmed in Figure 10, and the color and bold highlighting have been removed. Also, the extracted item 1102 has changed to the item selected in Figure 9, and the highlighting has been removed. This is because the user's confirmation and correction process was completed when the confirmation button 1005 was pressed in Figure 10, and the display returned to normal. In this way, the CPU 211 functions as a position correction means, correcting the position of the string to be extracted from the image data based on instructions to correct the position of the string to be extracted from the image data.
[0117] Through the above series of processes, the accuracy of the extraction position of the extracted items can be directly obtained, allowing for the determination of extraction position errors for each extracted item, thus reducing the likelihood of missing extraction position errors. Furthermore, by selecting an icon, the user can transition to a correction screen corresponding to the error, reducing the time required for correction.
[0118] This concludes the description of this embodiment.
[0119] (Second Embodiment) In the first embodiment, the user made corrections based on their own judgment on the text correction screen 701 and the position correction screen 901. In the second embodiment, appropriate correction candidates are presented according to the correction method, making it possible for the user to make corrections more efficiently. When describing the second embodiment, the explanation of parts that are the same as the first embodiment in terms of configuration and processing procedure will be omitted, and only the parts that differ will be explained.
[0120] Figure 12(a) shows an example of a text correction screen in the second embodiment. The text string has not yet been corrected. Figure 12(b) shows a text correction screen with text correction candidates displayed. When pulldown 1201 is pressed, the text correction candidate group 1202 is displayed. The text correction candidates may be arranged from the top of the list as the first candidate. In this embodiment, the text correction candidate group 1202 is a string composed of combinations of the second and subsequent candidates from the OCR engine.
[0121] However, the means for generating string candidates are not particularly limited and other methods may be used. For example, string candidates may be generated from the OCR results of a second OCR engine or from strings that are similar to the extracted results among strings previously approved by the user. Alternatively, string correction candidates may be generated by combining the above methods. If there is no correct string in the string correction candidate group 1202, the user can directly input characters by selecting the character input field 1203, as in the first embodiment.
[0122] Figure 13 shows an example of the position correction screen in the second embodiment. Selected extraction item 1301 is the item selected as the file name. Position correction target candidates 1302 to 1304 can be selected by enclosing the text area with a rectangle. Unlike Figure 9, the number of candidate destinations with text areas enclosed by rectangles is limited, but it is also possible to select text areas that are not enclosed. Furthermore, among these, position correction target candidate 1304, which is the first candidate, has its area highlighted with color and thickness compared to position correction target candidates 1302 and 1303.
[0123] In this embodiment, in step S403, the confidence level obtained by taking a continuous value for each item calculated by the entity extraction performed by the character extraction result acquisition means 321 exceeds an arbitrary threshold, and these items are designated as candidates for position correction. Since the item with the highest confidence level has already been presented to the user as the selected extraction item 1301, the remaining items with a confidence level exceeding the threshold are presented as candidates for position correction, and among these, the item with the highest confidence level becomes the first candidate.
[0124] The method for generating candidate positions for correction is not particularly limited and other methods may be used. For example, this could include items that match the attributes of the extracted items, or items that match the position information of items previously selected by the user in image data with the same layout. Furthermore, the display method may not only involve highlighting with color or thickness, but also relatively highlighting candidate positions for correction by masking items other than the candidate positions with semi-transparency.
[0125] By presenting suggested corrections on each correction screen, users can efficiently perform corrections according to their needs, thus reducing the time required for the task.
[0126] This concludes the description of this embodiment.
[0127] According to the first and second embodiments, the accuracy of the extraction position of the extracted items is directly obtained, thereby reducing the likelihood of missed detections. Furthermore, prompting the user to a correction screen corresponding to each error can shorten the time required for correction.
[0128] (Other embodiments) This disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.
[0129] Furthermore, the embodiments described above are merely examples illustrating how to implement this disclosure, and they should not be interpreted as limiting the technical scope of this disclosure. In other words, this disclosure can be implemented in various ways without departing from its technical concept or its main features.
[0130] This embodiment includes the following configuration. (Item 1) A display control means that controls the display of information for correcting the position of a string extracted from image data, or information for correcting a string extracted from image data, based on a position confidence score indicating the certainty of the position of a string extracted from image data and an OCR confidence score indicating the certainty of the extracted string. A position correction means for correcting the position of a string to be extracted from the image data based on instructions for correcting the position of the string to be extracted from the image data, A string correction means for correcting the string extracted from the image data, based on instructions for correcting the string extracted from the image data. An image forming apparatus characterized by having the following features. (Item 2) The display control means is If the position confidence falls below the first threshold, the system controls the display of information to correct the position of the string extracted from the image data. The image forming apparatus according to item 1, characterized in that, if the OCR confidence level falls below a second threshold, it is controlled to display information for correcting the string extracted from the image data. (Item 3) The image forming apparatus according to item 2, characterized in that the display control means displays information for correcting the position of a string extracted from the image data, and controls the display of information for correcting a string extracted from the image data, when the position confidence falls below a first threshold and the OCR confidence falls below a second threshold. (Item 4) The image forming apparatus according to any one of items 1 to 3, characterized in that the display control means controls the display of string correction candidates when the string correction means makes corrections. (Item 5) The image forming apparatus according to any one of items 1 to 4, characterized in that the display control means controls to display a candidate for the position correction destination when the position correction means corrects the position. (Item 6) The image forming apparatus according to any one of items 1 to 5, characterized in that the display control means displays the name of a combination of multiple strings extracted from the image data, and controls the display to display information for correcting the position of the strings or information for correcting the strings based on the position confidence and OCR confidence of the strings in the name. (Item 7) The image forming apparatus according to item 6, characterized in that the name is the file name of the image data. (Item 8) The image forming apparatus according to any one of items 1 to 7, further comprising an image reading means for reading a document placed on a scanner and generating the image data. (Item 9) The image forming apparatus according to item 7, further comprising a transmission means for transmitting the aforementioned image data. (Item 10) An image forming apparatus described in any one of items 1 to 9, A means for obtaining character extraction results, which performs character extraction on the aforementioned image data and obtains the extraction results. A position confidence acquisition means that calculates the position confidence for the extraction result and obtains the position confidence, OCR confidence acquisition means for calculating the OCR confidence score based on the extraction results and obtaining the OCR confidence score. An information processing system characterized by having the following features. (Item 11) A display control step that controls the display to show information for correcting the position of a string extracted from image data, or information for correcting a string extracted from image data, based on a position confidence score indicating the likelihood of the position of a string extracted from image data and an OCR confidence score indicating the likelihood of the extracted string. A position correction step to correct the position of the string to be extracted from the image data based on instructions for correcting the position of the string to be extracted from the image data, A string correction step is performed to correct the string extracted from the image data based on instructions for correcting the string extracted from the image data. A control method for an image forming apparatus, characterized by having the following features. (Item 12) A program to cause a computer to function as an image forming apparatus as described in any one of items 1 to 9. [Explanation of symbols]
[0131] 110 image forming apparatus, 120 image processing servers, 130 storage servers
Claims
1. A display control means that controls the display of information for correcting the position of a string extracted from image data, or information for correcting a string extracted from image data, based on a position confidence score indicating the certainty of the position of a string extracted from image data and an OCR confidence score indicating the certainty of the extracted string. A position correction means for correcting the position of a string to be extracted from the image data based on instructions for correcting the position of the string to be extracted from the image data, A string correction means for correcting the string extracted from the image data, based on instructions for correcting the string extracted from the image data. An image forming apparatus characterized by having the following features.
2. The display control means is If the position confidence falls below the first threshold, the system controls the display of information to correct the position of the string extracted from the image data. The image forming apparatus according to claim 1, characterized in that, if the OCR confidence level falls below a second threshold, it is controlled to display information for correcting the string extracted from the image data.
3. The image forming apparatus according to claim 2, characterized in that the display control means displays information for correcting the position of a string extracted from the image data, and does not display information for correcting a string extracted from the image data, when the position confidence falls below a first threshold and the OCR confidence falls below a second threshold.
4. The image forming apparatus according to claim 1, characterized in that the display control means controls the display of string correction candidates when the string correction means makes corrections.
5. The image forming apparatus according to claim 1, characterized in that the display control means controls to display a candidate for the position correction destination when the position correction means corrects the position.
6. The image forming apparatus according to claim 1, characterized in that the display control means displays the name of a combination of multiple strings extracted from the image data, and controls the display of information for correcting the position of the strings or information for correcting the strings based on the position confidence and OCR confidence of the strings in the name.
7. The image forming apparatus according to claim 6, characterized in that the aforementioned name is the file name of the image data.
8. The image forming apparatus according to claim 1, further comprising an image reading means for reading a document placed on a scanner and generating the image data.
9. The image forming apparatus according to claim 7, further comprising a transmission means for transmitting the aforementioned image data.
10. An image forming apparatus according to any one of claims 1 to 9, A means for obtaining character extraction results, which performs character extraction on the aforementioned image data and obtains the extraction results. A position confidence acquisition means that calculates the position confidence for the extraction result and obtains the position confidence, An OCR confidence acquisition means that calculates the OCR confidence score for the extraction result and obtains the OCR confidence score. An information processing system characterized by having the following features.
11. A display control step that controls the display to show information for correcting the position of a string extracted from image data, or information for correcting a string extracted from image data, based on a position confidence score indicating the likelihood of the position of a string extracted from image data and an OCR confidence score indicating the likelihood of the extracted string. A position correction step to correct the position of the string to be extracted from the image data based on instructions for correcting the position of the string to be extracted from the image data, A string correction step is performed to correct the string extracted from the image data based on instructions for correcting the string extracted from the image data. A control method for an image forming apparatus, characterized by having the following features.
12. A program for causing a computer to function as an image forming apparatus according to any one of claims 1 to 9.
Citation Information
Patent Citations
Information processing device and information processing method
JP2021196686A