Information processing apparatus, information processing method, and program

The information processing apparatus addresses the issue of incorrect user corrections being learned in OCR systems by incorporating a learning unit that verifies the reliability of character recognition and ensures that only accurate corrections are learned, thereby enhancing the accuracy of OCR results.

JP7696730B2Active Publication Date: 2025-06-23CANON KK
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
JP2021037211
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-09
Publication Date
2025-06-23
Estimated Expiration
2041-03-09

AI Technical Summary

Technical Problem

Existing technologies for optical character recognition (OCR) assistive data input do not effectively prevent incorrect user corrections from being learned and reused, leading to potential misrecognition in similar documents.

Method used

An information processing apparatus that includes a character recognition unit, a display control unit, an acquisition unit, and a learning unit. The learning unit learns error patterns based on differences between the extracted character string and the corrected character string, but only if the character recognition reliability meets a predetermined threshold, and ensures that incorrect corrections are not learned by verifying the confidence level of character recognition.

Benefits of technology

Prevents incorrect user corrections from being learned and applied to OCR results, thereby improving the accuracy of character recognition and reducing the burden on users correcting multiple documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696730000001
    Figure 0007696730000001
  • Figure 0007696730000002
    Figure 0007696730000002
  • Figure 0007696730000003
    Figure 0007696730000003
Patent Text Reader

Abstract

To prevent wrong correction of an OCR result made by a user from being learned.SOLUTION: An information processing apparatus comprises: character recognition means that extracts a character string including at least one character through character recognition on an image obtained by reading a document; display control means for displaying the extracted character string on a display; acquisition means that receives user correction of the extracted character string and acquires a character string corrected based on the user correction; and learning means that learns correction characteristics in a difference part between the extracted character string and the corrected character string. When the reliability in the character recognition of the difference part is equal to or more than a predetermined threshold, the learning means does not learn the correction characteristics for characters corresponding to the difference part of the extracted character string.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology for assisting data input for optical character recognition results.

Background Art

[0002] A technology for converting a document image (read image) obtained by reading a document or form including handwritten characters and printed characters into a character code for a computer, that is, optical character recognition (OCR), is widely known. For obtaining a read image, a scanner provided in a multi-function printer (MFP) or the like, or a camera function provided in a smartphone or the like can be used.

[0003] Generally, since there are variations in the image quality and character state of the read document image, the correct recognition rate of OCR does not reach 100%, and misrecognition occurs not infrequently. For characters misrecognized by OCR, it is necessary for the user to correct them by directly inputting the correct characters, etc. However, the operations of discovering, specifying, and correcting misrecognized portions in a character string are complicated, and the user's burden is high when processing many documents.

[0004] On the other hand, Patent Document 1 discloses a technology for learning a user's correction result for an OCR result and automatically correcting the OCR result based on the learned content. Here, when the OCR result and the user's input result are different, the character string before correction and the character string after correction are associated and registered as learning data, and thereafter, the OCR result is corrected based on the learning data. According to this technology, it is possible to reduce the user's input load particularly when processing documents including the same description content and the errors in the OCR results for the description content have high reproducibility. However, in the technology of Patent Document 1, even if the correction by the user is an incorrect correction due to the user's misinput, it is registered as learning data, so incorrect correction may be performed.

[0005] Therefore, in Patent Document 2, a technique for verifying the correction result by the user is disclosed in order to prevent incorrect input by the user when correcting the OCR result. Based on the string after correction by the user for the OCR result and the probability (confidence level) of each character of the OCR result, a location highly likely to be an incorrect input by the user is identified, and the user is notified that an incorrect correction is suspected. Thereby, incorrect corrections can be suppressed in the user's string correction work.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] However, with the technique described in Patent Document 2, it is not possible to prevent incorrect corrections by the input user from being learned with respect to the OCR result.

[0008] An object of the present invention is to prevent incorrect corrections by the user with respect to the OCR result from being learned.

Means for Solving the Problems

[0009] The information processing apparatus according to the present invention includes: a character recognition unit that extracts a character string including at least one character by performing character recognition on an image obtained by reading a document; a display control unit for displaying the extracted character string on a display device; an acquisition unit that receives a user correction for the extracted character string and acquires a character string corrected based on the user correction; and a learning unit that learns an error pattern specified based on a difference portion between the extracted character string and the corrected character string, wherein the learning unit Previous the reliability of the character recognition is equal to or higher than a predetermined thresholdpreviously, and when the character corresponding to the different part of the extracted character string has not been set as a learning target in advance, the extracted between the character string and the corrected character string the differences for the character corresponding to the error pattern without performing learning even if the confidence level of the character recognition is equal to or higher than the predetermined threshold value, when the character corresponding to the different part of the extracted character string has been set as a learning target in advance, learn the error pattern for the character corresponding to the different part characterized by

Advantages of the Invention

[0010] According to the present invention, it is possible to prevent the user's incorrect correction of the OCR result from being learned.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Best Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments do not limit the present invention, and not all the configurations described in the embodiments are essential means for solving the problems of the present invention.

[0013] <First Embodiment> [System Configuration] FIG. 1 is a diagram showing an information processing system according to the first embodiment. The information processing system includes a reading device 100 and an information processing device 110. The reading device 100 includes a scanner 101 and a communication unit 102. The scanner 101 reads a document and generates a document image obtained by scanning. The communication unit 102 communicates with an external device via a network.

[0014] The information processing device 110 includes a system control unit 111 such as a CPU, a ROM 112, a RAM 113, an HDD 114, a display unit 115, an input unit 116, and a communication unit 117. The system control unit 111 reads a control program stored in the ROM 112 and executes various processes. The RAM 113 is used as a temporary storage area such as a main memory and a work area of the system control unit 111. The HDD 114 stores various data and various programs. Note that the functions and processes of the information processing device 110 described later are realized by the system control unit 111 reading a program stored in the ROM 112 or the HDD 114 and executing this program.

[0015] The communication unit 117 performs communication processing with an external device via a network. The display unit 115 is controlled to display various information by the system control unit 111. As a specific device for realizing the display unit 115, it may be an operation panel 201 as shown in FIG. 2 mounted on the MFP, or a display connected to a PC capable of communicating with the information processing device 110 via a network. The input unit 116 has a keyboard and a mouse, and accepts various operations by the user. Note that the display unit 115 and the input unit 116 may be realized as an integrated device such as a touch panel. Further, the display unit 115 may perform projection by a projector, and the input unit 116 may recognize the position of a fingertip with respect to the projected image by a camera.

[0016] The data management device 120 has a database 122 and manages information used in the information processing device 110 and information generated by the information processing device 110. The data management device 120 may be a server managed on-premises or managed in the cloud. The data management device 120 communicates with the information processing device 110 via the communication unit 121 and transmits and receives data managed in the database 122. The information processing device 110 manages the data received from the data management device 120 in the RAM 113 or the HDD 114 and uses it for processing performed in the information processing device 110.

[0017] In the present embodiment, the scanner 101 of the reading device 100 reads a paper document such as a form and generates a document image obtained by scanning. The generated document image is transmitted to the information processing device 110 by the communication unit 102 on the reading device side. In the information processing device 110, the communication unit 117 on the information processing device side receives the document image obtained by scanning and stores the image in a storage device such as the HDD 114. Note that some functions of the display unit 115 and the input unit 116 may be in the reading device 100.

[0018] [UI] FIG. 2 is a diagram showing an example of an operation panel 201 that realizes a user interface (UI) composed of a display unit 115 and an input unit 116 of the information processing apparatus 110 in the present embodiment. The operation panel 201 includes a touch panel 202 and a numeric keypad 203 that function as the display unit 115 and the input unit 116. On the touch panel 202, the user ID during login, the main menu, and the like are displayed.

[0019] In the present embodiment, the UI provides the user with information extraction results from a document to be processed, OCR results, and character correction results obtained by correcting the OCR results, as shown in FIG. 6. Further, the UI provides the user with correction character string candidates displayed in a list as shown in FIG. 10, and an incorrect correction pointing screen 1200 as shown in FIG. 12. Although the UI is displayed on the touch panel 202, various information displayed by the UI may be displayed on a display connected to a PC capable of communicating with the information processing apparatus 110, and various operations by the user may be received from the keyboard and mouse of the PC.

[0020] [Software Configuration] FIG. 3 is a diagram showing a software configuration that realizes the data input support unit 300. The data input support unit 300 is composed of each means (304 to 305) of the processing result providing unit 301, each means (306 to 308) of the character recognition result generation unit 302, and various stages (309 to 312) of the character string correction unit 303.

[0021] The processing result providing unit 301 includes a character string display means 304 that displays the character recognition result which is the OCR result or the character string correction result obtained by correcting the character recognition result. Further, the processing result providing unit 301 includes a corrected character string input means 305 for the user to input a corrected character string for the character string obtained as the character recognition result. The processing result providing unit 301 provides the processing results of the character recognition result generation unit 302 and the character string correction unit 303 to the user, and transmits the input received from the user to the character string correction unit 303. The processing result providing unit 301 functions as the above-described UI, and is realized, for example, by the touch panel 202 provided in the information processing apparatus 110 or a PC capable of communicating with the information processing apparatus 110.

[0022] The character recognition result generation unit 302 is composed of an image processing means 306, a character recognition means 307, and an item extraction means 308. The image processing means 306 performs pre-processing of the image so that character recognition can be executed on the input document image. The character recognition means 307 performs conversion to a character code (OCR processing) on the character string in the image. The item extraction means 308 extracts the character string of the area required by the user from the document image. With these configurations, the character recognition result generation unit 302 executes character recognition on the document image input to the information processing apparatus 110 to generate a character recognition result, and extracts an item value corresponding to the item name to be extracted from the character recognition result.

[0023] The character string correction unit 303 is composed of a character string comparison means 309, a character string correction means 310, a confidence determination means 311, and a correction result management means 312. The character string comparison means 309 compares the OCR result obtained by the character recognition means 307 or the corrected character string generated by the subsequent character string correction means 310 with the correction result of the character string by the user input by the corrected character string input means 305. The character string correction means 310 generates a corrected result character string of the OCR result. The confidence determination means 311 determines the certainty of recognition for each character obtained by the character recognition means 307. The correction result management means 312 manages the comparison result generated by the character string comparison means 309 and the character string obtained by the corrected character string input means 305 or an expression abstracted therefrom. Here, the comparison result is information that associates the character before replacement and the character after replacement with the character string before correction and the character string after correction. These comparison results and the character string or an expression abstracted therefrom are used in the item value correction in S407 of FIG. 4. The correction result management means 311 is realized by the HDD 114 of the information processing apparatus 110 in FIG. 1 or the database 122 of the data management apparatus 120. With these configurations, the character string correction unit 303 receives the user's correction result for the item value from the processing result providing unit 301, detects, registers, and learns the correction content, and generates a correction candidate for the character recognition result based on the registered learning content.

[0024] The specific operations of the above various means will be specifically described with reference to the flowchart shown in FIG. 4.

[0025] [Processing Flow] FIG. 4 is a flowchart showing the processing flow from when a document image obtained by scanning with the reading device 100 is input to the information processing apparatus 110, the user checks and corrects the OCR result, and registers the correction content in the system.

[0026] First, in S400, the reading device 100 reads a document with the scanner 101 and generates a document image. The document to be read is, for example, a claim as shown in 500 of FIG. 5.

[0027] In S401, the image processing means 306 corrects the inclination by detecting the description direction of the character string in the document for the document image converted into an image.

[0028] In S402, the image processing means 306 sets a predetermined threshold value for the grayscale document image and performs binarization processing.

[0029] In S403, the image processing means 306 removes the ruled lines unnecessary for character recognition from the binarized document image.

[0030] The processing performed in these steps from S401 to S403 is positioned as preprocessing for accurately executing the identification of the character area in S404 and the character recognition in S405 to be performed later.

[0031] In S404, the character recognition means 307 identifies the character area for the preprocessed document image.

[0032] In S405, the character recognition means 307 performs character recognition processing (OCR processing) on the document image based on the result of identifying the character area, and acquires a character code as an OCR result for each character area in the document image detected in S404.

[0033] In S406, the item extraction means 308 extracts the item values required by the user from the document image for each predetermined item name. Taking the invoice in FIG. 5 as an example, for item names such as "phone number" and "total amount", the corresponding item values are "03-123-4567" and "¥11,286". The item names and item values extracted here are registered in the system in S409 described later. 600 in FIG. 6 shows an example of the screen provided to the user during item extraction. 601 is a preview screen of the document image to be processed, 602 is an item name display column for displaying the item names to be extracted, and 603 is an item value display column for displaying the item values extracted from the document image 601.

[0034] The extraction process of item values is realized, for example, when the user designates the area in which the item values to be extracted are described for the document image 601. As another method for extracting item values, a string related to the item name displayed in the item name display column 602 is searched in the document image 601, and based on the predefined positional relationship between the string and the item value, the necessary item value is extracted from the position of the searched string. Thus, it is also possible to automatically extract the item value corresponding to the item name without an instruction from the user. For example, when extracting the item value "¥11,232" for the item name "Total Amount", it is searched whether a string such as "Claim Amount" related to "Total Amount" exists in the document for the document type to be processed (in this case, the claim form). As a result, if the string "Requested Amount" is found in the document, based on the position of the searched string and the predefined positional relationship between "Requested Amount" held by the information processing apparatus 110 and the position of the item value of the amount corresponding thereto, the item value "¥11,232" is extracted. In this case, the item value "¥11,232" is extracted based on the rule that "the item value of the amount exists on the right side of the string 'Requested Amount'". The item value extraction methods shown above are merely examples, and any method may be used as long as it can extract the information desired by the user, and other methods are not limited to these.

[0035] In S407, the string correction means 310 performs string correction of the item value extracted as necessary. The OCR result of the extraction target image area, which is the extraction result of the item value, may include misrecognition depending on the quality of the document image 601. Therefore, in S407, string correction of the item value extracted based on the information obtained in the past by the correction character string input means 305 and the string comparison means 309 in FIG. 3 is performed. The specific process of item value correction will be described using the flowchart shown in FIG. 7.

[0036] Thereafter, the corrected character string of the item value is displayed by the character string display means 304. 603 in FIG. 6 shows one such state, and one character string that is the corresponding item value for each item name is displayed. In addition, for example, as shown in FIG. 10, the OCR result 1002 for the image 1001 to be character-recognized and the corrected character string candidate 1003 of the item value by S407 may be displayed for each extraction item. When the OCR result and a plurality of corrected character string candidates are displayed as shown in FIG. 10, the user can select the correct character string or the corrected character string determined to be the closest to it from among the OCR result or the list-displayed corrected character strings. When the desired character string does not exist in the list, the user can erase the displayed corrected character string candidates by pressing the clear button 1004 and input the desired character string from the input unit 116. The correction by the user for the character string obtained by this OCR is performed by the corrected character string input means 305 in FIG. 3. That is, the user can either correct the character string after once selecting a corrected character string close to the desired character string from the corrected character string candidates 1003 or directly correct the character string of the OCR result without selecting from the corrected character string candidates 1003. In this embodiment, when the user has completed the confirmation and correction of each item, the user can check the corresponding check box 604 to determine the item value for each item name.

[0037] In S408, when it is detected that all items have been checked, the "Next" button 605 is enabled. When the user presses the next button, that is, when the confirmation and correction of all items have been completed (when S408 is true), the process proceeds to S409.

[0038] In S409, corrections made by the user for each item are detected, and the correction characteristics are learned by registering the correction contents. Here, the user's correction contents are identified and registered from the difference between the character string before correction and the character string after correction by the user. Specifically, the correction contents are the character string confirmed after correction by the user or an abstract character string (expression) thereof, and an error pattern indicating the characteristics of the user's correction (correction characteristics) described later. The error pattern is generated by the character string comparison means 309. The specific processing of this step will be described with reference to the flowchart shown in FIG. 8.

[0039] In S410, the data is registered in the system, and all processing ends.

[0040] [Text correction for item values] The character string correction of the item value in S407 of FIG. 4 will be described with reference to FIG.

[0041] First, in S700, the confidence level determination means 311 performs confidence level determination for each character in the OCR result. The confidence level determination is performed based on whether or not the score indicating the likelihood of the OCR result for each character obtained during OCR processing, that is, the confidence level value, is equal to or higher than a preset threshold. For example, consider a case where the item value "¥112,82" of the item name "total amount" is recognized as "¥1,2B2" as the OCR result. FIG. 11 shows the distribution of the confidence level values ​​for each character in the OCR result at that time, and the dashed line on the distribution 1101 indicates the threshold for determining the confidence level of the OCR result. In this example, the confidence levels of "" and "B" are less than the threshold, and high confidence levels equal to or higher than the threshold are obtained for the other characters. In this case, the confidence level determination means 311 determines "" and "B" to have low confidence levels, and the other characters to have high confidence levels.

[0042] Here, confidence determination is performed using a common confidence threshold for the characters in the OCR result, but different thresholds may be used for each character. Also, the confidence determination may be performed by any means as long as it can evaluate the reliability of the OCR result, and is not limited to the above method. For example, a voting-type confidence determination means using a plurality of character recognition means (OCR engines) may be used.

[0043] In S701, a lattice for generating candidate correction strings for the OCR result is generated by referring to the confidence determination result obtained in S700 and the character error pattern. A lattice is a structure for creating candidate correction strings as shown in FIG. 9. Also, the error pattern is obtained by associating the characters misrecognized by the OCR engine (OCR error) with their correct characters and registering them. For example, when the OCR result is "Kiyan0 Co., Ltd." and the correct character string is "Kiyano Co., Ltd.", the error pattern obtained therefrom is the relationship between the OCR error character: "0" - correct character: "o". The error pattern is managed by the correction result management means 312 realized in the HDD 114 or the database 122. The error pattern may be obtained from the user's character string correction in S409, may be generated in advance and registered in the HDD 114 of the information processing apparatus 110 or the database 122 of the data management apparatus 120, or may be a combination thereof.

[0044] In S702, candidate correction strings are generated based on the lattice generated in S701.

[0045] In S703, the candidate correction strings generated in S702 are narrowed down. The methods for generating and narrowing down the candidate correction strings will be specifically described below.

[0046] 900 shown in FIG. 9 is a lattice of a character string to be corrected, in this case, “han1, 2B2” which is an OCR result. The lattice is defined based on lower candidates of the OCR result or error patterns for characters or character strings managed by the information processing device 110 or the data management device 120. A lower candidate is an OCR result that is next most likely to one OCR result. For example, “刊” in 903 in FIG. 9 is a lower candidate for “han” obtained as an OCR result. An error pattern is a candidate registered in advance for an image in which “han” is obtained as an OCR result, and “¥リ”, “¥1”, and “¥7” in 904 are character strings of other registered lower candidates. As described above, the error pattern may be one registered in the information processing device in advance, or one acquired based on the user's correction history for the OCR result. In this embodiment, the latter error pattern acquisition will be described. The error pattern acquisition is performed in S409 of the flow in FIG. 4, and a specific means will be described later in [Detection and registration of correction contents].

[0047] A lattice 900 is formed using lower candidates and error patterns of the OCR result, but the locations where lattice nodes are added are limited to locations determined to have low confidence in S700, as shown in 903 to 905. This is because it is not necessary to consider replacing characters that are not actually OCR errors with other characters, or it is better not to consider this. This is because if lattice nodes are allowed to be added to each character of the OCR result, including characters with high confidence, rather than being limited to characters with low confidence, the number of subsequent correction character string candidates will increase accordingly, and the processing time will increase. Furthermore, there is a concern that the accuracy of correction will decrease due to the increased number of options when narrowing down subsequent correction character string candidates, which reduces the probability of selecting the correct character string.

[0048] In this embodiment, since the parts determined to have low confidence in S700 are "HAN" and "B", correction candidates are generated in the lattice 900 of Fig. 9 by limiting them to the shaded "HAN" in 901 and "B" in 902. This eliminates the need to generate candidates for other parts in the character string, and reduces the number of patterns of generated correction character string candidates.

[0049] The correction string candidates are generated by selecting all possible paths of the generated lattice 900. Multiple correction string candidates are generated from the lattice 900, and there are six patterns starting with a kanji character: "Han1,2B2", "Han1,282", "Han1,2132", "Kan1,2B2", "Kan1,282", and "Kan1,2132". There are also nine patterns starting with a yen sign: "¥Li1,2B2", "¥Li1,282", "¥Li1,2132", "¥11,2B2", "¥11,282", "¥11,2132", "¥71,2B2", "¥71,282", and "¥71,2132". In other words, there are a total of 15 patterns of correction string candidates generated from the lattice 900.

[0050] In S703, the candidates are further narrowed down based on the knowledge set in the item value. Here, by using the knowledge that "amounts cannot contain characters other than numbers," the candidates are narrowed down to "¥11,282," "¥11,2132," "¥71,282," and "¥71,2132." Furthermore, by using the knowledge that "commas are placed every three digits" as a rule for expressing amounts, the candidates are narrowed down to "¥11,282" and "¥71,282."

[0051] Also, as another piece of knowledge, it is also possible to use the abstracted number "¥00,000" or the regular expression "¥¥¥d{2},¥d{3}", which is the format of the item value for the "total amount" obtained from the past user's correction history. As a condition that the correction string candidate should satisfy, by setting that it conforms to such a format, it is possible to narrow down to the above two candidates. The format set for the item value as described above is managed by the correction result management means 312. The step of obtaining a character string or a format from the user's correction history is performed in S409, and the specific means will be described later in [Detection and Registration of Correction Content]. The acquisition of such a character string obtained from the user's correction result or the format abstracted therefrom is managed by the correction result management means 312 realized in the HDD 114 or the database 122. Alternatively, a format defined in advance may be registered in the HDD 114 of the information processing apparatus 110 or the database 122 of the data management apparatus 120 and used. For example, for an amount, a format as described above can be registered and used, and for a phone number, a format such as ¥d{2}―¥d{4}―¥d{4} can be registered and used. Or, both the format obtained from the user and the format defined in advance may be used.

[0052] Note that the means of narrowing down candidates from the character string generated from the lattice used above is just an example, and various other means can be considered. For example, when the item name is "name", the knowledge that it "does not contain numbers" may be used. Also, for general item values, candidates can be narrowed down by taking the means of selecting the path with the highest likelihood by assigning the occurrence probability of characters and the transition probability of adjacent characters to the lattice.

[0053] In S704, if there are multiple candidates, the priority of the candidates is determined based on information obtained from the user's correction history. Even after the candidates are narrowed down in S703, there may be two correction string candidates, "¥11,282" and "¥71,282", such as when the character string to be corrected is "判1,2B2". In such a case, the priority of each correction string candidate is determined based on, for example, information obtained from the frequency of use of the character string previously confirmed by the user and the range of values. For example, if there is information that only item values ​​in the range of ¥0 to ¥50,000 have been input in the "total amount" field in the past, the priority of the two candidates, "¥11,282" and "¥71,282", is set to be higher than that of "¥11,282". As another example, if "Kiyano Co., Ltd." and "kiyano Co., Ltd." are obtained as the correction results, the priority can be determined based on the frequency of use of the character string previously confirmed by the user. Specifically, if there is information that a user has previously entered and confirmed "Kiyano Co., Ltd." ten times and "kiyano Co., Ltd." twice in the "Company Name" field, the priority of "Kiyano Co., Ltd." will be set high.

[0054] Furthermore, the priority may be determined based on the most recent time of use, rather than the frequency of use. For example, consider a case where a user has entered "Kiyano Co., Ltd." 10 times and "kiyano Co., Ltd." twice in the past. Looking at the frequency of use, "Kiyano Co., Ltd." is higher, but if "kiyano Co., Ltd." is confirmed as the most recent "company name" item, the priority of "kiyano Co., Ltd." is raised as a candidate. This allows the user's most recent correction results to be reflected without being bound by the frequency of past use. The above priority can also be determined using known technology, such as kana-kanji conversion.

[0055] All steps in Fig. 7 are completed above, and the string correction process for the item value, which is the process of S407 in Fig. 4, is completed. When only one correction result is displayed as in 603 of Fig. 6, the correction string candidate with the highest priority is displayed. Also, when the correction string candidates are listed as in Fig. 10, the correction string candidates are displayed on list 1003 in order of priority.

[0056] [Obtaining Modifications by the User and Registering Modification Contents] Regarding the obtaining of modifications by the user and the registration of modification contents in S409 of Fig. 4, it will be described with reference to Fig. 8. The following processes are executed for each item value.

[0057] First, in S800, the string determined by the user is obtained. The determined string is the string input to 603 in Fig. 6 when the confirmation button 605 in Fig. 6 is pressed, that is, when S408 becomes true.

[0058] In S801, it is determined whether there is a string modification by the user. This is determined that there is a string modification by the user when the string before the modification by the user, that is, the string output as the result of S407, is different from the string determined by the user obtained in S800.

[0059] In S802, the confidence level of the OCR result is obtained. The confidence level to be obtained is based on the determination result of the reliability of the OCR result obtained in S700. Regarding the characters for which the OCR result has been replaced by the item value correction in S407, since the characters of the original OCR result have been overwritten by the string correction and the confidence level determination result cannot be used, the confidence level determination is obtained for the other characters.

[0060] In S803, the correction result in S407 is compared with the character string after being corrected by the user, and the matching positions are searched for. In the following, consider the case where the OCR result including misrecognition cannot be corrected to the correct character string, and "judgment 1,2B2" is obtained as the correction result in S407. In this case, since the character string correction is not performed by S407, the confidence determination results of all characters can be used. When the character string before being corrected by the user is "judgment 1,2B2" and the character string after being corrected is "¥11.282", the matching positions are "1", "2", "2" from the front. The search for these matching characters can be obtained, for example, from a graph obtained by calculating the edit distance between two character strings. Here, it is assumed that the user made an input error and input a "." instead of a "," which should originally be there.

[0061] In S804, the different positions other than the matching positions obtained in S803 are detected. The different positions in the above example are "judgment", ",", "B". Although "," is not an error, it is detected as a different position here because it is different from the user input. The identification of the different positions is performed by a different position identification means (not shown).

[0062] In S805, an error pattern is generated for the identified different positions. By the processes of S803 and S804, the matching positions and different positions between the corrected character string before correction and the character string after being corrected by the user are detected. Therefore, an error pattern that defines the character pairs before and after replacement is generated using that information. The error pattern generated in the above example is "judgment" → "¥1", ",", "→", ".", "B" → "8".

[0063] However, if the character detected as an error in the process of S804 is a character resulting from correction in S407, an error pattern for that character is not generated in S805 even if the user corrects that character. This is because the detected error should be considered an error in the correction result, not an error in the OCR result. In this case, the user corrects the result of character string correction, so the character string correction is denied by the user. As a process in this case, for example, the relative importance or priority of the information used in the current character string correction in S407 that was held up to S804 and S805 among the information managed by the correction result management means 312 may be lowered or deleted.

[0064] In S806, the user is notified that a character with high character confidence is about to be corrected. The error patterns generated in S805 are (1) "判" → "¥1", (2) "," → ".", and (3) "B" → "8", of which (2) "," → "." is a correction to the character "," that was determined to have high confidence using the confidence determination result acquired in S802. Therefore, the correction by the user of (2) "," → "." is an erroneous correction, and it is highly likely that the OCR result before the correction was correct, so a notification is displayed to encourage the user to confirm the correction content.

[0065] 1200 shown in FIG. 12 is an example of a notification to the user displayed during the process of S806, and notifies the user that the correction from “,” to “.” in the form as shown in 1201 may be an incorrect correction. This can prevent the user from making an incorrect correction without noticing an input error. At the same time, a button for selecting whether to register the correction content while ignoring the control of the registered content based on the above confidence determination result is displayed. This is assumed for the case when the confidence determination is incorrect. For example, in the case where the confidence determination for the characters at the OCR error location is misjudged as a high confidence level, it is to prevent the user correction for the OCR error characters from becoming unable to be registered. When registering the correction for the character notified with the high confidence determination, in FIG. 12, by the user pressing the OK button 1203, it is possible to register the correction content while ignoring the high confidence determination result. At this time, when there are a plurality of items in the table of 1201, it may be configured such that the user can select a desired item from among these plurality of items. On the other hand, when not registering the correction content, the user can end the process without registering the content by pressing the cancel button 1202.

[0066] As described above, by the error pattern generation process using the confidence determination, even if the user makes an incorrect correction, the generation and registration of the accompanying error pattern can be prevented. As a result, when generating a corrected string candidate using the error pattern in the process of S407, the possibility of generating an incorrect candidate is reduced, and an improvement in the accuracy of string correction can be expected.

[0067] In S807, the confirmed character string by the user obtained in S806 and the error pattern generated in S805 are registered in the correction result management means 312. When registering the confirmed character string here, in the case of a character string that does not contain numbers, for example, in the case of "Kiyano Co., Ltd.", the character string "Kiyano Co., Ltd." is registered. Alternatively, it may be registered by splitting it into words such as "Kiyano" and "Co., Ltd.". On the other hand, in the case of a character string composed of numbers, for example, in the case of "¥11,282", a character string obtained by abstracting each digit is registered. Abstraction means, for example, replacing a number with an arbitrary number representing the number such as "0", and in the case of "¥11,282", it becomes "¥00,000". This abstraction may be a replacement with a regular expression. In the case of "¥11,282", it becomes "¥¥¥d{2},¥d{3}". This operation is a process based on the premise that it is easily assumed that the numbers described in character strings composed of numbers are different each time even if the types of forms to be handled and the description locations are the same or different forms.

[0068] The registered error pattern is used for generating the lattice 900 after the next time. Also, the confirmed character string by the registered user is used for narrowing down the correction character string candidates in S703. The specific method is as described in the item value correction of S407. If the error pattern to be registered and the confirmed character string by the user are already registered in the correction result management means 314, for example, by counting its usage frequency, it can be used for determining the weight and priority when generating correction character string candidates during correction.

[0069] <The First Variant of the First Embodiment> In the item value correction of S407 in the first embodiment, the generation of sub-candidates for each character or replacement candidates based on error patterns for generating correction string candidates is limited to the locations determined by the confidence determination means 311 to have low confidence. However, it is conceivable that the confidence determination is incorrect, that is, characters with OCR errors are determined to have high confidence. In this case, for the characters with incorrect confidence determination, sub-candidates are not considered and replacement candidates based on error patterns are not generated, and correction string candidates with a high possibility of being the correct character string cannot be generated.

[0070] In view of such a situation, it is conceivable to add a determination on whether to generate sub-candidates and replacement candidates based on error patterns according to the character type. Specifically, when characters that are likely to have incorrect confidence determination and are set as learning targets in advance appear, replacement candidates are generated for those characters regardless of the result of the confidence determination. For example, "l" is likely to be misrecognized as "1", and depending on the font, the glyphs are very similar, so in an OCR engine that recognizes by glyph similarity, even if misrecognized, the confidence may be high. For character types that are difficult to distinguish from other character codes due to the characteristics of such OCR engines, there is a high tendency for incorrect confidence determination, so they are set as learning targets in advance. Therefore, the information processing apparatus 110 stores in the HDD 114 or the database 122 provided in the data management apparatus 120 data in which such character types that are likely to have incorrect confidence determination are listed in advance. When a character type existing in the list appears in the OCR result, it is permitted to generate recognition sub-candidates and replacement candidates using error patterns on the lattice 900 regardless of the result of the confidence determination. Thereby, it is possible to reduce the decrease in the accuracy of character string correction due to incorrect confidence determination. In addition, it is possible to reduce the display of notifications prompting the user to confirm the correction content, even though the possibility of incorrect correction by the user is not high.

[0071] <Second Variant of the First Embodiment> The error patterns and the character strings confirmed by the user managed by the correction result management means 312 in the first embodiment may be configured to be editable by the user. For example, the configuration may be such that the edits can be made using the touch panel 202 provided in the information processing device 110 or a display connected to a PC, and the managed information can be added and deleted. The editing work is performed, for example, by a system administrator in each tenant. With this configuration, if a user has registered an incorrect correction, it can be deleted. As a result, it becomes possible to correct the character string in accordance with the user's wishes.

[0072] <Second embodiment> In the first embodiment, as shown in FIG. 12, after the user corrects the character string, the correction of the high confidence character is detected and the erroneous correction is indicated. However, a confidence judgment result display means (not shown) may be provided and the confidence judgment result may be added to the character string before correction and displayed. In this configuration, the user can check the part of the character string of the OCR result or the item value correction result by S407 that is likely to require correction with low confidence before correction. Specifically, as shown in 1300 of FIG. 13, the high confidence character and the low confidence character are displayed in different modes by being superimposed on the character string so that they can be distinguished. This is displayed, for example, by the touch panel 202 provided in the information processing device 110 or a display connected to a PC. In 1301 of FIG. 13, the low confidence characters "判" and "B" are displayed in a shaded manner, so that the user can know the part to be corrected before correction. This can prevent erroneous correction by the user, and the accuracy of the registered error pattern and character string is improved, so that the accuracy of character string correction can be expected to be improved.

[0073] (Other Examples) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

Explanation of Symbols

[0074] 100 Reader 110 Information Processing Device 120 Data Management Device 201 Operation Panel 301 Processing Result Providing Device 302 Character Recognition Result Generation Device 303 String Correction Device

Claims

1. Character recognition means for extracting a character string including at least one character by character recognition for an image obtained by reading a document, Display control means for causing the display device to display the extracted character string, Acquisition means for receiving a user correction for the extracted character string and obtaining a corrected character string based on the user correction, Learning means for learning an error pattern specified based on a difference portion between the extracted character string and the corrected character string, comprising The learning means When the reliability of the character recognition is equal to or higher than a predetermined threshold and a character corresponding to the difference portion of the extracted character string is not set as a learning target in advance, the learning of the error pattern for the character corresponding to the difference portion between the extracted character string and the corrected character string is not performed, Even when the reliability of the character recognition is equal to or higher than the predetermined threshold, if a character corresponding to the difference portion of the extracted character string is set as a learning target in advance, the error pattern for the character corresponding to the difference portion is learned. An information processing apparatus characterized by the above.

2. The learning means learns an error pattern specified based on a difference portion between a character string in which the reliability of the character recognition is less than the predetermined threshold and the corrected character string among the difference portions. The information processing apparatus according to claim 1, characterized in that.

3. The character corresponding to the difference portion is at least one pair of characters before and after replacement replaced between the extracted character string and the corrected character string, The error pattern is an error pattern associating a character before replacement and a character after replacement. The information processing apparatus according to claim 1 or 2, characterized in that.

4. The learning means registers, as the error pattern, a character pair before and after substitution in which the confidence of the character before substitution in the character pair before and after each substitution is less than the predetermined threshold, and does not register, as the error pattern, a character pair before and after substitution in which the confidence of the character before substitution in the character pair before and after each substitution is equal to or greater than the predetermined threshold. The information processing apparatus according to claim 3, characterized in that.

5. The error pattern includes an expression obtained by abstracting numbers included in the corrected character string. The information processing apparatus according to any one of claims 1 to 4, characterized in that.

6. Correction means for generating a corrected character string by performing character substitution on the character string extracted by the character recognition means based on the error pattern learned by the learning means. further comprising The display control means causes the display device to display the corrected character string. The information processing apparatus according to any one of claims 1 to 5, characterized in that.

7. The acquisition means acquires a corrected character string based on the user correction for the corrected character string. The information processing apparatus according to claim 6, characterized in that.

8. The correction means generates the corrected character string by replacing, among the characters included in the extracted character string, the characters registered as the characters before substitution in the error pattern with the characters after substitution in the error pattern. The information processing apparatus according to claim 6 or 7, characterized in that.

9. The correction means determines the priority of the corrected character string based on the usage frequency of at least some of the character pairs before and after substitution used in the corrected character string. The information processing apparatus according to claim 8, characterized in that.

10. The correction means determines the priority of the corrected character string based on the proximity of the times of use of at least some of the character pairs before and after substitution used in the corrected character string. The information processing apparatus according to claim 8 or 9, characterized in that.

11. The correction means generates the corrected character string by replacing, among the characters included in the extracted character string, the characters for which there are lower candidates in the character recognition by the character recognition means with the lower candidates. The information processing apparatus according to any one of claims 8 to 10, characterized in that.

12. The learning means a determination means for determining whether or not the reliability of the characters extracted by the character recognition is equal to or higher than a predetermined threshold value; a comparison means for comparing the extracted character string with the corrected character string and specifying the different part; The information processing apparatus according to any one of claims 1 to 11, characterized in that it includes.

13. The display control means further causes the display device to display a display representing the reliability of the characters included in the extracted character string. The information processing apparatus according to any one of claims 1 to 12, characterized in that.

14. When the reliability of the different part is equal to or higher than the predetermined threshold value, the learning means notifies the user whether or not to learn the error pattern for the character corresponding to the different part of the extracted character string, receives a user input corresponding to the notification, and determines whether or not to learn the error pattern based on the user input. The information processing apparatus according to any one of claims 1 to 13, characterized in that.

15. The error pattern is editable by the user. The information processing apparatus according to any one of claims 1 to 14, characterized in that.

16. A step of extracting a character string including at least one character by performing character recognition on an image obtained by the information processing apparatus reading a document; A step of the information processing apparatus causing the display device to display the extracted character string; A step of the information processing apparatus receiving a user correction for the extracted character string and obtaining a corrected character string based on the user correction; A step of the information processing apparatus learning an error pattern specified based on a difference portion between the extracted character string and the corrected character string; comprising: The step of learning: When the reliability of the character recognition is equal to or higher than a predetermined threshold value and a character corresponding to the difference portion of the extracted character string has not been set as a learning target in advance, learning of the error pattern for the character corresponding to the difference portion between the extracted character string and the corrected character string is not performed; Even when the reliability of the character recognition is equal to or higher than the predetermined threshold value, if a character corresponding to the difference portion of the extracted character string has been set as a learning target in advance, the error pattern for the character corresponding to the difference portion is learned. An information processing method characterized by the above.

17. A program for causing a computer to function as the information processing apparatus according to any one of Claims 1 to 15.

Citation Information

Patent Citations

  • Tesseract engine based character recognition method and device

    CN105825214A

  • Character recognition system

    JP1996016730A

  • Dictionary learning system and character recognition device

    JP1999184976A

  • Character recognizing device

    JP2001297306A

  • Information processing system and character correction method

    JP2004046388A