Information processing device, information processing method, and program

The information processing apparatus addresses the inefficiencies in updating character strings from document images by automating the process when document types are modified, enhancing user efficiency and reducing manual intervention.

JP2025087172APending Publication Date: 2025-06-10CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023201636
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing technologies for processing document images face challenges in efficiently updating character strings when the document type is modified, leading to slow response times and the need for manual updates for non-type modifications.

Method used

An information processing apparatus that determines document types, extracts character strings, and presents document information, allowing for automatic updates of character strings when document types are modified, and enabling quick display of matching document information based on user modifications.

Benefits of technology

This solution allows users to easily modify character strings extracted from document images, reducing the time and effort required for updates and improving work efficiency by automating the update process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087172000001_ABST
    Figure 2025087172000001_ABST
Patent Text Reader

Abstract

To enable a user to further easily correct character strings extracted from a document image.SOLUTION: An information processing server acquires a document image, performs OCR processing to acquire a group of character strings contained in the document image, and uses a document type determination model to determine multiple document types that may be a document type of the document image. The information processing server also uses an item value extraction model to extract character strings that are item values of specified detailed items from the group of character strings contained in the document image. Based on the extracted character strings, the information processing server estimates character strings that are item values of detailed items corresponding to common items for each document type for all possible document types. The information processing server performs confirmation / correction processing of document information including one document type determined based on document type determination results and item value extraction results and the item values of the corresponding detailed items. Based on the confirmed document information, the information processing server generates a file with a file name for classifying and storing the document image.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing technology for extracting information from document images.

Background Art

[0002] Conventionally, there has been a technology for determining the document type (e.g., claim, estimate, delivery note) of an input document image and extracting character strings corresponding to predetermined items (e.g., title, document number, issue date, company name, amount) common to all document types described in the document image.

[0003] Patent Document 1 discloses a technology for recognizing the document type of an input document image and extracting a character string that is an item value of a predetermined item according to the recognized document type. For example, individual documents may have an identification number assigned to identify those documents. In the case of a claim, the claim number is the identification number of the document, and in the case of a delivery note, the delivery note number is the identification number. On the other hand, there may be some other identification number described in the document image in addition to the identification number of the document. For example, inside a delivery note, in addition to the delivery note number, the claim number of the claim related to that delivery note may be described. In such a case, the claim number does not correspond to the identification number of the document as a delivery note. Similarly, even if the document issue date, which is the date when the document image was generated, is described, if the document is a claim, the filing date is the issue date of the document, and if the document is a delivery note, the delivery date is the issue date of the document, and the document issue date does not correspond to the issue date of a claim or a delivery note. Here, items common to all document types, such as identification numbers and issue dates, are referred to as common items, and items such as claim numbers, delivery note numbers, filing dates, delivery dates, and document issue dates, which are set as items to be used as common items for each document type, are referred to as detailed items. In Patent Document 1, detailed items to be used as common items for each document type are determined in advance, and the character strings of the detailed items corresponding to the determined document type are extracted. In Patent Document 1, when the user modifies the document type, all the character strings of the detailed items can be automatically updated by re-extracting the character strings of the detailed items corresponding to the modified document type based on the modified document type.

Prior Art Documents

Patent Document

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the technology of Patent Document 1, when a user modifies the document type, it is necessary to re-extract the character string from the document image for automatic update of the character string, and there is a problem that the response to the user's modification operation is slow and time-consuming. In addition, the automatic update of the character string is limited to the case where the document type is modified. When the document type is not modified, modifications other than the document type need to be made manually, which is troublesome.

Means for Solving the Problems

[0006] The present invention is an information processing apparatus, comprising: determination means for determining the document type of an input document image; extraction means for extracting, from the document image, a character string as an item value of a predetermined item for the document type determined by the determination means; presentation means for causing a display means to display document information including the determined document type and the character string extracted by the extraction means for the items determined for the document type; and reception means for receiving a modification to the document information, wherein when the determination means determines a plurality of document types, the extraction means extracts character strings of predetermined items for each of the plurality of document types, and there are a plurality of pieces of the document information corresponding to the plurality of document types, the presentation means causes the display means to display the document information that matches the modification received by the reception means from among the plurality of pieces of document information.

Effects of the Invention

[0007] According to the present invention, a user can more easily modify the character string extracted from a document image.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. Note that the components described in this embodiment are examples and are not intended to limit the scope of the present invention.

[0010] [First Embodiment] <Information Processing System> FIG. 1 is a diagram showing a configuration example of an information processing system. As shown in FIG. 1, the information processing system 100 is composed of, for example, an information processing device 101, a learning device 102, and an information processing server 103, and is connected to each other via a network 104. In the information processing system 100, the information processing device 101, the learning device 102, and the information processing server 103 may be configured to be connected to the network 104 in a multiple connection manner instead of a single connection. For example, the information processing server 103 may be composed of a first server device having high-speed computing resources and a second server device having a large-capacity storage, and may be configured to be connected to each other via the network 104.

[0011] The information processing apparatus 101 generates a document image 113 and transmits it to the information processing server 103. For example, the information processing apparatus 101 is realized by an MFP (Multi-Function Peripheral) having a plurality of functions such as printing, scanning, and FAX, and functions as an image acquisition unit 151. The image acquisition unit 151 of the information processing apparatus 101 optically reads a document 111 printed on a recording medium such as paper, performs predetermined scan image processing to generate a document image 113, and transmits it to the information processing server 103. Further, the image acquisition unit 151 of the information processing apparatus 101 receives FAX data 112 transmitted from a FAX transmitter (not shown), performs predetermined FAX image processing to generate a document image 113, and transmits it to the information processing server 103. Note that the information processing apparatus 101 may be configured to be realized by a PC (Personal Computer) or the like in addition to the MFP having the above-described scanning and FAX functions. Specifically, for example, a document image 113 such as a PDF or JPEG generated using a document creation application operating on a PC as the information processing apparatus 101 may be transmitted to the information processing server 103.

[0012] The learning device 102 functions as a generation unit 152 that generates learning data based on a plurality of document image samples 114, and a learning unit 153 that generates a machine learning model by learning the learning data generated by the generation unit 152 through machine learning. Specifically, the generation unit 152 generates learning data including a character string group included in each of the plurality of document image samples 114 and a correct label indicating the document type of the document image sample 114 (claim, estimate, order form, delivery note, etc.). The learning unit 153 generates a document type determination model 121 for estimating the document type of the document image 113 based on the learning data including the correct label of this document type. In addition, the generation unit 152 generates learning data by assigning a correct label to a character string that is an item value of a detailed item classified into each of predetermined common items among the character string group included in the document image 113. Here, the common item is an item (title, document number, issue date, company name, amount, etc.) that is common to all document types used for the file name or folder name for classifying and storing the document image 113.

[0013] It was stated that a plurality of detailed items are classified into each common item. For example, for the document number, detailed items such as claim number, estimate number, order number, delivery number, etc. are classified. The document number is an identification number for identifying the document in which it is described, and it is necessary to distinguish it from an identification number for identifying a related document different from the described document. Therefore, for each document type, one of the detailed items classified into the document number is defined as the item to be used as the document number, so that an appropriate document number can be specified according to the document type in combination with the document type determination result. As another example, when the common item is the company name, for each document type, one of the detailed items classified into the company name (such as the selling-side company name, purchasing-side company name, consignee company name, etc.) is defined as the item to be used as the company name. This enables an appropriate company name to be specified according to the document type in combination with the document type determination result. The learning unit 153 generates an item value extraction model 122 for extracting a character string of a common item (a candidate character string of a detailed item defined for each document type) based on the learning data generated by the generation unit 152. In the following description, the document type determination model 121 and the item value extraction model 122 are collectively referred to as the machine learning model 115.

[0014] After that, the learning device 102 transmits the generated document type determination model 121 and item value extraction model 122 to the information processing server 103 via the network 104.

[0015] The information processing server 103 functions as an information processing unit 154 and a storage unit 155. First, the information processing unit 154 executes OCR processing on the document image 113 received from the information processing device 101 and acquires a plurality of character string groups included in the document image 113.

[0016] The information processing unit 154 acquires the document type determination model 121 and the item value extraction model 122 from the learning device 102. The information processing unit 154 uses the document type determination model 121 for performing document type determination, and acquires a document type determination result 117 indicating which document type among invoices, estimates, purchase orders, delivery notes, etc. the document image 113 is. Further, the information processing unit 154 uses the item value extraction model 122 for performing item value extraction, and extracts a character string that is an item value of a predetermined common item from among the character string group included in the document image 113. Then, from among the character strings that are the item values of the extracted common items, a character string that is an item value of a detailed item determined for each document type is estimated. Next, the information processing unit 154 acquires an item value extraction result 118 for the acquired document type determination result 117 from the character strings that are the item values of the detailed items determined for each estimated document type.

[0017] When storing the file of the document image 116 that is the processing target in the information processing unit 154 among the document images 113 received from the information processing device 101, the storage unit 155 stores the character strings of the document type determination result 117 and the item value extraction result 118 using them as part of the file name.

[0018] The network 104 is realized by a LAN, a WAN, etc., and is a communication unit that mutually connects the information processing device 101, the learning device 102, and the information processing server 103, and transmits and receives data between the devices.

[0019] <Device Configuration> FIG. 2 is a diagram showing a configuration example of the information processing device 101, the learning device 102, and the information processing server 103 for realizing the information processing system 100 of FIG. 1.

[0020] FIG. 2(a) is a diagram showing the configuration of the information processing device 101. As shown in FIG. 2(a), the information processing device 101 includes a CPU 201, a ROM 202, a RAM 204, a printer device 205, a scanner device 206, a storage 208, an external interface 211, etc., and is mutually connected via a data bus 203.

[0021] The CPU 201 is a control unit for controlling the overall operation in the information processing apparatus 101. By executing the startup program stored in the ROM 202, the CPU 201 starts up the system of the information processing apparatus 101, and by executing the control program stored in the storage 208, the CPU 201 realizes functions such as printing, scanning, and FAX of the information processing apparatus 101.

[0022] The ROM 202 is realized by a non-volatile memory and is a storage unit for storing the startup program for starting up the information processing apparatus 101.

[0023] The data bus 203 is a communication unit for mutually transmitting and receiving data between the devices constituting the information processing apparatus 101.

[0024] The RAM 204 is realized by a volatile memory and is a storage unit used as a work memory when the CPU 201 executes the control program.

[0025] The printer device 205 is an image output device and is a processing unit for printing and outputting the document image inside the information processing apparatus 101 on a recording medium such as paper.

[0026] The scanner device 206 is an image input device and is a processing unit for optically reading a recording medium such as paper on which characters, charts, etc. are printed and acquiring it as a document image.

[0027] The original document conveyance device 207 is realized by an ADF (Auto Document Feeder) or the like and is a processing unit for detecting the original document placed on the document table and conveying the detected original document one by one to the scanner device 206.

[0028] The storage 208 is realized by an HDD (Hard Disk Drive) or the like and is a storage unit for storing the aforementioned control program and document images.

[0029] The input device 209 is realized by a touch panel, hard keys, etc., and is a processing unit for receiving operation inputs from a user who uses the information processing apparatus 101.

[0030] The display device 210 is realized by a liquid crystal display or the like, and is a display unit for displaying and outputting the setting screen of the information processing apparatus 101 to the user.

[0031] The external interface 211 connects the information processing apparatus 101 and the network 104, and is an interface unit for receiving FAX data from a FAX transmitter (not shown) or transmitting a document image to the information processing server 103.

[0032] Figure 2(b) is a diagram showing the configuration of the learning apparatus 102. As shown in Figure 2(b), the learning apparatus 102 is composed of a CPU 231, a ROM 232, a RAM 234, a storage 235, an input device 236, a display device 237, an external interface 238, and a GPU 239, and is mutually connected via a data bus 233.

[0033] The CPU 231 is a control unit for controlling the overall operation in the learning apparatus 102. The CPU 231 starts the system of the learning apparatus 102 by executing the boot program stored in the ROM 232, and generates a machine learning model 115 for performing document type determination and item value extraction by executing the learning program stored in the storage 235.

[0034] The ROM 232 is realized by a non-volatile memory, and is a storage unit for storing a boot program for starting the learning apparatus 102.

[0035] The data bus 233 is a communication unit for mutually transmitting and receiving data between the devices constituting the learning apparatus 102.

[0036] The RAM 234 is realized by a volatile memory and is a storage unit used as a work memory when the CPU 231 executes a learning program.

[0037] The storage 235 is realized by an HDD or the like and is a storage unit for storing the aforementioned learning program and document images.

[0038] The input device 236 is realized by a mouse, a keyboard, or the like and is a processing unit for receiving operation inputs from an engineer who controls the learning device 102.

[0039] The display device 237 is realized by a liquid crystal display or the like and is a display unit for displaying and outputting the setting screen of the learning device 102 to the engineer.

[0040] The external interface 238 connects the learning device 102 and the network 104 and is an interface unit for receiving a document image sample 114 from the outside or transmitting a machine learning model 115 to the information processing server 103.

[0041] The GPU 239 is an arithmetic unit composed of an image processing processor. The GPU 239 executes, for example, an operation for generating a machine learning model 115 using learning data generated from a character string group included in a given document image according to a control command given from the CPU 231.

[0042] Figure 2(c) is a diagram showing the configuration of the information processing server 103. As shown in Figure 2(c), the information processing server 103 is composed of a CPU 261, a ROM 262, a RAM 264, a storage 265, an input device 266, a display device 267, and an external interface 268, and is connected to each other via a data bus 263.

[0043] The CPU 261 is a control unit for controlling the overall operation in the information processing server 103. By executing the boot program stored in the ROM 262, the CPU 261 boots the system of the information processing server 103, and by executing the information processing program stored in the storage 265, the CPU 261 executes information processing such as OCR processing and information extraction.

[0044] The ROM 262 is realized by a non-volatile memory and is a storage unit for storing the boot program for booting the information processing server 103.

[0045] The data bus 263 is a communication unit for mutually transmitting and receiving data between the devices constituting the information processing server 103.

[0046] The RAM 264 is realized by a volatile memory and is a storage unit used as a work memory when the CPU 261 executes the information processing program.

[0047] The storage 265 is realized by an HDD or the like and is a storage unit for storing the aforementioned information processing program, machine learning model 115, document image 116, document type determination result 117, item value extraction result 118, and the like.

[0048] The input device 266 is realized by a mouse, a keyboard, or the like and is a processing unit for receiving an operation input to the information processing server 103 from a user who uses the information processing server 103 or an engineer who controls the information processing server 103.

[0049] The display device 267 is realized by a liquid crystal display or the like and is a display unit for displaying and outputting the setting screen of the information processing server 103 to a user who uses the information processing server 103 or an engineer who controls the information processing server 103.

[0050] The external interface 268 connects the information processing server 103 and the network 104, and is an interface unit for receiving the machine learning model 115 from the learning device 102 or receiving the document image 113 from the information processing device 101.

[0051] <Usage sequence> FIG. 3 is a diagram showing the usage sequence of the information processing system 100 in FIG. 1.

[0052] FIG. 3(a) is a diagram for explaining the processing flow when an engineer who manages the information processing system 100 registers the generation of the machine learning model 115 and the setting information of the information processing system.

[0053] In S301, the learning device 102 acquires a plurality of document image samples 114 prepared by the engineer.

[0054] In S302, the learning device 102 performs OCR processing on the document image sample 114 acquired in S301 to obtain a character string group.

[0055] In S303, the learning device 102 acquires the document type of each document image sample 114 input by the engineer as the correct label assigned to the corresponding document image sample 114. Examples of document types include claim forms, estimates, purchase orders, delivery notes, etc. Then, the learning device 102 uses the character string group acquired in S302 and the correct label indicating the document type assigned to each document image sample 114 as learning data to perform learning of the machine learning model 115 and generate a document type determination model 121.

[0056] In S304, the learning device 102 transmits the document type determination model 121 generated in S303 to the information processing server 103. The information processing server 103 stores the received document type determination model 121 in the storage unit 155.

[0057] In S305, the learning device 102 obtains, as correct labels, predetermined common items assigned to the strings to be extracted by the engineer from among the group of strings acquired in S302. The common items are items common to all document types such as, for example, title, claim number, estimate number, order number, delivery number, billing date, estimate date, order date, shipping date, selling company name, purchasing company name, consignee company name, total amount, etc. Then, the learning device 102 uses the group of strings acquired in S302 and the correct labels indicating the common items assigned to the strings to be extracted as learning data to perform learning of the machine learning model 115 and generate the item value extraction model 122.

[0058] In S306, the learning device 102 transmits the item value extraction model 122 generated in S305 to the information processing server 103. The information processing server 103 stores the received item value extraction model 122 in the storage unit 155.

[0059] In S307, the information processing apparatus 101 displays a setting screen (not shown) on the display device 210, obtains a file storage rule from an engineer via the input device 209, and sets the obtained file storage rule in the information processing server 103. The file storage rule includes a rule for assigning a file name using the document type determination result 117 and the item value extraction result 118 when saving the document image 113, and a definition of the storage destination of the file. For example, as the rule for assigning a file name, information such as "<document type name>_<issue date>_<document number>.pdf" is described. Also, as the storage destination of the file, information such as the folder path of the file storage destination (for example, the root directory of the storage 208) and the folder name "<company name>" for classifying and storing under the folder path is described. Here, the part indicated by <> is replaced with the document type or the character string of the item value in the document type determination result 117 and the item value extraction result 118 according to the content of the character string in <>, and the file name and the folder name are generated. For example, it is assumed that "invoice" is obtained as the document type determination result 117 of the document image 113, and "issue date: January 1, 2000, document number: 100, company name: ABC Co., Ltd." is obtained as the item value extraction result 118. At this time, according to the file storage rule, a folder named "ABC Co., Ltd." is created, and it is saved with the file name "invoice_20000101_100.pdf" under it. As another example, for example, for the document image 113, it is assumed that "purchase order" is obtained as the document type determination result 117, and "issue date: January 1, 2000, document number: 999, company name: XYZ Co., Ltd." is obtained as the item value extraction result 118. Similarly, a folder named "XYZ Co., Ltd." is created and saved with the file name "purchase order_20000101_999.pdf". Here, the file storage rule set by the engineer is the default file storage rule in the information processing system 100 that is applied when the user uses the information processing system 100 without previously setting the file storage rule. As will be described later, this default file storage rule can be updated by the file storage rule given by the user.

[0060] FIG. 3(b) is a diagram for explaining the process flow in which a user who uses the information processing system 100 performs document type determination and item value extraction on the document image 113 obtained using the information processing apparatus 101, confirms / corrects the recognition result, and then saves the file.

[0061] In S311, the information processing apparatus 101 displays a setting screen (not shown) on the display device 210, acquires a file storage rule as a user custom setting from the user via the input device 209, and sets the acquired file storage rule in the information processing server 103. Note that this setting does not need to be performed every time, and it may be performed only when the file storage rule is changed.

[0062] In S312, the information processing apparatus 101 executes scanning of the document 111 placed on the scanner device 206 or the document transport device 207 of the information processing apparatus 101.

[0063] In S313, the information processing apparatus 101 transmits the document image 113 acquired by scanning the document 111 to the information processing server 103.

[0064] In S314, the information processing server 103 executes OCR processing on the document image 113 transmitted in S313, and acquires a character string group included in the document image 113. In S315, the information processing server 103 inputs the character string group acquired in S314 into the document type determination model 121 stored in S303, and determines the document type of the document image 113.

[0065] In S316, the information processing server 103 inputs the character string group included in the document image 113 acquired in S314 into the item value extraction model 122 stored in S305, and extracts a character string that is an item value of a predetermined common item from the character string group included in the document image 113. Subsequently, a character string that is an item value of a detailed item defined for each possible document type is extracted, and the extraction result is stored in the storage 235.

[0066] In S317, the information processing server 103 causes the display device 237 to display, as a recognition result, a character string that is the document type acquired in S315 and S316 and the item value of the detailed item corresponding to the document type.

[0067] In S318, after the information processing server 103 displays the recognition result on the display device 237, when it receives an input for correcting the recognition result from the user via the input device 236, it corrects the recognition result based on the user input.

[0068] In S319, when the information processing server 103 detects a correction by the user, it refers to the extraction result of the item value for the entire document type stored in S316, and identifies the document type and item value that need to be changed in relation to the user's correction. Details of this process will be described later.

[0069] In S320, the information processing server 103 updates the recognition result displayed on the display device 237 based on the document type identified as needing change in S319 and the item value of the detailed item corresponding to the document type.

[0070] In S321, when the information processing server 103 receives an input from the user via the input device 236 to confirm the recognition result displayed on the display device 237, it confirms the recognition result. In S322, the information processing server 103 generates a file name for the document image 113 in accordance with the file storage rules acquired in S306 and S311, using the document type and the character string of each item value confirmed in S321. Then, the information processing server 103 stores the document image 113 with the generated file name in the storage destination specified by the file storage rules.

[0071] <Generation of Machine Learning Model> FIG. 4 is a flowchart for explaining the process by which the learning device 102 generates the document type determination model 121 and the item value extraction model 122, which are the machine learning models 115. Note that the execution program for each step shown in FIG. 4 is stored in any one of the ROM 232, RAM 234, and storage 235 of the learning device 102 and is executed by either the CPU 231 or GPU 239 of the learning device 102.

[0072] In S401, the CPU 231 acquires a plurality of document image samples 114 input from the engineer in S301 of FIG. 3. Specifically, for example, document image samples 114 created with different layouts for each issuing company, which are generally called quasi-standard forms such as invoices, estimates, and purchase orders, are acquired.

[0073] In S402, the CPU 231 executes block selection (BS) processing and OCR processing on the document image sample 114 acquired in S401, and acquires a character string group included in the document image sample 114 as the OCR result. Here, the character string group to be acquired may be handled in units of word separations obtained by dividing the document image by blank spaces or ruled lines using a technique such as block selection (BS), which is a known technique for identifying object units constituting a document image of a quasi-standard form. Further, the character string group to be acquired may be handled, for example, in units of word separations obtained by dividing the text included in the document image using a morphological analysis method.

[0074] In S403, the CPU 231 acquires a first correct label indicating the document type assigned to the document image sample 114 by the engineer via the input device 236.

[0075] In S404, the CPU 231 acquires a second correct label assigned to a character string that is an item value of a common item to be extracted by the engineer via the input device 236 from among the character string group acquired in S402.

[0076] In S405, the CPU 231 uses the character string group acquired in S402 and the first and second correct labels assigned in S403 and S404 to train a document type determination model 121 and an item value extraction model 122 for realizing document type determination and item value extraction. This process is executed by the CPU 231 controlling the GPU 239. Here, both the document type determination model 121 and the item value extraction model 122 are composed of a feature vector conversion processing unit that converts learning data into feature vectors, and an inference processing unit that performs inference based on the converted feature vectors. For the feature vector conversion processing unit, known techniques such as Word2Vec, fastText, BERT, XLNet, ALBERT, etc. can be used. Specifically, for example, by using a pre-trained BERT language model for general articles (e.g., the full text of Wikipedia articles), single character string data can be converted into feature vectors represented by 768-dimensional numerical values. Also, in the inference processing unit, the correct label indicating the document type or common item assigned in S403 and S404 is inferred from the converted feature vectors. Here, for the learning model used in the inference processing unit, algorithms generally known as machine learning algorithms, such as logistic regression, decision tree, random forest, support vector machine, neural network, etc. can be used. Specifically, any machine learning model that inputs the feature vectors converted by the feature vector conversion unit and outputs an inference result indicating a predetermined document type or common item label as the output value of the fully connected layer of the neural network may be used.

[0077] In S406, the CPU 231 transmits the document type determination model 121 generated in S405 to the information processing server 103. The information processing server 103 stores the received document type determination model 121 in the storage unit 155.

[0078] In S407, the CPU 231 transmits the item value extraction model 122 generated in S405 to the information processing server 103. The information processing server 103 stores the received item value extraction model 122 in the storage unit 155.

[0079] <Storage of Document Image Files> FIG. 5 is a flowchart for explaining a process in which the information processing server 103 assigns a file name generated using the character strings of the document type determination result 117 and the item value extraction result 118 to the received document image 113 and stores the document image 113 under the assigned file name. Note that the execution program for each step shown in FIG. 5 is stored in any one of the ROM 262, the RAM 264, and the storage 265 of the information processing server 103 and is executed by the CPU 261 of the information processing server 103.

[0080] In S501, the CPU 261 acquires the document type determination model 121 transmitted from the learning device 102.

[0081] In S502, the CPU 261 acquires the item value extraction model 122 transmitted from the learning device 102.

[0082] In S503, the CPU 261 acquires the file storage rules available in the information processing system 100. The file storage rules acquired here are the default settings acquired from the engineer in S307 or the user custom settings acquired from the user in S311.

[0083] In S504, the CPU 261 acquires the document image 113 from the information processing device 101. As specific examples of the document image 113, FIG. 6(a) shows a document image without a title, and FIG. 6(b) shows a document image in which the document type cannot be specified uniquely from the title. As shown in FIG. 6, the document image 113 includes character strings corresponding to predetermined detailed items. The document image 113 includes detailed items such as titles 601 and 612, a claim date 604 and a claim number 605 corresponding to the issue date and document number of the claim, an order date 602, an order number 603, and a delivery number 606 which are information of the order form and delivery note which are related documents. The document image 113 also includes detailed items such as a selling company name 607, a purchasing company name 608, total amounts 609 and 610, and a payment due date 611. In the following description, it is assumed that the document image 113 shown in FIG. 6(a) is acquired in S504 for explanation.

[0084] In S505, the CPU 261 executes the block selection (BS) process and the OCR process described in S402 on the document image 113, and acquires the character string group included in the document image 113.

[0085] In S506, the CPU 261 determines a plurality of document types that can be the document type of the document image 113 by using the document type determination model 121 acquired in S501. When there is no character string corresponding to the title as shown in FIG. 6(a), or when it is a document that combines "invoice" and "delivery note" as shown in FIG. 6(b), a plurality of document types that can be obtained based on the character strings extracted from the document image 113 are determined. In the present embodiment, it is assumed that the document types that can be an invoice, a quotation, a purchase order, and a delivery note, and among them, the probability of being a purchase order is determined to be the highest.

[0086] Here, with reference to FIG. 7, the method for determining the document type in S506 will be described. For example, when a document image with the description "invoice" in the title is input, as shown in FIG. 7(a), as the output result of the document type determination model 121, probability values (0 to 1) for each document type of invoice, quotation, purchase order, delivery note, and contract are calculated. The CPU 261 determines that the document type is "invoice" according to these probability values. In the present embodiment, as the document type determination result 117, a document type with a probability value equal to or higher than a predetermined threshold, for example, 0.9 or higher, may be adopted. As another example, when a document image with the description "delivery note" in the title is input, as shown in FIG. 7(b), as the output result of the document type determination model 121, since the probability value of "delivery note" is 0.9 or higher than the predetermined threshold, it is determined that the document type is "delivery note". When determining a plurality of possible document types, for example, the threshold may be set to 0.001, or the top four probability values may be selected.

[0087] FIG. 8(a) is a diagram for explaining the process of estimating and outputting the label of the document type determination result 117 for the input of the document image of the "claim" by the document type determination model 121. As shown in FIG. 8(a), the input of the character string group is generally input in order according to a predetermined reading order in the document image in units of character strings called tokens, and is input as data combined with commands such as [CLS] at the head and [SEP] at the delimiter position. Here, the [CLS] at the head can generally be treated as a feature vector indicating the characteristics of the entire character string constituting the document image. Therefore, the document type determination result 117 can be obtained by using the output result of the multi-class classifier with the feature vector as the input.

[0088] In S507, the CPU 261 uses the item value extraction model 122 obtained in S502 to extract a character string that is the item value of a predetermined detailed item from the character string group included in the document image 113. FIG. 8(b) is a conceptual diagram explaining the flow of outputting any label indicating a detailed item for each character string unit from the input of the character string group included in the document image of the "claim form" using the item value extraction model 122. As shown in FIG. 8(b), the input of the character string group is generally input in the order of a predetermined reading order in the document image in units of character strings called tokens, and is input as data combined with commands such as [CLS] at the beginning and [SEP] at the delimiter position. For this input, it is converted into a 768-dimensional feature vector (distributional representation) representing the feature amount of the character string using the above-mentioned BERT, and a multi-class classifier with the converted feature vector as input is used to extract a character string having a label for each detailed item. The example shown in FIG. 8(b) shows an example of obtaining the output result by a multi-class classifier using the BIO tag format generally used in the task of named entity recognition. This multi-class classification method using the BIO tag format is a method of expressing so as to be able to specify the net range of a character string having a label for each detailed item by B (Begin), I (Inside), and O (Outside). Specifically, the B-TITLE tag is given to the leading character string of the title, the I-TITLE tag is given to the character string included within the character string range of the title, and the O tag is given to a character string having no label. Through a series of processes by the item value extraction model 122 shown in FIG. 8(b), a character string corresponding to the detailed item is obtained from the character strings included in the document image of the "claim form". That is, the character string range of I-TITLE following B-TITLE can be extracted, and the character string "claim form" corresponding to the title can be obtained. Similarly, the character string range of I-DATEorder following B-DATEorder can be extracted, and the character string "July 15, 2007" corresponding to the order date (the issue date of the order form) can be obtained. Furthermore, the character string range of I-DATEinvoice following B-DATEinvoice (not shown) can be extracted, and the character string "August 5, 2007" of the invoice date (the issue date of the claim form) can be obtained.

[0089] In S508, based on the character string that is the item value of the detailed item acquired in S507, for all possible document types, the CPU 261 estimates the character string that is the item value of the detailed item corresponding to the common item for each document type. FIG. 9 is a diagram showing a correspondence table between the common items and the detailed items defined for each possible document type in the document image shown in FIG. 6(a). FIG. 10 is a diagram showing, for each document type, the recognition result of the character string that is the item value of the detailed item extracted from the document image 113 shown in FIG. 6 based on the correspondence table shown in FIG. 9. For example, when estimating the character string corresponding to the issue date as the common item, from FIG. 9, when the document type is "invoice", it is the "invoice date", and when the document type is "purchase order", it is the "order date", which are the detailed items classified under the issue date for each document type. Therefore, as the issue date of the invoice, the character string "August 5, 2007" of the invoice date 604 is estimated, and as the issue date of the purchase order, the character string "July 15, 2007" of the order date 602 is estimated. On the other hand, when the document type is a quotation and a delivery note, since the character strings corresponding to the "quotation date" or "shipping date" of the detailed items classified under the issue date are not included in the document image 113, it is assumed that there is no character string for the issue date. The CPU 261 associates the estimation result of the item value of the detailed item corresponding to each document type shown in FIG. 10 obtained as described above with the document image 113 as document information and stores it in the storage unit 155 of the information processing server 103.

[0090] In S509, the CPU 261 performs a confirmation / correction process on the document information including one document type determined based on the document type determination result 117 and the item value extraction result 118 and the item value of the detailed item corresponding thereto. This is performed in response to a confirmation / correction operation by the user on the recognition result of the character string of the document type determination result 117 acquired in S506 and the item value extraction result 118 acquired in S508. Details will be described later with reference to FIG. 5(b).

[0091] In S510, the CPU 261 generates a file with a file name for classifying and storing the document image 113. Here, based on the file storage rule acquired in S503, a file name is generated based on the document type determination result 117 confirmed / modified in S509 and the item value extraction result 118. For example, assume that the file name generation rule of the acquired file storage rule is set as "<issue date>_<document type>_<document number>_<issuing company name>_<total amount>.pdf". At this time, using the character strings such as "invoice" which is the document type determination result 117, and "20070805", "BN0037", "ABC Co., Ltd.", "41,800 yen", etc. which are the item value extraction results 118, a file name 1102 is generated as shown in Fig. 11(a). At this time, the character string which is the item value of the detailed item determined for each document type acquired in S508 may be given as metadata for the document image 113. Specifically, for example, as shown in Fig. 11(b), the character string which is the item value of each detailed item corresponding to all the document types shown in Fig. 10 may be given as metadata 1103. By giving metadata to the file in this way, it becomes possible to correct the recognition result only from the file based on the metadata of the file.

[0092] In S511, the CPU 261 stores the file of the document image with the file name 1102 in a predetermined storage destination folder "ABC Co., Ltd." 1101 based on the file storage rule acquired in S503.

[0093] Note that the result shown in Fig. 10 stored in the storage unit 155 of the information processing server 103 in association with the document image 113 in S508 may be deleted after the storage of the document is completed in S511. Alternatively, the result shown in Fig. 10 may be stored in the storage unit 155 by further associating the user information and the execution date and time. In this case, deletion may be performed at the timing when the user gives a deletion instruction, or automatically after a certain period has elapsed (for example, after one month, or at the end of the following month).

[0094] <Recognition result confirmation / modification process> FIG. 5(b) is a flowchart for explaining the confirmation / correction process of the recognition result of S509 in FIG. 5(a).

[0095] In S521, the CPU 261 causes the display device 237 to display a confirmation screen for presenting document information including one document type for the document image 113 and the item values of the corresponding detailed items based on the document type determination result 117 and the item value extraction result 118. FIG. 12 shows an example of the confirmation screen 1200. The confirmation screen 1200 includes a preview image display screen 1201, a result display screen 1202, and an OK button 1203. On the preview image display screen 1201, a preview image of the document image 113 and rectangles 1211 and 1212 for emphasizing the character strings extracted as the item values of each detailed item are displayed. On the result display screen 1202, character strings representing the document type information and the item values of each detailed item are displayed based on the document type determination result 117 and the item value extraction result 118. For the document type in the document type determination result 117, the character string 1221 of "purchase order" obtained in S506 is displayed, and a pull-down button 1231 for selecting and correcting the document type is displayed beside it. Also, for the detailed items in the item value extraction result 118, character strings 1222 to 1224, which are the item values of the corresponding detailed items for the "purchase order" from the extraction results estimated in S508, are displayed. Also, edit buttons 1232 to 1234 are displayed beside the character strings, and by pressing the edit buttons, corrections can be made when the OCR results of the character strings 1222 to 1224 are incorrect or when the extracted character strings are incorrect. The OK button 1203 can be pressed to confirm the recognition result for the document image 113 when the confirmation and correction of the extraction results are completed.

[0096] In S522, the CPU 261 determines whether there is a user operation received via the input device 236. If it is determined that there is no user operation (in the case of No), the process of S522 is repeated. On the other hand, if an instruction to press the correction button by the user is detected (in the case of YES / correction), the process proceeds to S523, and if an instruction to press the OK button is detected (in the case of YES / OK), the process proceeds to S525.

[0097] In S523, when the CPU 261 detects that the user has modified the recognition result in S522, it identifies the items that need to be modified in relation to the content of the user's modification. FIG. 13(a) shows an example in which the document type is changed from "purchase order" to "invoice" using the pull-down button 1231 on the result display screen 1202 of the confirmation screen 1200. When the CPU 261 detects that the document type has been changed from "purchase order" to "invoice", it obtains, from the recognition results for the detailed items for each document type shown in FIG. 13 stored in the storage unit 155, the character strings that are the item values for each detailed item where the document type is "invoice". Then, it identifies the item values of the document number, issue date, and issuer company name for which modification is necessary. Similarly, when the CPU 261 detects a modification to the item value of a detailed item, it identifies, according to the content of the modification, the item values of the document type or the character strings of the item values other than the modified item value for which modification is necessary. FIG. 13(b) shows an example of editing the character string 1322 for the document number as an example of modifying the item value. First, the CPU 261 detects that the document number has been changed from "QT0037" to "BN0037". Next, from the recognition results for the detailed items for each document type shown in FIG. 10, it identifies the document type that includes "BN0037" as the character string of the document number, and obtains the character strings of each detailed item at that time. Then, it identifies that modification of the character strings for the document type, issue date, and issuer company name is necessary.

[0098] In S524, the CPU 261 updates the recognition result for the item identified as needing correction in S523. For example, when the document type is corrected in Fig. 13(a), as shown in Fig. 13(c), the character string 1321 of the document type on the result display screen 1202, and the character strings 1322 to 1324 of the document number and issue date that are identified as needing correction in relation are updated. Alternatively, for each item value identified as needing correction, a screen for determining whether to apply the correction may be displayed on the result display screen 1202 so that the user can select whether to apply the correction. For example, by displaying the candidate display screen 1340 as shown in Fig. 13(d), the item values that need to be corrected in relation to the correction by the user are presented to the user, and the user may be able to determine whether to apply the correction collectively / individually. In this example, the user can determine whether to apply the correction candidates collectively by pressing the collective application button 1341 and the cancel button 1342. Alternatively, the user can determine whether to apply the correction candidates individually by pressing the application buttons 1343 and 1344 for the individual item values.

[0099] In S525, when the CPU 261 detects that the user has pressed the confirmation button in S522, it confirms the result presented on the confirmation screen 1200 as the recognition result.

[0100] As described above, according to the first embodiment of the present invention, as recognition results of the document type determination result 117 and the item value extraction result 118 for the document image, one or more item values are extracted in advance for each of a plurality of document types. Therefore, it is possible to identify the locations that need to be changed according to the correction content by the user only by searching the already saved recognition results, and it is possible to automatically update or present them to the user as correction candidates with a short response time. Also, even when the document type is not corrected, it is possible to automatically update or present correction candidates for the locations that need to be changed according to the correction content. Thereby, the labor and time required for the user to correct the recognition result can be reduced, and the work efficiency can be improved.

[0101] [Second Embodiment] In the first embodiment, in relation to a single correction of the recognition result by the user, items that required correction were identified and automatically updated or presented to the user as correction candidates. In this embodiment, based on successive corrections by the user, it is possible to narrow down candidates that require correction while presenting correction candidates to the user.

[0102] <Recognition result confirmation / correction processing> The flow of the processing in the second embodiment is shown in FIG. 14. For the same processing as the flowchart of FIG. 5(b) described in the first embodiment, the same numbers are assigned, and only the different parts will be described.

[0103] S521 to S522 are the same as in the first embodiment, and thus the description thereof will be omitted.

[0104] In S1423, when the CPU 261 detects in S522 that the user has corrected the recognition result via the input device 236, the CPU 261 identifies the items and item values that need to be corrected in relation to the user's correction. FIG. 15(a) shows an example in which the issuer company name is edited in the result display screen 1202 of the confirmation screen 1200 in FIG. 12. First, the CPU 261 detects that the issuer company name has been changed from "XYZ Co., Ltd." 1224 to "ABC Co., Ltd." 1504. Subsequently, from among the item value extraction results 118 of S508 stored in the storage unit 155 shown in FIG. 10, the CPU 261 identifies the document types for which the issuer company name is "ABC Co., Ltd." and the item values that need to be changed at that time. Here, it is specified that the document types of invoice, estimate, and delivery note, and the document number and issue date as the item values are the items that need to be changed, respectively. Furthermore, when the correction candidates cannot be narrowed down to one, it is also possible to identify the correction locations necessary for narrowing down the correction candidates. In the example shown in FIG. 15(a), if the document type or the document number is corrected, it becomes possible to uniquely identify the content of all the item values. Therefore, the document type and the document number are identified as the correction locations.

[0105] In S1424, the CPU 261 updates the confirmation screen 1200 based on the multiple correction candidates identified in S1423. For example, as shown in FIG. 15(b), on the result display screen 1202 of the confirmation screen 1200, icons 1511 to 1513 indicating that the items are those requiring correction and a message 1510 notifying that the correction candidates are in the multiple items indicated by the icons are displayed. Alternatively, as shown in FIG. 15(c), on the result display screen 1202, icons 1521 to 1522 indicating that the items are those requiring correction and a message 1520 notifying that the narrowing-down will be completed if the items indicated by the icons are corrected are displayed. Further, based on the information identified in S523, it is also possible to perform a display that enables easy correction of the document type and item values. For example, as the result display screen 1202 for correcting the document type, as shown in FIG. 15(d), the document types narrowed down in S523 may be displayed at the top as candidates in the pull-down menu 1561. Also, as shown in FIG. 15(e), for correcting the document number, a list screen 1570 of correction candidates is displayed on the result display screen 1202, and the character strings 1571 to 1573 of the correction candidates and the buttons 1581 to 1583 for applying them are displayed. Thereby, in addition to the operation of directly correcting the character string 1552 for the document number, the user can perform the correction by a simple operation of pressing the button corresponding to the appropriate character string from the list screen 1570 of the correction candidates.

[0106] Thereafter, when corrections made by the user for the second and subsequent times are detected, the CPU 261 similarly repeats the processes of S1423 to 1424. Thereby, other item values that further require changes are identified and updated. Since the subsequent processes are the same as those of the first embodiment, the description thereof is omitted.

[0107] As described above, according to the second embodiment, when corrections of the document type and item values are required multiple times, by performing narrowing-down of the correction candidates that match the correction each time one correction is made and presenting them to the user, the user can efficiently correct all the items.

[0108] [Third Embodiment] In the first and second embodiments, a mechanism is provided that identifies correction candidates according to user corrections and automatically updates or presents the correction candidates. In this embodiment, when no correction candidate that matches the user's correction is found, the user is prompted to confirm the correction content, and if there is an error in the correction content, the error can be corrected.

[0109] <Correction processing of recognition results> The flow of the process in the third embodiment is shown in FIG. 16. Note that the same processes as those in the flowchart of FIG. 5(b) described in the first embodiment are given the same numbers, and only the parts with differences will be described.

[0110] Since S521 to S522 are the same as those in the first embodiment, the description thereof is omitted.

[0111] In S1623, when the CPU 261 detects in S522 that the user has modified the recognition result via the input device 236, it identifies the items that need to be modified in relation to the user's modification. Figure 17(a) shows an example of editing the issue date for the confirmation screen in Figure 11. When the CPU 261 detects that the issue date 1703 has been changed from "July 15, 2007" to "July 5, 2007", it identifies the document type that matches the issue date from among the item value extraction results 118 of S508 stored in the storage unit 155 shown in Figure 10. Here, since no document type that matches the issue date is found, it is determined that there may be an error in the user's modified content. Therefore, in the example shown in this Figure 17(a), the modified issue date 1703 is identified as the item that needs to be modified in S1623. On the other hand, Figure 17(b) shows an example of editing the document number 1702 and the issue date 1703 for the result display screen 1202 in Figure 11. In response to the change of the document number 1702 from "QT0037" to "DL0037", the CPU 261 identifies "claim form" as the document type that matches the document type. Also, in response to the change of the issue date 1703 from "July 15, 2007" to "August 5, 2007", it identifies "delivery note" as the document type that matches the issue date. Here, since the document types identified from each change do not match, the CPU 261 determines that the user's modified content is inconsistent. Therefore, in the example shown in this Figure 17(b), the modified document number 1702 and issue date 1703 are identified as the items that need to be modified in S1623.

[0112] In S1624, based on the result of S1623, the CPU 261 updates the result display screen 1202 of the confirmation screen 1200 so as to indicate the possibility that the user's correction content is incorrect. FIGS. 17(c) and 17(d) respectively show the result display screen 1202 for the operations in FIGS. 17(a) and 17(b). As shown in FIG. 17(c), an icon 1721 indicating an item that needs correction is displayed on the result display screen 1202, and an error message 1720 notifying that the correction content may be incorrect is displayed below the result display screen 1202. Also, as shown in FIG. 17(d), icons 1731 to 1732 indicating items that need correction are displayed on the result display screen 1202, and an error message 1730 notifying that the correction content may be inconsistent is displayed below the result display screen 1202.

[0113] Thereafter, when a correction by the user for the second time and subsequent times is detected, the CPU 261 similarly repeats the processes of S1623 to 1624. As a result, other item values that further need to be changed are specified and updated. Since the subsequent processes are the same as those in the first embodiment, the description thereof is omitted.

[0114] As described above, according to the third embodiment, when the document type and item values that need to be changed cannot be specified based on the document type and the correction content of the item values, it is possible to notify the user that the correction content is incorrect or that inconsistent corrections are being made. Thereby, the user can easily review the correction content and perform the correction work efficiently. (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or device via a network or a storage medium, and having one or more processors in a computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions. The present disclosure includes the following configurations and methods. [Configuration 1] Determination means for determining the document type of the input document image, Extraction means for extracting, from the document image, a character string as an item value of a predetermined item for the document type determined by the determination means, Presentation means for causing a display means to display document information including the determined document type and the character string extracted by the extraction means for the items determined for the document type, Receiving means for receiving a correction to the document information, Comprising, When the determination means determines a plurality of document types, and the extraction means extracts character strings of predetermined items for each of the plurality of document types, and has a plurality of pieces of the document information corresponding to the plurality of document types, the presentation means causes the display means to display the document information that matches the correction received by the receiving means among the plurality of pieces of document information, An information processing apparatus characterized by the above. [Configuration 2] At least one of the document type included in the document information that matches the correction and the character string of the item determined for the document type is included in the correction, The information processing apparatus according to Configuration 1, characterized by the above. [Configuration 3] The presentation means replaces and presents the document information that matches the correction with the document information presented before receiving the correction, The information processing apparatus according to Configuration 1 or 2, characterized by the above. [Configuration 4] The presentation means notifies that there is a correction candidate for a part different from the document information that matches the correction among the document information presented before receiving the correction, The information processing apparatus according to Configuration 1 or 2, characterized by the above. [Configuration 5] When the receiving means newly receives a correction, the presentation means narrows down and notifies the correction candidates that match the newly received correction among the correction candidates, The information processing apparatus according to Configuration 4, characterized by the above. [Configuration 6] When there is no document information in the determined document type and the extracted character string that matches the correction, the prompting means notifies that there is an error in the correction. The information processing apparatus according to any one of Configurations 1 to 5, characterized in that. [Configuration 7] When the correction includes corrections to a plurality of locations in the document information and there is no document information in the determined document type and the extracted character string that matches all of the corrections to the plurality of locations, the prompting means notifies that there is a contradiction in the corrections to the plurality of locations. The information processing apparatus according to any one of Configurations 1 to 5, characterized in that. [Configuration 8] The determination means calculates a probability value determined as the document type of the document image for a predetermined document type based on the character string included in the document image, and determines the document type for which the probability value is equal to or greater than a predetermined threshold value as the document type of the document image. The information processing apparatus according to any one of Configurations 1 to 7, characterized in that. [Configuration 9] The predetermined items for each of the plurality of document types are items determined to be used as items corresponding to the common items for each document type among the items classified into common items common to all document types. The information processing apparatus according to any one of Configurations 1 to 8, characterized in that. [Configuration 10] A step of determining the document type of the input document image; A step of extracting, from the document image, a character string as an item value of a predetermined item for the document type determined in the determining step; A step of causing a display means to display document information including the determined document type and the character string extracted in the extracting step for the item determined for the document type; A step of accepting a correction to the document information; Comprising The step of causing the display is to cause the display means to display document information that matches the correction received in the receiving step from among a plurality of pieces of the document information, when the determining step determines a plurality of document types, and the extracting step extracts character strings of items predetermined for each of the plurality of document types, and there are a plurality of pieces of the document information corresponding to the plurality of document types. An information processing method characterized by the above. [Configuration 11] A program for causing a computer to function as the information processing apparatus according to any one of Configurations 1 to 9.

Claims

1. Determination means for determining the document type of the input document image, Extraction means for extracting, from the document image, a character string as an item value of a predetermined item for the document type determined by the determination means, Presentation means for causing a display means to display document information including the determined document type and the character string extracted by the extraction means for the items determined for the document type, Acceptance means for accepting a correction to the document information, comprising: When the determination means determines a plurality of document types, and the extraction means extracts character strings of predetermined items for each of the plurality of document types, and has a plurality of pieces of the document information corresponding to the plurality of document types, the presentation means causes the display means to display the document information that matches the correction accepted by the acceptance means among the plurality of pieces of document information. An information processing apparatus characterized by the above.

2. At least one of the document type included in the document information that matches the correction and the character string of the item determined for the document type is included in the correction. The information processing apparatus according to claim 1, characterized by the above.

3. The presentation means replaces and presents the document information that matches the correction with the document information presented before accepting the correction. The information processing apparatus according to claim 1, characterized by the above.

4. The presentation means notifies that there is a correction candidate for a part different from the document information that matches the correction among the document information presented before accepting the correction. The information processing apparatus according to claim 1, characterized by the above.

5. When the acceptance means newly accepts a correction, the presentation means narrows down and notifies the correction candidates that match the newly accepted correction among the correction candidates. The information processing apparatus according to claim 4, characterized by the above.

6. When there is no document information that matches the correction in the determined document type and the extracted character string, the presentation means notifies that there is an error in the correction. The information processing apparatus according to claim 1, characterized by the above.

7. When the correction includes corrections to a plurality of parts of the document information, and there is no document information that matches all of the corrections to the plurality of parts in the determined document type and the extracted character string, the presentation means notifies that there is a contradiction in the corrections to the plurality of parts. The information processing apparatus according to claim 1, characterized by the above.

8. The determination means calculates a probability value determined as the document type of the document image for a predetermined document type based on the character string included in the document image, and determines the document type whose probability value is equal to or greater than a predetermined threshold value as the document type of the document image. The information processing apparatus according to claim 1, characterized in that.

9. The items predetermined for each of the plurality of document types are items determined to be used as items corresponding to the common items for each document type among the items classified into common items common to all document types. The information processing apparatus according to any one of claims 1 to 8, characterized in that.

10. A step of determining the document type of the input document image; A step of extracting, from the document image, a character string as an item value of an item predetermined for the document type determined in the determining step; A step of causing a display means to display document information including the determined document type and the character string extracted in the extracting step for the item determined for the document type; A step of receiving a correction to the document information; Comprising: In the step of causing the display, when the determining step determines a plurality of document types, and the extracting step extracts character strings of items predetermined for each of the plurality of document types, and has a plurality of the document information corresponding to the plurality of document types, the display means displays the document information that matches the correction received in the receiving step from among the plurality of the document information. An information processing method, characterized in that.

11. A program for causing a computer to function as the information processing apparatus according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for filing document

    JP1996221558A