Information processing apparatus, information processing method, and computer readable medium

By extracting the title from the document's read image to determine the category, and using the rules corresponding to the category to extract the item value, the problem of needing to prepare rules in advance for each document type in the existing technology is solved, and more flexible item value extraction is achieved.

CN112580414BActive Publication Date: 2025-12-05FUJIFILM BUSINESS INNOVATION CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010126738.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-30
Filing Date
2020-02-28
Publication Date
2025-12-05
Estimated Expiration
2040-02-28

AI Technical Summary

Technical Problem

Existing technologies require pre-prepared rules for each document type to extract project values, lacking flexibility and adaptability.

Method used

The document category is determined by extracting the title of the category from the read image of the document, and the item value is extracted using the pre-prepared rules corresponding to the category.

Benefits of technology

Even without rules prepared for each document type, it can effectively extract project values, improving the system's flexibility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112580414B_ABST
    Figure CN112580414B_ABST
Patent Text Reader

Abstract

An information processing apparatus, an information processing method, and a computer readable medium. The information processing apparatus has a processor that determines a category of a document to which a document of a kind is classified using a title representing the kind of the document extracted from a read image of the document, extracts a project name from the document using definition information prepared for each category of a document and in which rules for extracting a project value from a document are defined, and for which definition information corresponding to the determined category of the document is prepared.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a computer-readable medium. BACKGROUND

[0002] Sometimes, it is desired to automatically extract an item value for a specific item from a document. For example, in a case where a document is a bill or the like, there are many cases where a form of the bill is determined in advance by a business that issues the bill or the like. Therefore, if the form of the bill is analyzed, and it is determined in which part of the bill an item value is written, then in a case where a bill of the same form is obtained later, the desired item value can be automatically extracted from the bill.

[0003] In addition, a desired item value is generally written on a bill in the vicinity of an item name corresponding to the item. For example, the possibility is high that an item value for an item such as a total amount, that is, a number indicating the total amount, is located immediately below or to the right of a character string indicating an item name such as "total amount" on a bill. Therefore, by finding the character string such as "total amount" from a read image of the bill, the item value can be automatically extracted.

[0004] In any of the above cases, in the related art, information in which rules and the like for extracting an item value are defined is prepared in advance for each type of document.

[0005] As related art documents, for example, Japanese Patent Application Publication No. 2001-202466 and Japanese Patent Application Publication No. 2013-142955 can be cited. SUMMARY

[0006] An object of the present disclosure is to enable extraction of an item value even if definition information in which rules for extracting an item value from a document are defined is not prepared for each type of document.

[0007] According to a first aspect of the present disclosure, there is provided an information processing apparatus having a processor that determines a category of a document to which a document of a type is classified using a title indicating the type of the document extracted from a read image of the document, and extracts an item name from the document using definition information prepared in advance for each category of document and defining rules for extracting an item value from a document, among the definition information, definition information prepared in correspondence with the determined category of the document.

[0008] According to a second aspect of the present disclosure, when a title of the document is extracted, the category of the document is determined by referring to category classification information in which titles classified into a predetermined category of document are included in correspondence with the category.

[0009] According to a third aspect of the present disclosure, the title included in the category classification information is a title of a document classified into a category of the kind, the category including at least an order form, a bill, or a receipt.

[0010] According to a fourth aspect of the present disclosure, in a case where read images of a plurality of documents are continuously acquired, the processor determines a category of a document for each document, and performs a classification process of the plurality of documents in accordance with the determined category of each document.

[0011] According to a fifth aspect of the present disclosure, in a case where a category of a document is specified when a classification process is performed, for a plurality of documents in succession, the documents are classified in a manner that a group is constituted from a document belonging to the specified category of the document to a document immediately before a next document belonging to the specified category of the document or a final document.

[0012] According to a sixth aspect of the present disclosure, each document classified into each group is subjected to a process corresponding to a category of the document.

[0013] According to a seventh aspect of the present disclosure, there is provided a computer-readable medium storing a program causing a computer to execute a process having steps of determining a category of a document to which a document of a kind is classified using a title indicating the kind of the document extracted from a read image of the document, and extracting a project name from the document using definition information corresponding to the determined category of the document, the definition information being prepared in advance for each category of the document and defining a rule of extracting a project value from the document.

[0014] According to an eighth aspect of the present disclosure, there is provided an information processing method including steps of determining a category of a document to which a document of a kind is classified using a title indicating the kind of the document extracted from a read image of the document, and extracting a project name from the document using definition information corresponding to the determined category of the document, the definition information being prepared in advance for each category of the document and defining a rule of extracting a project value from the document.

[0015] Effects of Invention

[0016] According to the above first aspect, even if definition information defining a rule of extracting a project value from a document is not prepared for each kind of the document, the project value can be extracted.

[0017] According to the above second aspect, a document category can be determined using category classification information.

[0018] According to the above third aspect, a document can be determined as a document category corresponding to a kind of the document.

[0019] According to the above-described 4th aspect, the determined document category can be used as a reference to classify a document.

[0020] According to the above-described 5th aspect, a document not belonging to the specified document category is handled as a document attached to a document belonging to the specified document category, thereby enabling formation of a group of documents.

[0021] According to the above-described 6th aspect, a document not belonging to the specified document category can be subjected to a process corresponding to the document category.

[0022] According to the above-described 7th aspect, even if definition information defining a rule for extracting a value of an item from a document is not prepared for each kind of document, the value of the item can be extracted.

[0023] According to the above-described 8th aspect, even if definition information defining a rule for extracting a value of an item from a document is not prepared for each kind of document, the value of the item can be extracted. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 is a structural block diagram of an image forming apparatus in Embodiment 1.

[0025] Figure 2 is a hardware structure diagram of the image forming apparatus in Embodiment 1.

[0026] Figure 3 is a diagram showing an example of a data structure of document category information stored in a document category information storage section in Embodiment 1.

[0027] Figure 4 is a flowchart showing an item value extraction process in Embodiment 1.

[0028] Figure 5 is a structural block diagram of an image forming apparatus in Embodiment 2.

[0029] Figure 6 is a flowchart showing a document classification process in Embodiment 2.

[0030] Figure 7 is a diagram showing a plurality of documents read by a scanner and information associated with each document in Embodiment 2.

[0031] Figure 8 is a conceptual diagram showing a case where a document classification is stored in a folder in Embodiment 2. DETAILED DESCRIPTION

[0032] Hereinafter, a preferred embodiment of the present disclosure will be described with reference to the accompanying drawings. In each of the embodiments described later, a document will be described as an example of a document.

[0033] Embodiment 1

[0034] Figure 1 is a structural block diagram of the image forming apparatus 10 in the present embodiment. Figure 2 is a hardware structure diagram of the image forming apparatus 10 in the present embodiment. The image forming apparatus 10 of the present embodiment mounts an information processing apparatus relating to the present disclosure, and can be realized by a multifunction peripheral mounting various functions such as a copy function, a scan function, and the like. In the present embodiment, the image forming apparatus 10 is realized as a multifunction peripheral mounting a copy function, a scan function, and the like. Figure 2 In the present embodiment, various programs for realizing control of the present apparatus and the characteristic processing functions of the present embodiment described later are stored in the ROM 2. The CPU 1 realizes the action control and various functions of various mechanisms mounted on the present apparatus such as the scanner 6, the printer 7, and the like, in accordance with the various programs stored in the ROM 2. The RAM 3 is used as a work memory at the time of program execution, a communication buffer. The HDD (Hard Disk Drive) 4 stores electronic documents and the like read using the scanner 6. The operation panel 5 performs reception of instructions from a user, display of information. The scanner 6 reads a document set by a user, and accumulates it as an electronic document in the HDD 4 and the like. The printer 7 prints an image on an output sheet in accordance with an instruction from the control program executed by the CPU 1. The network interface (IF) 8 connects to a network, and is used for transmission and reception of electronic data with an external apparatus, or access to the present apparatus via a browser, and the like. The address data bus 9 is connected to various mechanisms as control targets of the CPU 1, and performs communication of data.

[0035] Figure 1 The cloud 30 connected to the image forming apparatus 10 in a communicable manner via a network (not illustrated) such as the Internet is illustrated. The image forming apparatus 10 in the present embodiment has a read image acquisition section 11, an image analysis section 12, a document category determination section 13, a project value extraction section 14, an information provision section 15, a document category information storage section 16, a definition information storage section 17, and a document information storage section 18. In addition, components not used for explanation in the present embodiment are omitted from the drawing.

[0036] The read image acquisition section 11 acquires a read image of a document read by the scanner 6. The image analysis section 12 analyzes the read image acquired by the read image acquisition section 11 and extracts a character string recorded in the document. The document category determination section 13 extracts a title (hereinafter referred to as "title") indicating the kind of the document from the character string extracted by the image analysis section 12 and determines the category of the document based on the extracted title. The item value extraction section 14 extracts an item value from the read image of the document using the definition information prepared in correspondence with the category of the document determined by the document category determination section 13 among the definition information stored in the definition information storage section 17. Further, the document information including the extracted item value is saved in the document information storage section 18. The information provision section 15 provides the user with the document information.

[0037] Here, the "kind of document" and the "category of document" are explained.

[0038] The kind of document is determined based on the providing source (also referred to as "issuing source") and the providing destination (also referred to as "destination") of the document and the category of the document. The category of document (hereinafter also referred to as "document category") is generally also referred to as the kind of document, but indicates each group classified by kind. The document category is determined to some extent restrictively by a manager or the like. In the case of a document, a bill, an estimate, an order form, a receipt, a contract, and the like correspond to the document category. For example, a bill received by a company A from a company B and a bill received by the company A from a company C are bills different in the issuing source, and thus constitute different kinds of documents. However, they respectively constitute documents classified into the same document category of a bill. In the present embodiment, the "kind of document" and the "document category" are thus used to be distinguished explicitly.

[0039] Figure 3 is a drawing showing an example of a data structure of the document category information stored in the document category information storage section 16 in the present embodiment. The document category information is category classification information in which the document category and the title of the document classified into the document category correspond. The document category information is set in advance by a manager of the document or the like. However, even if the document category is the same, the expression of the title in each document is sometimes different if the kind of document is different. For example, if the case where the document category is an estimate is taken as an example, the title of each estimate made by each company can be basically decided freely as the issuing source. Therefore, the character string indicating the title of the estimate made by each company is not necessarily consistent as "estimate", "offer", "preliminary estimate", and the like, and fluctuation in the expression is possible. That is, even if the document category is the same, the title is different if the kind of document is different. Therefore, it is preferable that the title of the document conforming to the document category, particularly the title set to the document by each company, be set in the title set to the document category information.

[0040] The definition information storage section 17 stores definition information that is set in advance for each type of document. In the definition information, a rule for extracting one or more item values from a document classified as the type of document is defined in advance. In the present embodiment, the definition information is prepared for each type of document, rather than for each type of document. The item value extraction section 14 extracts an item value that is an extraction target from a read image of a document that is a processing target, using definition information corresponding to the type of document of the document.

[0041] The document information storage section 18 stores item value information that is generated by the item value extraction section 14 for each document. With respect to the item value information, a group of an item value extracted by the item value extraction section 14 and an item name corresponding to the item value is generated in correspondence with identification information (for example, "document ID") of a document that is a processing target and the type of document of the document.

[0042] The respective components 11 to 15 in the image forming apparatus 10 are realized by a computer mounted on the image forming apparatus 10 and a program that operates using the CPU 1 mounted on the computer. In addition, the respective storage sections 16 to 18 are realized by the HDD 4 mounted on the image forming apparatus 10. Alternatively, the RAM 3 or a storage unit located outside can be used via a network.

[0043] In addition, the program used in the present embodiment can of course be provided from a communication unit or stored in a computer-readable recording medium such as a CD-ROM or a USB memory. The program provided from the communication unit or the recording medium is installed in a computer, and various processes are realized by the CPU of the computer sequentially executing the program.

[0044] Next, the process of extracting an item value from a read image of a document in the present embodiment will be described using the flowchart shown in FIG. 6. Figure 4

[0045] When the user causes the scanner 6 to read a document, the read image acquisition section 11 acquires a read image of the document (step 101). Next, the image analysis section 12 analyzes the acquired read image and extracts a character string recorded in the document (step 102). Specifically, a character string is extracted from the read image of the document using an OCR (Optical Character Recognition) technique. In addition, the "character string" refers to a collection of characters, but sometimes only one character is included in the collection.

[0046] ​Next, the document category determining section 13 extracts a character string that meets a prescribed extraction condition from among the character strings extracted by the image analyzing section 12 as a candidate for the title of the document (step 103). Generally, the title of a document is located at the top of the document and is a character string having a font size of a certain size or more. Therefore, a condition related to the position of the title on such a document and the attribute of the expression word of the title is set in advance as the prescribed extraction condition, and a character string that meets the extraction condition is extracted as a candidate for the title. Also, the document category determining section 13 refers to the document category information storage section 16 and collates the character string extracted as a candidate for the title with each title set in the document category information. If there is a title that coincides with the character string constituting the title candidate, the coinciding title is determined as the title of the document (step 104), and the document category corresponding to the title of the document in the document category information is determined as the document category of the document (step 105). In the present embodiment, the document category of the document is thus determined from the expression of the title in the document.

[0047] In addition, in a case where the document does not belong to any document category, the document is classified into a document category such as "Others".

[0048] When the document category of the document is determined, the item value extracting section 14 acquires the definition information set in correspondence with the document category from the definition information storage section 17 (step 106), and extracts the item value of the item designated by the definition information from the read image of the document (step 107). In a case where the definition information defines the position and region of each item value of the extraction target on the document, the item value extracting section 14 extracts the item value from the designated position and the like of the read image of the document with reference to the definition information. In a case where the definition information does not define the position and region of each item value of the extraction target on the document, but defines the item name corresponding to the item value of the extraction target, the item value extracting section 14 determines the position of the item name from the read image of the document with reference to the definition information, and extracts a character string in the vicinity of the item name as the item value. Further, in a case where the definition information defines the pattern of each item value of the extraction target, such as the data type representing the item value, the item value extracting section 14 extracts a character string belonging to the defined data type from the read image of the document as the item value with reference to the definition information. As for the data type representing the item value, for example, in a case where the item value is a date, it is "YYYY / MM / DD", and the item value extracting section 14 extracts a character string that meets the "YYYY / MM / DD" type as the item value. Also, for example, in a case where the item value is an amount of money, it is a numeric string with a "¥" attached at the beginning, and the item value extracting section 14 extracts the numeric string with the "¥" attached as the item value. The extraction processing of the item value by the item value extracting section 14 can also be performed using a related art technique.

[0049] The item value extraction section 14 generates item value information corresponding to the item name of the item for which the item value was extracted as described above, and stores it in the document information storage section 18 (step 108). More specifically, item value information including the document identification information and the document category to which the document is classified, and the item name of the item extracted from the document and the item value of the item is generated and stored.

[0050] The information providing section 15 provides the generated item value information to, for example, a subsequent step of processing the document, or to the cloud 30 for data management. The method of providing is not particularly limited. For example, it can be sent in file form via a network, or provided using the function of an email or the like.

[0051] Embodiment 2.

[0052] In the above-described Embodiment 1, a case in which the documents are processed one by one is assumed, but in business, sometimes a plurality of documents are processed collectively at the end of the month or the like. In the present embodiment, it is characterized that in a case in which the user causes the scanner 6 to read a plurality of documents continuously, each document having relevance is stored in classified storage.

[0053] Figure 5 is a structural block diagram of the image forming apparatus 10 in the present embodiment. The same symbols are given to the same members as in Embodiment 1, and the description is appropriately omitted. The image forming apparatus 10 in the present embodiment has a structure obtained by adding the document classification processing section 19 to the structure of Embodiment 1.

[0054] The document classification processing section 19, in a case in which the read images of a plurality of documents have been acquired continuously, performs classification processing of the plurality of documents in accordance with the category of each document determined by the document category determination section 13 when the category of each document is determined. The document classification processing section 19 is realized by the coordinated action of the computer mounted on the image forming apparatus 10 and the program that acts using the CPU 1 mounted on the computer.

[0055] Next, the processing of classifying the documents that are the processing targets in the present embodiment will be described using the flowchart shown in Figure 6

[0056] In a case in which the user wants to cause the scanner 6 to read a plurality of documents to perform the classification of the documents described later, the user performs a prescribed operation to cause the operation panel 5 to display a designation screen of the document category. Also, the user inputs and designates the document category that is the basis of the classification from the designation screen. Thus, when the user designates the document category, the image forming apparatus 10 accepts the designated document category (hereinafter, also referred to as "designated document category") (step 201).

[0057] ​Next, the user sets a plurality of documents to be processed to the ADF (Auto Document Feeder) of the image forming apparatus 10 and reads them in order. First, when one document is read, the image forming apparatus 10 performs the item value extraction processing explained in Embodiment 1 (step 202). The contents of the item value extraction processing can be the same as those using Figure 3 The contents are omitted from explanation.

[0058] Here, in a case where the document category of the document to be processed coincides with the specified document category (YES in step 203), the document classification processing section 19 newly generates a group in order to classify and manage the documents (step 204), and registers the document to be processed in the newly generated group (step 205). Also, in a case where there is still an unprocessed document (YES in step 206), the processing returns to step 202, and the item value extraction processing is performed on the document read next from the ADF.

[0059] Here, in a case where the document category of the document to be processed does not coincide with the specified document category (NO in step 203), the document classification processing section 19 registers the document to be processed in the same group as the group in which the immediately preceding document has been generated and registered (step 205). Thereby, the document to be processed is assigned to the same group as the most recent document belonging to the specified document category.

[0060] Further, in a case where the document category of the document to be processed coincides with the specified document category (YES in step 203), the document classification processing section 19 newly generates a group as explained above (step 204). That is, a group different from the already generated group is generated, and the document to be processed is registered in the newly generated group (step 205).

[0061] The above processing is repeated, and when the above processing has been performed on all the documents (NO in step 206), the document classification processing section 19 stores each document in the folder of the corresponding group (step 207). In addition, each folder is provided in the document information storage section 18.

[0062] As explained above, in the present embodiment, the documents are classified in a manner that, for a plurality of consecutive documents, the documents from the document belonging to the specified document category to the immediately preceding document of the next document belonging to the specified document category, or the final document (i.e., the last read document among the plurality of read documents) become a group.

[0063] Further, even if the documents are classified into the same group, the processing of the category to which the document belongs is performed. That is, the item value extraction section 14 extracts from the read image of the document using the definition information set in correspondence with the category of the document, instead of using the definition information set in correspondence with the specified document category, for the document not belonging to the specified document category.

[0064] The above-described document classification processing will be described using a specific example.

[0065] Figure 7 The documents 31a to 31f read continuously are shown. Further, the titles extracted from the documents are shown as "title extraction results" in correspondence with the documents 31a to 31f. Also, the document categories determined from the documents are shown together. For example, the title of the document 31b is "Attachment 1" which is not set in the title of the document category information, and therefore the document category is "other". The same applies to the document 31c. The documents 31a, 31d, 31e, and 31f are determined the document categories in accordance with the set contents of the document category information.

[0066] Here, it is assumed that the user wants to group the plurality of documents based on the bill, and therefore "bill" is specified from the document category specification screen. In this case, since the document category of the document 31a is "bill", the document 31a is processed, and therefore a new group (for example, "group A") is generated in step 204, and the document 31a is registered in the group A. Further, the group A becomes a group to which a later document is registered at the current time point.

[0067] Since the document category of the next document 31b is "other" instead of bill, the document 31b is assigned to the same group A as the document (i.e., the document of the bill which is the most recently processed) 31a processed immediately before in step 205. The same applies to the document 31c.

[0068] Since the document category of the next document 31d is "bill", the document 31d is processed, and therefore a new group (for example, "group B") is generated in step 204, and the document 31d is registered in the group B. Thus, the group B becomes a group to which a later document is registered at the current time point. Since the document category of the next document 31e is "estimate", instead of bill, the document 31e is assigned to the same group B as the document 31d which is the most recently processed bill in step 205.

[0069] As described above, even documents classified into the same group are processed according to their category. That is, for example, for documents 31b and 31c classified into group A, the item value extraction unit 14 extracts the item value based on the corresponding definition information rather than the invoice. Furthermore, for document 31e classified into group B, the item value extraction unit 14 extracts the item value according to the definition information corresponding to the valuation form rather than the invoice.

[0070] Additionally, since document 31f is classified as "Bill", a new group (e.g., "Group C") is generated in step 204. Thus, group B is determined to consist of documents 31d and 31e.

[0071] Figure 8 It is shown Figure 7 A conceptual diagram illustrating a scenario where the illustrated documents are stored in a folder. (Example:) Figure 8 As illustrated, each document 31a to 31f is categorized into a corresponding group and stored. Furthermore, as indicated by the additional label attached to document 31a, each document 31a to 31f is associated with corresponding item value information generated by the item value extraction unit 14.

[0072] According to this implementation method, when a specified document category is specified, multiple documents can be classified by referring to the document category of each document read.

[0073] In this embodiment, the user specifies the document category as the basis for classification in step 201 ("invoice" in the example above). However, assuming no document category is specified, the document classification processing unit 19 may also classify and store documents according to each document category. That is, groups are generated according to each document category such as invoice, estimate, and others, and each document is classified into its respective group.

[0074] In the above embodiments, a document is described as an example of a document, but it can be applied as long as there are multiple types of documents, and is not limited to a document.

[0075] Furthermore, in the above embodiments, the information processing apparatus disclosed herein is described as being mounted on the image forming apparatus 10. However, by configuring it to obtain a read image of a document from the image forming apparatus 10, the information processing apparatus can also be configured as a different device from the image forming apparatus 10. Alternatively, it can be configured to be implemented via the cloud 30. Additionally, it can be configured to utilize other information processing apparatuses to perform part of the processing functions of the image forming apparatus 10, for example... Figure 1 , 5 The image analysis unit 12, etc., shown in the processing functions.

[0076] In addition, in the above-described embodiments, the processor refers to a broad processor, and includes a general-purpose processor (for example, CPU: Central Processing Unit, and the like), a dedicated processor (for example, GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, and the like).

[0077] Furthermore, the operations of the processor in each of the above-described embodiments can be performed not only by one processor but also by a plurality of processors located at physically separate positions in cooperation. Furthermore, the order of the operations of the processor is not limited to the order described in the above-described embodiments, and can be changed as appropriate.

Claims

1. An information processing device, wherein, The information processing device has a processor. When extracting the title indicating the document's category from the read image of the document, the processor determines the document category to which the document of that category belongs by referring to category classification information, which includes the title indicating the document's category corresponding to the pre-determined document category. The project name is extracted from the document by using the definition information prepared in advance according to each category of the document, which contains the rules for extracting project values ​​from the document.

2. The information processing apparatus according to claim 1, wherein, The titles included in the category classification information are the titles of documents classified into that category. The document categories must include at least order forms, invoices, or receipts.

3. The information processing apparatus according to claim 1, wherein, When multiple document images are acquired consecutively, the processor determines the category of each document and performs classification processing on the multiple documents based on the determined categories of each document.

4. The information processing apparatus according to claim 1, wherein, When a document category is specified during classification, for multiple consecutive documents, the documents are classified in groups from the document belonging to the specified document category up to the document immediately preceding the next document belonging to the specified document category or the last document.

5. The information processing apparatus according to claim 3 or 4, wherein, Each document categorized into groups is processed according to its category.

6. A computer-readable medium storing a program that causes a computer to perform processing, wherein... The processing procedure has the following steps: When extracting the title indicating the document's category from the read image of the document, the category of the document to which that category belongs is determined by referring to category classification information, which includes a title indicating the document's category corresponding to a pre-determined document category; and The project name is extracted from the document by using the definition information prepared in advance according to each category of the document, which contains the rules for extracting project values ​​from the document.

7. An information processing method, wherein, The information processing method includes the following steps: When extracting the title indicating the document's category from the read image of the document, the category of the document to which that category belongs is determined by referring to category classification information, which includes a title indicating the document's category corresponding to a pre-determined document category; and The project name is extracted from the document by using the definition information prepared in advance according to each category of the document, which contains the rules for extracting project values ​​from the document.

8. A computer program product comprising a program that causes a computer to perform processing, wherein, The processing procedure has the following steps: When extracting the title indicating the document's category from the read image of the document, the category of the document to which that category belongs is determined by referring to category classification information, which includes a title indicating the document's category corresponding to a pre-determined document category; and The project name is extracted from the document by using the definition information prepared in advance according to each category of the document, which contains the rules for extracting project values ​​from the document.

Citation Information

Patent Citations

  • Slip type discriminator

    JP2001202466A

  • Document processing device and program

    JP2013142955A

  • Form recognition apparatus and form recognition method

    JP2014016762A

  • Image forming apparatus, mobile device, method for classifying document, and computer readable recording medium

    US20170155783A1