Voucher information processing system, voucher information processing method, and voucher information processing program

The system addresses the limitation of existing models by using a learning model for general items and a language model for additional items, enabling accurate extraction of user-specific information from supporting documents, particularly in the transportation industry.

JP2026011945AActive Publication Date: 2026-01-23FAST ACCOUNTING INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024112957
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2026-01-23
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

Existing learning models are not fully trained to estimate user-specific usage items, such as 'shipping name' and 'shipper,' which are specific to the transportation industry, limiting their ability to accurately extract relevant information from supporting documents.

Method used

A system comprising a learning model for general items and a language model for additional items, where the learning model outputs general items and the language model outputs additional items based on input data and extraction instructions, allowing for the extraction of user-specific information using a general item extraction unit and an additional item extraction unit.

Benefits of technology

Enables accurate extraction of user-specific information from supporting documents, including items not previously learned by the model, enhancing the system's capability to handle user-specific usage items with high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026011945000001_ABST
    Figure 2026011945000001_ABST
Patent Text Reader

Abstract

To provide a voucher information processing system, a voucher information processing method and a voucher information processing program for obtaining information corresponding to items even when there are use items peculiar to a user who is not sufficiently learned in a learning model.SOLUTION: The voucher information processing system 1 includes the learning model 11 that has machine-learned the general items and machine-learned to output information corresponding to the general items when the voucher data is input. The information processing apparatus includes the general item extraction unit 13 that inputs the voucher data of the voucher to be extracted to the learning model and outputs the information corresponding to the general item, and the additional item extraction unit 14 that includes the language model 12 that outputs the information corresponding to the item when the item, the extraction instruction of the item, and the voucher data are input, inputs the additional item to be extracted and not machine-learned by the learning model, the extraction instruction of the additional item, and the voucher data of the voucher to be extracted to the language model, and outputs the information corresponding to the additional item.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voucher information processing system, a voucher information processing method, and a voucher information processing program. [Background technology]

[0002] A technology has been proposed for reading bills using an AI-OCR function (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] https: / / www.ricoh.co.jp / service / cloud-ocr / special / invoice Summary of the Invention [Problem to be solved by the invention]

[0004] By machine learning image data of supporting documents such as invoices and training data containing a set of information corresponding to common items on supporting documents, a learning model is created that outputs information corresponding to the learned items when supporting document data is input, and when supporting document data is input into the learning model, it is possible to output and estimate with high accuracy information corresponding to the items that were previously machine-learned. For example, using invoices as an example, a learning model is constructed to output information corresponding to items such as "billing company name" and "billing amount" that are commonly used in invoice processing, so it is possible to output information corresponding to these items with high accuracy.

[0005] However, the learning model has not been fully trained to estimate user-specific usage items. For example, information corresponding to items such as "shipping name" and "shipper," which are specific to the transportation industry, cannot be obtained by inputting supporting data into the learning model.

[0006] Therefore, the present invention has been made in consideration of these points, and aims to provide a supporting information processing device, a supporting information processing method, and a supporting information processing program supporting information processing device that can obtain information corresponding to user-specific usage items that have not been sufficiently learned in the learning model. [Means for solving the problem]

[0007] The information processing system of the first aspect of the present invention comprises a learning model that has been machine-learned to learn general items and that has been machine-learned to output information corresponding to the general items when supporting document data is input; a language model that outputs information corresponding to the items when an item, an extraction instruction for the item, and supporting document data are input; a general item extraction unit that inputs supporting document data of the supporting document to be extracted into the learning model and outputs information corresponding to the general items; and an additional item extraction unit that inputs additional items that are the subject of extraction and have not been machine-learned by the learning model, extraction instructions for the additional items, and supporting document data of the supporting document to be extracted into the language model and outputs information corresponding to the additional items.

[0008] The learning model is a learning model that has been machine-learned to output information corresponding to general items when evidence image data is input, and the language model is a language model that outputs information corresponding to the items when an item, an extraction instruction for the item, and evidence image data are input, and the general item extraction unit outputs information corresponding to the general items when evidence image data of the evidence to be extracted is input into the learning model, and the additional item extraction unit outputs information corresponding to the additional items when an additional item that is to be extracted and has not been machine-learned in the learning model, an extraction instruction for the additional item, and evidence image data to be extracted are input into the language model.

[0009] The learning model is a learning model that has been machine-learned to output information corresponding to general items when document image data is input, and the language model is a language model that outputs information corresponding to the items when an item, extraction instructions for the item, and document text data are input, and the general item extraction unit outputs information corresponding to the general items when document image data of the document to be extracted is input into the learning model, and the additional item extraction unit outputs information corresponding to the additional items when an additional item that is to be extracted and has not been machine-learned in the learning model, extraction instructions for the additional items, and document text data to be extracted are input into the language model.

[0010] In the additional item extraction unit, the evidence text data to be input to the language model may be evidence text data obtained as a result of applying a character recognition technique to evidence image data of the evidence to be extracted.

[0011] The voucher may be an estimate, an order, an invoice, a receipt, a delivery note, an inspection slip, a contract, or a bankbook.

[0012] The additional items to be extracted and input to the additional item extraction unit may be additional items designated by the user.

[0013] Specific examples of the additional items may be input to the additional item extraction unit.

[0014] The information processing method of the second aspect of the present invention includes a general item extraction step executed by a computer, in which general items have been machine-learned and supporting data to be extracted is input into the learning model that has been machine-learned to output information corresponding to the general items when supporting data is input, and the method outputs information corresponding to the general items; and an additional item extraction step, in which an additional item that has not been machine-learned in the learning model to be extracted, an instruction to extract the additional item, and supporting data to be extracted are input into a language model that outputs information corresponding to the item when an item, an extraction instruction for the additional item, and supporting data are input, and the method outputs information corresponding to the additional item.

[0015] In the third aspect of the document information processing program of the present invention, a computer is caused to execute a general item extraction step in which document image data or document text data to be extracted is input into a learning model that has already machine-learned general items and has been machine-learned to output information corresponding to the general items when document image data or document text data is input, and the computer outputs information corresponding to the general items; and an additional item extraction step in which an additional item that is the subject of extraction and has not been machine-learned in the learning model, an extraction instruction for the additional item, and the document data to be extracted are input into a language model that, when an item, an extraction instruction for the additional item, and the document data to be extracted, output information corresponding to the additional item, and the additional item is input into a language model that, when an item, an extraction instruction for the additional item, and the document data to be extracted, outputs information corresponding to the additional item. [Effects of the Invention]

[0016] According to the present invention, it is possible to provide a supporting document information processing system, a supporting document information processing method, and a supporting document information processing program supporting document information processing device that can obtain information corresponding to user-specific usage items that have not been sufficiently learned in the learning model. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a diagram showing a configuration of a voucher information processing system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating a configuration of a computer according to an embodiment of the present invention. [Figure 3] FIG. 10 is a diagram showing an example of an invoice image. [Figure 4] FIG. 10 is a diagram illustrating an example of an order sheet image. [Figure 5] FIG. 10 is a flowchart of the evidence information processing according to the first embodiment. [Figure 6] FIG. 10 is a flowchart of the evidence information processing according to the second embodiment. [Figure 7] FIG. 10 is a diagram showing item-related information extracted according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] An embodiment of the present invention will be described below. Fig. 1 is a diagram showing the configuration of a documented evidence information processing system 1 according to an embodiment of the present invention. The documented evidence information processing system 1 includes a learning model 11 that has been machine-learned to learn general items (sometimes referred to as "existing items") and that outputs information corresponding to the general items when documented evidence data (documented evidence image data and documented evidence text data are collectively referred to as "documented evidence data") is input; a language model 12 that outputs information corresponding to the items when an item, an extraction instruction for the item, and documented evidence data are input; a general item extraction unit 13 that inputs documented evidence data 15 of a documented evidence to be extracted into the learning model 11 and outputs information 18 corresponding to the general items; and an additional item extraction unit 14 that inputs additional items 16 that are to be extracted but have not been machine-learned by the learning model, extraction instructions 17 for the additional items, and documented evidence data 15 of a documented evidence to be extracted into the language model and outputs information 19 corresponding to the additional items.

[0019] In this specification, "evidence" refers to documents related to the movement of money or changes in rights and obligations, such as contracts, approval documents, estimates, purchase orders, order forms, delivery notes, inspection slips, acceptance slips, invoices, bankbooks, and other passbooks. Also, in this specification, "accounting evidence" refers to evidence related to payment for goods and services, such as estimates, purchase orders, order forms, and invoices. The present invention is particularly effective when applied to accounting evidence, as it allows users to obtain the information they want with a high degree of accuracy.

[0020] In this specification, "general items" refer to typical items used by most users. For example, in the case of an invoice, these are items such as "billing company name" and "billing amount." "Additional items" refer to items that are used only by specific users and are not machine-learned in the learning model. For example, these are items such as "shipper" and "shipping name" on purchase orders used by transportation companies.

[0021] In this specification, "evidence image data" refers to image data of a voucher. Evidence image data is expressed, for example, by the position coordinates of each pixel of the evidence image, the color and intensity of the pixels, etc. "Evidence text data" refers to text data that represents the content of the evidence. Evidence text data is typically expressed as a single line of continuous text, such as letters, numbers, and symbols, such as "AAA Corporation S1000224 Management Department July 1, 2024..." "Evidence data" refers to data related to evidence, and collectively refers to evidence image data and evidence text data. Evidence text data can be obtained by a person viewing and inputting evidence image data, or by applying character recognition technology such as OCR (Optical Character Recognition) to evidence image data.

[0022] The learning model in this embodiment is constructed by storing a large amount of teacher data (correct answer data) that pairs supporting data with information corresponding to general items in the supporting data, and then performing machine learning.When supporting data is input into the learning model, information corresponding to the items is output.

[0023] The language model in this embodiment is a model of human language using the probability of word occurrence. When an instruction is given to the language model in language, an answer is output. In this embodiment, the instruction is, for example, "Extract information corresponding to item XX from the evidence data of the evidence to be extracted." The larger the model size (number of parameters) of the language model, the more accurate the answer, but when the number of parameters exceeds 10 million, the accuracy becomes practical. Language models include vision-LLM (Large Language Models) that can process image data, and language models that can only process text data.

[0024] The functions of the evidential information processing system 1 according to this embodiment are realized by a computer 2. FIG. 2 is a diagram showing the configuration of the computer. The input unit 21 is a component that inputs information to the storage unit, and includes, for example, a keyboard, a mouse, a digital camera, an Internet connection, an intranet connection, etc. The storage unit 23 is a component that stores (records) information. Methods for inputting evidential data to the storage unit 23 using the input unit 21 include taking a photo of the evidential data with a digital camera or uploading evidential data stored on a server from the server via an intranet connection. The storage unit 23 stores a learning model 11 and a language model 12. The calculation unit 24 is a component that processes information. The calculation unit 24 outputs general item information by processing information based on the evidential data and learning model stored in the storage unit 23, and outputs additional item information based on the evidential data, additional items, and additional item extraction instructions stored in the storage unit 23. The output unit 22 is, for example, a display. The display displays information corresponding to the general items and additional items of the evidential data to be extracted output by the calculation unit 24.

[0025] Figure 3 shows an example of a typical invoice image. Invoice image 301 contains the following fields: "Document Title" (302), "Destination" (303), "Billing Source" (304), "Invoice Number" (305), "Billing Date" (306), "Billing Period" (307), "Item Name," "Quantity," "Unit Price," "Amount," and "Billed Amount" (308), and "Payment Account" (309). All of these fields are typical of invoices and fall under the category of general fields. Since the learning model has already learned about general fields, simply inputting invoice image data 301 into the learning model allows the user to obtain all of these necessary fields with high accuracy. FIG. 4 shows an example of an image of a purchase order used by a shipping company. Item 402 is a general item, so if purchase order image data 401 is input into the learning model, "AAA Transport Co., Ltd." can be extracted with high accuracy as information corresponding to "destination." However, "Shipper Name" (403) and "Ship Name" (404) are items specific to purchase orders used by shipping companies and are not machine-learned in the learning model. Therefore, even if order image data 401 is input into the learning model, information corresponding to "Shipper Name" and "Ship Name" cannot be extracted. In this embodiment, by inputting the purchase order image data 401 and additional items "Shipper Name" and "Ship Name" into the language model, and inputting instructions for extracting additional items such as "Shipper Name" and "Ship Name" into the language model, such as "Extract Shipper Name" and "Ship Name" from the order image data 401," information corresponding to the additional items "Shipper Name" and "Ship Name" can be extracted.

[0026] By inputting specific examples such as "shipper name" and "ship name" into the learning model, it is possible to estimate the information corresponding to the additional items with even greater accuracy.

[0027] Figure 7 shows the general item information and additional item information extracted from the purchase order image data 401. The extracted item information can be displayed on a screen as needed, and can also be linked to a deadline management system to manage deadlines, display alerts when the port entry date approaches, and search for required evidence by specifying an item or keyword.

[0028] As described above, according to this embodiment, even if there are additional items that are user-specific usage items that have not been sufficiently learned in the learning model, information corresponding to these items can be obtained. [Example]

[0029] Hereinafter, the evidence information processing method according to the first embodiment will be described with reference to FIG.

[0030] First, the user inputs additional items, which are non-general items used specifically by the user, into the system (S50).

[0031] Next, the image data of the document to be extracted is acquired (S51). Methods for acquiring the image data of the document include, for example, taking a picture of the paper document with a digital camera, saving the image data of the document attached to an email in a specified folder, or downloading the image data of the document saved on a server.

[0032] The acquired document image data is input into a learning model that has been trained to output general items when document image data is input (S52). Then, information corresponding to the general items is extracted and output (S54).

[0033] Meanwhile, the document image data, additional items, and instructions for extracting the additional items are input to a language model capable of processing image information (S55). Then, information corresponding to the additional items is output from the language model (S56).

[0034] The information corresponding to the output general items and additional items is stored in the memory unit, and the information corresponding to the required items can be displayed on the display as needed, or the required evidence can be found by searching using keywords.

[0035] In this embodiment, since information corresponding to the additional items is extracted using a language model capable of processing image information, it is possible to extract information corresponding to the additional items taking into account the layout of the items in the voucher, and it is possible to extract information corresponding to the additional items with high accuracy. In particular, since the layout of items in accounting vouchers such as estimates, invoices, and receipts is standardized, this embodiment makes it possible to extract information corresponding to the additional items with high accuracy. [Example]

[0036] The method for processing evidence information according to the second embodiment will be described below with reference to Fig. 6. Steps S60 to S64 are the same as those in the first embodiment, and therefore will not be described here.

[0037] This embodiment differs from embodiment 1 in steps S65 to S66. In this embodiment, the document image data is converted into text data by character recognition technology such as OCR (S65).

[0038] The converted document text data, additional items, and additional item extraction instructions are input to a language model that cannot process image data and can only process text data (S66).The language model then outputs the additional information.

[0039] Generally, it takes a long time and a great deal of cost to build a language model that can process image information. According to this embodiment, a language model that can process only text data can be used, so the time and cost required to build a language model that can process images can be reduced.

[0040] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]

[0041] 1. Voucher information processing system 11 Learning Model 12 Language Models 13 General item extraction part 14 Additional item extraction part 15. Evidence data 17 Instructions for extracting additional items 18 General information 19. Information applicable to additional items 21 Input section 22 Output section 23 Memory section 24 Arithmetic section 301 Invoice Image 401 Purchase Order Image

Claims

1. A learning model that has been machine-learned to learn general items and that outputs information corresponding to the general items when supporting document data is input; a language model that outputs information corresponding to an item when an item, an instruction to extract the item, and evidence data are input; a general item extraction unit that inputs evidence data of the evidence to be extracted into the learning model and outputs information corresponding to the general items; an additional item extraction unit that inputs additional items to be extracted in the language model and not machine-learned by the learning model, instructions to extract the additional items, and evidence data of the evidence to be extracted, and outputs information corresponding to the additional items; Evidence information processing system.

2. The learning model is a learning model that has been machine-trained to output information corresponding to general items when evidence image data is input, the language model is a language model that outputs information corresponding to an item when an item, an instruction to extract the item, and evidence image data are input; the general item extraction unit outputs information corresponding to general items when the document image data of the document to be extracted is input to the learning model; the additional item extraction unit, when receiving additional items to be extracted in the language model and not machine-learned in the learning model, an instruction to extract the additional items, and evidence image data to be extracted, outputs information corresponding to the additional items; The evidence information processing system according to claim 1.

3. The learning model is a learning model that has been machine-trained to output information corresponding to general items when evidence image data is input, the language model is a language model that outputs information corresponding to an item when an item, an extraction instruction for the item, and evidence text data are input; the general item extraction unit outputs information corresponding to general items when the document image data of the document to be extracted is input to the learning model; the additional item extraction unit, upon receiving input of an additional item to be extracted in the language model and not yet machine-learned in the learning model, an instruction to extract the additional item, and evidence text data to be extracted, outputs information corresponding to the additional item; The evidence information processing system according to claim 1.

4. In the additional item extraction unit, the evidence text data input to the language model is evidence text data obtained as a result of applying character recognition technology to evidence image data of the evidence to be extracted. The evidence information processing system according to claim 3.

5. The voucher is an estimate, purchase order, invoice, receipt, delivery note, inspection slip, contract, or bankbook.

4. The evidence information processing system according to claim 1.

6. The additional items to be extracted and input to the additional item extraction unit are additional items designated by a user.

4. The evidence information processing system according to claim 1.

7. inputting specific examples of the additional items into the additional item extraction unit; 4. The evidence information processing system according to claim 1.

8. The computer executes a general item extraction step in which the general items have been machine-learned and the learning model has been trained to output information corresponding to the general items when supporting data is input, and the supporting data to be extracted is input into the learning model, and the information corresponding to the general items is output; an additional item extraction step of inputting an additional item that has not been machine-learned in the learning model to be extracted, an extraction instruction for the additional item, and supporting data to be extracted into a language model that outputs information corresponding to the item when the item, an extraction instruction for the additional item, and supporting data to be extracted are input, and outputting information corresponding to the additional item; Method for processing evidentiary information.

9. On the computer, a general item extraction step in which the learning model has been machine-learned to learn general items, and when voucher image data or voucher text data is input, the learning model is machine-learned to output information corresponding to the general items, and the learning model inputs the voucher image data or voucher text data to output information corresponding to the general items; an additional item extraction step of inputting an item, an extraction instruction for the item, and supporting data to be extracted, and outputting information corresponding to the additional item to a language model that outputs information corresponding to the item, an additional item that has not been machine-learned in the learning model, an extraction instruction for the additional item, and supporting data to be extracted; A document information processing program that executes the above.

Citation Information

Patent Citations

  • Recognition device and recognition method

    JP2018005462A

  • Information processing device, information processing method, program, and document reading system

    JP2020016946A

  • Information processing device and program

    JP2022032831A

  • Information processing system, item value extraction method, model generation method, and program

    WO2023062798A1