Information processing device, information processing method, and program
The information processing apparatus uses machine learning to cluster and translate transaction documents, addressing inefficiencies in manual document verification by automating the association of relevant documents, thereby reducing labor and enhancing accounting process efficiency.
Patent Information
- Application Number
- PCT/JP2024/000846
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-24
AI Technical Summary
Existing accounting processing systems require manual verification of document matches, which is time-consuming and inefficient due to the need for accounting staff to confirm document relevance and correctness, leading to increased labor costs.
An information processing apparatus that utilizes machine learning to generate clusters of related transaction documents by extracting and associating summary data from various document types, enabling automatic linking and translation of relevant documents.
Facilitates easy association and translation of related documents, reducing the labor required for document verification and improving efficiency in accounting processes.
Smart Images

Figure JP2024000846_24072025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present invention relates to an information processing device, an information processing method, and a program.
[0002] There is known an accounting processing device that performs journal entries using AI (artificial intelligence) that has previously performed machine learning based on training data and learned to select combinations of account items corresponding to transaction detail information (for example, Patent Document 1).
[0003] Japanese Patent Application Laid-Open No. 2021-165967
[0004] However, in the accounting procedures covered by the prior art, accounting staff had to refer to related documents to confirm the fact of delivery as a pre-processing step for journal entries, as well as confirm that the contract or order details, delivery details, and invoice details match, and check for any improper accounting procedures, which was a time-consuming task of checking documents scattered throughout the company.
[0005] The present invention has been made in consideration of these points, and aims to make it possible to easily link related documents in a transaction so as to reduce the effort required to check transaction documents.
[0006] The information processing device of the first aspect of the present invention includes an acquisition unit that acquires document data, which is electronic data of documents related to transactions; a summary generation unit that generates a plurality of summary data representing summaries extracted from each of a plurality of document data including document data of a plurality of different types of documents, with information for associating each document; and a cluster generation unit that generates clusters by dividing the plurality of summary data into related cases.
[0007] The summary generation unit may generate summary data by extracting information indicating transaction details contained in the document data, which information is used to associate the transaction details with transaction details in other document data indicating the same transaction.
[0008] The summary generation unit generates the summary data including different types of information for each type of document, and the information processing device may further have a memory unit that stores a machine learning model that generates the cluster when the type of document indicated by the document data and the summary data extracted from the document data are input, and the cluster generation unit may input the multiple summary data into the machine learning model and generate the cluster.
[0009] The acquisition unit may further acquire target document data, which is document data that is the target for determining related cases, the summary generation unit may generate target summary data, which is the summary data corresponding to the target document data, and the information processing device may further have an identification unit that identifies cases to which the document corresponding to the target summary data is related based on the cluster generated by the cluster generation unit, and an output unit that outputs document data related to the cases identified by the identification unit.
[0010] The information processing device may further have a reception unit that receives a selection of document data to be displayed, and may further have an identification unit that identifies document data that has a predetermined relationship with the cluster to which the document data received by the reception unit belongs as candidate document data that is a candidate related to the document data, and an output unit that associates the document data to be displayed with the candidate document data and outputs them.
[0011] The system may further include a reception unit that receives a selection of document data to be journalized, and a journal data generation unit that generates journal data that journalizes transactions indicated by the document data based on information contained in the document data to be journalized and other document data corresponding to the cluster to which the document data belongs.
[0012] The journal data generation unit may generate journal data that journalizes transactions indicated by the document data based on items contained in the document data to be journalized and specified items contained in other document data that correspond to the cluster to which the document data belongs, and that correspond to the document type of the document data.
[0013] The system may further include a memory unit that associates and stores document data with journal entry data indicating the journal entries for the transactions indicated by the document data, a reception unit that accepts a selection of document data to be journalized, and a journal entry data generation unit that generates the journal entry data for the document data to be journalized based on the document data to be journalized and the journal entry data associated with document data that belongs to a cluster that has a predetermined relationship with the cluster to which the document data belongs.
[0014] An information processing method of a second aspect of the present invention includes the steps of: acquiring document data, which is electronic data of documents related to a transaction, executed by a computer; generating a plurality of summary data representing summaries obtained by extracting information for associating each document from each of a plurality of document data, which includes document data of a plurality of different types of documents; and generating clusters by dividing the plurality of summary data into related cases.
[0015] In a third aspect of the program of the present invention, a computer is caused to execute the steps of acquiring document data, which is electronic data of documents related to a transaction; generating a plurality of summary data representing summaries obtained by extracting information for associating each document from each of a plurality of document data, which includes document data of a plurality of different types of documents; and generating clusters by dividing the plurality of summary data into related cases.
[0016] According to the present invention, documents related to a transaction can be easily linked.
[0017] It is a diagram for explaining an overview of an information processing system S. It is a block diagram showing the configuration of an information processing device 1. It is a diagram showing an example of summary data. It is a diagram showing an example of the data structure of cluster information. It is a flowchart showing the flow of processing in the information processing device 1.
[0018] [Outline of Information Processing System S] Fig. 1 is a diagram for explaining an outline of the information processing system S. The configuration of the information processing system S will be explained with reference to Fig. 1(a). The information processing system S is a system for exchanging transaction documents and processing the exchanged transaction documents. The information processing system S has an information processing device 1 and an information terminal 2.
[0019] The information processing device 1 is a device for electronically exchanging and processing transaction documents. The information processing device 1 is, for example, a server. The information processing device 1 accumulates transaction documents and associates documents related to the same transaction by clustering the accumulated transaction documents. The information processing device 1 may journalize received invoices, etc., based on information on other transaction documents belonging to the cluster.
[0020] The information terminal 2 is a terminal used by an accounting processing staff member. As an example, the information terminal 2 is a smartphone, a tablet, or a personal computer. As an example, the information terminal 2 transmits documents to be processed to the information processing device 1. As an example, the information terminal 2 displays a screen based on control from the information processing device 1 and transmits the contents of user operations to the information processing device 1.
[0021] The processing in the information processing system S will be described with reference to Fig. 1(b). The information processing device 1 acquires document data. The document data is electronic data obtained by digitizing documents related to transactions. Examples of the document data include digitized estimates, approval documents, contracts, purchase orders, delivery notes, and invoices.
[0022] Document data includes information that is not necessary for linking transaction documents. For example, a quote may include information such as the person in charge of the quote, the quote number, and the quote's validity period, but this information is not necessarily required for linking transaction documents. Furthermore, a contract may include various clauses that do not affect the linking of transaction documents (e.g., clauses regarding compensation for damages, late fees, risk of loss, and termination). Therefore, by removing such information and determining the relationship between documents, the accuracy of linking documents can be improved.
[0023] The information processing device 1 generates summary data D2 based on the acquired document data. Summary data D2 is information extracted from the information contained in the document data, which is the content of the transaction and is used to associate the transaction with other document data. Summary data D2 is data generated by extracting information used to associate different types of document data issued for the same transaction. In other words, summary data D2 includes information used to associate the transaction content of the document data with other document data issued for the same transaction. Summary data D2 includes information used to associate each document from multiple document data sets, including document data for multiple different types of documents. Summary data D2 includes, for example, information such as the transaction counterparty, transaction date, document issue date, transaction object, and quantity.
[0024] The information processing device 1 generates clusters (C1, C2, C3, C4) divided for each related case based on the generated summary data. The information processing device 1 generates clusters based on the similarity between the summary data generated from the transaction documents. As a result, documents that have the same transaction parties and objects and are dated closely are presumed to be related transaction documents, so the information processing device 1 can generate clusters divided from the transaction documents (summary data) for each series of transactions (also called cases).
[0025] Based on the selection of document data accepted from information terminal 2, information processing device 1 may display document data belonging to the same cluster as the selected document data on information terminal 2 in association with the selected document data.
[0026] By configuring the information processing system S in this way, it is possible to easily link documents related to a transaction.
[0027] 2 is a block diagram showing the configuration of the information processing device 1. The information processing device 1 has a communication unit 11, a storage unit 12, and a control unit 13. The control unit 13 has an acquisition unit 131, a summary generation unit 132, a cluster generation unit 133, an identification unit 134, an output unit 135, a reception unit 136, and a journal data generation unit 137.
[0028] The communication unit 11 is a communication interface for transmitting and receiving data to and from other devices via a network. The storage unit 12 is a storage medium including a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD), a hard disk drive, etc. The storage unit 12 pre-stores a program to be executed by the control unit 13. The storage unit 12 stores a machine learning model that, when inputted with summary data extracted from each of a plurality of document data, generates a cluster to which each of the plurality of summary data belongs. The machine learning model may further receive as input the document type indicated by the document data.
[0029] The control unit 13 is a processor such as a CPU (Central Processing Unit), etc. The control unit 13 executes the programs stored in the storage unit 12 to function as an acquisition unit 131, a summary generation unit 132, a cluster generation unit 133, an identification unit 134, an output unit 135, a reception unit 136, and a journal data generation unit 137.
[0030] The acquisition unit 131 acquires document data, which is electronic data of documents related to a transaction. For example, the acquisition unit 131 may acquire the document data from an external device (not shown), or may acquire document data stored in the storage unit 12. The acquisition unit 131 may acquire the document data by uploading the document data by the information terminal 2. The acquisition unit 131 may acquire the document data in association with the document type, or may identify the type of the acquired document based on items included in the document data.
[0031] The summary generation unit 132 generates multiple summary data representing summaries extracted from multiple document data, including document data for multiple different types of documents, for associating each document. FIG. 3 illustrates an example of the summary data. The summary data includes, for example, information indicating the 5W2H (Five Whats and Means) of the transaction, such as the purpose of the transaction (Why), document number, expected delivery date (When), transaction counterparty (Who), transaction target (What), transaction volume (How many), and transaction amount (How much). The summary data may include different items depending on the type of document received. The summary data may also include information for associating documents related to the transaction, such as an estimate number, contract number, draft number, settlement number, delivery number, and invoice number.
[0032] For example, information to be extracted as summary data for each document type may be predetermined. In this case, the storage unit 12 stores item information that associates document types with items to be extracted in the summary data for each document type. The summary generation unit 132 references the item information, extracts items corresponding to the document data acquired by the acquisition unit 131, and generates summary data.
[0033] The storage unit 12 may also store a trained model that has been trained using document data including document types and items to be extracted for each document type as training data, and that generates summary data when the document data is input. In this case, the summary generation unit 132 generates summary data by inputting the document data into the trained model.
[0034] The cluster generation unit 133 generates clusters by dividing multiple pieces of summary data into related cases. For example, the cluster generation unit 133 extracts features from each piece of summary data, performs clustering based on the similarity between the extracted features, and generates clusters divided into cases. For example, the cluster generation unit 133 generates clusters based on known cosine similarity. The features are vector data that indicate the characteristics of each piece of summary data. For example, the cluster generation unit 133 divides the clusters so that the distance of the feature of the document from the center of the cluster in the feature space is within a predetermined threshold. For example, when multiple pieces of summary data are input, the cluster generation unit 133 performs clustering using a machine learning model that divides each piece of summary data into clusters. The cluster generation unit 133 may output document data in association with the cluster to which the document data belongs.
[0035] The cluster generating unit 133 may store cluster information in the storage unit 12, which associates document data with the cluster to which the document data belongs. FIG. 4 is a diagram showing an example of the data structure of cluster information. The cluster information includes a "document data ID," a "cluster ID," and a "document type." The cluster information may also include information contained in the summary data of each document. The "document data ID" is an ID (identification) for identifying document data. The "cluster ID" is an ID for identifying the cluster to which the document data belongs. The "document type" is the type of the document data.
[0036] By configuring the information processing device 1 in this manner, it is possible to easily link related documents in a transaction.
[0037] If the input document data lacks information necessary to generate summary data, the summary generation unit 132 may generate data for the missing items based on other items. In this case, the item information stored in the storage unit 12 includes definitions (e.g., calculation formulas) of the relationships between items included in the summary data. For example, the definition of the relationships between items defines the consumption tax amount as the subtotal amount multiplied by the consumption tax rate. If the acquired document data lacks an item, the summary generation unit 132 generates data to supplement the missing item based on the definitions of the relationships between items included in the item information. For example, if the document data lacks an item for the consumption tax amount, the summary generation unit 132 calculates the consumption tax amount by multiplying the subtotal amount by the consumption tax rate, and then generates summary data by adding an item for the calculated consumption tax amount.
[0038] The cluster generation unit 133 may filter the summary data based on conditions such as the transaction counterparty, transaction amount, and transaction time. As an example, the cluster generation unit 133 extracts summary data corresponding to the same transaction counterparty from the target summary data, and performs clustering on the extracted summary data. Configuring the cluster generation unit 133 in this way improves the accuracy of linking related cases.
[0039] Since different types of documents contain different information, the summary data may contain different information for each document type. The summary generation unit 132 generates summary data containing different types of information for each document type. For example, summary data extracted from a "contract" includes the transaction counterparty, the object of the transaction, the contract unit price, the contract period, etc. Summary data extracted from a "purchase order" includes the order date, the transaction counterparty, the object of the transaction, and the order quantity. Summary data extracted from an "invoice" includes the order date, delivery date, the transaction counterparty, the object of the transaction, the transaction quantity, the unit price, the invoice amount, etc. The cluster generation unit 133 inputs the summary data including the document type into a machine learning model and generates clusters.
[0040] The information processing device 1 may be configured to output document data related to the document data obtained after the clusters are generated.
[0041] The acquisition unit 131 acquires target document data, which is document data to be subjected to a related case determination, and the summary generation unit 132 generates target summary data, which is summary data corresponding to the target document data, as described above.
[0042] The identification unit 134 identifies cases to which the document corresponding to the target summary data is related, based on the clusters generated by the cluster generation unit 133. As an example, the identification unit 134 inputs the target summary data into the machine learning model that generated the clusters, thereby outputting the cluster to which the summary data belongs, and identifies the cluster to which the target summary data belongs.
[0043] The output unit 135 outputs document data related to the case identified by the identification unit 134. As an example, the output unit 135 causes the information terminal 2 to display a screen for displaying other document data belonging to the cluster identified by the identification unit 134 as a cluster belonging to the target summary data.
[0044] The information processing device 1 may be configured so that the user can view document data that belong to similar clusters.
[0045] Receiving unit 136 receives a selection of document data to be displayed. Receiving unit 136 causes information terminal 2 to display a screen for receiving the selection of document data to be displayed. The user operates information terminal 2 to select the document data to be displayed. Receiving unit 136 obtains information indicating the document data selected by the user from information terminal 2.
[0046] The identifying unit 134 identifies, as candidate document data that is a candidate related to the document data, document data that has a predetermined relationship with the cluster to which the document data accepted by the accepting unit 136 belongs. Specifically, the identifying unit 134 identifies, as candidate document data, document data that belongs to a cluster within a predetermined distance from the selected document to be displayed.
[0047] The output unit 135 outputs the document data to be displayed and the candidate document data in association with each other. The output unit 135 causes the information terminal 2 to display a screen for displaying the document data to be displayed and the candidate document data identified by the identification unit 134.
[0048] The reception unit 136 receives the selection of document data to be journalized. The reception unit 136 displays a screen for receiving the selection of documents to be journalized on the information terminal 2 and receives the user's selection operation. The journal data generation unit 137 generates journal data by journalizing transactions indicated by the document data based on information contained in the document data to be journalized and other document data corresponding to the cluster to which the document data belongs. Specifically, the journal data generation unit 137 generates journal data by journalizing transactions indicated by the document data based on items contained in the document data to be journalized and predetermined items contained in other document data corresponding to the cluster to which the document data belongs, which items correspond to the document type of the document data. As an example, the memory unit 12 stores items to be extracted from other documents when creating journal data. As an example, the memory unit 12 stores information to acquire accounting codes and contract purposes contained in a request form. The journal data generation unit 133 refers to the storage unit 12, identifies items to extract from other documents, and references document data belonging to the same cluster as the document data to be journalized to acquire those items. The journal data generation unit 137 then creates journal data based on items such as amounts included in the document data to be journalized and items such as accounting codes acquired from the other documents. The output unit 135 may display the created journal data on the information terminal 2.
[0049] As an example, the storage unit 12 may store a trained model that has been trained using multiple different types of document data and correct journalization data as training data, and that has been trained to output journalization data using multiple different types of document data as input. The journalization data generation unit 137 may input the document data to be journalized and document data that belongs to the same cluster as the document data to the trained model, and output the journalization data.
[0050] In the case of consecutive transactions, invoices that have already been journalized may belong to the same cluster. Therefore, the information processing device 1 may be configured to journalize document data based on other document data that belong to the same cluster.
[0051] In this case, the storage unit 12 stores document data and journal data indicating the journalization of the transaction indicated by the document data in association with each other. As an example, the storage unit 12 stores a document data ID and journal data created for the document data indicated by the document data ID in association with each other. The journal data includes, for example, a summary, an account title, various accounting codes, etc.
[0052] The journal data generation unit 137 generates journal data for the document data to be journalized based on the document data to be journalized and journal data associated with document data belonging to a cluster that has a predetermined relationship with the cluster to which the document data belongs. The journal data generation unit 137 references the memory unit 12 and determines whether journal data has already been created for document data belonging to the same cluster as the selected document data. If journal data has been created for document data belonging to the same cluster, the journal data creation unit 137 creates journal data for the transaction indicated by the selected document data to be journalized based on the journal data already created for document data belonging to the same cluster stored in the memory unit 12. The journal data generation unit 137 may create journal data for the transaction indicated by the document data to be journalized based on, in addition to the document data belonging to the same cluster, journal data created for document data belonging to clusters whose clusters are within a predetermined threshold distance from each other.
[0053] By configuring the information processing device 1 in this way, it is possible to reduce the effort required for journalizing work.
[0054] [Processing Flow in Information Processing Device 1] Fig. 5 is a flowchart showing the processing flow in the information processing device 1. The flowchart shown in Fig. 5 starts from the point in time when an instruction to perform clustering is received.
[0055] The acquisition unit 131 acquires a plurality of document data (S01). The summary generation unit 132 generates summary data for each of the acquired document data (S02). The cluster generation unit 133 divides the summary data generated by the summary generation unit 132 into clusters for each related case (S03).
[0056] Accepting unit 136 accepts the selection of document data to be displayed (S04). Output unit 135 associates the selected document data with document data belonging to the same cluster and displays them on information terminal 2 (S05). Information processing device 1 then ends the process.
[0057] [Effects of the Present Embodiment] As described above, the information processing device 1 provides the effect of easily linking documents related to a transaction.
[0058] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments.
[0059] REFERENCE SIGNS LIST 1 Information processing device 2 Information terminal 11 Communication unit 12 Storage unit 13 Control unit 131 Acquisition unit 132 Summary generation unit 133 Cluster generation unit 134 Identification unit 135 Output unit 136 Reception unit 137 Journal data generation unit
Claims
1. An information processing apparatus comprising: an acquisition unit that acquires document data which is electronic data of documents related to transactions; a summary generation unit that generates a plurality of summary data indicating summaries obtained by extracting information for associating each document from each of the plurality of document data including document data of a plurality of different types of documents; and a cluster generation unit that generates clusters obtained by dividing the plurality of summary data for each related case.
2. The information processing apparatus according to claim 1, wherein the summary generation unit generates summary data which is information indicating the transaction content included in the document data and is obtained by extracting information for associating with the transaction content in other document data indicating the same transaction.
3. The summary generation unit generates the summary data including different types of information for each type of the documents. The information processing apparatus further includes a storage unit that stores a machine learning model for generating the clusters. When the type of the document indicated by the document data and the summary data extracted from the document data are input, the cluster generation unit inputs the plurality of summary data into the machine learning model to generate the clusters. The information processing apparatus according to claim 1.
4. The acquisition unit further acquires target document data which is document data to be determined for related cases. The summary generation unit generates target summary data which is the summary data corresponding to the target document data. The information processing apparatus further includes: a specifying unit that specifies a case related to the document corresponding to the target summary data based on the clusters generated by the cluster generation unit; and an output unit that outputs document data related to the case specified by the specifying unit. The information processing apparatus according to any one of claims 1 to 3.
5. The information processing apparatus according to any one of claims 1 to 3, further comprising a reception unit that receives a selection of document data to be displayed. The information processing apparatus further includes: a specifying unit that specifies, as candidate document data which is a candidate related to the document data, document data having a predetermined relationship with the cluster to which the document data received by the reception unit belongs; and an output unit that outputs the document data to be displayed and the candidate document data in association with each other.
6. A reception unit that receives a selection of document data to be translated, and a translation data generation unit that generates translation data obtained by translating the transaction indicated by the document data based on the information included in the document data to be translated and other document data corresponding to the cluster to which the document data belongs. The information processing apparatus according to any one of claims 1 to 3, further comprising:
7. The translation data generation unit generates translation data obtained by translating the transaction indicated by the document data based on an item included in the document data to be translated and a predetermined item included in other document data corresponding to the cluster to which the document data belongs, the predetermined item corresponding to the type of the document of the document data. The information processing apparatus according to claim 6.
8. A storage unit that stores in association document data and translation data indicating a translation of a transaction indicated by the document data, a reception unit that receives a selection of document data to be translated, and based on the document data to be translated and the translation data associated with document data belonging to a cluster having a predetermined relationship with the cluster to which the document data belongs, a translation data generation unit that generates the translation data for the document data to be translated. The information processing apparatus according to any one of claims 1 to 3, further comprising:
9. An information processing method comprising: a step of obtaining document data which is electronic data of a document related to a transaction; a step of generating a plurality of summary data indicating a summary obtained by extracting information for associating each document from each of the plurality of document data including document data of a plurality of different types of documents; and a step of generating a cluster obtained by dividing the plurality of summary data for each related case.
10. A program for causing a computer to execute: a step of obtaining document data which is electronic data of a document related to a transaction; a step of generating a plurality of summary data indicating a summary obtained by extracting information for associating each document from each of the plurality of document data including document data of a plurality of different types of documents; and a step of generating a cluster obtained by dividing the plurality of summary data for each related case.
Citation Information
Patent Citations
Information processing apparatus, program, and information processing system
JP2020017140A
File management device and file management program
JP2021157589A
Data registration processing method, data registration processing program, and data registration processing apparatus
JP2022101276A
Estimate evaluation support device and estimate evaluation support program
JP2023146850A
Data processing device, data processing method, and program
WO2022091354A1