Document Verification System
The document verification method and system address the inconsistency challenge by using a trained decision model to automatically verify accounting data and certificates, enhancing accuracy and efficiency beyond OCR limitations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SEMICON ENERGY LAB CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-29
AI Technical Summary
The challenge of checking the consistency between accounting data and paper-based certificates is hindered by the variability in certificate formats and the inaccuracies of optical character recognition (OCR), requiring manual verification for reconciliation.
A document verification method and system that uses a processing unit to extract local image features and text data from certificates, applying a trained decision model to determine matching consistency with accounting data, independent of OCR accuracy.
Enables automatic and accurate verification of consistency between supporting documents and accounting data, reducing reliance on manual checks and improving efficiency.
Smart Images

Figure 2026123183000001_ABST
Abstract
Description
Technical Field
[0001] One aspect of the present invention relates to a certificate verification method. Another aspect of the present invention relates to a certificate verification system.
Background Art
[0002] As accounting processes for certificates, there are tasks such as data entry work for accounting data and auditing work for accounting reports. Accounting software that supports the accounting processes of certificates is becoming widespread. However, since certificates are often on paper media, these tasks need to be performed while checking each certificate one by one even when using accounting software. Therefore, the work efficiency is low. In recent years, technological developments have been carried out to improve the work efficiency of data entry for accounting data by using optical character recognition (OCR). Patent Document 1 discloses accounting software in which OCR is performed on a certificate to obtain a completed electronic form, and the accounting data is registered in an accounting ledger.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] As one of the auditing tasks for accounting reports, there is a task (reconciliation) of checking the consistency between the accounting data registered in accounting software and the certificates. Since many certificates are on paper media and furthermore, the formats of certificates differ for each company, it is difficult to mechanically read the certificates. Therefore, the confirmation of the consistency between the accounting data and the certificates has to rely on human eyes.
[0005] It is possible to capture paper documents as images using devices such as image scanners and extract text from the images using OCR. However, the text extracted by OCR contains various kinds of noise and may not be complete. In other words, the accuracy of the text extracted by OCR may be insufficient for checking accounting data.
[0006] Furthermore, when numerical data is extracted using OCR, it is difficult to determine what the numbers represent solely through OCR. Therefore, a human needs to determine which accounting data registered in the accounting software the numerical data extracted by OCR should be compared against.
[0007] Therefore, one aspect of the present invention aims to provide a document verification method that automatically checks the consistency between supporting documents and accounting data. Another aspect of the present invention aims to provide a document verification method that checks the consistency between supporting documents and accounting data regardless of the performance of the OCR. Another aspect of the present invention aims to provide a document verification system that automatically checks the consistency between supporting documents and accounting data. Another aspect of the present invention aims to provide a document verification system that checks the consistency between supporting documents and accounting data regardless of the performance of the OCR. Note that "automatically checking" refers to performing some or all of the consistency check between supporting documents and accounting data using a system. Therefore, "automatically" as described herein can be rephrased as "systematically".
[0008] Furthermore, the description of these problems does not preclude the existence of other problems. Moreover, one aspect of the present invention does not need to solve all of these problems. Other problems will naturally become apparent from the description in the specification, drawings, and claims, and it is possible to extract other problems from the description in the specification, drawings, and claims. [Means for solving the problem]
[0009] One aspect of the present invention is a document verification method that uses a processing unit to check the consistency between accounting data and supporting documents, wherein the processing unit receives first accounting data, image data of a first supporting document, and first text data, extracts first local image features from the image data of the first supporting document, and uses a trained decision model to determine whether the first supporting document and the first accounting data match based on the first local image features, the first text data, and the first accounting data, and outputs the result of the determination. The first text data is data extracted from the image data of the first supporting document by optical character recognition.
[0010] Another aspect of the present invention is a document verification method that uses a processing unit to check the consistency between accounting data and supporting documents, wherein the processing unit receives first accounting data, image data of first supporting documents, and first text data, extracts first local image features from the image data of first supporting documents, generates a first vector based on the first accounting data using a trained decision model, generates a second vector based on the first local image features and first text data using the trained decision model, calculates the similarity between the first vector and the second vector, determines whether the first supporting documents and first accounting data match using the calculated similarity, and outputs the result of the determination. The first text data is data extracted from the image data of first supporting documents by optical character recognition.
[0011] Another aspect of the present invention is a document verification method that uses a processing unit to check the consistency between accounting data and supporting documents, wherein the processing unit receives first accounting data and image data of first supporting documents, extracts first text data from the image data of first supporting documents by optical character recognition, extracts first local image features from the image data of first supporting documents, generates a first vector based on the first accounting data using a trained decision model, generates a second vector based on the first local image features and first text data using a trained decision model, calculates the similarity between the first vector and the second vector, determines whether the first supporting documents and first accounting data match using the calculated similarity, and outputs the result of the determination.
[0012] In the above document verification method, it is preferable that the trained judgment model undergoes a first training to generate a vector using the second text data, and after the first training, undergoes a second training to generate a vector using the second local image features, the third text data, and the second accounting data. Furthermore, it is preferable that the second accounting data is data corresponding to the second document, the second local image features are extracted from the image data of the second document, and the third text data is data extracted from the image data of the second document.
[0013] Furthermore, in the above evidence verification method, it is preferable that the second learning is supervised learning.
[0014] Furthermore, in the above-described document verification method, it is preferable that the first accounting data includes data manually entered by the user with reference to the first document.
[0015] Furthermore, in the above-described document verification method, it is preferable that the first accounting data includes machine-entered data based on the first document.
[0016] Another aspect of the present invention is a document verification system comprising a storage unit, a receiving unit, and a processing unit. The storage unit stores a trained judgment model. The receiving unit has the function of receiving first accounting data, image data of a first document, and first text data. The processing unit has the function of extracting first local image features from the image data of the first document, generating a first vector based on the first accounting data using the trained judgment model, generating a second vector based on the first local image features and the first text data using the trained judgment model, calculating the similarity between the first vector and the second vector, and determining whether the first document and the first accounting data match using the calculated similarity. The first text data is data extracted from the image data of the first document by optical character recognition.
[0017] In the above-described document verification system, it is preferable that the trained judgment model undergoes a first training to generate a vector using the second text data, and after the first training, undergoes a second training to generate a vector using the second local image features, the third text data, and the second accounting data. Furthermore, it is preferable that the second accounting data is data corresponding to the second document, the second local image features are extracted from the image data of the second document, and the third text data is data extracted from the image data of the second document.
[0018] Furthermore, in the above-mentioned evidence verification system, it is preferable that the second learning process is supervised learning.
[0019] Furthermore, it is preferable that the above-mentioned document verification system includes a display unit, the display unit having a function to display the result of the determination. [Effects of the Invention]
[0020] According to one aspect of the present invention, it is possible to provide a voucher verification method for automatically checking the consistency between vouchers and accounting data. Also, according to one aspect of the present invention, it is possible to provide a voucher verification method for checking the consistency between vouchers and accounting data regardless of the performance of OCR. Also, according to one aspect of the present invention, it is possible to provide a voucher verification system for automatically checking the consistency between vouchers and accounting data. Also, according to one aspect of the present invention, it is possible to provide a voucher verification system for checking the consistency between vouchers and accounting data regardless of the performance of OCR.
[0021] Note that the effects of one aspect of the present invention are not limited to the effects listed above. The effects listed above do not prevent the existence of other effects. Note that other effects are the effects not mentioned in this item as described below. Effects not mentioned in this item can be derived by those skilled in the art from the descriptions in the specification, drawings, etc., and can be appropriately extracted from these descriptions. Note that one aspect of the present invention has at least one of the effects listed above and / or other effects. Therefore, one aspect of the present invention may, in some cases, not have the effects listed above.
Brief Description of Drawings
[0022] [[ID=e10]] [Figure 1] FIG. 1A is a diagram for explaining an example of a voucher verification system according to one aspect of the present invention. FIG. 1B is a diagram for explaining an example of a processing unit according to one aspect of the present invention. [Figure 2] FIG. 2A is a diagram for explaining an example of a voucher verification system according to one aspect of the present invention. FIG. e2B is a diagram for explaining an example of a processing unit according to one aspect of the present invention. [Figure 3] FIG. 3 is a flowchart showing an example of a voucher verification method according to one aspect of the present invention. [Figure 4] FIG. 4 is a flowchart showing an example of a voucher verification method according to one aspect of the present invention. [Figure 5] FIG. 5 is a flowchart showing an example of a voucher verification method according to one aspect of the present invention. [Figure 6]FIG. 6 is a flowchart showing an example of a method for learning a determination model according to an aspect of the present invention. [Figure 7] FIG. 7 is a diagram for explaining an example of a processing unit according to an aspect of the present invention. [Figure 8] FIG. 8 is a flowchart showing an example of steps according to a processing verification method which is an aspect of the present invention [Figure 9] FIG. 9 is a diagram for explaining an example of a processing unit according to an aspect of the present invention. [Figure 10] FIG. 10 is a flowchart showing an example of steps according to a processing verification method which is an aspect of the present invention. [Figure 11] FIG. 11 is a diagram showing an example of the hardware of a credential verification system. [Figure 12] FIG. 12 is a diagram showing an example of the hardware of a credential verification system.
MODE FOR CARRYING OUT THE INVENTION
[0023] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and those skilled in the art can easily understand that the form and details thereof can be variously changed without departing from the spirit and scope of the present invention. Therefore, the present invention should not be construed as being limited to the description of the embodiments shown below.
[0024] [[ID=!29]] In the configuration of the invention described below, the same reference numerals are commonly used between different drawings for the same part or parts having the same or similar functions, and the repeated description thereof will be omitted. Also, when referring to similar functions, the hatch patterns may be the same and may not be particularly labeled.
[0025] In addition, the positions, sizes, ranges, etc. of each configuration shown in the drawings may not represent the actual positions, sizes, ranges, etc. for the sake of easy understanding. For this reason, the disclosed invention is not necessarily limited to the positions, sizes, ranges, etc. disclosed in the drawings.
[0026] <! Furthermore, it should be noted that the ordinal numbers "1st," "2nd," and "3rd" used in this specification are added to avoid confusion of constituent elements and do not imply any numerical limitation.
[0027] In this specification and other documents, the string information contained in a document may be simply referred to as "document." In other words, when simply referred to as "document," it may refer to the string information contained in the document.
[0028] Furthermore, in this specification, accounting data registered in accounting software that has not been verified for consistency with supporting documents (accounting data that has not been reconciled) may be referred to as unaudited accounting data. Similarly, supporting documents that have not been verified for consistency with accounting data registered in accounting software (supporting documents that have not been reconciled) may be referred to as unaudited supporting documents. In other words, unaudited supporting documents can also be considered supporting documents that are subject to audit, or supporting documents that are the subject of audit.
[0029] Furthermore, in this specification, accounting data registered in the accounting software that has been verified for consistency with supporting documents (accounting data that has been reconciled) may be referred to as audited accounting data. Similarly, supporting documents that have been verified for consistency with accounting data registered in the accounting software (supporting documents that have been reconciled) may be referred to as audited documents.
[0030] In this specification, the conversion of a string (words, numbers, etc., and combinations thereof) into a vector is referred to as "generating a vector based on a string." This vector includes low-dimensional compressed (dimensionally reduced) vectors, vectors using distributed representations, and so on. Therefore, "generating a vector based on a string" as described in this specification can be rephrased as "generating a low-dimensional compressed (dimensionally reduced) vector based on a string," or "generating a vector based on a string using a model that has previously learned distributed representations," etc.
[0031] Optical Character Recognition (OCR) is a mechanism that extracts text data from image data by identifying characters within an image. Note that devices and software equipped with OCR functionality are sometimes simply referred to as OCR. Therefore, the term OCR used in this specification may be rephrased as either a device equipped with OCR functionality or software equipped with OCR functionality.
[0032] (Embodiment 1) In this embodiment, a document verification system and document verification method according to one aspect of the present invention will be described with reference to Figures 1 to 10.
[0033] <Document Verification System> First, one embodiment of the present invention, a document verification system, will be described with reference to Figures 1 and 2. One embodiment of the present invention, a document verification system, is a system for checking the consistency between unaudited documents and accounting data entered based on said unaudited documents.
[0034] Figure 1A shows the configuration of the document verification system 100.
[0035] The document verification system 100 includes at least a processing unit 101. The document verification system 100 shown in Figure 1A includes a processing unit 101, a storage unit 102, and a receiving unit 103.
[0036] The document verification system 100 can be installed on an information processing device such as a personal computer used by the user. Alternatively, the processing unit 101 can be installed on a server and accessed from a client PC via a network.
[0037] The reception unit 103 has the function of receiving data. This data includes accounting data, image data, and text data. The reception unit 103 receives at least accounting data and image data.
[0038] In this embodiment, accounting data consists of string information contained in the document, such as the transaction date, item name, payment amount, and the name of the trading partner company. Image data is the image data of the document. Text data is character data (also called string data) extracted from the image data by optical character recognition (OCR). The image data may contain embedded information such as font name, font size, coordinates, and borders. Hereafter, the information such as font name, font size, coordinates, and borders embedded in the image data will be referred to as supplementary information.
[0039] For example, reception unit 103 accepts unaudited accounting data. Reception unit 103 also accepts image data of unaudited supporting documents. In addition, reception unit 103 may accept text data extracted from the image data of unaudited supporting documents.
[0040] The memory unit 102 stores the trained judgment model. The memory unit 102 may also store data received by the reception unit 103 (such as accounting data and image data).
[0041] The processing unit 101 has the following functions: extracting local image features from image data; generating vectors based on accounting data using a trained decision model; generating vectors based on local image features and text data using a trained decision model; calculating the similarity between two vectors; and determining whether the supporting documents and accounting data match using the calculated similarity.
[0042] Furthermore, the processing unit 101 may have a function to extract supplementary information from image data in addition to a function to extract local image features from image data. In this case, the processing unit 101 has a function to generate vectors based on accounting data using a trained decision model, a function to generate vectors based on either or both of the local image features and supplementary information, as well as text data, using a trained decision model, a function to calculate the similarity between the two vectors, and a function to determine whether the supporting documents and accounting data match using the calculated similarity.
[0043] Furthermore, since the supplementary information is embedded in the image data, it can be extracted from the image data without analyzing local image features. Therefore, the processing unit 101 may have a function to extract supplementary information from image data, a function to generate a vector based on accounting data using a trained decision model, a function to generate a vector based on supplementary information and text data using a trained decision model, a function to calculate the similarity between the two vectors, and a function to determine whether the supporting documents and accounting data match using the calculated similarity.
[0044] Furthermore, the processing unit 101 may have a function to output the result of the determination of whether or not the supporting documents and accounting data match.
[0045] Local image features are features extracted from a specific region of image data. Examples of local image features include SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), and HOG (Histograms of Oriented Gradients).
[0046] As a method for extracting local image features, the computational algorithms for feature extraction described above can be applied. For example, computational algorithms such as SIFT, SURF, and HOG can be used.
[0047] Furthermore, the extraction of local image features may be performed by inference using a neural network. For example, it may be done using a convolutional neural network (CNN).
[0048] The similarity between two vectors can be calculated using, for example, cosine similarity, covariance, unbiased covariance, Pearson's product-moment correlation coefficient, or deviation pattern similarity.
[0049] Figure 1B shows the configuration of the processing unit 101. The processing unit 101 may include a feature extraction unit 101a, a vector generation unit 101b, a calculation unit 101c, and a determination unit 101d, as shown in Figure 1B.
[0050] The feature extraction unit 101a has the function of extracting local image features from image data. In addition to the function of extracting local image features from image data, the feature extraction unit 101a may also have the function of extracting supplementary information from image data. Alternatively, the feature extraction unit 101a may have the function of extracting supplementary information from image data instead of extracting local image features from image data. Furthermore, the feature extraction unit 101a outputs the extracted local image features to, for example, the vector generation unit 101b.
[0051] The vector generation unit 101b has the function of generating vectors based on accounting data. The vector generation unit 101b also has the function of generating vectors based on local image features and text data. Furthermore, the vector generation unit 101b may have the function of generating vectors based on local image features and / or accompanying information and text data. The vector generation unit 101b also may have the function of generating vectors based on accompanying information and text data. The vector generation unit 101b outputs the generated vectors to, for example, the calculation unit 101c.
[0052] The vector is generated using a pre-trained decision model. This pre-trained decision model is stored in the memory unit 102. Therefore, the vector generation unit 101b receives the pre-trained decision model from the memory unit 102 and generates the vector. This pre-trained decision model may also be stored in the memory unit of the processing unit 101.
[0053] The calculation unit 101c has the function of calculating the similarity between two vectors. One of the two vectors is a vector generated based on accounting data, and the other of the two vectors is a vector generated based on local image features and text data, a vector generated based on local image features and / or accompanying information and text data, or a vector generated based on accompanying information and text data. The calculation unit 101c also outputs the calculated similarity to, for example, the determination unit 101d.
[0054] The determination unit 101d has the function of determining whether the supporting documents and accounting data match using similarity scores. For example, the determination unit 101d determines whether the similarity score is greater than a preset threshold. The determination unit 101d also has the function of outputting the determination result.
[0055] By using a processing unit 101 that includes a feature extraction unit 101a, a vector generation unit 101b, a calculation unit 101c, and a determination unit 101d, the consistency between supporting documents and accounting data can be automatically checked.
[0056] In the processing unit 101 shown in Figure 1B, the functions of the processing unit 101 are classified and independent of each other, however, some or all of the functions of the processing unit 101 do not have to be independent. For example, the determination unit 101d may have the functions of the calculation unit 101c. Alternatively, the determination unit 101d may have the functions of the vector generation unit 101b and the functions of the calculation unit 101c.
[0057] Furthermore, when a CNN is used as the trained decision model, the vector generation unit 101b may use the trained decision model to calculate the similarity between two vectors, or to determine whether the supporting documents and accounting data match. In other words, the vector generation unit 101b may have the functions of the calculation unit 101c, or the functions of the decision unit 101d. In this case, the processing unit 101 may be configured without the calculation unit 101c and / or the decision unit 101d.
[0058] The configuration of the processing unit according to one aspect of the present invention is not limited to the configuration of the processing unit 101 shown in Figure 1B. For example, the configuration of the processing unit 101A shown in Figure 2B may also be used.
[0059] The processing unit 101A shown in Figure 2B includes a feature extraction unit 101a, a vector generation unit 101b, a calculation unit 101c, and a determination unit 101d, in addition to an OCR unit 101e.
[0060] The OCR unit 101e has OCR functionality. By including the OCR unit 101e, text data can be extracted from image data. Therefore, the amount of data received by the reception unit 103 can be reduced.
[0061] As shown in Figure 1A, the document verification system 100 may be connected to an optical character reader 110, an input device 130, an output device 140, and a storage device 150, etc., via a network 120.
[0062] Network 120 is a computer network that forms the basis of the World Wide Web (WWW), including the Internet, intranets, extranets, PANs (Personal Area Networks), LANs (Local Area Networks), CANs (Campus Area Networks), MANs (Metropolitan Area Networks), WANs (Wide Area Networks), and GANs (Global Area Networks). Network 120 includes wired and wireless communications.
[0063] The optical character recognition device 110 has the function of extracting strings of characters contained in an image (strings of characters that have been converted into an image) as text data from image data using OCR.
[0064] The input device 130 has the function of reading a paper document and generating an electronic document. For example, an image scanner or a digital camera can be used as the input device 130. In this embodiment, the document is, for example, a document of evidence. The electronic document can be in image file format. In this case, the electronic document can be rephrased as image data.
[0065] Furthermore, the input device 130 may be a device for inputting data. For example, the input device 130 could be a keyboard, a pointing device, or a touch panel. Using the input device 130, the user can input accounting data, etc.
[0066] The output device 140 has the function of outputting data output from the processing unit 101. For example, the output device 140 can be a display, projector, printer, audio output device, or memory.
[0067] The storage device 150 stores accounting data and image data. The storage device 150 may also store text data. The storage device 150 can also be referred to as a database.
[0068] The accounting data stored in the storage device 150 may be audited accounting data, or it may be audited accounting data and unaudited accounting data. Furthermore, the image data stored in the storage device 150 may be image data of audited documents, or it may be image data of audited documents and unaudited documents.
[0069] Furthermore, some or all of the accounting data, image data, etc., stored in the storage device 150 may be stored in the storage unit 102.
[0070] The above describes the configuration of the document verification system 100. Note that the configuration of the document verification system according to one aspect of the present invention is not limited to the configuration of the document verification system 100 shown in Figure 1A. For example, the configuration of the document verification system 100A shown in Figure 2A may also be used.
[0071] The document verification system 100A shown in Figure 2A includes a processing unit 101, a storage unit 102, and a reception unit 103, as well as a display unit 105.
[0072] The display unit 105 has the function of displaying the results of the determination made by the processing unit 101. For example, a display, projector, or printer can be used as the display unit 105. This allows the user to quickly identify accounting data that does not match the supporting documents, or to quickly identify accounting data with low similarity.
[0073] According to one aspect of the present invention, a document verification system can be provided that automatically checks the consistency between supporting documents and accounting data. Furthermore, according to another aspect of the present invention, a document verification system can be provided that checks the consistency between supporting documents and accounting data regardless of the performance of the OCR. Also, according to another aspect of the present invention, a document verification system can be provided that checks the consistency between supporting documents and accounting data using an existing optical character recognition device.
[0074] <Method of verifying evidence> Next, a document verification method according to one aspect of the present invention will be described with reference to Figures 3 to 5. This document verification method according to one aspect of the present invention is a method for checking the consistency between unaudited documents and accounting data entered based on said unaudited documents.
[0075] Before performing the document verification method, prepare accounting data 11, image data 12, and text data 13 based on document 10. Note that the data prepared before performing the document verification method may be accounting data 11 and image data 12 based on document 10. Here, document 10 is assumed to be an unaudited document.
[0076] Accounting data 11 is data registered in the accounting software based on supporting document 10.
[0077] The accounting data 11 may include data manually entered by the user, referring to the supporting document 10. Alternatively, the accounting data 11 may include machine-entered data based on the supporting document 10. In other words, the accounting data 11 may consist only of data manually entered by the user, or only of machine-entered data, or a combination of manually entered and machine-entered data.
[0078] Image data 12 is image data of document 10. If document 10 is in paper format, image data 12 should be created by capturing document 10 using an image scanner and / or a digital camera. If document 10 is electronic data (especially image data), the electronic data itself should be used as image data 12.
[0079] Text data 13 is data extracted from the image data of document 10 using optical character recognition (OCR).
[0080] Figure 3 is a flowchart illustrating an example of a document verification method according to one aspect of the present invention. Figure 3 is also a flowchart illustrating the processing flow performed by a document verification system according to one aspect of the present invention. The document verification method according to one aspect of the present invention is performed using the processing unit described above.
[0081] In the document verification method explained using Figure 3, accounting data 11, image data 12, and text data 13 are prepared before performing the document verification method.
[0082] The evidence verification method has steps S001 to S004, as shown in Figure 3.
[0083] Step S001 is the process in which the processing unit receives accounting data 11, image data 12, and text data 13.
[0084] Step S002 is the process by which the processing unit extracts local image features 14. The processing unit may have a function to extract local image features from image data. This allows the local image features 14 to be extracted from image data 12.
[0085] Local image feature extraction can be performed using computational algorithms such as SIFT, SURF, and HOG, as described above. Alternatively, local image feature extraction may be performed by inference using a neural network. For example, it may be performed using a CNN.
[0086] Step S003 is the process of checking the consistency between the supporting document 10 and the accounting data 11. In other words, step S003 is the process of determining whether or not the supporting document 10 and the accounting data 11 match.
[0087] Step S003 includes steps S011 to S013 shown in Figure 3. Here, steps S011 to S013 will be described in order to explain step S003.
[0088] Step S011 is the process in which the processing unit generates vectors 15 and 16. Each of vectors 15 and 16 is generated using a pre-trained decision model. The pre-trained decision model will be described later.
[0089] Vector 15 is generated based on accounting data 11. Vector 16 is generated based on local image features 14 and text data 13.
[0090] Step S012 is the process in which the processing unit calculates the similarity between vector 15 and vector 16. The similarity can be calculated using cosine similarity, Pearson correlation coefficient, or deviation pattern similarity, as described above.
[0091] Step S013 is the process in which the processing unit determines whether the similarity calculated in step S012 is greater than or equal to a pre-set threshold. If the calculated similarity is greater than or equal to the threshold, the processing unit determines that the document 10 and the accounting data 11 are a match. If the calculated similarity is less than the threshold, the processing unit determines that the document 10 and the accounting data 11 are not a match.
[0092] Users may set or change the above thresholds based on factors such as the accuracy of the trained classification model.
[0093] Steps S011 to S013 allow for a determination of whether the supporting document 10 and the accounting data 11 match. Therefore, the consistency between the supporting document 10 and the accounting data 11 can be checked.
[0094] Step S004 is the process in which the processing unit outputs the result of the determination obtained in step S003.
[0095] Therefore, it is possible to check the consistency between unaudited evidence and the accounting data entered based on that evidence.
[0096] Furthermore, if the image data 12 includes supplementary information, a step in which the processing unit extracts the supplementary information may be included between step S002 and step S003. The processing unit may have a function to extract supplementary information from the image data in addition to a function to extract local image features from the image data. In this case, in step S011, the vector 16 may be generated based on the local image features 14 and the supplementary information, or both, and the text data 13.
[0097] Furthermore, if the image data 12 includes supplementary information, the step in step S002 in which the processing unit extracts local image features 14 may be replaced with a step in which the processing unit extracts supplementary information. The processing unit should preferably have a function to extract supplementary information from the image data. In this case, in step S011, the vector 16 should preferably be generated based on the supplementary information and the text data 13.
[0098] The above is a description of one example of a document verification method. Note that the document verification method according to one aspect of the present invention is not limited to the method described using Figure 3. For example, the document verification method may be performed according to the flow shown in Figure 4.
[0099] Figure 4 is a flowchart illustrating another example of a document verification method according to one aspect of the present invention. Figure 4 is also a flowchart illustrating the processing flow performed by a document verification system according to one aspect of the present invention. The document verification method according to one aspect of the present invention is performed using the processing unit described above.
[0100] In the document verification method explained using Figure 4, accounting data 11 and image data 12 are prepared before performing the document verification method.
[0101] The document verification method described using Figure 4 includes steps S021, S022, S002, S003, and S004. In other words, the document verification method described using Figure 4 differs from the document verification method described using Figure 3 in that it includes steps S021 and S022 instead of step S001.
[0102] Step S021 is the process in which the processing unit receives accounting data 11 and image data 12.
[0103] Step S022 is the process in which the processing unit extracts text data 13 from image data 12. The processing unit may preferably have optical character recognition (OCR) functionality. This allows the text data 13 to be extracted from image data 12.
[0104] After performing step S022, proceed to steps S002, S003, and S004 in order. Note that for steps S002, S003, and S004, refer to the previous explanation.
[0105] Based on the above, it is possible to check the consistency between unaudited evidence and the accounting data entered based on said unaudited evidence. Note that the order in which steps S022 and S002 are executed may be reversed. Also, if the processing unit has a feature extraction unit and an OCR unit (see Figure 2B), steps S022 and S002 may be executed simultaneously.
[0106] In addition, another aspect of the present invention is a document verification method in which step S022 of the document verification method described with reference to Figure 4 is replaced with another step.
[0107] Figure 5 is a flowchart illustrating another example of a document verification method according to one aspect of the present invention. Figure 5 is also a flowchart illustrating the processing flow performed by a document verification system according to one aspect of the present invention. The document verification method according to one aspect of the present invention is performed using the processing unit described above.
[0108] The evidence verification method described using Figure 5 comprises steps S021, S023, S024, S002, S003, and S004. In other words, the evidence verification method described using Figure 5 differs from the evidence verification method described using Figure 4 in that it includes steps S023 and S024 instead of step S022. Steps S021, S002, S003, and S004 can be explained in the previous explanation.
[0109] Step S023 is the process in which the processing unit transmits the image data 12 to the optical character recognition device. The optical character recognition device receives the image data 12 and extracts text data 13 from the image data 12.
[0110] Step S024 is the process in which the processing unit receives the text data 13 extracted in step S023 from the optical character recognition device.
[0111] The above is another example of a document verification method. The document verification method described using Figure 5 is effective when the processing unit does not have OCR functionality. Note that steps S023 and S024 and step S002 may be executed in any order, or they may be executed simultaneously.
[0112] <Pre-trained decision model> Next, we will explain the trained classification model using Figure 6.
[0113] As described above, by using a pre-trained classification model, vectors can be generated based on text data. Furthermore, vectors can be generated based on local image features and text data.
[0114] Figure 6 is a flowchart showing the method for creating a pre-trained classification model. Note that the method for creating a pre-trained classification model can also be described as the method for training the classification model.
[0115] As shown in Figure 6, the method for creating a trained decision model comprises steps S101 and S102. In other words, the decision model is trained by performing steps S101 and S102 in order.
[0116] Step S101 is the process of performing a first training on the decision model. By performing the first training, the decision model becomes able to generate vectors based on text data. Preferably, the text data used for the first training includes audited accounting data. However, the text data used for the first training is not limited to audited accounting data and may also include text data extracted from documents other than supporting documents, such as general documents.
[0117] Step S102 is the process of performing a second training on the judgment model that has undergone the first training. By performing the second training, the judgment model can generate vectors based on local image features and text data.
[0118] The second learning step is preferably supervised learning. For example, it is preferable to perform supervised learning using a training dataset as the second learning step. Here, it is preferable that the training dataset consists of multiple accounting data (the first to the nth accounting data (where n is an integer of 2 or more)), multiple text data (the first to the nth text data), and multiple local image features (the first to the nth local image features). Specifically, it is preferable to use the multiple text data and the multiple local image features as input data, and to use the vectors generated from the multiple accounting data as training data (labels). Through the first learning step, the decision model can generate vectors based on the accounting data, and therefore the vectors generated from the multiple accounting data can be used as training data (labels). Thus, the multiple accounting data may be considered as training data (labels).
[0119] Furthermore, the input of the i-th accounting data (where i is an integer between 1 and n), the extraction of the i-th text data using OCR, and the extraction of the i-th local image features are all performed using the same document. Specifically, the i-th accounting data is registered in the accounting software based on the document. The i-th text data is extracted from the image data of the document using OCR. The i-th local image features are also extracted from the image data of the document.
[0120] Furthermore, the second learning process is preferably performed using audited documents. For example, the above-mentioned multiple accounting data preferably includes audited accounting data. Furthermore, the above-mentioned multiple text data preferably includes text data extracted by OCR from the image data of the audited documents. Furthermore, the above-mentioned multiple local image features preferably include local image features extracted from the image data of the audited documents.
[0121] In the second learning exercise, it is recommended to assign correct labels to audited accounting data and incorrect labels to data obtained from sources other than audited accounting data, ensuring that the number of correctly labeled data points is roughly equal to the number of incorrectly labeled data points.
[0122] Audited accounting data is already registered in the accounting software. Furthermore, when creating ledgers using accounting software, the electronic storage of paper documents is permitted. Therefore, image data of audited documents is often stored in a storage device connected to the device on which the accounting software can run. This means that extracting text data and local image features from the image data of audited documents is straightforward. Consequently, a training dataset consisting of the aforementioned multiple accounting data, multiple text data, and multiple local image features can be easily created.
[0123] The above-mentioned multiple accounting data, multiple text data, and multiple local image features may be stored in a storage unit of the document verification system (for example, the storage unit 102 shown in Figure 1A), or in a storage device connected to the document verification system via a network 120 (for example, the storage device 150 shown in Figure 1A).
[0124] The second type of learning is not limited to supervised learning; it may also be semi-supervised learning. Compared to supervised learning, semi-supervised learning requires fewer training data points in the training dataset, allowing for high-accuracy vector generation even with a limited number of audited documents. Semi-supervised learning is particularly effective in the initial stages of accounting software implementation, when the number of audited documents is small.
[0125] A pre-trained classification model can be created by performing the first and second training steps. In other words, the classification model is trained by performing the first and second training steps. As a result, the pre-trained classification model can generate vectors based on text data. It can also generate vectors based on local image features and text data. Since the first training step is performed before the second training step, the first training step can be called pre-training.
[0126] Furthermore, the decision model may undergo a second training process to enable it to generate vectors based on either or both of the local image features and associated information, and text data. When supervised learning is performed using a training dataset as the second training process, it is preferable that the training dataset consists of multiple accounting data, multiple text data, and either or both of the multiple local image features and multiple associated information.
[0127] Furthermore, the decision model may undergo a second learning process to enable it to generate vectors based on supplementary information and text data. When supervised learning is performed using a training dataset as the second learning process, it is preferable that the training dataset consists of multiple accounting data, multiple text data, and multiple supplementary information.
[0128] The above describes how to create a pre-trained judgment model. Note that the pre-trained judgment model may be created by the processing unit of the document verification system, or by a device other than the document verification system.
[0129] By using this pre-trained judgment model, even if the OCR fails to accurately extract the text information contained in the document, it is possible to generate a vector corresponding to the text information contained in the document. Therefore, the consistency between the document and the accounting data can be checked regardless of the performance of the OCR used to extract the text data. In other words, there is no need to improve the accuracy of the OCR itself. To put it another way, existing OCR can be used.
[0130] <Example 1> This section describes variations of the document verification system. Since the document verification system is related to the document verification method and the trained decision model, this section also includes explanations of variations of the document verification method and the trained decision model.
[0131] First, as an example of a modified document verification system, we will explain, using Figure 7, a processing unit 101B that has a different configuration from the processing unit 101 shown in Figure 1B and the processing unit 101A shown in Figure 2B, among the components of the document verification system.
[0132] Figure 7 illustrates the configuration of the processing unit 101B. The processing unit 101B comprises a feature extraction unit 101a, an inference unit 101f, and a determination unit 101g. The feature extraction unit 101a can be described in the previous explanation.
[0133] The inference unit 101f has the function of performing accounting data inference based on text data and local image features. Alternatively, the inference unit 101f has the function of performing accounting data inference based on one or more selected from image data, local image features, and text data. This function allows the acquisition of string information described in the document and the generation of accounting data. The inference unit 101f also outputs the accounting data generated by this inference to, for example, the determination unit 101g.
[0134] The determination unit 101g has the function of determining whether the supporting documents and accounting data match. For example, the determination unit 101g determines whether the accounting data received by the reception unit 103 matches the accounting data generated by the above inference. The determination unit 101g also has the function of outputting the result of the determination.
[0135] Inference of accounting data is preferably performed using a neural network. For example, it is preferable to use a CNN. Therefore, it is preferable to use a CNN as the trained decision model.
[0136] Furthermore, when a CNN is used as the trained decision model, the trained decision model may be used to determine whether the supporting documents and accounting data match. In other words, the inference unit 101f may have the functions of the decision unit 101g, or the decision unit 101g may have the functions of the inference unit 101f. In this case, the processing unit 101B may be configured not to have the inference unit 101f or the decision unit 101g.
[0137] By using a document verification system equipped with processing unit 101B, the consistency between documents and accounting data can be automatically checked.
[0138] The processing unit 101B has the function of extracting local image features from image data, the function of performing accounting data inference based on the local image features and text data using a trained decision model, and the function of determining whether the supporting documents and accounting data match. The processing unit 101B may also have the function of outputting the result of the determination of whether the supporting documents and accounting data match.
[0139] The processing unit 101B may also have a function to extract local image features from image data, a function to extract supplementary information from image data, a function to perform accounting data inference based on one or more selected from image data, local image features, and text data using a trained decision model, and a function to determine whether or not the supporting documents and accounting data match. Furthermore, the processing unit 101B may also have a function to output the result of the determination of whether or not the supporting documents and accounting data match.
[0140] The above is an explanation of the modified version of the document verification system.
[0141] Next, a modified example of the document verification method will be explained using Figure 8. The document verification method explained using Figure 8 will be performed using a processing verification system equipped with the processing unit 101B shown in Figure 7.
[0142] Figure 8 is a flowchart showing an example of a document verification method. The document verification method described using Figure 8 differs from the document verification method described using Figure 3 in that step S003 has steps S014 and S015 instead of steps S011 to S013.
[0143] The evidence verification method shown in Figure 8 comprises steps S001 to S004. Step S003 also comprises steps S014 and S015. Steps S001, S002, and S004 can be described in the previous explanation.
[0144] Step S014 is the process in which the processing unit 101B performs accounting data inference based on the text data 13 and local image features 14. The accounting data inference is performed using a pre-trained decision model, which will be described later. Hereafter, the accounting data generated by the inference will be referred to as accounting data 11A.
[0145] Step S015 is the process in which the processing unit 101B determines whether accounting data 11A and accounting data 11 match.
[0146] Steps S014 and S015 allow for a determination of whether the document 10 and the accounting data 11 match. Therefore, by using a document verification system that includes step S003 having steps S014 and S015, the consistency between the document and the accounting data can be automatically checked.
[0147] The above is an explanation of variations in the evidence verification method.
[0148] Next, a modified version of the trained decision model will be described. This trained decision model will be used in a document verification method having the steps shown in Figure 8.
[0149] As described above, by using a pre-trained decision model, accounting data inference can be performed based on one or more selected from image data, local image features, and text data. Note that this pre-trained decision model differs from the pre-trained decision model described in the previous section on <Pre-trained decision model> in that the data it generates (output data) is different, and therefore the method for creating the pre-trained decision model (the method for training the decision model) is also different.
[0150] As mentioned above, it is preferable to use a neural network as the decision model. For example, it is preferable to use a CNN, a recurrent neural network (RNN), long short-term memory (LSTM), or an attention mechanism.
[0151] The training of the decision model is preferably supervised learning. For example, it is preferable to perform supervised learning using a training dataset. Here, it is preferable that the training dataset consists of multiple accounting data (first to nth accounting data), multiple text data (first to nth text data), and multiple local image features (first to nth local image features). Specifically, it is preferable to use the multiple text data and the multiple local image features as input data, and the multiple accounting data as training data (labels). The training dataset may also consist of one or more selected from the multiple accounting data, multiple image data, multiple local image features, and multiple text data.
[0152] As explained in the previous section on the <trained judgment model>, the input of the i-th accounting data, the OCR extraction of the i-th text data, and the extraction of the i-th local image features are all performed using the same document.
[0153] Furthermore, it is preferable that the learning is performed using audited documents. For example, the above-mentioned multiple accounting data preferably include audited accounting data. Furthermore, the above-mentioned multiple text data preferably include text data extracted by OCR from the image data of the audited documents. Furthermore, the above-mentioned multiple local image features preferably include local image features extracted from the image data of the audited documents.
[0154] Audited accounting data is already registered in the accounting software. Furthermore, when creating ledgers using accounting software, the electronic storage of paper documents is permitted. Therefore, image data of audited documents is often stored in a storage device connected to the device on which the accounting software can run. This means that extracting text data and local image features from the image data of audited documents is straightforward. Consequently, a training dataset consisting of the aforementioned multiple accounting data, multiple text data, and multiple local image features can be easily created.
[0155] The above-mentioned multiple accounting data, multiple text data, and multiple local image features may be stored in a storage unit of the document verification system (for example, the storage unit 102 shown in Figure 1A), or in a storage device connected to the document verification system via a network 120 (for example, the storage device 150 shown in Figure 1A).
[0156] This learning method is not limited to supervised learning; it may also be semi-supervised learning. Compared to supervised learning, semi-supervised learning requires a smaller amount of training data in the training dataset, allowing for the generation of accounting data with high accuracy even with a small number of audited documents. This is particularly effective in the initial stages of accounting software implementation, when the number of audited documents is limited.
[0157] By performing this learning process, a pre-trained decision model can be created. In other words, the decision model is trained through this learning process. As a result, the pre-trained decision model can perform inference on accounting data based on local image features and text data.
[0158] The above is an explanation of variations of the trained decision model.
[0159] Furthermore, when using the document verification system described in <Modification 1>, the determination of whether the document and accounting data match is made based on the accounting data. When the inferred accounting data is output as a result of the determination, the user can intuitively understand the result of the determination compared to when a similarity score is output. Therefore, the user can quickly identify accounting data that does not match the document.
[0160] <Modification 2> This section describes other variations of the document verification system. Since the document verification system is also related to the document verification method and the trained decision model, this section also includes descriptions of other variations of the document verification method and the trained decision model.
[0161] First, as another variation of the document verification system, we will explain, using Figure 9, a processing unit 101C that has a different configuration from the processing unit 101 shown in Figure 1B, the processing unit 101A shown in Figure 2B, and the processing unit 101B shown in Figure 7, which are all components of the document verification system.
[0162] Figure 9 is a diagram illustrating the configuration of the processing unit 101C. The processing unit 101C comprises a feature extraction unit 101a, a determination unit 101g, an estimation unit 101h, and an inference unit 101i. The feature extraction unit 101a and the determination unit 101g can be described in the previous explanation.
[0163] The estimation unit 101h has the function of estimating the name of a trading partner company based on local image features. This estimation is preferably performed using a pre-trained estimation model. The estimation unit 101h also outputs the name of the trading partner company obtained through this estimation to, for example, the inference unit 101i.
[0164] The inference unit 101i has the function of inferring accounting data based on the company name and text data of the trading partner obtained by the above estimation. This function allows it to obtain string information described in the supporting documents and generate accounting data. The inference unit 101i also outputs the accounting data generated by the inference to, for example, the determination unit 101g.
[0165] Inference of accounting data is preferably performed using a neural network. For example, it is preferable to use a CNN. Therefore, it is preferable to use a CNN as the trained decision model.
[0166] Furthermore, when a CNN is used as the trained decision model, the trained decision model may be used to determine whether the supporting documents and accounting data match. In other words, the inference unit 101i may have the functions of the decision unit 101g, or the decision unit 101g may have the functions of the inference unit 101i. In this case, the processing unit 101C may be configured not to have the inference unit 101i or the decision unit 101g.
[0167] By using a document verification system equipped with processing unit 101C, the consistency between documents and accounting data can be automatically checked.
[0168] The processing unit 101C has the following functions: extracting local image features from image data; estimating the name of a trading partner company based on the local image features using a pre-trained estimation model; inferring accounting data based on the trading partner company name and text data using a pre-trained judgment model; and determining whether the supporting documents and accounting data match. The processing unit 101C may also have a function to output the result of the determination of whether the supporting documents and accounting data match.
[0169] The above describes other variations of the document verification system.
[0170] Next, other variations of the document verification method will be explained using Figure 10. The document verification method explained using Figure 10 will be performed using a processing verification system equipped with the processing unit 101C shown in Figure 9.
[0171] Figure 10 is a flowchart showing an example of a document verification method. The document verification method described using Figure 10 differs from the document verification method described using Figure 8 in that step S003 has steps S016 and S017 instead of step S014.
[0172] The evidence verification method shown in Figure 10 comprises steps S001 to S004. Step S003 also comprises steps S016, S017, and S015. Steps S001, S002, S004, and S015 can be explained in the preceding description.
[0173] Step S016 is the process in which the processing unit 101C estimates the name of the trading partner company based on the local image features 14. It is preferable to use a pre-trained estimation model for estimating the trading partner company name. Hereafter, the trading partner company name obtained through estimation will be referred to as company name 17.
[0174] Step S017 is the process in which the processing unit 101C performs accounting data inference based on the text data 13 and the company name 17. The accounting data inference is performed using a pre-trained decision model. Hereafter, the accounting data generated by the inference will be referred to as accounting data 11A.
[0175] Steps S016, S017, and S015 allow for a determination of whether the document 10 and the accounting data 11 match. Therefore, by using a document verification system that includes step S003 having steps S016, S017, and S015, the consistency between the document and the accounting data can be automatically checked.
[0176] The above is a description of other variations of the evidence verification method.
[0177] Furthermore, the estimation of the trading partner's company name may be performed based on the image data of the document. In this case, local image features do not need to be extracted from the image data of the document. Therefore, the processing unit 101C may not need to include the feature extraction unit 101a. Also, in the document verification method described using Figure 8, step S002 may be omitted.
[0178] <Pre-trained estimation model> Next, we will describe the trained estimation model. This trained estimation model will be used in the evidence verification method having the steps shown in Figure 10.
[0179] As described above, by using a pre-trained estimation model, it is possible to estimate the names of trading partners based on local image features.
[0180] It is preferable to use a neural network as the estimation model. For example, it is preferable to use an RNN, LSTM, or attention mechanism.
[0181] The training of the estimation model is preferably supervised learning. For example, it is preferable to perform supervised learning using a training dataset. Here, it is preferable that the training dataset consists of multiple local image features (the first local image feature to the mth local image feature (where m is an integer of 2 or more)) and multiple business partner company names (the name of the first business partner to the mth business partner). Specifically, it is preferable to use the multiple local image features as input data and the multiple business partner company names as training data (labels).
[0182] Furthermore, the extraction of the j-th local image feature (where j is an integer between 1 and m) and the acquisition of the j-th trading partner company name are performed using the same document.
[0183] Furthermore, it is preferable that the learning is performed using audited documents. For example, it is preferable that the multiple local image features include local image features extracted from the image data of the audited documents. It is also preferable that the names of the multiple trading partners include the names of trading partners obtained from the image data of the audited documents, or from local image features extracted from the image data of the audited documents. The names of the multiple trading partners may also include the names of trading partners obtained from audited accounting data.
[0184] Audited accounting data is already registered in the accounting software. Furthermore, when creating ledgers using accounting software, the electronic storage of paper documents is permitted. Therefore, image data of audited documents is often stored in a storage device connected to a device capable of running the accounting software. This means that extracting local image features from the image data of audited documents is easy. Consequently, a training dataset consisting of the above-mentioned local image features and the names of the above-mentioned trading partners can be easily created.
[0185] The above-mentioned multiple local image features and the names of the above-mentioned multiple trading partners may be stored in a memory unit of the document verification system (for example, the memory unit 102 shown in Figure 1A), or they may be stored in a storage device connected to the document verification system via a network 120 (for example, the storage device 150 shown in Figure 1A).
[0186] This learning method is not limited to supervised learning; it may also be semi-supervised learning. Compared to supervised learning, semi-supervised learning requires a smaller amount of training data in the training dataset, allowing for high-accuracy estimation of trading partner company names even with a limited number of audited documents. This is particularly effective in the initial stages of accounting software implementation, when the number of audited documents is typically small.
[0187] The format of supporting documents varies from company to company. Therefore, local image features extracted from the image data of these documents are effective in estimating the names of trading partners. Consequently, by performing this training, it is possible to estimate the names of trading partners with high accuracy.
[0188] By performing this learning process, a pre-trained estimation model can be created. In other words, by performing this learning process, the estimation model is trained. As a result, the pre-trained estimation model can estimate the names of trading partners based on local image features.
[0189] The above is an explanation of the pre-trained estimation model.
[0190] Next, other variations of the trained decision model will be described. This trained decision model will be used in a document verification method having the steps shown in Figure 10.
[0191] As mentioned above, by using a pre-trained decision model, it is possible to perform accounting data inference based on the names of trading partners and text data.
[0192] It is preferable to use a neural network as the decision model. For example, it is preferable to use an RNN, LSTM, or attention mechanism.
[0193] The training of the decision model is preferably supervised learning. For example, it is preferable to perform supervised learning using a training dataset. Here, it is preferable that the training dataset consists of multiple accounting data (the first to the nth accounting data), multiple text data (the first to the nth text data), and multiple names of trading partners (the names of the first to the nth trading partners). Specifically, it is preferable to use the multiple text data and the multiple names of trading partners as input data, and the multiple accounting data as training data (labels).
[0194] As explained in the previous section on the <trained judgment model>, the input of the i-th accounting data, the extraction of the i-th text data using OCR, and the acquisition of the i-th trading partner's company name are all performed using the same supporting document.
[0195] Furthermore, it is preferable that the learning is performed using audited documents. For example, the above-mentioned multiple accounting data preferably includes audited accounting data. Also, the above-mentioned multiple text data preferably includes text data extracted by OCR from the image data of the audited documents. Furthermore, the above-mentioned multiple business partner company names preferably include the business partner company names obtained from the image data of the audited documents or from local image features extracted from the image data of the audited documents. Note that the above-mentioned multiple business partner company names may also include the business partner company names obtained from audited accounting data.
[0196] Audited accounting data is already registered in the accounting software. Furthermore, when creating ledgers using accounting software, electronic storage of paper documents is permitted. Therefore, image data of audited documents is often stored in a storage device connected to a device capable of running the accounting software. This means that extracting text data and local image features from the image data of audited documents is straightforward. Additionally, the processing unit 101C can be used to obtain the names of trading partners from the local image features. Thus, a training dataset consisting of the above-mentioned accounting data, text data, and the names of the above-mentioned trading partners can be easily created.
[0197] The above-mentioned multiple accounting data, multiple text data, and multiple business partner company names may be stored in the storage unit of the document verification system (for example, the storage unit 102 shown in Figure 1A), or they may be stored in a storage device connected to the document verification system via the network 120 (for example, the storage device 150 shown in Figure 1A).
[0198] This learning method is not limited to supervised learning; it may also be semi-supervised learning. Compared to supervised learning, semi-supervised learning requires a smaller amount of training data in the training dataset, allowing for the generation of accounting data with high accuracy even with a small number of audited documents. This is particularly effective in the initial stages of accounting software implementation, when the number of audited documents is limited.
[0199] By performing this learning process, a trained decision model can be created. In other words, by performing this learning process, the decision model is trained. As a result, the trained decision model can perform accounting data inference based on the names of trading partners and text data.
[0200] The format of supporting documents varies from company to company. In other words, the name of the trading partner company is useful for identifying transaction dates, product names, and payment amounts contained in text data. Therefore, by performing this training, accounting data can be inferred with high accuracy.
[0201] The above is a description of other variations of the trained decision model.
[0202] According to one aspect of the present invention, a document verification method can be provided that automatically checks the consistency between supporting documents and accounting data. Furthermore, according to one aspect of the present invention, a document verification method can be provided that checks the consistency between supporting documents and accounting data regardless of the performance of the OCR. Furthermore, according to one aspect of the present invention, a document verification method can be provided that checks the consistency between supporting documents and accounting data using existing OCR.
[0203] This embodiment can be implemented in appropriate combination with other embodiments described herein, at least in part.
[0204] (Embodiment 2) In this embodiment, the hardware configuration of a document verification system according to one aspect of the present invention will be described with reference to Figures 11 and 12.
[0205] The document verification system of this embodiment can automatically check the consistency between documents and accounting data using the document verification method shown in Embodiment 1.
[0206] <Example Configuration of a Document Verification System 1> Figure 11 shows a block diagram of the document verification system 200. In the diagrams attached to this specification, the components are classified by function and shown as independent blocks. However, in reality, it is difficult to completely separate the components by function, and one component may be involved in multiple functions. Also, one function may be involved in multiple components; for example, the processing performed by the processing unit 202 may be executed on different servers depending on the processing.
[0207] The document verification system 200 has at least a processing unit 202. The document verification system 200 shown in Figure 11 further has a reception unit 201, a storage unit 203, a database 204, a display unit 205, and a transmission line 206.
[0208] [Reception Desk 201] The reception unit 201 receives image data from outside the document verification system 200. This image data is, for example, image data of an unaudited document and corresponds to image data 12 shown in Embodiment 1. The reception unit 201 may also receive text data from outside the document verification system 200. This text data is, for example, text data extracted from the image data of an unaudited document and corresponds to text data 13 shown in Embodiment 1. The image data, text data, etc., received by the reception unit 201 are supplied to the processing unit 202, the storage unit 203, or the database 204, respectively, via the transmission line 206.
[0209] Input methods for image data and text data include, for example, key input using a keyboard or touch panel, voice input using a microphone, image input using a scanner or camera, reading from a recording medium, and acquisition using communication.
[0210] The document verification system 200 may have an optical character recognition (OCR) function. This allows it to recognize characters contained in image data and extract text data. For example, the processing unit 202 may have this function. Alternatively, the document verification system 200 may further have an OCR unit that has this function.
[0211] [Processing 202] The processing unit 202 has the function of processing data supplied from the reception unit 201, the storage unit 203, or the database 204, etc. The processing unit 202 can supply the processing results to the storage unit 203, the database 204, or the display unit 205, etc.
[0212] The processing unit 202 includes the processing unit 101 shown in Embodiment 1. Specifically, the processing unit 202 has the functions of: extracting local image features from image data; generating vectors based on accounting data using a trained decision model; generating vectors based on local image features and text data using a trained decision model; calculating the similarity between two vectors; and determining whether the supporting documents and accounting data match using the calculated similarity. The processing unit 202 may also have a function to output the result of the determination of whether the supporting documents and accounting data match.
[0213] The processing unit 202 may use a transistor having a metal oxide in its channel formation region. Because this transistor has an extremely small off-current, using it as a switch to hold the charge (data) flowing into a capacitive element that functions as a memory element ensures that the data can be retained for a long period of time. By using this characteristic in at least one of the registers and cache memory of the processing unit 202, the processing unit 202 can be operated only when necessary, and the information from the previous processing can be saved to the memory element in other cases, thereby turning off the processing unit 202. In other words, normally-off computing becomes possible, and the power consumption of the evidence verification system can be reduced.
[0214] In this specification, a transistor using an oxide semiconductor in the channel formation region is referred to as an Oxide Semiconductor transistor (OS transistor). The channel formation region of an OS transistor preferably contains a metal oxide.
[0215] The metal oxide in the channel-forming region preferably contains indium (In). When the metal oxide in the channel-forming region contains indium, the carrier mobility (electron mobility) of the OS transistor increases. Furthermore, the metal oxide in the channel-forming region preferably contains element M. Element M is preferably aluminum (Al), gallium (Ga), or tin (Sn). Other elements applicable to element M include boron (B), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, it is also possible to combine multiple of the aforementioned elements as element M. Element M is, for example, an element with a high bond energy with oxygen. For example, an element with a higher bond energy with oxygen than indium. Furthermore, the metal oxide in the channel-forming region preferably contains zinc (Zn). Metal oxides containing zinc may be more prone to crystallization.
[0216] The metal oxides present in the channel-forming regions are not limited to indium-containing metal oxides. For example, the metal oxides present in the channel-forming regions may be zinc-tin oxides, gallium-tin oxides, or other metal oxides that do not contain indium but contain zinc, gallium, or tin.
[0217] Furthermore, the processing unit 202 may use a transistor that includes silicon in its channel formation region.
[0218] Furthermore, the processing unit 202 may use a combination of a transistor containing an oxide semiconductor in its channel formation region and a transistor containing silicon in its channel formation region.
[0219] The processing unit 202 includes, for example, an arithmetic circuit or a central processing unit (CPU).
[0220] The processing unit 202 may have a microprocessor such as a DSP (Digital Signal Processor) and a GPU (Graphics Processing Unit). The microprocessor may be implemented using a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) and an FPAA (Field Programmable Analog Array). The processing unit 202 can perform various data processing and program control by interpreting and executing instructions from various programs by the processor. Programs that can be executed by the processor are stored in at least one of the processor's memory area and the storage unit 203.
[0221] The processing unit 202 may have main memory. The main memory includes at least one of volatile memory such as RAM and non-volatile memory such as ROM.
[0222] For RAM, for example, DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory) are used, and a memory space is virtually allocated and used as the workspace for the processing unit 202. The operating system, application programs, program modules, program data, and lookup tables stored in the storage unit 203 are loaded into RAM for execution. These data, programs, and program modules loaded into RAM are directly accessed and manipulated by the processing unit 202.
[0223] ROM can store BIOS (Basic Input / Output System) and firmware, etc., which do not require rewriting. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory). Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows data to be erased by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), and flash memory.
[0224] [Storage section 203] The memory unit 203 has the function of storing the program executed by the processing unit 202. The memory unit 203 also has the function of storing the trained judgment model shown in Embodiment 1. Furthermore, the memory unit 203 may also have the function of storing, for example, data received by the reception unit 201, processing results generated by the processing unit 202, etc.
[0225] The storage unit 203 has at least one of volatile memory and non-volatile memory. The storage unit 203 may have volatile memory such as DRAM and SRAM. The storage unit 203 may have non-volatile memory such as ReRAM (Resistive Random Access Memory), PRAM (Phase Change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory), or flash memory. The storage unit 203 may also have recording media drives such as hard disk drives (HDD) and solid state drives (SSD).
[0226] [Database 204] The document verification system 200 may have a database 204. For example, the database 204 has the function of storing data relating to audited documents (for example, the first to nth accounting data, the first to nth text data, and the first to nth local image features as shown in Embodiment 1). The database 204 may also store data relating to documents for which integrity checks have been completed.
[0227] Note that the storage unit 203 and the database 204 do not necessarily have to be separate from each other. For example, the document verification system 200 may have a storage unit that has the functions of both the storage unit 203 and the database 204.
[0228] Furthermore, the memory in the processing unit 202, the storage unit 203, and the database 204 can each be considered an example of a non-temporary computer-readable storage medium.
[0229] [Display section 205] The display unit 205 has a function to display the processing results from the processing unit 202. For example, the display unit 205 has a function to display the result of a determination made in the processing unit 202. The display unit 205 may also have a function to display image data of supporting documents and accounting data.
[0230] The document verification system 200 may also have an output unit. The output unit has the function of supplying data to an external source.
[0231] [Transmission path 206] The transmission line 206 has the function of transmitting various types of data. Data can be transmitted and received between the receiving unit 201, processing unit 202, storage unit 203, database 204, and display unit 205 via the transmission line 206. For example, image data of evidence and text data extracted from the image data of evidence are transmitted and received via the transmission line 206.
[0232] <Example Configuration of a Document Verification System 2> Figure 12 shows a block diagram of the document verification system 210. The document verification system 210 includes a server 220 and terminals 230 (such as personal computers).
[0233] The server 220 includes a processing unit 222, a storage unit 223, a transmission line 226, and a communication unit 227. Although not shown in Figure 12, the server 220 may further include a receiving unit, an output unit, an OCR unit, and a database.
[0234] Terminal 230 includes a reception unit 231, a processing unit 232, a storage unit 233, a display unit 235, a transmission line 236, and a communication unit 237. Although not shown in Figure 12, terminal 230 may further include a database and an OCR unit.
[0235] In the document verification system 210, the receiving unit 231 of the terminal 230 receives image data. This image data is image data of an unaudited document and corresponds to image data 12 shown in Embodiment 1. The receiving unit 231 of the terminal 230 may also receive text data. This text data is, for example, text data extracted from the image data of an unaudited document and corresponds to text data 13 shown in Embodiment 1. This image data and text data are transmitted from the communication unit 237 of the terminal 230 to the communication unit 227 of the server 220.
[0236] The image data and text data received by the communication unit 227 are stored in the storage unit 223 via the transmission line 226. Alternatively, the image data and text data may be supplied directly from the communication unit 227 to the processing unit 222.
[0237] The extraction of local image features and the determination of whether the supporting documents and accounting data match, as described in Embodiment 1, require high processing power. The processing unit 222 of the server 220 has higher processing power than the processing unit 232 of the terminal 230. Therefore, it is preferable that the extraction of local image features and the determination of whether the supporting documents and accounting data match are performed by the processing unit 222.
[0238] The processing unit 222 then outputs the result of the determination. The result of the determination is supplied directly from the processing unit 222 to the communication unit 227. The result of the determination is transmitted from the communication unit 227 of the server 220 to the communication unit 237 of the terminal 230. The result of the determination is displayed on the display unit 235 of the terminal 230. The result of the determination may also be stored in the storage unit 223 or storage unit 233.
[0239] [Processing Unit 222 and Processing Unit 232] The processing unit 222 has the function of processing data supplied from the storage unit 223 and the communication unit 227, etc. The processing unit 232 has the function of processing data supplied from the receiving unit 231, the storage unit 233, the display unit 235, and the communication unit 237, etc. The processing units 222 and 232 can be described by referring to the description of the processing unit 202. It is preferable that the processing unit 222 has higher processing capacity than the processing unit 232.
[0240] [Storage section 223] The storage unit 223 has the function of storing programs executed by the processing unit 222. The storage unit 223 also has the function of storing data related to audited evidence (e.g., accounting data, image data, and text data), processing results generated by the processing unit 222, and data input to the communication unit 227. The storage unit 223 can refer to the description of the storage unit 203.
[0241] [Storage section 233] The memory unit 233 has the function of storing the program executed by the processing unit 232. The memory unit 233 also has the function of storing calculation results generated by the processing unit 232, data received by the receiving unit 231, and data input to the communication unit 237. The memory unit 233 can refer to the description of the memory unit 203.
[0242] [Transmission lines 226 and 236] Transmission lines 226 and 236 have the function of transmitting data. Data transmission and reception between the processing unit 222, the storage unit 223, and the communication unit 227 can be performed via transmission line 226. Data transmission and reception between the receiving unit 231, the processing unit 232, the storage unit 233, the display unit 235, and the communication unit 237 can be performed via transmission line 236.
[0243] [Communications Units 227 and 237] Communication units 227 and 237 can be used to send and receive data between the server 220 and the terminal 230. Hubs, routers, or modems can be used as communication units 227 and 237. Data transmission and reception can be done via wired or wireless methods (e.g., radio waves and infrared rays).
[0244] Communication between server 220 and terminal 230 may be performed by connecting to computer networks such as the Internet, intranet, extranet, PAN (Personal Area Network), LAN (Local Area Network), CAN (Campus Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), and GAN (Global Area Network), which form the basis of the World Wide Web (WWW).
[0245] [Reception Desk 231] The reception unit 231 can refer to the description of the reception unit 201.
[0246] [Display section 235] Display unit 235 can refer to the description of display unit 205.
[0247] This embodiment can be implemented in appropriate combination with other embodiments described herein, at least in part. [Explanation of Symbols]
[0248] 10: Proof document, 11: Accounting data, 11A: Accounting data, 12: Image data, 13: Text data, 14: Local image features, 15: Vector, 16: Vector, 17: Company name, 100: Proof document verification system, 100A: Proof document verification system, 101: Processing unit, 101a: Feature extraction unit, 101A: Processing unit, 101b: Vector generation unit, 101B: Processing unit, 101c: Calculation unit, 101C: Processing unit, 101d: Judgment unit, 101e: OCR unit, 101f: Inference unit, 101g: Judgment unit, 101h: Estimation unit, 101i: Inference unit, 102: Memory unit, 103: Reception Unit, 105: Display Unit, 110: Optical Character Recognition Device, 120: Network, 130: Input Device, 140: Output Device, 150: Storage Device, 200: Document Verification System, 201: Reception Unit, 202: Processing Unit, 203: Memory Unit, 204: Database, 205: Display Unit, 206: Transmission Line, 210: Document Verification System, 220: Server, 222: Processing Unit, 223: Memory Unit, 226: Transmission Line, 227: Communication Unit, 230: Terminal, 231: Reception Unit, 232: Processing Unit, 233: Memory Unit, 235: Display Unit, 236: Transmission Line, 237: Communication Unit
Claims
1. Having a server, The server comprises a processing unit, a storage unit, a transmission line, and a communication unit. The memory unit stores the trained decision model, The aforementioned communication unit receives the image data and text data of the evidence, The processing unit receives the image data and text data of the evidence from the communication unit via the transmission line. The processing unit extracts local image features from the image data of the evidence, The processing unit generates a first vector based on the accounting data using the trained decision model, The processing unit generates a second vector using the trained decision model based on the local image features and the text data. The processing unit calculates the similarity between the first vector and the second vector, The processing unit uses the similarity score to determine whether the supporting document and the accounting data match. The communication unit receives the result of the determination from the processing unit. Document verification system.
2. In claim 1, Having a device, The result of the aforementioned determination is transmitted from the communication unit to the terminal by the document verification system.
3. In claim 1 or claim 2, The aforementioned text data is extracted from the image data of the document by optical character recognition in the document verification system.