Document correction support program, document correction support method, and information processing apparatus

The document correction support program uses machine learning to identify the correct document by calculating differences and relationships between documents, enhancing the efficiency of document correction processes.

JP7704045B2Active Publication Date: 2025-07-08FUJITSU LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022023005
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-07-08
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

Existing document correction systems struggle to identify which document to correct among multiple documents of the same type, as they do not consider the relationship between them.

Method used

A document correction support program that utilizes machine learning to generate explanatory variables based on predefined restoration operations, calculating differences between document values and their relationships, and uses these variables to determine which document requires correction.

Benefits of technology

Facilitates easy identification of the document to be corrected by considering the relationship between multiple documents, improving the efficiency of document correction processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704045000001
    Figure 0007704045000001
  • Figure 0007704045000002
    Figure 0007704045000002
  • Figure 0007704045000003
    Figure 0007704045000003
Patent Text Reader

Abstract

To provide a document correction support program, a document correction support method, and an information processing apparatus which can specify a document to be corrected.SOLUTION: A document correction support program according to the present invention causes a computer to execute an input processing and an output processing. In the input processing, an explanatory variable generated based on each of a determination target documents is input to a model obtained by machine learning based on an explanatory variable including a difference from a value obtained by calculation, based on learning data of a plurality of cases including correction history of each document, using a value of an item included in a specific first document and a value of an item included in a specific second document based on pre-defined restoration calculation for each case, and a value of an item included in a first document, and an objective variable including a value corresponding to the correction history in each of documents for a case. In the output processing based on an output from the model, an output specifies whether or not there is any correction in each of the determination target documents.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a document correction support program, a document correction support method, and an information processing apparatus.

Background Art

[0002] Conventionally, for various documents (hereinafter also referred to as materials) submitted at a counter or the like, the consistency between the documents is checked, and items with inconsistencies are corrected by staff. For example, at the tax service counter, many tax return documents are submitted every year. The submitted documents are checked by staff against the basic information of residents and the documents submitted by employers to ensure there are no mistakes.

[0003] FIG. 8 is an explanatory diagram for explaining an example of correcting document deficiencies. As shown in FIG. 8, resident H1 submits a final tax return D1 and a resident tax return D2 to the city hall. Also, employers K1 and K2 of resident H1 submit salary payment reports D3 and D4 regarding resident H1. Also, a pension institution K3 submits a pension payment report D5 regarding resident H1. Staff H2 at the city hall compares the entries in the submitted final tax return D1, resident tax return D2, salary payment reports D3 and D4, and pension payment report D5. Then, staff H2 detects items with inconsistencies and corrects the data of those items.

[0004] Regarding support for such document correction work, there is a prior art in which the description contents of invoices and specifications received from each medical institution server are checked according to existing rules. Also, there is a prior art in which the validity of item contents is determined based on whether the semantic vector of each item content in a document deviates from the centroid of the semantic vector of the corresponding item in a normative document in which the semantic vectors of each item content are registered in advance. Also, there is a prior art in which an evaluation result obtained by inputting data of a document to be evaluated into a learned model obtained by machine learning of corrected material data is displayed. Also, there is a prior art in which a model is updated by using a new document set as learning data.

Prior Art Documents

Patent Documents

[0005] Patent Document 1 Japanese Unexamined Patent Application Publication No. 2007-241986 Patent Document 2 Japanese Unexamined Patent Application Publication No. 2020-140442 Patent Document 3 Japanese Unexamined Patent Application Publication No. 2018-147280 Patent Document 4 Japanese Unexamined Patent Application Publication No. 2021-89473 Summary of the Invention Problems to be Solved by the Invention

[0006] In some cases, among the submitted documents, there may be multiple documents of the same type, such as salary payment reports D3 and D4. In such cases, the municipal office staff H2 determines which document to correct among the documents of the same type after considering the relationship between the documents.

[0007] However, in the above prior art, the relationship between the multiple submitted documents is not considered, and even if inconsistent items can be identified, it is difficult to identify which document to correct.

[0008] One aspect aims to provide a document correction support program, a document correction support method, and an information processing apparatus that can easily identify a document to be corrected. Means for Solving the Problems

[0009] In one aspect, the document correction support program causes a computer to execute an input process and an output process. The input process is based on learning data of a plurality of cases including the correction history of each document. For each case, based on a predefined restoration operation, the difference between the value of an item included in a specific first document and the value calculated from the value of an item included in a specific second document, and the value of the item included in the first document, and the explanatory variable including the value corresponding to the correction history in each document of the case, the explanatory variable generated based on each document to be determined is input to the model learned based on the objective variable including the value corresponding to the correction history in each document of the case. The output process outputs whether there is a correction in each document to be determined based on the output from the model.

Advantages of the Invention

[0010] The document to be corrected can be easily identified.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0012] Hereinafter, with reference to the drawings, a document correction support program, a document correction support method, and an information processing apparatus according to an embodiment will be described. In the embodiment, components having the same function are denoted by the same reference numerals, and redundant descriptions are omitted. Note that the document correction support program, the document correction support method, and the information processing apparatus described in the following embodiments are merely examples and do not limit the embodiments. Also, the following embodiments may be appropriately combined within a non - conflicting range.

[0013] FIG. 1 is a block diagram showing a functional configuration example of an information processing apparatus according to an embodiment. As shown in FIG. 1, the information processing apparatus 1 includes an input unit 10, an objective variable generation unit 20, an explanatory variable generation unit 30, a storage unit 40, a learning unit 50, and a determination unit 60.

[0014] The information processing apparatus 1 is an example of a learning apparatus that generates a learning model 71 by performing machine learning using the corrected document data 11 as learning data. Also, the information processing apparatus 1 is an example of a determination apparatus that determines and outputs whether there is a correction for each document included in the uncorrected document data 12 to be determined, using the generated learning model 71.

[0015] Note that the above - mentioned learning apparatus and determination apparatus may be realized by a single information processing apparatus 1, or may be realized separately. For example, the information processing apparatus 1 may be a learning apparatus having an input unit 10, an objective variable generation unit 20, an explanatory variable generation unit 30, a storage unit 40, and a learning unit 50. Also, the information processing apparatus 1 may be a determination apparatus having an input unit 10, an objective variable generation unit 20, an explanatory variable generation unit 30, a storage unit 40, and a determination unit 60.

[0016] The input unit 10 is a processing unit that receives input of various data such as the corrected document data 11 and the uncorrected document data 12.

[0017] The corrected document data 11 is data that describes the corrected content of each document in each case. For example, the corrected document data 11 is data that includes, for each case, the content of each document (material) (for example, the entered value in each item) and the presence or absence of corrections in each document (for example, the item with a defect and the corrected value in that item). As an example, the corrected document data 11 includes the values of the items of each document submitted by resident H1 in the tax business and corrected by staff H2, and the presence or absence of corrections.

[0018] In machine learning, based on the corrected document data 11, for each case, the feature amount corresponding to the content of each document is used as an explanatory variable, and the value corresponding to the presence or absence of correction is used as a target variable to generate. Next, in machine learning, when the generated explanatory variable is input to the input layer of a model such as a neural network, the model parameters are obtained so that the target variable can be obtained from the output layer of the model.

[0019] The uncorrected document data 12 is data that describes the content of each document related to the case to be determined. For example, the uncorrected document data 12 is data that includes the content of each document (material) (for example, the entered value in each item) in the case to be determined. As an example, the uncorrected document data 12 includes the values of the items of each document submitted by resident H1 in the tax business.

[0020] The target variable generation unit 20 is a processing unit that generates a target variable based on the corrected document data 11. Specifically, the target variable generation unit 20 generates a target variable in which the presence or absence of correction in each document of the case is quantified for each case included in the corrected document data 11 during machine learning.

[0021] The explanatory variable generation unit 30 is a processing unit that generates an explanatory variable based on the corrected document data 11 or the uncorrected document data 12. Specifically, the explanatory variable generation unit 30 generates an explanatory variable in which the content of each document (material) (for example, the entered value in each item) is quantified for each case of the corrected document data 11 or the uncorrected document data 12.

[0022] During machine learning, the information processing apparatus 1 uses, for each case of the corrected document data 11, data obtained by combining, for each document, the objective variable generated by the objective variable generation unit 20 and the explanatory variable generated by the explanatory variable generation unit 30 as, for example, an array with each document as a row, as the learning data 70 for machine learning. Further, during determination using the learning model 71, the information processing apparatus 1 uses, for each document based on the uncorrected document data 12, data obtained by using, as, for example, an array with each document as a row, the explanatory variable generated by the explanatory variable generation unit 30 as the evaluation data 80.

[0023] The explanatory variables generated by the explanatory variable generation unit 30 include a value obtained by performing an operation based on a predefined restoration operation on a specific document and a value obtained by performing an operation on the values of items of a specific document (referred to as another document) separate from the specific document, that is, a difference between the value obtained by the operation and the value of an item of the specific document, that is, a value corresponding to the relationship between the specific document and another document.

[0024] Here, the specific document and the other document are those specified in advance in the definition content of the restoration operation. For example, by specifying the document type in the definition content, documents of the same type are specified. For example, by specifying a salary payment report in the definition content of the restoration operation, a plurality of salary payment reports (for example, salary payment reports D3 and D4) submitted in each case are used as the specific document and the other document.

[0025] Further, the restoration operation defines a typical "way of dividing documents" between a specific document and another document and is an operation for restoring it. As an example, the restoration operation is an operation that a staff member H2 performs in his / her head when estimating correct values in items of a plurality of documents (for example, salary payment reports D3 and D4) submitted by a resident H1 for each document type. That is, the result of the restoration operation is a candidate for the correct value for each document type. Further, the result of the restoration operation represents the characteristics of the entire document of the corresponding document type (for example, salary payment reports D3 and D4). Therefore, the difference between the operation result and the item of a specific document indirectly represents the relationship between the specific document and another document.

[0026] As a configuration for generating the above-described explanatory variables, the explanatory variable generation unit 30 includes a restoration operation calculation unit 31 and a difference calculation unit 32.

[0027] Based on the restoration operation definition information 41 stored in the storage unit 40, the restoration operation calculation unit 31 is a processing unit that performs a restoration operation on a specific document included in each case of the corrected document data 11 or the uncorrected document data 12 and the value of an item of another document separate from the document.

[0028] Based on the restoration operation definition information 41, the difference calculation unit 32 is a processing unit that obtains the difference between the result of the restoration operation by the restoration operation calculation unit 31 and the value of an item of a specific document.

[0029] FIG. 2 is an explanatory diagram for explaining an example of the restoration operation definition information 41. As shown in FIG. 2, the restoration operation definition information 41 defines a stereotypical "way of dividing documents" between a specific document and another document, and shows a "restoration operation" for restoring it. Note that the restoration operation definition information 41 is prepared for each document type such as the final tax return, D2, D3, etc., and for each document type, the content of the restoration operation for each item between documents is defined. In the following example, the value of the document targeted by the restoration operation is limited to 0 and positive for explanation, but the value of the document targeted by the restoration operation is not limited to these.

[0030] For example, as a "way of dividing documents" in the "restoration operation" defined in the restoration operation definition information 41, there is a case where the value of a certain document (for example, the salary payment report D3) is included in the value of another document (for example, the salary payment report D4). In the "restoration operation" for such a case, the maximum value of the values of the two documents is obtained. Therefore, the meaning of the difference between the result of the restoration operation for this case and the value of a certain document is that if it is 0, this document is the maximum value, and if it is other than 0, one of the other documents is the maximum value.

[0031] In addition, as for the "way of dividing materials" in the "restoration operation" defined in the restoration operation definition information 41, there is a case where values are recorded in separate materials and there is no duplication of values between materials. In the "restoration operation" in such a case, the sum of the values of all materials is obtained. Therefore, the meaning of the difference between the result of the restoration operation in this case and the value of a certain material is that if it is 0, there is a positive value only in this material, and if it is other than 0, there is also a positive value in other materials.

[0032] In addition, as for the "way of dividing materials" in the "restoration operation" defined in the restoration operation definition information 41, there is a case where the value of another item is misrecorded in the item of a specific material. In the "restoration operation" in such a case, the value of the corresponding material is subtracted from the sum of the values of all materials. Therefore, the meaning of the difference between the result of the restoration operation in this case and the value of a certain material is that if it is 0, there is a positive value only in this material and the corresponding material, and if it is other than 0, there is also a positive value in other materials, or this material is the corresponding material or the value is 0.

[0033] The storage unit 40 is a storage device such as a memory or an HDD (Hard Disk Drive), and stores the restoration operation definition information 41.

[0034] The learning unit 50 is a processing unit that generates a learning model 71 by performing known machine learning processing based on the explanatory variables generated by the explanatory variable generation unit 30 and the target variables generated by the target variable generation unit 20 for each case of the learning data 70. Examples of the machine learning processing performed by the learning unit 50 include decision trees, random forests, deep learning, etc. For example, in the case of deep learning, the learning unit 50 generates the learning model 71 by obtaining the parameters of the hidden layer so as to output the corresponding output for the target variable when the explanatory variables are input.

[0035] The determination unit 60 is a processing unit that inputs evaluation data 80 including explanatory variables generated based on the uncorrected document data 12 for the case to be determined into the learning model 71 to obtain the discrimination result of the case to be determined. Specifically, the determination unit 60 reads the parameters obtained by the machine learning of the learning unit 50 to construct the learning model 71. Next, the determination unit 60 inputs the explanatory variables generated based on the uncorrected document data 12 to the learning model 71, that is, the feature amounts of each document to be determined. Next, the determination unit 60 obtains the accuracy (evaluation value) indicating the presence or absence of correction in each document to be determined from the output of the learning model 71. The determination unit 60 outputs a determination result 81 indicating the presence or absence of correction in each document by comparing, for example, with a predetermined threshold value based on the obtained accuracy.

[0036] Figure 3 is a flowchart showing an operation example of the information processing apparatus according to the embodiment. As shown in Figure 3, in the information processing apparatus 1, the corrected document data 11 is used as learning data, and a learning process including the result of the restoration calculation based on the restoration operation definition information 41 in the explanatory variables is performed (S1), and the learning model 71 is generated.

[0037] Figure 4 is a flowchart showing an example of the learning process of the information processing apparatus 1 according to the embodiment. As shown in Figure 4, when the learning process is started, the restoration operation calculation unit 31 acquires the values of each submitted document (material) for each case (resident) included in the corrected document data 11. Next, the restoration operation calculation unit 31 performs a restoration operation based on the defined content of the restoration operation definition information 41 from the values of the items of a specific document and another document separate from that document (S10).

[0038] Next, the difference calculation unit 32 calculates the difference between the restoration operation result by the restoration operation calculation unit 31 and the value of the material (S11), and stores this calculation result as an explanatory variable in the learning data 70 (S12).

[0039] Next, the explanatory variable generation unit 30 determines whether or not the process of generating explanatory variables for all the materials included in the case has been completed (S13). If the process has not been completed for all the materials included in the case (S13: No), the explanatory variable generation unit 30 returns the process to S10.

[0040] If the process has been completed for all the materials included in the case (S13: Yes), the explanatory variable generation unit 30 determines whether or not all the restoration operations defined in the restoration operation definition information 41 have been processed (S14). If not all the restoration operations have been processed (S14: No), the explanatory variable generation unit 30 returns the process to S10.

[0041] If all the restoration operations have been processed (S14: Yes), the explanatory variable generation unit 30 determines whether or not the process of generating explanatory variables has been performed for all the items of the materials (S15). If the process has not been performed for all the items (S15: No), the explanatory variable generation unit 30 returns the process to S10.

[0042] If all the items have been processed (S15: Yes), the explanatory variable generation unit 30 determines whether or not all the residents (cases) have been processed (S16). If not all the residents have been processed (S16: No), the explanatory variable generation unit 30 returns the process to S10.

[0043] If all the residents have been processed (S16: Yes), the target variable generation unit 20 calculates whether or not each submitted document (material) has been corrected for each case (resident) included in the corrected material data 11. Next, the target variable generation unit 20 stores the calculation result regarding whether or not each material has been corrected as a target variable in the learning data 70.

[0044] Next, the target variable generation unit 20 determines whether or not the process of generating target variables has been performed for all the materials (S19). If the process has not been performed for all the materials (S19: No), the target variable generation unit 20 returns the process to S17.

[0045] When processing has been performed on all the materials (S19: Yes), the learning unit 50 performs machine learning based on the learning data 70 to generate a learning model 71 (S20), and ends the processing.

[0046] FIG. 5 is an explanatory diagram for explaining an example of the learning data 70. As shown in FIG. 5, when performing machine learning, for each case (Resident A, Resident B) included in the corrected material data 11, based on the corrected materials (1), (2), (3)... the feature amount corresponding to the content of each material is used as an explanatory variable, and learning data 70 is generated with the value corresponding to the presence or absence of correction as the objective variable.

[0047] Here, the explanatory variables in the learning data 70 include, based on the restoration operation defined by the restoration operation definition information 41, the difference between a specific document and the value calculated from the values of items in another document separate from that document, and the value of the item in the specific document, that is, the value corresponding to the relationship between the specific document and other documents.

[0048] For example, the explanatory variables related to the material (1) of Resident A include the difference between the items of the material (1) and the values obtained by the restoration operations a, b, c..., that is, the values corresponding to the relationship between the material (1) and other documents (for example, material (2), material (3)...). Similarly, the explanatory variables related to the material (2) include the difference between the items of the material (2) and the values obtained by the restoration operations a, b, c..., that is, the values corresponding to the relationship between the material (2) and other documents (for example, material (1), material (3)...). Also, the explanatory variables related to the material (3) include the difference between the items of the material (3) and the values obtained by the restoration operations a, b, c..., that is, the values corresponding to the relationship between the material (3) and other documents (for example, material (1), material (2)...). Note that the same applies to each case (Resident B...).

[0049] In the information processing apparatus 1, by performing machine learning based on the learning data 70 including such explanatory variables, a learning model 71 that has learned the relationship between each document can be generated.

[0050] Returning to FIG. 3, following S1, in the information processing apparatus 1, the generated learning model 71 is applied to the uncorrected document data 12, and a determination process is performed to determine whether there are any corrections regarding each document included in the uncorrected document data 12 (S2). As a result, the information processing apparatus 1 obtains a determination result 81 indicating whether there are any corrections regarding each document, and ends the process.

[0051] FIG. 6 is a flowchart showing an example of the determination process of the information processing apparatus 1 according to the embodiment. As shown in FIG. 6, when the determination process is started, the restoration operation calculation unit 31 acquires the values of each document (material) submitted for the case (resident) to be determined included in the uncorrected document data 12. Next, the restoration operation calculation unit 31 performs a restoration operation based on the defined content of the restoration operation definition information 41 from the values of a specific document and the items of another document separate from that document (S20).

[0052] Next, the difference calculation unit 32 calculates the difference between the restoration operation result by the restoration operation calculation unit 31 and the value of the material (S31), and stores this calculation result as an explanatory variable in the evaluation data 80 (S32).

[0053] Next, the explanatory variable generation unit 30 determines whether the process of generating explanatory variables for all the materials included in the case to be determined has ended (S33). If the process has not ended for all the materials included in the case to be determined (S33: No), the explanatory variable generation unit 30 returns the process to S30.

[0054] If the process has ended for all the materials included in the case to be determined (S33: Yes), the explanatory variable generation unit 30 determines whether all the restoration operations defined in the restoration operation definition information 41 have been processed (S34). If not all the restoration operations have been processed (S34: No), the explanatory variable generation unit 30 returns the process to S30.

[0055] When all restoration operations have been processed (S34: Yes), the explanatory variable generation unit 30 determines whether processing for generating explanatory variables has been performed for all items in the document (S35). If processing has not been performed for all items (S35: No), the explanatory variable generation unit 30 returns the processing to S30.

[0056] When all items have been processed (S35: Yes), the explanatory variable generation unit 30 determines whether all residents (cases) to be determined have been processed (S36). If processing has not been performed for all residents (S36: No), the explanatory variable generation unit 30 returns the processing to S30.

[0057] When all residents have been processed (S36: Yes), the determination unit 60 inputs the evaluation data 80 into the learning model 71 to determine the presence or absence of corrections in each document of the case to be determined and obtains the determination result 81 (S37), and ends the processing.

[0058] As described above, based on the corrected document data 11 of a plurality of cases including the correction history of each document, the information processing apparatus 1, for each case, based on a predefined restoration operation, calculates the difference between the value of an item included in a specific first document and the value calculated from the value of an item included in a specific second document, and generates an explanatory variable including the difference. Then, the information processing apparatus 1 generates a learning model 71 by machine learning based on the generated explanatory variable and an objective variable including a value corresponding to the correction history in each document of the case. The information processing apparatus 1 inputs the explanatory variable generated in the same manner as during learning based on each document to be determined into this learning model 71, and outputs the presence or absence of corrections in each document to be determined based on the output from the learning model 71.

[0059] As described above, the explanatory variables input to the learning model 71 used for determination include the values restored from the values of the items included in a specific first document and the values of the items included in a specific second document, and the relationship between the documents based on the difference between the items included in the first document. Therefore, in the information processing apparatus 1, since the learning model 71 is used to determine whether there is a correction in each document to be determined, it is possible to identify the document to be corrected corresponding to the relationship between the documents.

[0060] In addition, the restoration operation in the information processing apparatus 1 is an operation for obtaining the maximum value of the values of items between different specific documents. Thereby, in the information processing apparatus 1, based on the difference between the result of the restoration operation for obtaining the maximum value of the values of items between different specific documents and the value of the first document, if it is 0, the value of the first document is the maximum value, and if it is other than 0, any of the values of other documents is the maximum value. A feature amount indicating the relationship between the documents can be included in the explanatory variable. For example, as an example of a typical way of dividing documents, there is a case where the value of a certain document (for example, the first document) is included in the value of another document (for example, the second document). In the information processing apparatus 1, a feature amount related to such a case can be included in the explanatory variable by the above-described restoration operation.

[0061] In addition, the restoration operation in the information processing apparatus 1 is an operation for obtaining the total value of the values of items between different specific documents. Thereby, in the information processing apparatus 1, based on the difference between the result of the restoration operation for obtaining the total value of the values of items between different specific documents and the value of the first document, if it is 0, only the first document has a positive value, and if it is other than 0, there are also positive values in other documents. A feature amount indicating the relationship between the documents can be included in the explanatory variable. For example, as an example of a typical way of dividing documents, there is a case where values are formed in separate documents and the values do not overlap. In the information processing apparatus 1, a feature amount related to such a case can be included in the explanatory variable by the above-described restoration operation.

[0062] In addition, the restoration operation in the information processing apparatus 1 is an operation of subtracting the value of an item of a specific third document included in either the first document or the second document from the total value of the items between different specific documents. Thereby, in the information processing apparatus 1, based on the difference between the result of the restoration operation of subtracting the value of the item of the third document from the total value of the items between different specific documents and the value of the first document, if it is 0, only the values of the first document and the third document to be subtracted are positive, and if it is other than 0, there is a positive value in other documents or this document is either the third document or has a value of 0. A feature amount indicating the relationship between documents can be included as an explanatory variable. For example, as an example of a typical way of dividing documents, there is a case where the value of another item is misdescribed in a specific document. In the information processing apparatus 1, a feature amount related to such a case can be included as an explanatory variable by the above restoration operation.

[0063] In addition, the first document and the second document in the information processing apparatus 1 are documents of the same type. Thereby, in the information processing apparatus 1, the document to be corrected can be specified from the relationship between documents of the same type.

[0064] (Others) Note that each component of each illustrated apparatus does not necessarily have to be physically configured as shown in the figure. That is, the specific form of dispersion / integration of each apparatus is not limited to that shown in the figure, and all or part of it can be functionally or physically dispersed / integrated in any unit according to various loads, usage situations, etc. For example, regarding the information processing apparatus 1, the configuration for generating the learning model 71 and the configuration for making a determination based on the generated learning model 71 may be dispersed.

[0065] In addition, various processing functions of the information processing apparatus 1 (input unit 10, target variable generation unit 20, explanatory variable generation unit 30, learning unit 50, and determination unit 60) may be executed in whole or in any part thereof on a CPU (or a microcomputer such as an MPU or an MCU (Micro Controller Unit)). Needless to say, various processing functions may also be executed in whole or in any part thereof on a program analyzed and executed by a CPU (or a microcomputer such as an MPU or an MCU), or on hardware based on wired logic. Further, various processing functions performed by the information processing apparatus 1 may be executed by a plurality of computers in cooperation by cloud computing.

[0066] (Computer configuration example) Incidentally, various processes described in the above embodiments can be realized by executing a program prepared in advance on a computer. Therefore, hereinafter, an example of a computer configuration (hardware) that executes a program having the same functions as those in the above embodiments will be described. FIG. 7 is a block diagram showing an example of a computer configuration.

[0067] As shown in FIG. 7, the computer 200 includes a CPU 201 that executes various arithmetic processes, an input device 202 that receives data input, a monitor 203, and a speaker 204. The computer 200 also includes a medium reading device 205 that reads a program and the like from a storage medium, an interface device 206 for connecting to various devices, and a communication device 207 for communicating with external devices by wire or wirelessly. The information processing apparatus 1 also includes a RAM 208 that temporarily stores various information, and a hard disk device 209. Further, each unit (201 to 209) in the computer 200 is connected to a bus 210.

[0068] The hard disk device 209 stores a program 211 for executing various processes in the functional configurations (e.g., the input unit 10, the target variable generation unit 20, the explanatory variable generation unit 30, the learning unit 50, and the determination unit 60) described in the above embodiments. The hard disk device 209 also stores various data 212 referred to by the program 211. The input device 202 receives, for example, input of operation information from an operator. The monitor 203 displays, for example, various screens operated by the operator. The interface device 206 is connected to, for example, a printing device or the like. The communication device 207 is connected to a communication network such as a LAN (Local Area Network) and exchanges various information with external devices via the communication network.

[0069] The CPU 201 reads out the program 211 stored in the hard disk device 209, expands it in the RAM 208, and executes it, thereby performing various processes related to the above functional configurations (e.g., the input unit 10, the target variable generation unit 20, the explanatory variable generation unit 30, the learning unit 50, and the determination unit 60). Note that the program 211 may not be stored in the hard disk device 209. For example, the program 211 stored in a computer-readable storage medium may be read out and executed. The computer-readable storage medium corresponds to, for example, a portable recording medium such as a CD-ROM, a DVD disk, a USB (Universal Serial Bus) memory, a semiconductor memory such as a flash memory, a hard disk drive, or the like. Also, the program 211 may be stored in a device connected to a public line, the Internet, a LAN, or the like, and the computer 200 may read out and execute the program 211 from these.

[0070] Regarding the above embodiments, the following additional remarks are further disclosed.

[0071] (Appendix 1) Based on the learning data of multiple cases including the revision history of each document, for each case, based on a predefined restoration operation, the difference between the value of the item included in a specific first document and the value obtained by calculating from the value of the item included in a specific second document and the value of the item included in the first document, for a model trained by machine learning based on an explanatory variable including the difference and an objective variable including the value corresponding to the revision history in each document of the case, input the explanatory variable generated based on each document to be determined, Output whether there is a revision in each document to be determined based on the output from the model. A document revision support program characterized by causing a computer to execute the process.

[0072] (Appendix 2) The restoration operation is an operation for obtaining the maximum value of the item values between different specific documents. The document revision support program according to Appendix 1, characterized in that.

[0073] (Appendix 3) The restoration operation is an operation for obtaining the total value of the item values between different specific documents. The document revision support program according to Appendix 1, characterized in that.

[0074] (Appendix 4) The restoration operation is an operation for subtracting the value of the item of a specific third document included in either the first document or the second document from the total value of the item values between different specific documents. The document revision support program according to Appendix 1, characterized in that.

[0075] (Appendix 5) The first document and the second document are documents of the same type. The document revision support program according to any one of Appendices 1 to 4, characterized in that.

[0076] (Appendix 6) Based on the learning data of multiple cases including the revision history of each document, for each of the cases, based on a predefined restoration operation, the value of the difference between the value calculated from the values of the items included in a specific first document and the values of the items included in a specific second document and the value of the item included in the first document, for a model trained by machine learning based on an explanatory variable including the difference and an objective variable including the value corresponding to the revision history in each document of the case, input the explanatory variable generated based on each document to be determined, Output whether there is a revision in each document to be determined based on the output from the model. A document revision support method characterized in that a computer executes the process.

[0077] (Appendix 7) The restoration operation is an operation for obtaining the maximum value of the values of items between different specific documents. The document revision support method according to Appendix 6, characterized in that.

[0078] (Appendix 8) The restoration operation is an operation for obtaining the total value of the values of items between different specific documents. The document revision support method according to Appendix 6, characterized in that.

[0079] (Appendix 9) The restoration operation is an operation for subtracting the value of an item of a specific third document included in either the first document or the second document from the total value of the values of items between different specific documents. The document revision support method according to Appendix 6, characterized in that.

[0080] (Appendix 10) The first document and the second document are documents of the same type. The document revision support method according to any one of Appendices 6 to 9, characterized in that.

[0081] (Appendix 11) Based on the learning data of multiple cases including the revision history of each document, for each case, based on a predefined restoration operation, the difference between the value of the operation calculated from the values of the items included in a specific first document and the values of the items included in a specific second document, and the value of the items included in the first document, is used as an explanatory variable, and for a model trained based on the objective variable including the value corresponding to the revision history in each document of the case, the explanatory variable generated based on each document to be determined is input, Based on the output from the model, output whether there is a revision in each document to be determined, An information processing apparatus, characterized in that a control unit executes the processing.

[0082] (Appendix 12) The restoration operation is an operation for obtaining the maximum value of the values of items between different specific documents. The information processing apparatus according to Appendix 11, characterized in that.

[0083] (Appendix 13) The restoration operation is an operation for obtaining the total value of the values of items between different specific documents. The information processing apparatus according to Appendix 11, characterized in that.

[0084] (Appendix 14) The restoration operation is an operation for subtracting the value of the items of a specific third document included in either the first document or the second document from the total value of the values of the items between different specific documents. The information processing apparatus according to Appendix 11, characterized in that.

[0085] (Appendix 15) The first document and the second document are documents of the same type. The information processing apparatus according to any one of Appendices 11 to 14, characterized in that.

Explanation of Reference Numerals

[0086] 1... Information processing apparatus 10... Input unit 11... Revised material data 12... Unrevised material data 20... Objective variable generation unit 30…Explanation variable generation unit 31…Restoration operation calculation unit 32…Difference calculation unit 40…Memory unit 41…Restoration operation definition information 50…Learning unit 60…Judgment unit 70…Learning data 71…Learning model 80…Evaluation data 81…Judgment result 200…Computer 201…CPU 202…Input device 203…Monitor 204…Speaker 205…Media reader 206…Interface device 207…Communication device 208…RAM 209…Hard disk device 210…Bus 211…Program 212…Various data D1…Final tax return form D2…Resident tax return form D3, D4…Salary payment report D5…Pension payment report H1…Resident H2…Employee K1, K2…Workplace K3…Pension institution

Claims

1. Based on learning data of multiple cases including the revision history of each document, for each case, based on a predefined restoration operation, the difference between the value of an item included in a specific first document and the value calculated from the values of items included in a specific second document and the value of the item included in the first document, and for a model trained by machine learning based on an explanatory variable including the difference and a target variable including a value corresponding to the revision history in each document of the case, input the explanatory variable generated based on each document to be determined, Based on the output from the model, output whether there is a revision in each document to be determined. A document revision support program characterized by causing a computer to execute the process.

2. The restoration operation is an operation for obtaining the maximum value of the values of items between different specific documents. The document revision support program according to claim 1, characterized in that.

3. The restoration operation is an operation for obtaining the total value of the values of items between different specific documents. The document revision support program according to claim 1, characterized in that.

4. The restoration operation is an operation for subtracting the value of an item of a specific third document included in either the first document or the second document from the total value of the values of items between different specific documents. The document revision support program according to claim 1, characterized in that.

5. The first document and the second document are documents of the same type. The document revision support program according to any one of claims 1 to 4, characterized in that.

6. Based on learning data of multiple cases including the revision history of each document, for each case, based on a predefined restoration operation, the difference between the value of an item included in a specific first document and the value calculated from the values of items included in a specific second document and the value of the item included in the first document, and for a model trained by machine learning based on an explanatory variable including the difference and a target variable including a value corresponding to the revision history in each document of the case, input the explanatory variable generated based on each document to be determined, Based on the output from the model, output whether there is a revision in each document to be determined. A document revision support method characterized by a computer executing the process.

7. Based on the learning data of multiple cases including the revision history of each document, for each case, based on a predefined restoration operation, the difference between the value of the operation calculated from the value of the item included in a specific first document and the value of the item included in a specific second document and the value of the item included in the first document, for the model machine-learned based on the explanatory variable including the difference and the objective variable including the value corresponding to the revision history in each document of the case, input the explanatory variable generated based on each document to be determined, Based on the output from the model, output whether there is a revision in each document to be determined, An information processing apparatus, wherein a control unit executes the processing.

Citation Information

Patent Citations

  • Data input system, data input program

    JP2006344012A

  • Method for electronic examination of medical fees

    JP2007241986A

  • Document data processing apparatus and program thereof

    JP2010134766A

  • Data analysis device and data analysis method

    JP2018147280A

  • Document correction method and document correction device

    JP2020140442A