Document analysis apparatus and document analysis method

The document analysis apparatus effectively extracts and processes structured information from tabular documents by aligning cell identification with item values, enabling accurate document analysis and evaluation using large language models.

JP7709583B1Active Publication Date: 2025-07-16HITACHI INDUSTRY & CONTROL SOLUTIONS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024184406
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-07-16
Estimated Expiration
2044-10-18

Smart Images

  • Figure 0007709583000001_ABST
    Figure 0007709583000001_ABST
Patent Text Reader

Abstract

Enables extraction of text-form information that reflects the structure from tabular-form documents. 【Solution means】 A document analysis device (document processing device 100) extracts document information indicating the correspondence between the identification information of each cell and the item value of the cell from a tabular-form document including cells in which item values, which are the values of items, are entered, and forms information indicating the correspondence between the identification information of the cells in which the item values to be acquired are entered among the cells and the item identification information of the items, and creates an item value acquisition prompt that requests an item value acquisition result indicating the correspondence between the item identification information and the item value entered in the cell corresponding to the item identification information by referring to the document information. A document analysis unit 111, and an answer acquisition unit 118 that transmits the item value acquisition prompt to the language model server 210 and acquires the item value acquisition result directly or indirectly from the language model server 210.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a document analysis apparatus and a document analysis method for reading item values from tabular documents.

Background Art

[0002] Natural language processing using large language models is becoming widespread, and attention is focused on automating and streamlining document processing such as evaluation / review that has conventionally been performed by humans using this technology. The input to a large language model (also simply referred to as a language model) is text. Therefore, it is necessary to extract text-formatted information from documents.

[0003] As a technique for extracting information in consideration of the structure and content of a document, there is a system described in Patent Document 1. This system enables a user to select and submit one or more text segments for analysis and to select one or more citations that match the text segments and the profile data included in the document from a set of recommended citations. The system also queries one or more citation libraries or source databases to find citations for recommendations that best match the text segments selected and submitted by the author. Further, the system automatically processes the data submitted by the author while the document is presented by a document rendering application and generates a set of recommended citations for consideration and inclusion in the document.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] When using a language model for tabular documents, it is necessary to convert them into text (text information) that reflects the document structure for processing. Especially for tabular documents with templates (formats, styles), there are many forms for filling in item values, and it is required to extract the correspondence between items (item identification information, item names) and item values. The technology of Patent Document 1 is not suitable for information extraction from tabular documents.

[0006] The present invention has been made in view of such a background, and an object thereof is to provide a document analysis apparatus and a document analysis method that enable extraction of text-formatted information reflecting the structure from tabular documents.

Means for Solving the Problems

[0007] In order to solve the above problems, a document analysis apparatus according to the present invention extracts document information indicating the correspondence between the identification information of each cell and the item value of the cell from a tabular document including cells in which item values, which are values of items, are entered, and refers to the identification information of the cells in which the item values to be acquired among the cells are entered and the form information indicating the correspondence between the item identification information of the items and the identification information of the cells, and requests an item value acquisition result indicating the correspondence between the item identification information and the item value entered in the cell corresponding to the item identification information. Merge the text, the format information, and the document information Create an item value acquisition prompt and the previous Send the item value acquisition prompt to a language model server and obtain the item value acquisition result directly or indirectly from the language model server. document analysis unit It is provided with.

Effects of the Invention

[0008] According to the present invention, it is possible to provide a document analysis apparatus and a document analysis method that enable extraction of text-formatted information reflecting the structure from tabular documents. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Modes for Carrying Out the Invention

[0010] ≪Outline of Document Processing Apparatus≫ The following describes the outline of a document processing apparatus as a document analysis apparatus, a document evaluation apparatus, a document evaluation criterion extraction apparatus, and a document correction apparatus in an embodiment (embodiment) for carrying out the present invention. First, the document processing apparatus functions as a document analysis apparatus that extracts information considering the document structure from a tabular document such as a spreadsheet. More specifically, the document processing apparatus refers to the format information of the document, extracts the item value corresponding to the item identification information (item name), and outputs it as text.

[0011] Second, the document processing apparatus functions as a document evaluation apparatus that evaluates / reviews a document. More specifically, the document processing apparatus evaluates the document with reference to the text extracted from the document and the evaluation criteria of the document, and outputs a report (evaluation report). Third, the document processing apparatus functions as a document evaluation criterion extraction apparatus that extracts the evaluation criteria (check contents) of a document from a document checklist or the minutes of a review meeting. Fourth, the document processing apparatus functions as a document correction apparatus that corrects fluctuations in the notation of a document. More specifically, the document processing apparatus detects synonyms from the document and corrects fluctuations in the notation. After explaining the basic configuration of the document processing apparatus, the functional configurations of the document analysis apparatus, document evaluation apparatus, document evaluation criterion extraction apparatus, and document correction apparatus will be described in this order.

[0012] ≪Configuration of Document Processing Apparatus≫ FIG. 1 is a functional block diagram of a document processing apparatus 100 according to the present embodiment. The document processing apparatus 100 is a computer and includes a control unit 110, a storage unit 130, and an input / output unit 180. User interface devices such as a display, a keyboard, and a mouse are connected to the input / output unit 180. The input / output unit 180 includes a communication device and can transmit and receive data to and from a language model server 210 and other devices. Note that the language model server 210 is a server of a service (dialogue artificial intelligence service) that receives a prompt including an instruction and returns an answer to the prompt.

[0013] ≪Document Processing Apparatus: Storage Unit≫ The storage unit 130 includes storage devices such as a ROM (Read Only Memory), a RAM (Random Access Memory), and an SSD (Solid State Drive). A document database 131, a document format database 140, a document evaluation report rule database 150, an evaluation report database 160, and a program 138 are stored in the storage unit 130. Note that various storage contents of the storage unit 130 may be read as needed from an external storage device such as a cloud server.

[0014] Documents to be processed by the document processing apparatus 100 are stored in the document database 131 in association with document identification information. The program 138 includes descriptions of processes executed by functional units provided in the control unit 110 described later. Other configurations of the storage unit 130 will be described later.

[0015] ≪Document Processing Apparatus: Control Unit≫ The control unit 110 is configured to include a CPU (Central Processing Unit), and is provided with a document analysis unit 111, a document format generation unit 112, a document evaluation unit 113, an evaluation report unit 114, an evaluation criterion extraction unit 115, a dictionary generation unit 116, a document correction unit 117, and a response acquisition unit 118. The response acquisition unit 118 sends the prompts generated by other functional units provided in the control unit 110 described below to the language model server 210, and directly or indirectly receives responses. The control unit 110 may be configured to include a GPU (Graphics Processing Unit), an NPU (Neural (network) Processing Unit), an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), etc.

[0016] ≪Document processing device: Document analysis device≫ Hereinafter, the document processing device 100 as a document analysis device will be described. The document processing device 100 refers to the format information (template information) of a tabular document, and extracts pairs of item identification information (item name) and item values for the items included in the tabular document.

[0017] ≪Document analysis device: Tabular document · Format information≫ FIG. 2 is a diagram showing an example of a tabular document 410 according to the present embodiment. The tabular document 410 is a spreadsheet-formatted document, and is a document in which item values corresponding to item names are entered / described in the cells adjacent to the cells (columns) in which the item names are described. In FIG. 2, the item name is described in the cell immediately to the right of the cell in which the item name is described. For example, the item value corresponding to the item name "Required delivery date" is "2020 / 10 / 10". Note that the item value may be described in the cell below the item name. Note that an example of a tabular document / spreadsheet-formatted document is an Excel (registered trademark) document, but it is not limited thereto.

[0018] In a spreadsheet-formatted document, cells are identified by identification information indicating position information. For example, the identification information of the cell in which the item name "Required delivery date" is described is "B6:C7". The tabular document 410 is stored in the document database 131 in association with document identification information. The document database 131 may store documents in text format.

[0019] FIG. 3 is a diagram showing a format information document 420 (template document) of the tabular document according to the present embodiment. The cells of the item names are the same as those of the tabular document 410. In the cells where the item values are entered, a character string with "#" added to the head of the item name (item identification information) is described. Hereinafter, unless there is a need for particular distinction, the item identification information and the item identification information with "#" added to the head will be described without distinction. For example, "request delivery date" as item identification information and "#request delivery date" are not distinguished.

[0020] ≪Document analysis device: Document format database≫ In the document format database 140, the format information document 420 is stored in association with format information identification information. Also, the identification information (position information) and values of the cells extracted from the format information document 420 by the document analysis unit 111 described later are stored in the document format database 140 in association with the format information document 420. In other words, the document format database 140 stores format information identification, the format information document 420 (see FIG. 3), and the correspondence between the identification information and values of the cells of the format information document 420 (see element 612 in FIG. 4 described later). Hereinafter, the correspondence between the identification information and values of the cells of the format information document 420 will also be referred to as format information.

[0021] ≪Document analysis device: Document analysis unit≫ Before starting the processing of the tabular document 410, the document analysis unit 111 extracts format information (the correspondence between the identification information and values of the cells) from the format information document 420 and stores it in the document format database 140. The document analysis unit 111 extracts the item value corresponding to the item name (item identification information) from the tabular document 410 with reference to the format information. More specifically, the document analysis unit 111 first extracts document information (see element 613 in FIG. 4 described later), which is the correspondence between the identification information (position information) and values of the cells, from the tabular document 410.

[0022] Next, the document analysis unit 111 generates an item value acquisition prompt 610 (see FIG. 4 described later) based on the extracted information. Subsequently, the document analysis unit 111 instructs the response acquisition unit 118 to send the item value acquisition prompt 610 to the language model server 210 and receives an item value acquisition result 430 (see FIG. 5 described later) as a response.

[0023] ≪Document Analysis Apparatus: Item Value Acquisition Prompt≫ FIG. 4 is a diagram showing the configuration of an item value acquisition prompt 610 according to the present embodiment. Generally, a prompt is configured to include one or more elements. The item value acquisition prompt 610 includes elements 611 to 613. Element 611 indicates an instruction / direction to extract an item value corresponding to item identification information starting with "#" from the document shown in element 613 with reference to the format (format information) shown in element 612.

[0024] Element 612 is the correspondence (format information) of the cell identification information and value extracted from the format information document 420 (see FIG. 3). Each row shows the correspondence of the identification information and value. Element 613 is the correspondence (document information) of the cell identification information and value extracted from the tabular document 410 (see FIG. 2). Each row shows the correspondence of the identification information and value.

[0025] ≪Document Analysis Apparatus: Item Value Acquisition Result≫ FIG. 5 is a diagram showing an item value acquisition result 430 according to the present embodiment. The item value acquisition result 430 is data in JSON format specified by element 611 of the item value acquisition prompt 610 and shows the correspondence between the item identification information (item name) and the item value. Note that JSON is an abbreviation for JavaScript Object Notation and is the format of an object in the programming language JavaScript.

[0026] Taking the "Business Order Certainty" which is item identification information as an example, the correspondence will be explained. The identification information of the cell in which the document format information "Business Order Certainty" is described is "E2:F3" (refer to element 612 described in FIGS. 3 and 4). The language model server 210 acquires "A" as the corresponding value at "E2:F3" in element 613 and includes it in the item value acquisition result 430. The document analysis unit 111 associates the item value acquisition result 430 with the tabular document 410 and stores it in the document database 131.

[0027] As described above, the document processing apparatus 100 as a document analysis apparatus includes a document analysis unit 111 that extracts document information (refer to element 613) indicating the correspondence between the identification information of each cell and the item value of the cell from the tabular document 410 including the cell in which the item value, which is the value of the item, is entered.

[0028] The document analysis unit 111 creates an item value acquisition prompt 610 (refer to FIG. 4) that requests format information (refer to element 612) indicating the correspondence between the identification information of the cell in which the item value to be acquired is entered among the cells and the item identification information of the item, and the item value acquisition result 430 (refer to FIG. 5) indicating the correspondence between the item identification information and the item value entered in the cell corresponding to the item identification information, with reference to the document information. The document processing apparatus 100 as a document analysis apparatus includes a response acquisition unit 118 that transmits the item value acquisition prompt 610 to the language model server 210 and directly or indirectly acquires the item value acquisition result 430 from the language model server 210.

[0029] ≪Document Analysis Apparatus: Document Analysis Processing≫ FIG. 6 is a flowchart of the document analysis processing according to the present embodiment. In step S11, the document analysis unit 111 extracts the identification information and value of the cell from the tabular document 410 which is the document to be processed (refer to element 613 described in FIG. 4).

[0030] Step S12 is a document format acquisition process that acquires format information corresponding to the tabular document 410 of the document to be processed from the document format database 140. Details will be described later with reference to FIG. 7. In step S13, the document analysis unit 111 generates an item value acquisition prompt 610 (see FIG. 4). Here, element 612 is the format information acquired in step S12. Element 613 is the correspondence between the cell identification information and the value acquired in step S11.

[0031] In step S14, the document analysis unit 111 instructs the response acquisition unit 118 to send the item value acquisition prompt 610 to the language model server 210. Next, the document analysis unit 111 stores the item value acquisition result 430 (see FIG. 5), which is the response, in the document database 131 in association with the tabular document 410.

[0032] ≪Document Analysis Apparatus: Document Format Acquisition Process≫ FIG. 7 is a flowchart of the document format acquisition process according to the present embodiment. The document format acquisition process of step S12 (see FIG. 6) will be described with reference to FIG. 7. In step S21, the document analysis unit 111 starts a process of repeating the process of step S22 for each format information stored in the document format database 140. Note that the format information is the correspondence between the cell identification information and the value of the format information document 420 (see element 612 in FIG. 4).

[0033] In step S22, the document analysis unit 111 calculates the degree of match between the format information and the tabular document 410 that is the processing target of the document analysis process (see FIG. 6). The degree of match is the ratio or number of matches between the values and identification information of the cells in the format information that do not start with "#" and the values and identification information of the cells included in the tabular document 410. In other words, it is the ratio or number of matches between the cells in the format information document 420 that do not start with "#" and have item names described, and the cells in the tabular document 410 that have the same item names described. Note that the cells with item names are, for example, the cells of "(Sales) Order Accuracy" in "B2:D3" and "Completion" in "F6:F7". In step S23, the document analysis unit 111 sets the format information with the maximum degree of match calculated in step S22 as the format information of the tabular document 410 of the document to be processed.

[0034] As described above, the cell identification information is information indicating the position of the cell in the tabular document 410. Among the format information in the document format database 140 that stores one or more format information, for the cells included in the format information and the cells included in the document information (refer to element 613), the document analysis unit 111 refers to the format information with the largest number of cells where the cell identification information and the cell value match, and creates the item value acquisition prompt 610.

[0035] ≪Features of the Document Analysis Device≫ The document processing device 100 as a document analysis device extracts the item value acquisition result 430 (refer to FIG. 5) from the tabular document 410 (refer to FIG. 2) with reference to the format information (refer to element 612 described in FIG. 4). Even for a document with a complex tabular format, it is possible to obtain the item value acquisition result that is the correspondence between the item identification information and the item value necessary for performing document processing such as review / evaluation. According to the document format acquisition process (refer to FIG. 7), the format information that most closely matches the format of the document to be processed is acquired. There is no need to specify the format (template document) for each document, and the item value acquisition result can be efficiently obtained.

[0036] ≪Modification Example of the Document Analysis Device: Item Value Acquisition Result≫ The item value acquisition result 430 (refer to FIG. 5) was in JSON format, but it is not limited to this. It may be in text format. FIG. 8 is a diagram showing the text format item value acquisition result 430A according to a modification example of the present embodiment. The item value acquisition result 430A is text format data with the same content as the item value acquisition result 430.

[0037] ≪Document Analysis Device: Document Format Generation Unit≫ In the above-described embodiment, the format information (refer to element 612 shown in FIG. 4) is extracted from a format information document 420 (refer to FIG. 3) in which item identification information starting with "#" is described. The document format generation unit 112 extracts format information from a format information document 420A (refer to FIG. 9 to be described later) in which the item identification information is blank. FIG. 9 is a diagram showing a format information document 420A (template document) in which the item identification information according to a modification example of the present embodiment is blank. Compared with the format information document 420, the item identification information is blank.

[0038] FIG. 10 is a flowchart of the format information acquisition process according to the present embodiment. With reference to FIG. 10, a process of acquiring format information from a format information document 420A in which an entry field (cell) for an item value is blank will be described. By the document format acquisition process, it becomes possible to acquire format information without preparing a format information document 420 in which item identification information starting with "#" is entered.

[0039] In step S31, the document format generation unit 112 extracts the correspondence (template document information) between the cell identification information and the value from the format information document 420A. In step S32, the document format generation unit 112 acquires a cell that is a blank cell, the cell above it is not blank, and the cell to its left is blank. Note that the cells above / below / to the left / right of a cell can be determined from the identification information indicating the position. Also, assume that the cell above has the same width and the cell to the left has the same height.

[0040] In step S33, the document format generation unit 112 starts a process of repeating the process of step S34 for each cell acquired in step S32. In step S34, the document format generation unit 112 uses, as format information, the item identification information obtained by adding "#" to the beginning of the value (character string) of the cell above and the cell identification information.

[0041] In step S35, the document format generation unit 112 acquires a cell that is a blank cell, the cell to its left is not blank, and the cell above it is blank. Note that the cell above has the same width and the cell to the left has the same height. In step S36, the document format generation unit 112 starts a process of repeating the process of step S37 for each cell acquired in step S35.

[0042] In step S37, the document format generation unit 112 uses the item identification information obtained by adding "#" to the head of the value (character string) of the left adjacent cell and the identification information of the cell as format information. In step S38, the document format generation unit 112 stores the format information acquired in steps S34 and S37 in the document format database 140 as the format information of the format information document 420A.

[0043] As described above, the document processing device 100 as a document analysis device includes a document format generation unit 112 that extracts template document information indicating the correspondence between the identification information indicating the position of each cell and the value entered in the cell from a template document (refer to the format information document 420A shown in FIG. 9) in which the item values are not entered and that serves as a template for the tabular document 410.

[0044] When the cell immediately above a blank cell (a blank cell) is not blank and the cell immediately to the left is blank, the document format generation unit 112 associates the identification information of the blank cell with the value of the cell immediately above as item identification information. Also, when the cell immediately to the left of a blank cell is not blank and the cell immediately above is blank, the document format generation unit 112 associates the identification information of the blank cell with the value of the cell immediately to the left as item identification information. The document format generation unit 112 uses the correspondence between the identification information of the associated blank cell and the item identification information as format information.

[0045] ≪Document Processing Device: Document Evaluation Device≫ Hereinafter, the document processing device 100 as a document evaluation device will be described. The document processing device 100 performs evaluation / review on a document including items, such as the tabular document 410, in accordance with the evaluation criteria for each item and reports the results. Note that even for a document in text format that includes structures such as chapters, sections, and items, it is also a processing target of the document evaluation device. In the following description, elements / components of the document structure such as chapters and sections are also referred to as items.

[0046] <<Document Evaluation Device: Overview of Processing>> FIG. 11 is a diagram for explaining the overview of document evaluation report processing (see FIG. 24 described later) by the document processing device 100 according to the present embodiment. The document evaluation unit 113 generates an evaluation acquisition prompt 620 (see FIG. 12 described later) based on the evaluation criteria stored in the document evaluation report rule database 150 and the item value acquisition result 430 (see FIG. 5). The evaluation acquisition prompt 620 is transmitted to the language model server 210, and an evaluation result 450 (see FIG. 13 described later) is returned. The evaluation result 450 is the result of the evaluation / review of the item value acquisition result 430 according to the evaluation criteria.

[0047] The evaluation report unit 114 generates a report acquisition prompt 630 (see FIG. 15 described later) based on the evaluation result 450 and the report examples stored in the document evaluation report rule database 150. The report acquisition prompt 630 is transmitted to the language model server 210, and a report 460 (see FIG. 17 described later) is returned. Hereinafter, the details of the document processing device 100 as the document evaluation device will be described.

[0048] <<Document Evaluation Device: Document Evaluation Report Rule Database>> In the document evaluation report rule database 150, the evaluation criteria for evaluating / reviewing a document are stored in association with format information (see the element 612 shown in FIG. 4 of the document format database 140). In the format information, not only the items (item identification information and item values) of the tabular document, but also the components such as chapters, sections, and charts included in the text document are regarded as items. The evaluation criteria (see the element 622 in FIG. 12 described later) are the evaluation criteria for each of these items. Since the format information includes one or more items, one or more evaluation criteria are associated with one format information.

[0049] Also, in the document evaluation report rule database 150, questions corresponding to the evaluation criteria (see the element 624 in FIG. 12 described later) are stored. The questions are questions for obtaining the evaluation results for each item or evaluation criterion. Since the format information includes one or more items, one or more questions are associated with one format information.

[0050] Furthermore, the document evaluation report rule database 150 also stores report examples (see elements 632 to 635 in FIG. 15 described later) regarding the evaluation results evaluated / reviewed based on the evaluation criteria. The report examples are report examples of the evaluation results for each item. Since the format information includes one or more items, one or more report examples are associated with one piece of format information. In addition, the document evaluation report rule database 150 may store a document (also referred to as a checklist) that describes the viewpoints, check contents, and check rules for the evaluation / review of documents.

[0051] ≪Document Evaluation Device: Evaluation Report Database≫ In the evaluation report database 160, the evaluation results of evaluating the document and the reports summarizing the evaluation results are stored in association with the document identification information of the document to be evaluated. Furthermore, in the evaluation report database 160, the minutes of the meeting for evaluating / reviewing the document while referring to the document and the report are stored in association with the document identification information of the document to be evaluated.

[0052] ≪Document Evaluation Device: Document Evaluation Unit≫ The document evaluation unit 113 evaluates / reviews the document by referring to the evaluation criteria related to the items constituting the document. The document evaluation unit 113 generates an evaluation acquisition prompt 620 (see FIG. 12 described later) based on the evaluation criteria and the document information. Subsequently, the document evaluation unit 113 instructs the response acquisition unit 118 to send the evaluation acquisition prompt 620 to the language model server 210 and receives the evaluation result (simply referred to as an evaluation) as the response.

[0053] ≪Document Evaluation Device: Evaluation Acquisition Prompt≫ FIG. 12 is a diagram showing the configuration of the evaluation acquisition prompt 620 according to the present embodiment. The evaluation acquisition prompt 620 includes elements 621 to 624. Element 621 indicates an instruction / direction to answer the question shown in element 624 regarding the document that is the evaluation target shown in element 623.

[0054] Element 622 is an evaluation criterion for items included in a document to be evaluated (see Element 623). The evaluation criteria are stored in the document evaluation report rule database 150. Element 623 is the document information to be evaluated (evaluated document information). In FIG. 12, the document information at Element 623 is the correspondence between the item identification information and the item values obtained by the document analysis unit 111 from a tabular document (see the item value acquisition result 430 shown in FIG. 5). The document information of Element 623 may be text-formatted document information created by a word processor. For example, it may be the text-formatted item value acquisition result 430A (see FIG. 8).

[0055] Element 624 is a question for obtaining an evaluation result according to the evaluation criteria in Element 622. The question may include a procedure for obtaining the evaluation result. The questions shown in FIG. 12 include (1) correctly counting the numbers of "high", "medium", and "low" in the "evaluation result", (2) obtaining a project risk rank according to the evaluation criteria, and (3) comparing the obtained project risk rank with the project risk rank included in the evaluation target. In this way, the evaluation acquisition prompt 620 is a prompt for obtaining an evaluation result by referring to the evaluation criteria related to the item of "project risk rank".

[0056] ≪Document Evaluation Device: Evaluation Result≫ FIG. 13 is a diagram showing the evaluation result 450 according to the present embodiment. FIG. 14 is a diagram showing the evaluation result 450A according to the present embodiment. The evaluation results 450 and 450A are answers along the procedures shown in the questions of Element 624 included in the evaluation acquisition prompt 620. Note that the evaluation results 450 and 450A have different evaluation results depending on whether they match or do not match in the comparison of the project risk rank, which is the last procedure. In this way, the document evaluation unit 113 uses the evaluation acquisition prompt 620 to inquire of the language model server 210 for each evaluation criterion for each item included in the document to be evaluated and the format information corresponding to the document, and obtains the evaluation results 450 and 450A.

[0057] As described above, the document processing apparatus 100 as a document evaluation apparatus includes a document evaluation unit 113 that creates an evaluation acquisition prompt 620 for requesting evaluation results 450, 450A of evaluation target document information (see element 623) including the correspondence between item identification information and item values for one or more items, and an evaluation criterion (see element 622) related to the item, based on the evaluation criterion. The document processing apparatus 100 as a document evaluation apparatus includes a response acquisition unit 118 that transmits the evaluation acquisition prompt 620 to the language model server 210 and acquires the evaluation results 450, 450A directly or indirectly from the language model server 210. The request for the evaluation result included in the evaluation acquisition prompt 620 is a request in a question format (see element 624) for obtaining the evaluation results 450, 450A.

[0058] ≪Document Evaluation Apparatus: Evaluation Report Section≫ The evaluation report unit 114 generates an evaluation report with reference to the evaluation results 450, 450A. Specifically, the evaluation report unit 114 generates a report acquisition prompt 630 (see FIG. 15 described later) based on the evaluation results and report examples. Subsequently, the evaluation report unit 114 instructs the response acquisition unit 118 to send the report acquisition prompt 630 to the language model server 210 and receives the report as the response. Similar to the document evaluation unit 113, the evaluation report unit 114 generates a report acquisition prompt 630 for each item constituting the document and receives a report 460 (see FIG. 17 described later).

[0059] ≪Document Evaluation Apparatus: Report Acquisition Prompt≫ FIGS. 15 and 16 are diagrams showing the configuration of the report acquisition prompt 630 according to the present embodiment. The report acquisition prompt 630 includes elements 631 to 637. Element 631 indicates an instruction / direction to report the question shown in element 636 and the answer (evaluation result) to the question shown in element 637 with reference to the report examples shown in elements 632 to 635.

[0060] Element 632 is an example of a question and an answer where the determination results do not match. Element 633 is an example of a report when the determination results shown in element 632 do not match. The report includes the presence or absence of pointed-out matters, the pointed-out content, and amendments. Element 634 is an example of a question and an answer where the determination results match. Element 635 is an example of a report when the determination results shown in element 634 match. The report includes the presence or absence of pointed-out matters, the pointed-out content, and amendments, but the content is equal to nothing.

[0061] Move on to FIG. 16 and continue the explanation of the report acquisition prompt 630. Element 636 is a question for obtaining the evaluation result in element 624 of the evaluation acquisition prompt 620 (see FIG. 12). Element 637 is the evaluation result 450 (see FIG. 13), which is the answer to the question.

[0062] ≪Document Evaluation Device: Report≫ FIG. 17 is a diagram showing the report 460 according to the present embodiment. The answer, which is element 637, is the evaluation result 450 where the determination results do not match, and corresponds to the question and answer in element 632. Therefore, the report 460 is a report along the lines of the report in element 633. As described above, the evaluation report unit 114 uses the report acquisition prompt 630 including examples of reports corresponding to the evaluation results 450, 450A for each item acquired by the document evaluation unit 113 to inquire of the language model server 210 and obtain the report 460.

[0063] As described above, the document processing device 100 as a document evaluation device includes an evaluation report unit 114 that creates a report acquisition prompt 630 for requesting a report 460 in a predetermined format based on the evaluation results 450, 450A. The answer acquisition unit 118 transmits the report acquisition prompt 630 to the language model server 210 and acquires the report 460 directly or indirectly from the language model server 210.

[0064] The report acquisition prompt 630 includes examples of questions included in a question-formatted request (see elements 632, 634), examples of evaluation results as answers to the questions (see elements 632, 634), examples of reports in a predetermined format for the examples of the questions and the examples of the answers (see elements 633, 635), a question-formatted request included in the evaluation acquisition prompt 620 (see element 636), and evaluation results acquired by the answer acquisition unit 118 (see element 637).

[0065] Also, the evaluation target document information (see element 623 described in FIG. 12) includes an evaluation item (for example, "project risk rank") whose item value is determined according to the number of items whose item values are within a predetermined range of one or more predetermined values (for example, "high", "medium", "low"). The evaluation criteria (see element 622) include criteria for determining the values of the evaluation items. The evaluation acquisition prompt 620 includes a request to count the number of items whose item values are within a predetermined range, a request to obtain a counted evaluation item value that is the value of an evaluation item determined according to the counted result, and a request to report whether the counted evaluation item value matches the described evaluation item value that is the value of the evaluation item included in the evaluation target document information (see element 624).

[0066] The predetermined format of the report (see element 633 described in FIG. 15) is a format for reporting a discrepancy when the counted evaluation item value and the described evaluation item value do not match. The report in the predetermined format (see element 633) is a format that includes an amendment for correcting the described evaluation item value to the counted evaluation item value when the counted evaluation item value and the described evaluation item value do not match.

[0067] <<Modification Examples of Evaluation Acquisition Prompt and Report Acquisition Prompt>> The evaluation criteria (see element 622) included in the evaluation acquisition prompt 620 (see FIG. 12) are evaluation criteria in which different item values ("AA", "A", "B", "C") are determined according to the number of item values ("high", "medium"). Some examples of evaluation criteria / check contents other than such a style are shown.

[0068] FIG. 18 is a diagram showing the evaluation acquisition prompt 620A according to the present embodiment. The evaluation criteria in the element 622A of the evaluation acquisition prompt 620A are evaluation criteria based on two item values ("project risk rank by scale" and "project risk rank by required quality") for determining another item value ("final project risk rank"). In other words, the evaluation acquisition prompt 620A is a prompt for acquiring an evaluation result related to the item of "final project risk rank".

[0069] FIG. 19 is a diagram showing the report acquisition prompt 630A according to the present embodiment. The report acquisition prompt 630A corresponds to the evaluation acquisition prompt 620A and is a prompt for acquiring a report of the evaluation result related to the item of "final project risk rank". Element 632A is an example where the determination results do not match. Element 633A is an example of a report when the determination results do not match as shown in element 632A. Element 634A is an example where the determination results match. Element 635A is an example of a report when the determination results match as shown in element 634A.

[0070] As described above, the evaluation target document information (see element 623A) includes an evaluation item (e.g., "final project risk rank") that is an item whose item value is determined according to the maximum or minimum item value among a plurality of item values of comparison target item values (e.g., "project risk rank by scale" and "project risk rank by required quality").

[0071] The evaluation criteria include criteria for determining the value of the evaluation item. The evaluation acquisition prompt 620A includes a request for obtaining a comparative evaluation item value that is the value of the evaluation item determined according to the result of comparing the comparison target item values, and a request for reporting whether the comparative evaluation item value matches the described evaluation item value that is the value of the evaluation item included in the evaluation target document information (see element 624A).

[0072] The prescribed format of the report (refer to element 633A) is a format that reports a discrepancy when the comparison evaluation item value and the described evaluation item value do not match. The prescribed format of the report (refer to element 633A) is a format that includes an amendment to correct the described evaluation item value to the comparison evaluation item value when the comparison evaluation item value and the described evaluation item value do not match.

[0073] Figure 20 is a diagram showing the evaluation acquisition prompt 620B according to this embodiment. The evaluation criteria in element 622B of the evaluation acquisition prompt 620B are determination criteria for determining whether the item value is a prescribed value (for example, blank) and determining the appropriateness (presence or absence of violation) of the item value. Figure 21 is a diagram showing the report acquisition prompt 630B according to this embodiment. Element 632B is an example of a violation. Element 633B is an example of a report when there is a violation as shown in element 632B. Element 634B is an example of no violation. Element 635B is an example of a report when there is no violation as shown in element 634B.

[0074] As described above, the evaluation criteria (element 622B) include prohibited item identification information, which is item identification information for which it is prohibited for the item value to be a prescribed value (for example, blank). The evaluation acquisition prompt 620B includes a request to obtain the item value of the prohibited item identification information included in the evaluation target document information, and a request to obtain prohibited value entry item identification information indicating the item identification information for which the obtained item value is a prescribed value. The prescribed format of the report (refer to element 633B) is a format that includes prohibited value entry item identification information.

[0075] Figure 22 is a diagram showing the evaluation acquisition prompt 620C according to this embodiment. The evaluation criteria in element 622C of the evaluation acquisition prompt 620C are evaluation criteria for determining the appropriateness of the item value. Figure 23 is a diagram showing the report acquisition prompt 630C according to this embodiment. Element 632C is an example where the item value is inappropriate (violating the evaluation criteria). Element 633C is an example of a report when the item value is inappropriate as shown in Element 632B. Element 634C is an example where the item value is appropriate. Element 635C is an example of a report when the item value is appropriate as shown in Element 634B.

[0076] ≪Document Evaluation Device: Document Evaluation Report Processing≫ FIG. 24 is a flowchart of the document evaluation report processing according to the present embodiment. The process of obtaining the evaluation result for a document and further generating a report will be described with reference to FIG. 24.

[0077] In step S41, the document evaluation unit 113 acquires the document information to be the target of the evaluation report. For example, for a tabular document, the document evaluation unit 113 acquires the item value acquisition result 430 (see FIG. 5) extracted by the document analysis unit 111 from the document database 131. Next, the document evaluation unit 113 acquires the evaluation criteria, questions, and report examples from the document evaluation report rule database 150 based on the format information corresponding to the document.

[0078] In step S42, the document evaluation unit 113 starts the process of repeating steps S43 to S44 for each item with an evaluation criterion. In step S43, the document evaluation unit 113 generates an evaluation acquisition prompt 620 (see FIG. 12). Here, element 622 is the evaluation criterion corresponding to the item among the evaluation criteria acquired in step S41. Also, element 624 is the question corresponding to the evaluation criterion. Element 623 is the document information acquired in step S41. Element 624 is the question corresponding to the evaluation criterion (corresponding to the item) among the questions acquired in step S41.

[0079] In step S44, the document evaluation unit 113 instructs the response acquisition unit 118 to send the evaluation acquisition prompt 620 to the language model server 210. Next, the document evaluation unit 113 stores the evaluation results 450, 450A (see FIGS. 13 and 14), which are the responses, in the evaluation report database 160 in association with the document identification information of the document.

[0080] In step S45, the evaluation report unit 114 starts a process of repeating steps S46 to S47 for each item with an evaluation criterion. In step S46, the evaluation report unit 114 generates a report acquisition prompt 630 (see FIGS. 15 to 16). Here, elements 632 to 635 are report examples corresponding to the item (evaluation criterion) among the report examples acquired in step S41. Further, element 636 is a question corresponding to the item, and is the question in element 624 of the evaluation acquisition prompt 620 generated in step S43. Element 637 is the evaluation result that is the response acquired in step S44.

[0081] In step S47, the evaluation report unit 114 instructs the response acquisition unit 118 to send the report acquisition prompt 630 to the language model server 210. Next, the evaluation report unit 114 stores the report 460 (see FIG. 17), which is the response, in the evaluation report database 160 in association with the document identification information of the document. In step S48, the evaluation report unit 114 merges the reports acquired in step S47 to create a report on the evaluation of the document and stores it in the evaluation report database 160.

[0082] ≪Features of the Document Evaluation Device≫ The document processing apparatus 100 as a document evaluation apparatus acquires evaluation results 450, 450A for each item with reference to evaluation criteria (see elements 622, 622A, 622B, 622C described in FIGS. 12, 18, 20, 22). Further, the document processing apparatus 100 acquires a report 460 for each item with reference to report examples (see elements 632 to 635, 632A to 635A, 632B to 635B, 632C to 635C described in FIGS. 15, 19, 21, 23). By evaluating / reviewing each item included in the document and generating a report, the document processing apparatus 100 can acquire an evaluation result and a report with high accuracy according to the item. Also, by preparing evaluation criteria and questions according to the item, it becomes possible to expect to obtain an evaluation result and a report with higher accuracy.

[0083] <<Document Processing Apparatus: Document Evaluation Criterion Extraction Apparatus>> Hereinafter, the document processing apparatus 100 as a document evaluation criterion extraction apparatus will be described. The document processing apparatus 100 collects evaluation criteria, check contents, and check rules for each item in a document including items, such as the tabular document 410. The document processing apparatus 100 extracts and collects them from the minutes of an evaluation / review meeting and a document describing evaluation criteria and check contents. Hereinafter, evaluation criteria, check contents, and check rules are collectively referred to as evaluation criteria.

[0084] <<Document Evaluation Criterion Extraction Apparatus: Evaluation Criterion Extraction Unit>> FIG. 25 is an example of a document 510 including the evaluation criteria according to the present embodiment. The document 510 is, for example, the minutes of a meeting in which the document was evaluated / reviewed. The minutes include evaluation criteria including review perspectives and rules. The evaluation criteria extraction unit 115 extracts evaluation criteria from a document containing the evaluation criteria. Specifically, the evaluation criteria extraction unit 115 generates an evaluation criteria acquisition prompt 640 (see FIG. 26 described later) that instructs / orders the extraction of evaluation criteria from the document. Subsequently, the evaluation criteria extraction unit 115 instructs the response acquisition unit 118 to send the evaluation criteria acquisition prompt 640 to the language model server 210 and receives the evaluation criteria as the response. The evaluation criteria extraction unit 115 stores the received evaluation criteria in the document evaluation report rule database 150 in association with the format information of the document to be evaluated / reviewed.

[0085] ≪Document Evaluation Criteria Extraction Device: Evaluation Criteria Acquisition Prompt≫ FIG. 26 is a diagram showing the configuration of the evaluation criteria acquisition prompt 640 according to the present embodiment. The evaluation criteria acquisition prompt 640 includes elements 641 to 642. Element 641 indicates an instruction / order to generate and acquire evaluation criteria with reference to the actual review points shown in element 642. Element 642 includes the text information contained in the document 510.

[0086] ≪Document Evaluation Criteria Extraction Device: Evaluation Criteria≫ FIG. 27 is a diagram showing the acquired evaluation criteria 520 according to the present embodiment. The evaluation criteria 520 are general evaluation criteria excluding the specific content of individual documents from the content of the review points included in element 642. In this way, the evaluation criteria extraction unit 115 extracts and acquires evaluation criteria from a document including the evaluation object.

[0087] As described above, the document processing device 100 as a document evaluation criteria extraction device includes an evaluation criteria extraction unit 115 that creates an evaluation criteria acquisition prompt 640 (see FIG. 26) for extracting the evaluation criteria of the document with reference to the minutes including the review results of the document. In addition, the document processing apparatus 100 as a document evaluation criterion extraction apparatus includes an answer acquisition unit 118 that transmits an evaluation criterion acquisition prompt 640 to the language model server 210 and acquires evaluation criteria directly or indirectly from the language model server 210. The evaluation criterion extraction unit 115 stores the evaluation criteria in the storage unit 130 (see the document evaluation report rule database 150).

[0088] ≪Document Evaluation Criterion Extraction Apparatus: Variation Example of Evaluation Criterion Acquisition Prompt≫ The evaluation criteria obtained using the evaluation criterion acquisition prompt 640 (see FIG. 26) are evaluation criteria related to the document that is the subject of review / evaluation, and by extension, documents in the same format as the said document. The evaluation criterion extraction unit 115 may acquire evaluation criteria for each item determined by the document format (format information).

[0089] FIG. 28 is a diagram showing an evaluation criterion acquisition prompt 640A according to the present embodiment. The evaluation criterion acquisition prompt 640A includes an element 643A that is not present in the evaluation criterion acquisition prompt 640. The element 641A indicates an instruction / order to generate and acquire evaluation criteria for each item of the document in the element 643A. FIG. 29 shows evaluation criteria 520A when the evaluation criterion acquisition prompt 640A according to the present embodiment is used. Evaluation criteria are extracted for each item (chapter).

[0090] As described above, the evaluation criterion extraction unit 115 creates an evaluation criterion acquisition prompt 640A that extracts evaluation criteria for each item of the said document (see the element 643A described in FIG. 28) by referring to the items described in the document and the minutes including the review results of the document. The evaluation criterion extraction unit 115 stores the evaluation criteria in the storage unit 130 (see the document evaluation report rule database 150) as evaluation criteria related to the items.

[0091] The document serving as the extraction source is not necessarily a text-formatted document such as minutes. Text information may be extracted from a checklist document in table format, and the evaluation criteria that are the contents of the check may be extracted. FIG. 30 is a diagram showing an evaluation criterion acquisition prompt 640B (check content acquisition prompt) according to the present embodiment. The evaluation criterion acquisition prompt 640B includes elements 641B and 642B.

[0092] Element 641B indicates an instruction / direction to generate and acquire check content (evaluation criteria) for each item (chapter) from the text information in element 642B. Element 642B indicates text information extracted from a checklist document (the "Function Specification Checklist" described in FIG. 30). FIG. 31 shows evaluation criteria 520B (check content) when the evaluation criterion acquisition prompt 640B according to the present embodiment is used. The evaluation criteria are extracted collectively for each item (chapter).

[0093] As described above, the evaluation criterion extraction unit 115 creates a check content acquisition prompt (see evaluation criterion acquisition prompt 640B) that extracts check content for each item (e.g., chapter) from a checklist document including the check content for each item of the document. The response acquisition unit 118 transmits the check content acquisition prompt to the language model server 210 and acquires the check content directly or indirectly from the language model server 210. The evaluation criterion extraction unit 115 stores the check content in the storage unit 130 (document evaluation report rule database 150) as evaluation criteria related to the items.

[0094] ≪Document Evaluation Criterion Extraction Device: Evaluation Criterion Extraction Process≫ FIG. 32 is a flowchart of the evaluation criterion extraction process according to the present embodiment. In step S51, the evaluation criterion extraction unit 115 acquires the minutes of the meeting in which the document was evaluated / reviewed from the evaluation report database 160. In step S52, the evaluation criterion extraction unit 115 starts a process of repeating steps S53 to S54 for each set of minutes.

[0095] In step S53, the evaluation criterion extraction unit 115 generates an evaluation criterion acquisition prompt 640A (see FIG. 28). Here, element 642A is the text information included in the minutes. Element 643A is an item (item identification information) included in the format information of the document to be evaluated / reviewed in the minutes.

[0096] In step S54, the evaluation criterion extraction unit 115 instructs the response acquisition unit 118 to send the evaluation criterion acquisition prompt 640A to the language model server 210. Next, the evaluation criterion extraction unit 115 stores the evaluation criterion 520A (see FIG. 29), which is the response, in the document evaluation report rule database 150 in association with the item of the format information. Note that the evaluation criterion extraction unit 115 may inquire the user of the document processing apparatus 100 as to whether or not to adopt the evaluation criterion before storing the evaluation criterion 520A in the document evaluation report rule database 150. Steps S55 to S58 are the same processing as steps S51 to S54, and the processing is performed for the checklist instead of the minutes.

[0097] ≪Features of the evaluation criterion extraction device≫ The document processing apparatus 100 as an evaluation criterion extraction device extracts evaluation criteria for each document format (format information) and for each item constituting the document from the minutes or the checklist. It becomes possible to acquire evaluation criteria without using manual labor. By collecting evaluation criteria and manually organizing them as necessary and referring to them in the document evaluation report process, it can be expected that the quality of document evaluation by the document processing apparatus 100 will be improved.

[0098] ≪Document processing apparatus: Document correction apparatus≫ Hereinafter, the document processing apparatus 100 as a document correction apparatus will be described. The document processing apparatus 100 performs document correction including elimination of fluctuations in word notation corresponding to the document or the business to which the document pertains.

[0099] ≪Document correction apparatus: Outline of processing≫ FIG. 33 is a diagram for explaining the outline of the document correction process (see FIG. 40 described later) by the document processing apparatus 100 according to the present embodiment. The dictionary generation unit 116 generates a synonym acquisition prompt 650 (see FIG. 34 described later) that refers to one or more related documents 540 such as documents related to the same business. The synonym acquisition prompt 650 is transmitted to the language model server 210, and a synonym dictionary 550 (see FIG. 35 described later) is returned. The synonym dictionary 550 is a synonym dictionary specialized for the context included in the document 540 and the business related to the document 540.

[0100] The document correction unit 117 generates a corrected document acquisition prompt 660 (see FIG. 36 described later) that refers to the synonym dictionary 550. The corrected document acquisition prompt 660 is transmitted to the language model server 210, and a corrected document 560 is returned. Hereinafter, the details of the document processing apparatus 100 as a document correction apparatus will be described.

[0101] ≪Document Correction Apparatus: Dictionary Generation Unit≫ The dictionary generation unit 116 generates a synonym acquisition prompt 650 (see FIG. 34 described later) that instructs / orders the generation of a synonym dictionary. Subsequently, the dictionary generation unit 116 instructs the response acquisition unit 118 to send the synonym acquisition prompt 650 to the language model server 210 and receives the synonym dictionary 550 (see FIG. 35 described later) as a response.

[0102] ≪Document Correction Apparatus: Synonym Acquisition Prompt≫ FIG. 34 is a diagram showing the configuration of the synonym acquisition prompt 650 according to the present embodiment. The synonym acquisition prompt 650 includes elements 651 to 652. Element 651 indicates an instruction / order to acquire a synonym dictionary (a group of synonyms) by referring to the text information of the document 540 shown in element 652. The instruction / order includes criteria for selecting which word as a headword among the synonyms. Examples of the criteria include katakana notation priority, kanji notation priority, abbreviation priority, etc. Element 652 includes the text information included in the document 540.

[0103] <<Document correction device: Synonym dictionary>> FIG. 35 is a diagram showing the acquired synonym dictionary 550 according to the present embodiment. In the synonym dictionary 550, words conforming to the heading word selection criteria instructed to the element 651 are used as heading words (keys in JSON format).

[0104] <<Document correction device: Document correction unit>> The document correction unit 117 generates a corrected document acquisition prompt 660 (see FIG. 36 described later) for instructing the correction of a document. Subsequently, the document correction unit 117 instructs the response acquisition unit 118 to send the corrected document acquisition prompt 660 to the language model server 210 and receives the corrected document as a response. The document correction unit 117 stores the corrected document in the document database 131 in association with the document to be configured.

[0105] <<Document correction device: Corrected document acquisition prompt>> FIG. 36 is a diagram showing the configuration of the corrected document acquisition prompt 660 according to the present embodiment. The corrected document acquisition prompt 660 includes elements 661 to 663. The element 661 indicates an instruction to correct and proofread the document shown in the element 663 with reference to the synonym dictionary 550 shown in the element 662. The element 661 includes instructions to correct not only the fluctuations of words but also the fluctuations of tenses and to correct the end of the sentence to the "desu / masu form". The synonym dictionary 550 is shown in the element 662. The text information of the document to be corrected is shown in the element 663.

[0106] As described above, the document processing device 100 as a document correction device includes a dictionary generation unit 116 that creates a synonym acquisition prompt 650 for requesting a synonym group in which words included in the text information (see the element 652 shown in FIG. 34) of a document group including one or more documents are grouped by synonyms. The document processing device 100 as a document correction device includes a response acquisition unit 118 that transmits the synonym acquisition prompt 650 to the language model server 210 and acquires a synonym group (see the synonym dictionary 550 shown in FIG. 35) directly or indirectly from the language model server 210.

[0107] The document processing apparatus 100 as a document correction apparatus includes a document correction unit 117 that creates a corrected document acquisition prompt 660 (see FIG. 36) for requesting a corrected document in which the fluctuations in the word notations of the documents included in the document group are corrected by referring to a synonym group. The response acquisition unit 118 transmits the corrected document acquisition prompt 660 to the language model server 210 and acquires a corrected document directly or indirectly from the language model server 210.

[0108] The synonym acquisition prompt 650 includes a request to use, as a headword, a word that satisfies a predetermined criterion for each synonym group (see element 651). The predetermined criterion is any one of katakana notation priority, kanji notation priority, and abbreviation priority. The corrected document acquisition prompt 660 requests a corrected document in which the fluctuations in the word notations are corrected to prioritize the headwords (see element 661).

[0109] Following the correction of the fluctuations in the word notations, the corrected document acquisition prompt 660 requests the correction of the fluctuations in the end notations (see element 661). The fluctuations in the end notations include at least one of the fluctuations in tense and the fluctuations between the desu / masu form and the da / de aru form.

[0110] ≪Document Correction Apparatus: Variants of Synonym Acquisition Prompt and Corrected Document Acquisition Prompt≫ FIG. 37 is a diagram showing a synonym acquisition prompt 650A according to the present embodiment. Compared with the synonym acquisition prompt 650 (see FIG. 34), it emphasizes summarizing synonyms in consideration of the context. FIG. 38 is a diagram showing a synonym dictionary 550A including meanings according to the present embodiment. FIG. 39 is a diagram showing a corrected document acquisition prompt 660A according to the present embodiment. Compared with the corrected document acquisition prompt 660 (see FIG. 36), it emphasizes correcting the fluctuations of words in consideration of the context.

[0111] As described above, the synonym acquisition prompt 650A requests a synonym group that takes into account the meaning in context. The synonym group includes the meaning in context. The corrected document acquisition prompt 660A requests a corrected document in which the fluctuations in word notation are corrected in consideration of the meaning in context.

[0112] ≪Document Correction Device: Document Correction Process≫ FIG. 40 is a flowchart of the document correction process according to the present embodiment. In step S61, the dictionary generation unit 116 acquires the related document 540 from the document database 131.

[0113] In step S62, the dictionary generation unit 116 generates the synonym acquisition prompt 650 (see FIG. 34). Here, the element 652 is the entire text information included in the document 540. In step S63, the dictionary generation unit 116 instructs the answer acquisition unit 118 to send the synonym acquisition prompt 650 to the language model server 210, and receives the synonym dictionary 550 as the answer.

[0114] In step S64, the document correction unit 117 starts a process of repeating steps S65 to S66 for each document acquired in step S61. In step S65, the document correction unit 117 generates the corrected document acquisition prompt 660 (see FIG. 36). Here, the element 662 is the text information included in the document.

[0115] In step S66, the document correction unit 117 instructs the answer acquisition unit 118 to send the corrected document acquisition prompt 660 to the language model server 210. The document correction unit 117 receives the configured corrected document as the answer and stores it in the document database 131 in association with the document before correction.

[0116] ≪Features of Document Correction Device≫ The document processing apparatus 100 as a document correction apparatus performs document correction to correct synonymous notations, tenses, and fluctuations in sentence endings. Regarding synonyms, corrections are made by referring to a synonym dictionary 550 generated based on related documents 540 such as documents of the same business (see element 662). Therefore, it is possible to correct words corresponding to the meanings specific to the business and the meanings according to the context in which the words are included.

[0117] ≪Variant Example of Document Correction Apparatus: Prompt for Obtaining Corrected Document≫ In the above-described embodiment, fluctuations in synonyms, fluctuations in tenses, and fluctuations in desu / masu style / darudesu style are corrected using one prompt for obtaining a corrected document 660, 660A. They may be corrected by dividing them into a plurality of prompts. By dividing them into a plurality of prompts, it can be expected that the accuracy of correction will be improved.

[0118] ≪Variant Example: Language Model≫ The above-described document processing apparatus 100 transmits a prompt to a language model server 210 that provides a language model service and obtains an answer. The document processing apparatus 100 itself may be equipped with a language model and query itself. FIG. 41 is a functional block diagram of a document processing apparatus 100A according to a variant example of the present embodiment. The document processing apparatus 100A includes a language model 132. The answer acquisition unit 118A outputs an answer using the language model 132 with the prompt as an input. By not using the language model server 210 of an external service, the risk of information leakage can be reduced. In addition, by using the language model 132 trained according to the business of the self-organization, the accuracy of the answer is improved, and an improvement in the accuracy of document evaluation / review can be expected.

[0119] As described above, the answer acquisition unit 118A provided in the document processing apparatus 100A as a document analysis apparatus obtains the item value acquisition result 430 for the item value acquisition prompt 610 using the language model 132 instead of the language model server 210.

[0120] The response acquisition unit 118A provided in the document processing apparatus 100A as a document evaluation apparatus acquires the evaluation result 450 for the evaluation acquisition prompt 620 and the report 460 for the report acquisition prompt 630 by using the language model 132 instead of the language model server 210.

[0121] The response acquisition unit 118A provided in the document processing apparatus 100 as a document evaluation criterion extraction apparatus acquires the check content for the check content acquisition prompt (see the evaluation criterion acquisition prompt 640B) by using the language model 132 instead of the language model server 210.

[0122] The response acquisition unit 118A provided in the document processing apparatus 100 as a document correction apparatus acquires the synonym group (synonym dictionary 550) for the synonym acquisition prompt 650 and the corrected document for the corrected document acquisition prompt 660 by using the language model 132 instead of the language model server 210.

[0123] ≪Other Variants≫ As described above, several embodiments and variants of the present invention have been explained. However, these embodiments are merely illustrative and do not limit the technical scope of the present invention. The present invention can take various other embodiments, and furthermore, various changes such as omissions and substitutions can be made without departing from the gist of the present invention. These embodiments and their variants are included in the scope and gist of the invention described in this specification and the like, and are also included in the invention described in the claims and its equivalent scope.

[0124] ≪Hardware Configuration≫ The document processing apparatuses 100 and 100A according to the above-described embodiments are realized by a computer 900 configured as shown in FIG. 42, for example. FIG. 42 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the document processing apparatuses 100 and 100A according to the above-described embodiments. The computer 900 includes a CPU 901, a ROM 902, a RAM 903, an SSD 904, and an input / output interface 905 (described as input / output I / F (Interface) in FIG. 42). Further, the computer 900 includes a communication interface 906 (described as communication I / F in FIG. 42) and a media interface 907 (described as media I / F in FIG. 42). The computer 900 may include an HDD (Hard Disc Drive) instead of the SSD 904, or may include an HDD in addition to the SSD 904.

[0125] The CPU 901 operates based on a program stored in the ROM 902 or the SSD 904, and performs control by the control unit 110 in FIG. 1. The ROM 902 stores a boot program executed by the CPU 901 when the computer 900 is started up, a program related to the hardware of the computer 900, and the like.

[0126] The CPU 901 controls an input device 910 such as a mouse or a keyboard, and an output device 911 such as a display or a printer via the input / output interface 905. The CPU 901 acquires data from the input device 910 via the input / output interface 905, and outputs the generated data to the output device 911.

[0127] The SSD 904 stores a program executed by the CPU 901 and data used by the program. The communication interface 906 receives data from another device (not shown), such as a language model server 210, via a communication network and outputs it to the CPU 901, and transmits data generated by the CPU 901 to another device via the communication network.

[0128] The media interface 907 reads a program or data stored in the recording medium 912 and outputs it to the CPU 901 via the RAM 903. The CPU 901 loads the program from the recording medium 912 onto the RAM 903 via the media interface 907 and executes the loaded program. The recording medium 912 is an optical recording medium such as a DVD (Digital Versatile Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductor memory tape medium, or a semiconductor memory, etc.

[0129] For example, when the computer 900 functions as the document processing apparatuses 100, 100A according to the above-described embodiment, the CPU 901 of the computer 900 realizes the functions of the document processing apparatuses 100, 100A by executing the program 138 (see FIG. 1) loaded onto the RAM 903. The CPU 901 reads and executes the program from the recording medium 912. In addition, the CPU 901 may read the program from another device via a communication network, or may install the program 138 from the recording medium 912 to the SSD 904 and execute it. Note that the document processing apparatuses 100, 100A are not limited to the hardware computer 900, and may be in the form of a virtual machine.

Explanation of Signs

[0130] 100 Document processing apparatus (document analysis apparatus, document evaluation apparatus, document evaluation criterion extraction apparatus, document correction apparatus) 111 Document analysis unit 112 Document format generation unit 113 Document evaluation unit 114 Evaluation report unit 115 Evaluation criterion extraction unit 116 Dictionary generation unit 117 Document correction unit 118, 118A Answer acquisition unit 131 Document database 132 Language model 140 Document format database 150 Document evaluation report rule database 160 Evaluation Report Database 210 Language Model Server 410 Tabular Document 420, 420A Formal Information Document 430, 430A Item Value Acquisition Result 450, 450A Evaluation Result 460 Report 510 Document 520, 520A Evaluation Criteria 520B Evaluation Criteria (Check Contents) 540 Document 550, 550A Thesaurus (Synonym Group) 560 Revised Document 610 Item Value Acquisition Prompt 612 Element (Formal Information) 613 Element (Document Information) 620, 620A, 620B, 620C Evaluation Acquisition Prompt 622 Element (Evaluation Criteria) 623 Element (Document Information to be Evaluated) 624 Element (Question) 630, 630A, 630B, 630C Report Acquisition Prompt 632~635 Elements (Report Example) 640, 640A Evaluation Criteria Acquisition Prompt 640B Evaluation Criteria Acquisition Prompt (Check Content Acquisition Prompt) 650, 650A Synonym Acquisition Prompt 660, 660A Revised Document Acquisition Prompt

Claims

1. Extract document information indicating the correspondence between the identification information of each cell and the item value of the cell from a tabular document composed of cells in which item values, which are values of items, are entered, Create an item value acquisition prompt by merging format information indicating the correspondence between the identification information of cells in which the item values to be acquired are entered among the cells and the item identification information of the item, and text, the format information, and the document information that request an item value acquisition result indicating the correspondence between the item identification information and the item value entered in the cell corresponding to the item identification information, A document analysis unit that transmits the item value acquisition prompt to a language model server and acquires the item value acquisition result directly or indirectly from the language model server Document analysis device.

2. The format information is Including the correspondence between the identification information of the cell and the item name, The identification information of the cell is Information indicating the position of the cell in the tabular document, The document analysis unit is Among the format information in the document format database storing one or more pieces of the format information, for a format information cell that is a cell indicated by the identification information of the cell included in the format information and a document information cell that is a cell indicated by the identification information of the cell included in the document information, create the item value acquisition prompt with the format information having the largest number of the format information cells and the document information cells in which the identification information of the format information cell and the identification information of the document information cell match and the item name of the format information cell and the item value of the document information cell match as the format information to be merged The document analysis device according to claim 1.

3. Extract template document information indicating the correspondence between the identification information indicating the position of each cell and the value entered in the cell from a template document in which item values are not entered and which serves as a template for the tabular document, When the cell above the blank cell, which is a blank cell, is not blank and the cell to the left is blank, associate the identification information of the blank cell with the value of the cell above and the item identification information, When the cell to the left of the blank cell is not blank and the cell above is blank, associate the identification information of the blank cell with the value of the cell to the left and the item identification information, Further include a document format generation unit that uses the correspondence between the identification information of the blank cell associated and the item identification information as the format information The document analysis device according to claim 1.

4. The document analysis unit is Instead of the language model server, using a language model, obtain the item value acquisition result for the item value acquisition prompt The document analysis device according to claim 1.

5. A document analysis device extracting document information indicating the correspondence between the identification information of each cell and the item value of the cell from a tabular document including cells in which item values, which are values of items, are entered; merging format information indicating the correspondence between the identification information of the cell in which the item value to be acquired is entered among the cells and the item identification information of the item, and text requesting an item value acquisition result indicating the correspondence between the item identification information and the item value entered in the cell corresponding to the item identification information, the format information, and the document information to create an item value acquisition prompt; sending the item value acquisition prompt to a language model server and obtaining the item value acquisition result directly or indirectly from the language model server Document analysis method.

Citation Information

Patent Citations

  • Systems, methods, and software for processing, presenting, and recommending citations.

    JP2015527641A