Document proofreading device and document proofreading method

The document proofreading device and method address inaccuracies in document correction by generating business-specific synonym groups through a language model server, enhancing precision and efficiency.

JP2026073853AActive Publication Date: 2026-05-01HITACHI INDUSTRY & CONTROL SOLUTIONS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HITACHI INDUSTRY & CONTROL SOLUTIONS LTD
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing document correction methods using general dictionaries often incorrectly treat words with different meanings as synonyms, leading to inaccuracies, and customizing dictionaries for each business or document is costly and labor-intensive.

Method used

A document proofreading device and method that utilizes a dictionary generation unit to create synonym groups based on document-specific prompts, leveraging a language model server for tailored corrections, including a document proofreading unit to refine word notation variations.

Benefits of technology

Enables document correction tailored to individual tasks and documents, improving accuracy and efficiency by directly utilizing a language model server for synonym acquisition and proofreading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073853000001_ABST
    Figure 2026073853000001_ABST
Patent Text Reader

Abstract

It enables document proofreading tailored to individual tasks and documents. [Solution] The document proofreading device (document processing device 100) includes a dictionary generation unit 116 that creates a synonym acquisition prompt requesting a synonym group, which is a collection of documents containing one or more documents, with words grouped by synonym; an answer acquisition unit 118 that sends the synonym acquisition prompt to a language model server 210 and obtains the synonym group directly or indirectly from the language model server 210; and a document proofreading unit 117 that refers to the synonym group and creates a proofreading document acquisition prompt requesting a proofreading document that corrects variations in word notation in the documents included in the document collection. The answer acquisition unit 118 sends the proofreading document acquisition prompt to the language model server 210 and obtains the proofreading document directly or indirectly from the language model server 210.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a document correction device and a document correction method for correcting documents.

Background Art

[0002] Natural language processing using large language models is becoming widespread, and attention is focused on automating and improving the efficiency of document correction that has conventionally been performed by humans using this technology. As content of document correction, there is sometimes an elimination and unification of synonymous expressions, tenses, and fluctuations at the end of sentences.

[0003] As an example of processing considering fluctuations in synonymous expressions, there is a log management device described in Patent Document 1. The log management device searches for logs using a word group database and a thesaurus dictionary that store synonyms.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Words included in documents used in business have meanings and contents specific to that business. Therefore, when using a general dictionary, there is a possibility of correcting (modifying) words used in different meanings as synonyms. Even if a dictionary is prepared for each business, the cost and labor are high. Such circumstances may occur not only in business but also in documents.

[0006] The present invention has been made in view of such a background, and an object thereof is to provide a document correction device and a document correction method that enable document correction according to individual businesses and documents.

Means for Solving the Problems

[0007] To solve the above-mentioned problems, the document proofreading apparatus according to the present invention comprises: a dictionary generation unit that creates a synonym acquisition prompt requesting a synonym group in which words contained in the text information of a group of documents including one or more documents are grouped by synonym; a response acquisition unit that transmits the synonym acquisition prompt to a language model server and obtains the synonym group directly or indirectly from the language model server; and a document proofreading unit that refers to the synonym group and creates a proofreading document acquisition prompt requesting a proofreading document in which variations in word notation in the documents included in the group of documents have been corrected, wherein the response acquisition unit transmits the proofreading document acquisition prompt to the language model server and obtains the proofreading document directly or indirectly from the language model server. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a document proofreading device and a document proofreading method that enable document proofreading tailored to individual tasks and documents. Other problems, configurations, and effects will be clarified by the following description of embodiments. [Brief explanation of the drawing]

[0009] [Figure 1] This is a functional block diagram of the document processing device according to this embodiment. [Figure 2] This figure shows an example of a tabular document according to this embodiment. [Figure 3] This figure shows a tabular document containing formal information. [Figure 4] This figure shows the configuration of the item value acquisition prompt according to this embodiment. [Figure 5] This figure shows the item value acquisition results according to this embodiment. [Figure 6] This is a flowchart of the document analysis process according to this embodiment. [Figure 7] This is a flowchart of the document format acquisition process according to this embodiment. [Figure 8] This figure shows the results of obtaining item values ​​in text format according to a modified version of this embodiment. [Figure 9] It is a diagram showing a format information document in which item identification information according to a modification example of this embodiment is blank. [Figure 10] It is a flowchart of the format information acquisition process according to this embodiment. [Figure 11] It is a diagram for explaining an overview of document evaluation report processing by the document processing apparatus according to this embodiment. [Figure 12] It is a diagram showing the configuration of an evaluation acquisition prompt according to this embodiment. [Figure 13] It is a diagram showing the evaluation result according to this embodiment. [Figure 14] It is a diagram showing the evaluation result according to this embodiment. [Figure 15] It is a diagram showing the configuration of a report acquisition prompt according to this embodiment. [Figure 16] It is a diagram showing the configuration of a report acquisition prompt according to this embodiment. [Figure 17] It is a diagram showing the report according to this embodiment. [Figure 18] It is a diagram showing the evaluation acquisition prompt according to this embodiment. [Figure 19] It is a diagram showing the report acquisition prompt according to this embodiment. [Figure 20] It is a diagram showing the evaluation acquisition prompt according to this embodiment. [Figure 21] It is a diagram showing the report acquisition prompt according to this embodiment. [Figure 22] It is a diagram showing the evaluation acquisition prompt according to this embodiment. [Figure 23] It is a diagram showing the report acquisition prompt according to this embodiment. [Figure 24] It is a flowchart of the document evaluation report processing according to this embodiment. [Figure 25] It is an example of a document including evaluation criteria according to this embodiment. [Figure 26] It is a diagram showing the configuration of an evaluation criteria acquisition prompt according to this embodiment. [Figure 27] It is a diagram showing the acquired evaluation criteria according to this embodiment. [Figure 28] This figure shows the evaluation criteria acquisition prompt according to this embodiment. [Figure 29] These are the evaluation criteria used when the evaluation criteria acquisition prompt according to this embodiment is used. [Figure 30] This figure shows the evaluation criteria acquisition prompt (check content acquisition prompt) according to this embodiment. [Figure 31] These are the evaluation criteria (check contents) when using the evaluation criteria acquisition prompt according to this embodiment. [Figure 32] This is a flowchart of the evaluation criteria extraction process according to this embodiment. [Figure 33] This diagram illustrates the outline of the document proofreading process performed by the document processing device according to this embodiment. [Figure 34] This figure shows the configuration of the synonym acquisition prompt according to this embodiment. [Figure 35] This figure shows the synonym dictionary obtained according to this embodiment. [Figure 36] This figure shows the configuration of the calibration document acquisition prompt according to this embodiment. [Figure 37] This figure shows the synonym acquisition prompt according to this embodiment. [Figure 38] This figure shows a synonym dictionary containing meaning, according to this embodiment. [Figure 39] This figure shows the calibration document acquisition prompt according to this embodiment. [Figure 40] This is a flowchart of the document proofreading process according to this embodiment. [Figure 41] This is a functional block diagram of a document processing device according to a modified example of this embodiment. [Figure 42] This is a hardware configuration diagram showing an example of a computer that implements the functions of the document processing device according to the above embodiment. [Modes for carrying out the invention]

[0010] ≪Overview of Document Processing Devices≫ The following describes the outline of a document processing device, which functions as a document analysis device, document evaluation device, document evaluation criterion extraction device, and document proofreading device in an embodiment for carrying out the present invention. Firstly, the document processing device functions as a document analysis device that extracts information from tabular documents such as spreadsheets, taking into account the document structure. More specifically, the document processing device refers to the format information of the document, extracts the item values ​​corresponding to the item identification information (item names), and outputs them as text.

[0011] Secondly, the document processing device functions as a document evaluation device that evaluates / reviews documents. More specifically, the document processing device evaluates documents by referring to text extracted from the documents and the document evaluation criteria, and outputs a report (evaluation report). Thirdly, the document processing device functions as a document evaluation criteria extraction device, which extracts document evaluation criteria (check items) from document checklists and review meeting minutes. Fourthly, the document processing device functions as a document proofreading device that corrects inconsistencies in the document's notation. More specifically, the document processing device detects synonyms in the document and corrects inconsistencies in notation. The following describes the basic configuration of the document processing device, followed by a description of the functional configurations of the document analysis device, document evaluation device, document evaluation criterion extraction device, and document proofreading device in that order.

[0012] ≪Configuration of Document Processing System≫ Figure 1 is a functional block diagram of the document processing device 100 according to this embodiment. The document processing device 100 is a computer and comprises a control unit 110, a storage unit 130, and an input / output unit 180. User interface devices such as a display, keyboard, and mouse are connected to the input / output unit 180. The input / output unit 180 is equipped with a communication device and is capable of sending and receiving data with the language model server 210 and other devices. The language model server 210 is a server for a service (interactive artificial intelligence service) that receives prompts including commands / instructions and returns answers to those prompts.

[0013] ≪Document Processing Device: Memory Unit≫ The memory unit 130 is comprised of memory devices such as ROM (Read Only Memory), RAM (Random Access Memory), and SSD (Solid State Drive). The memory unit 130 stores a document database 131, a document format database 140, a document evaluation report rule database 150, an evaluation report database 160, and a program 138. The various contents of the memory unit 130 may be stored in an external storage device such as a cloud server and read as needed.

[0014] The document database 131 stores documents to be processed by the document processing device 100, associated with document identification information. The program 138 includes a description of the processing to be performed by the functional unit provided in the control unit 110, which will be described later. The other components of the storage unit 130 will be described later.

[0015] Document Processing Unit: Control Section The control unit 110 includes a CPU (Central Processing Unit) and comprises a document analysis unit 111, a document format generation unit 112, a document evaluation unit 113, an evaluation reporting unit 114, an evaluation criteria extraction unit 115, a dictionary generation unit 116, a document proofreading unit 117, and an answer acquisition unit 118. The answer acquisition unit 118 sends prompts generated by other functional units of the control unit 110 (described later) to the language model server 210 and receives answers directly or indirectly. The control unit 110 may also include a GPU (Graphics Processing Unit), an NPU (Neural (network) Processing Unit), an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), etc.

[0016] Document Processing Devices: Document Analysis Devices The document processing device 100, which functions as a document analysis device, is described below. The document processing device 100 refers to the format information (template information) of a tabular document and extracts pairs of item identification information (item name) and item value for each item contained in the tabular document.

[0017] ≪Document analysis device: Tabular document / format information≫ Figure 2 shows an example of a tabular document 410 according to this embodiment. The tabular document 410 is a spreadsheet-format document in which the item value corresponding to the item name is entered in the cell (column) next to the cell in which the item name is entered. In Figure 2, the item name is entered in the cell to the right of the cell in which the item name is entered. For example, the item value corresponding to the item name "Requested Delivery Date" is "2020 / 10 / 10". Note that the item value may also be entered in the cell below the item name. Note that while an Excel® document is an example of a tabular document / spreadsheet-format document, it is not limited to this.

[0018] In a spreadsheet document, cells are identified by identifiers that show their location. For example, the identifier for the cell containing the item name "Required Delivery Date" is "B6:C7". The tabular document 410 is stored in the document database 131 in association with document identification information. The document database 131 may also store text-format documents.

[0019] Figure 3 shows a format information document 420 (template document) of a tabular document according to this embodiment. The cells for item names are the same as in the tabular document 410. The cells where item values ​​are entered contain a string with "#" added to the beginning of the item name (item identification information). In the following, unless there is a particular need to distinguish them, the item identification information and the item identification information with "#" added to the beginning will not be distinguished. For example, "requested delivery date" and "#requested delivery date" will not be distinguished as item identification information.

[0020] Document Analysis System: Document Format Database The document format database 140 stores the format information document 420, associated with its format information identification information. Furthermore, the cell identification information (location information) and values ​​extracted from the format information document 420 by the document analysis unit 111 (described later) are also stored in the document format database 140, associated with the format information document 420. In other words, the document format database 140 stores the format information identification, the format information document 420 (see Figure 3), and the correspondence between the cell identification information and values ​​of the format information document 420 (see element 612 in Figure 4, described later). Hereafter, the correspondence between the cell identification information and values ​​of the format information document 420 will also be referred to as format information.

[0021] ≪Document Analysis System: Document Analysis Department≫ Before starting to process the tabular document 410, the document analysis unit 111 extracts format information (cell identification information and value correspondence) from the format information document 420 and stores it in the document format database 140. The document analysis unit 111 then refers to the format information and extracts item values ​​corresponding to item names (item identification information) from the tabular document 410. More specifically, the document analysis unit 111 first extracts document information (see element 613 in Figure 4 below), which is the correspondence between cell identification information (location information) and values, from the tabular document 410.

[0022] Next, the document analysis unit 111 generates an item value acquisition prompt 610 (see Figure 4 below) based on the extracted information. Subsequently, the document analysis unit 111 instructs the response acquisition unit 118 to send the item value acquisition prompt 610 to the language model server 210 and receives the item value acquisition result 430 (see Figure 5 below), which is the response.

[0023] ≪Document Analysis Device: Item Value Acquisition Prompt≫ Figure 4 shows the configuration of the item value acquisition prompt 610 according to this embodiment. Generally, a prompt consists of one or more elements. The item value acquisition prompt 610 includes elements 611 to 613. Element 611 refers to the format information shown in Element 612 and indicates an instruction / command to extract item values ​​corresponding to item identifiers starting with "#" from the document shown in Element 613.

[0024] Element 612 is the correspondence between cell identification information and values ​​(formal information) extracted from the formal information document 420 (see Figure 3). Each row shows the correspondence between identification information and values. Element 613 is the correspondence between cell identifiers and values ​​(document information) extracted from tabular document 410 (see Figure 2). Each row shows the correspondence between identifier and value.

[0025] ≪Document Analysis Device: Item Value Acquisition Results≫ Figure 5 shows the item value acquisition result 430 according to this embodiment. The item value acquisition result 430 is JSON format data specified in element 611 of the item value acquisition prompt 610, and shows the correspondence between item identification information (item name) and item value. JSON is an abbreviation for JavaScript Object Notation, and is the format of objects in the programming language JavaScript.

[0026] Let's explain the correspondence using the item identification information "#Sales Order Probability" as an example. The identification information for the cell where the document format information "#Sales Order Probability" is written is "E2:F3" (see element 612 in Figures 3 and 4). The language model server 210 obtains "A" as the corresponding value for "E2:F3" in element 613 and includes it in the item value acquisition result 430. The document analysis unit 111 associates the item value acquisition result 430 with the tabular document 410 and stores it in the document database 131.

[0027] As described above, the document processing device 100, as a document analysis device, includes a document analysis unit 111 that extracts document information (see element 613) from a tabular document 410 which is composed of cells in which item values ​​are entered, showing the correspondence between the identification information of each cell and the item value of that cell.

[0028] The document analysis unit 111 refers to the identification information of the cell in which the item value to be acquired is entered within the cell, format information (see element 612) that shows the correspondence between the item identification information of the item, and the item identification information of the item, and the document information, and creates an item value acquisition prompt 610 (see Figure 4) that requests an item value acquisition result 430 (see Figure 5) that shows the correspondence between the item identification information and the item value entered in the cell corresponding to the item identification information. The document processing device 100, which functions as a document analysis device, includes a response acquisition unit 118 that sends an item value acquisition prompt 610 to the language model server 210 and acquires an item value acquisition result 430 directly or indirectly from the language model server 210.

[0029] Document Analysis Equipment: Document Analysis Processing Figure 6 is a flowchart of the document analysis process according to this embodiment. In step S11, the document analysis unit 111 extracts cell identification information and values ​​from the tabular document 410, which is the document to be processed (see element 613 shown in Figure 4).

[0030] Step S12 is a document format acquisition process that retrieves format information corresponding to the tabular document 410 of the document to be processed from the document format database 140. Details will be described later with reference to Figure 7. In step S13, the document analysis unit 111 generates an item value acquisition prompt 610 (see Figure 4). Here, element 612 is the format information acquired in step S12. Element 613 is the correspondence between the cell identification information and the value acquired in step S11.

[0031] In step S14, the document analysis unit 111 instructs the response acquisition unit 118 to send an item value acquisition prompt 610 to the language model server 210. Next, the document analysis unit 111 associates the item value acquisition result 430 (see Figure 5), which is the response, with the tabular document 410 and stores it in the document database 131.

[0032] Document Analysis System: Document Format Acquisition Processing Figure 7 is a flowchart of the document format acquisition process according to this embodiment. The document format acquisition process in step S12 (see Figure 6) will be explained with reference to Figure 7. In step S21, the document analysis unit 111 starts repeating the process in step S22 for each piece of format information stored in the document format database 140. The format information refers to the correspondence between the cell identification information and values ​​of the format information document 420 (see element 612 in Figure 4).

[0033] In step S22, the document analysis unit 111 calculates the degree of agreement between the formal information and the tabular document 410, which is the target of the document analysis process (see Figure 6). The degree of agreement is the percentage or number of cells in the formal information whose values ​​and identification information do not begin with "#" that match the values ​​and identification information of cells contained in the tabular document 410. In other words, it is the percentage or number of cells in the formal information document 420 that contain item names that do not begin with "#" that match the cells in the tabular document 410 that contain the same item names. Note that items in the cells are, for example, "B2:D3" for "(Sales) Order Probability" and "F6:F7" for "Completion of Work". In step S23, the document analysis unit 111 uses the format information that yields the highest degree of match calculated in step S22 as the format information for the tabular document 410 of the document to be processed.

[0034] As explained above, the cell identification information is information that indicates the location of the cell in the tabular document 410. The document analysis unit 111 creates an item value acquisition prompt 610 by referring to the format information in the document format database 140, which stores one or more format information items, that has the maximum number of cells in which the cell identification information and cell values ​​match between the cells in the format information and the cells in the document information (see element 613).

[0035] Features of the document analysis device The document processing device 100, acting as a document analysis device, extracts item value acquisition results 430 (see Figure 5) from a tabular document 410 (see Figure 2) by referring to format information (see element 612 shown in Figure 4). Even with complex tabular documents, it is possible to obtain item value acquisition results, which are the correspondence between item identification information and item values ​​necessary for document processing such as review / evaluation. According to the document format acquisition process (see Figure 7), the format information that best matches the format of the document being processed is obtained. This eliminates the need to specify a format (template document) for each document, allowing for efficient acquisition of item values.

[0036] <<Twisted example of a document analysis device: Item value acquisition results>> The item value acquisition result 430 (see Figure 5) was in JSON format, but is not limited to this. It may also be in text format. Figure 8 shows an item value acquisition result 430A in text format according to a modified example of this embodiment. Item value acquisition result 430A is text format data with the same content as item value acquisition result 430.

[0037] ≪Document analysis device: Document format generation unit≫ In the embodiment described above, format information (see element 612 shown in Figure 4) is extracted from a format information document 420 (see Figure 3) which contains item identification information starting with "#". The document format generation unit 112 extracts format information from a format information document 420A (see Figure 9 below) in which the item identification information is blank. Figure 9 shows a format information document 420A (template document) in which the item identification information is blank, according to a modified example of this embodiment. Compared to the format information document 420, the item identification information is blank.

[0038] Figure 10 is a flowchart of the format information acquisition process according to this embodiment. Referring to Figure 10, the process of acquiring format information from a format information document 420A in which the item value input fields (cells) are blank will be explained. The document format acquisition process makes it possible to acquire format information without having to prepare a format information document 420 in which item identification information starting with "#" has been entered.

[0039] In step S31, the document format generation unit 112 extracts cell identification information and value correspondence (template document information) from the format information document 420A. In step S32, the document format generation unit 112 obtains a cell that is blank, has a non-blank cell above it, and a blank cell to its left. The cells above, below, to the left, and to the right of a cell can be identified from the identification information indicating its position. It is also assumed that the cell above it has the same width and the cell to its left has the same height.

[0040] In step S33, the document format generation unit 112 starts repeating the process in step S34 for each cell acquired in step S32. In step S34, the document format generation unit 112 uses the item identification information, which is the value (string) of the cell above it with "#" added to the beginning, and the cell identification information as format information.

[0041] In step S35, the document format generation unit 112 obtains a cell that is blank, has a non-blank cell to its left, and a blank cell above it. It is assumed that the cell above it has the same width and the cell to its left has the same height. In step S36, the document format generation unit 112 starts repeating the process in step S37 for each cell acquired in step S35.

[0042] In step S37, the document format generation unit 112 uses the item identification information obtained by adding "#" to the beginning of the value (string) of the cell to its left, and the cell identification information, as format information. In step S38, the document format generation unit 112 stores the format information acquired in steps S34 and S37 in the document format database 140 as format information for the format information document 420A.

[0043] As described above, the document processing device 100, as a document analysis device, includes a document format generation unit 112 that extracts template document information from a template document (see format information document 420A shown in Figure 9), which is a template for a tabular document 410 and has no item values ​​entered, showing the correspondence between identification information indicating the position of each cell and the value entered in that cell.

[0044] The document format generation unit 112 associates the identification information of a blank cell with the value of the cell to its left if the cell to its left is not blank, and the cell to its left is blank. Furthermore, the document format generation unit 112 associates the identification information of a blank cell with the value of the cell to its left if the cell to its left is not blank, and the cell to its left is blank. The document format generation unit 112 uses the correspondence between the identification information of the associated blank cells and the item identification information as format information.

[0045] ≪Document Processing Device: Document Evaluation Device≫ The document processing device 100, which functions as a document evaluation device, is described below. The document processing device 100 evaluates / reviews documents containing items, such as tabular documents 410, according to evaluation criteria for each item, and reports the results. Documents containing structures such as chapters, sections, and items, even if in text format, are also subject to processing by the document evaluation device. In the following description, elements / components of document structure such as chapters and sections will also be referred to as items.

[0046] Document Evaluation System: Processing Overview Figure 11 is a diagram illustrating the overview of the document evaluation report processing by the document processing device 100 according to this embodiment (see Figure 24 below). The document evaluation unit 113 generates an evaluation acquisition prompt 620 (see Figure 12 below) based on the evaluation criteria stored in the document evaluation report rule database 150 and the item value acquisition results 430 (see Figure 5). The evaluation acquisition prompt 620 is sent to the language model server 210, and an evaluation result 450 (see Figure 13 below) is returned. The evaluation result 450 is the result of the evaluation / review of the item value acquisition results 430 according to the evaluation criteria.

[0047] The evaluation reporting unit 114 generates a report acquisition prompt 630 (see Figure 15 below) based on the evaluation results 450 and the report examples stored in the document evaluation report rule database 150. The report acquisition prompt 630 is sent to the language model server 210, and a report 460 (see Figure 17 below) is returned. The details of the document processing device 100 as a document evaluation device are described below.

[0048] ≪Document Evaluation System: Document Evaluation Reporting Rule Database≫ The document evaluation report rule database 150 stores evaluation criteria for evaluating / reviewing documents, associated with formal information (document format database 140, see element 612 in Figure 4). Formal information includes not only items in tabular documents (item identifiers and item values), but also constituent elements such as chapters, sections, and figures / tables included in text documents. The evaluation criteria (see element 622 in Figure 12 below) are the evaluation criteria related to each of these items. Since formal information contains one or more items, one or more evaluation criteria are associated with each piece of formal information.

[0049] The document evaluation report rule database 150 also stores questions corresponding to the evaluation criteria (see element 624 in Figure 12, described later). The questions are designed to obtain evaluation results for each item or evaluation criterion. Since the formal information contains one or more items, one or more questions are associated with each piece of formal information.

[0050] Furthermore, the document evaluation report rule database 150 also stores example reports of evaluation results evaluated / reviewed based on evaluation criteria (see elements 632-635 in Figure 15 below). The example reports are examples of evaluation results related to each item. Since the formal information contains one or more items, one or more example reports are associated with each piece of formal information. In addition, the document evaluation report rule database 150 may also store documents (also referred to as checklists) that describe the perspectives, check contents, and check rules for evaluating / reviewing documents.

[0051] Document Evaluation System: Evaluation Report Database The evaluation report database 160 stores evaluation results and reports summarizing those evaluation results, associated with the document identification information of the document being evaluated. Furthermore, the evaluation report database 160 also stores minutes of meetings in which the document was evaluated / reviewed while referring to the document and reports, associated with the document identification information of the document being evaluated.

[0052] Document Evaluation System: Document Evaluation Department The document evaluation unit 113 evaluates / reviews the document by referring to evaluation criteria related to the items that make up the document. Based on the evaluation criteria and document information, the document evaluation unit 113 generates an evaluation acquisition prompt 620 (see Figure 12 below). Subsequently, the document evaluation unit 113 instructs the response acquisition unit 118 to send the evaluation acquisition prompt 620 to the language model server 210 and receives the evaluation result (also simply referred to as evaluation), which is the response.

[0053] Document Evaluation Device: Evaluation Acquisition Prompt Figure 12 shows the configuration of the evaluation acquisition prompt 620 according to this embodiment. The evaluation acquisition prompt 620 includes elements 621 to 624. Element 621 provides instructions / directions for answering the questions shown in Element 624 regarding the document being evaluated, as shown in Element 623.

[0054] Element 622 is the evaluation criteria for items included in the document being evaluated (see Element 623). The evaluation criteria are stored in the document evaluation report rule database 150. Element 623 is the document information to be evaluated (evaluation target document information). In Figure 12, the document information in element 623 is the correspondence between item identification information and item values ​​obtained by the document analysis unit 111 from a tabular document (see item value acquisition result 430 shown in Figure 5). The document information of element 623 may also be text-format document information created with a word processor. For example, it may be the item value acquisition result 430A in text format (see Figure 8).

[0055] Element 624 is a question for obtaining an evaluation result according to the evaluation criteria in Element 622. The question may include a procedure for obtaining the evaluation result. The question shown in Figure 12 includes the following procedure: (1) correctly count the number of "High," "Medium," and "Low" values ​​in "Evaluation Result," (2) determine the project risk rank according to the evaluation criteria, and (3) compare the determined project risk rank with the project risk rank included in the evaluation subject. Thus, the evaluation acquisition prompt 620 is a prompt for obtaining an evaluation result by referring to the evaluation criteria related to the "Project Risk Rank" item.

[0056] Document Evaluation Device: Evaluation Results Figure 13 shows the evaluation result 450 according to this embodiment. Figure 14 shows the evaluation result 450A according to this embodiment. Evaluation results 450 and 450A are answers that follow the procedure shown in the questions of element 624 included in the evaluation acquisition prompt 620. Note that evaluation results 450 and 450A differ in the final step, the comparison of project risk ranks, if they do not match and if they do. In this manner, the document evaluation unit 113 queries the language model server 210 using the evaluation acquisition prompt 620 for each evaluation criterion for each item included in the formal information corresponding to the document that is the subject of evaluation, and obtains evaluation results 450 and 450A.

[0057] As described above, the document processing device 100, as a document evaluation device, includes a document evaluation unit 113 that creates an evaluation acquisition prompt 620 requesting evaluation results 450, 450A of the document information to be evaluated based on the evaluation criteria, by referring to the document information to be evaluated (see element 623), which includes the correspondence between item identification information and item values ​​for one or more items, and the evaluation criteria related to the items (see element 622). The document processing device 100, as a document evaluation device, includes a response acquisition unit 118 that sends an evaluation acquisition prompt 620 to the language model server 210 and acquires evaluation results 450, 450A directly or indirectly from the language model server 210. The request for evaluation results included in the evaluation acquisition prompt 620 is a question-based request (see element 624) that seeks evaluation results 450,450A.

[0058] Document Evaluation System: Evaluation Reporting Department The evaluation reporting unit 114 generates an evaluation report by referring to the evaluation results 450 and 450A. More specifically, the evaluation reporting unit 114 generates a report acquisition prompt 630 (see Figure 15 below) based on the evaluation results and example reports. Subsequently, the evaluation reporting unit 114 instructs the response acquisition unit 118 to send the report acquisition prompt 630 to the language model server 210 and receives the report, which is the response. Similar to the document evaluation unit 113, the evaluation reporting unit 114 generates a report acquisition prompt 630 for each item that makes up the document and receives the report 460 (see Figure 17 below).

[0059] ≪Document Evaluation Device: Report Acquisition Prompt≫ Figures 15 and 16 show the configuration of the report acquisition prompt 630 according to this embodiment. The report acquisition prompt 630 includes elements 631 to 637. Element 631 provides instructions / directions for reporting the question shown in Element 636 and the answer to that question (evaluation result) shown in Element 637, referring to the reporting examples shown in Elements 632-635.

[0060] Element 632 is an example of a question and answer where the judgment result does not match. Element 633 is an example of a report when the judgment results do not match those shown in Element 632. The report includes whether or not there are any issues, the nature of the issues, and proposed corrections. Element 634 is an example of a question and answer where the judgment result matches. Element 635 is an example of a report when the judgment result matches that shown in Element 634. The report includes whether or not there are any issues, the content of the issues, and proposed corrections, but the content is equivalent to nothing.

[0061] Moving on to Figure 16, we will continue the explanation of the report acquisition prompt 630. Element 636 is a question to obtain the evaluation result, which is located in element 624 of the evaluation acquisition prompt 620 (see Figure 12). Element 637 is the evaluation result 450 (see Figure 13), which is the answer to the question.

[0062] Document Evaluation Device: Report Figure 17 shows a report 460 according to this embodiment. The answer, which is element 637, is an evaluation result 450 in which the judgment result does not match, and corresponds to the question and answer in element 632. Therefore, report 460 is a report that conforms to the report in element 633. In this manner, the evaluation reporting unit 114 queries the language model server 210 using a report acquisition prompt 630 that includes a report example corresponding to the item-specific evaluation results 450, 450A acquired by the document evaluation unit 113, and obtains a report 460.

[0063] As described above, the document processing device 100, as a document evaluation device, includes an evaluation reporting unit 114 that generates a report acquisition prompt 630 that requests a report 460 in a predetermined format based on the evaluation results 450, 450A. The response acquisition unit 118 sends a report acquisition prompt 630 to the language model server 210 and acquires the report 460 directly or indirectly from the language model server 210.

[0064] The report acquisition prompt 630 includes an example of a question included in a question-format request (see elements 632, 634), an example of an evaluation result as an answer to the question (see elements 632, 634), an example of a report in a predetermined format for the example question and the example answer (see elements 633, 635), a question-format request included in the evaluation acquisition prompt 620 (see element 636), and an evaluation result acquired by the answer acquisition unit 118 (see element 637).

[0065] Furthermore, the document information to be evaluated (see element 623 in Figure 12) includes evaluation items (e.g., "project risk rank") whose item values ​​are determined according to the number of items that fall within a predetermined range (e.g., "high", "medium", "low") where the item value is one or more predetermined values. The evaluation criteria (see Element 622) include the criteria for determining the values ​​of the evaluation items. The evaluation acquisition prompt 620 includes a request to count the number of items whose item values ​​fall within a predetermined range, a request to obtain a count evaluation item value, which is the value of an evaluation item determined according to the count result, and a request to report whether the count evaluation item value matches the described evaluation item value, which is the value of an evaluation item included in the document information to be evaluated (see element 624).

[0066] The prescribed format of the report (see element 633 in Figure 15) is a format for reporting a discrepancy when the count evaluation item value and the recorded evaluation item value do not match. The prescribed report format (see element 633) includes a proposed amendment to correct the described evaluation item value to the count evaluation item value if the count evaluation item value and the described evaluation item value do not match.

[0067] <<Variations of the evaluation prompt and report prompt>> The evaluation criteria (see element 622) included in the evaluation acquisition prompt 620 (see Figure 12) are evaluation criteria in which other item values ​​("AA", "A", "B", "C") are determined according to the number of item values ​​("High", "Medium"). Several examples of evaluation criteria / check contents other than this format are shown below.

[0068] Figure 18 shows the evaluation acquisition prompt 620A according to this embodiment. The evaluation criteria in element 622A of the evaluation acquisition prompt 620A are evaluation criteria in which other item values ​​("final project risk rank") are determined based on two item values ​​("project risk rank by scale" and "project risk rank by required quality"). In other words, the evaluation acquisition prompt 620A is a prompt for acquiring the evaluation result related to the "final project risk rank" item.

[0069] Figure 19 shows the report acquisition prompt 630A according to this embodiment. The report acquisition prompt 630A corresponds to the evaluation acquisition prompt 620A and is a prompt for acquiring a report of the evaluation results related to the "final project risk rank" item. Element 632A is an example where the judgment results do not match. Element 633A is an example of reporting when the judgment results do not match, as shown in Element 632A. Element 634A is an example where the judgment result matches. Element 635A is an example of a report when the judgment results match, as shown in Element 634A.

[0070] As explained above, the document information to be evaluated (see element 623A) includes evaluation items (e.g., "final project risk rank") whose item value is determined by the maximum or minimum item value among multiple item values, such as "project risk rank by size" and "project risk rank by required quality".

[0071] The evaluation criteria include the standards for determining the values ​​of the evaluation items. The evaluation acquisition prompt 620A includes a request for a comparative evaluation item value, which is the value of an evaluation item determined according to the result of comparing the comparison target item values, and a request to report whether the comparative evaluation item value matches the described evaluation item value, which is the value of an evaluation item included in the evaluation target document information (see element 624A).

[0072] The prescribed format of the report (see element 633A) is a format for reporting discrepancies when the comparative evaluation item value and the stated evaluation item value do not match. The prescribed report format (see element 633A) includes a proposed revision that modifies the described evaluation item value to match the comparative evaluation item value if the comparative evaluation item value and the described evaluation item value do not match.

[0073] Figure 20 shows the evaluation acquisition prompt 620B according to this embodiment. The evaluation criterion in element 622B of the evaluation acquisition prompt 620B is a judgment criterion that determines whether the item value is a predetermined value (for example, blank) and whether the item value is appropriate (whether or not there is a violation). Figure 21 shows the report acquisition prompt 630B according to this embodiment. Element 632B is an example of a violation. Element 633B is an example of reporting a violation as shown in Element 632B. Element 634B is an example without a violation. Element 635B is an example of a report in the absence of a violation as shown in Element 634B.

[0074] As explained above, the evaluation criteria (element 622B) include prohibited item identification information, which is item identification information that prohibits an item value from being a predetermined value (e.g., blank). The evaluation acquisition prompt 620B includes a request to acquire the item value of the prohibited item identification information contained in the document information to be evaluated, and a request to acquire the prohibited value entry item identification information that indicates the item identification information for which the acquired item value is a predetermined value. The prescribed format of the report (see element 633B) includes identification information for prohibited value entries.

[0075] Figure 22 shows the evaluation acquisition prompt 620C according to this embodiment. The evaluation criterion in element 622C of the evaluation acquisition prompt 620C is an evaluation criterion for determining whether the item value is appropriate or not. Figure 23 shows the report acquisition prompt 630C according to this embodiment. Element 632C is an example of an item value that is inappropriate (violates the evaluation criteria). Element 633C is an example of reporting when an item value is inappropriate, as shown in Element 632B. Element 634C is an example of an item with an appropriate value. Element 635C is an example of reporting when the item value shown in Element 634B is appropriate.

[0076] Document Evaluation System: Document Evaluation Report Processing Figure 24 is a flowchart of the document evaluation report processing according to this embodiment. The process of obtaining evaluation results for a document and generating a report will be explained with reference to Figure 24.

[0077] In step S41, the document evaluation unit 113 acquires document information that is the subject of the evaluation report. For example, for tabular documents, the document evaluation unit 113 acquires the item value acquisition results 430 (see Figure 5) extracted by the document analysis unit 111 from the document database 131. Next, based on the format information corresponding to the document, the document evaluation unit 113 acquires evaluation criteria, questions, and report examples from the document evaluation report rule database 150.

[0078] In step S42, the document evaluation unit 113 starts the process of repeating steps S43 to S44 for each item for which there is an evaluation criterion. In step S43, the document evaluation unit 113 generates an evaluation acquisition prompt 620 (see Figure 12). Here, element 622 is an evaluation criterion corresponding to an item among the evaluation criteria acquired in step S41. Element 624 is a question corresponding to the said evaluation criterion. Element 623 is the document information acquired in step S41. Element 624 is a question among the questions acquired in step S41 that corresponds to the evaluation criterion (corresponding to an item).

[0079] In step S44, the document evaluation unit 113 instructs the response acquisition unit 118 to send an evaluation acquisition prompt 620 to the language model server 210. Next, the document evaluation unit 113 stores the evaluation results 450, 450A (see Figures 13 and 14), which are the responses, in the evaluation report database 160, associating them with the document identification information of the document.

[0080] In step S45, the evaluation reporting unit 114 starts the process of repeating steps S46 to S47 for each item for which there is an evaluation criterion. In step S46, the evaluation reporting unit 114 generates a report acquisition prompt 630 (see Figures 15-16). Here, elements 632-635 are report examples corresponding to the items (evaluation criteria) among the report examples obtained in step S41. Element 636 is a question corresponding to the item, and is the same question as element 624 of the evaluation acquisition prompt 620 generated in step S43. Element 637 is the evaluation result, which is the answer obtained in step S44.

[0081] In step S47, the evaluation reporting unit 114 instructs the response acquisition unit 118 to send a report acquisition prompt 630 to the language model server 210. Next, the evaluation reporting unit 114 stores the response report 460 (see Figure 17) in the evaluation report database 160, associating it with the document identification information of the document. In step S48, the evaluation reporting unit 114 merges the reports obtained in step S47 to create an evaluation report for the document and stores it in the evaluation report database 160.

[0082] Features of the document evaluation device The document processing device 100, acting as a document evaluation device, obtains evaluation results 450 and 450A for each item by referring to evaluation criteria (see elements 622, 622A, 622B, and 622C shown in Figures 12, 18, 20, and 22). Furthermore, the document processing device 100 obtains reports 460 for each item by referring to example reports (see elements 632-635, 632A-635A, 632B-635B, and 632C-635C shown in Figures 15, 19, 21, and 23). By evaluating / reviewing each item contained in a document and generating a report, the document processing device 100 can obtain highly accurate evaluation results and reports corresponding to each item. In addition, by preparing evaluation criteria and questions corresponding to each item, it is expected that even more accurate evaluation results and reports can be obtained.

[0083] ≪Document Processing Device: Document Evaluation Criteria Extraction Device≫ The following describes the document processing device 100 as a document evaluation criteria extraction device. The document processing device 100 collects evaluation criteria, check contents, and check rules for each item in documents containing items, such as tabular documents 410. The document processing device 100 extracts and collects this information from minutes of evaluation / review meetings and documents that describe evaluation criteria and check contents. Hereinafter, evaluation criteria, check contents, and check rules will be collectively referred to as evaluation criteria.

[0084] Document Evaluation Criteria Extraction Device: Evaluation Criteria Extraction Unit Figure 25 shows an example of a document 510 containing evaluation criteria according to this embodiment. Document 510 is, for example, the minutes of a meeting where a document was evaluated / reviewed. The minutes contain evaluation criteria, including review perspectives and rules. The evaluation criteria extraction unit 115 extracts evaluation criteria from the document containing the evaluation criteria. More specifically, the evaluation criteria extraction unit 115 generates an evaluation criteria acquisition prompt 640 (see Figure 26 below) that commands / instructs the extraction of evaluation criteria from the document. Subsequently, the evaluation criteria extraction unit 115 instructs the response acquisition unit 118 to send the evaluation criteria acquisition prompt 640 to the language model server 210 and receives the evaluation criteria as the response. The evaluation criteria extraction unit 115 stores the received evaluation criteria in the document evaluation report rule database 150, associating them with the format information of the document being evaluated / reviewed.

[0085] Document Evaluation Criteria Extraction Device: Evaluation Criteria Acquisition Prompt Figure 26 shows the configuration of the evaluation criteria acquisition prompt 640 according to this embodiment. The evaluation criteria acquisition prompt 640 includes elements 641 to 642. Element 641 provides instructions / directives for generating and obtaining evaluation criteria, referencing the actual review findings shown in Element 642. Element 642 contains text information included in document 510.

[0086] Document Evaluation Criteria Extraction Device: Evaluation Criteria Figure 27 shows the evaluation criteria 520 obtained according to this embodiment. The evaluation criteria 520 is a general evaluation criterion obtained by removing the specific content of individual documents from the content of review comments included in element 642. In this way, the evaluation criteria extraction unit 115 extracts and obtains evaluation criteria from documents containing the evaluation target.

[0087] As described above, the document processing device 100, which functions as a document evaluation criteria extraction device, includes an evaluation criteria extraction unit 115 that creates an evaluation criteria acquisition prompt 640 (see Figure 26) for extracting evaluation criteria for a document by referring to minutes containing the review results of the document. Furthermore, the document processing device 100, which functions as a document evaluation criteria extraction device, includes a response acquisition unit 118 that sends an evaluation criteria acquisition prompt 640 to the language model server 210 and acquires evaluation criteria directly or indirectly from the language model server 210. The evaluation criteria extraction unit 115 stores the evaluation criteria in the storage unit 130 (see document evaluation report rule database 150).

[0088] ≪Document Evaluation Criteria Extraction Device: Modified Version of Evaluation Criteria Acquisition Prompt≫ The evaluation criteria obtained using the evaluation criteria acquisition prompt 640 (see Figure 26) are the evaluation criteria for the document that was reviewed / evaluated, and by extension, for documents of the same format as that document. The evaluation criteria extraction unit 115 may also acquire evaluation criteria for each item determined by the document format (format information).

[0089] Figure 28 shows the evaluation criteria acquisition prompt 640A according to this embodiment. The evaluation criteria acquisition prompt 640A includes element 643A, which is not present in the evaluation criteria acquisition prompt 640. Element 641A indicates an instruction / command to generate and acquire evaluation criteria for each item of the document in element 643A. Figure 29 shows the evaluation criteria 520A when using the evaluation criteria acquisition prompt 640A according to this embodiment. Evaluation criteria are extracted for each item (chapter).

[0090] As described above, the evaluation criteria extraction unit 115 refers to the items described in the document and the minutes including the document review results, and creates an evaluation criteria acquisition prompt 640A that extracts evaluation criteria for each item of the document (see element 643A shown in Figure 28). The evaluation criteria extraction unit 115 stores the evaluation criteria as evaluation criteria for each item in the storage unit 130 (see document evaluation report rule database 150).

[0091] The source document for extraction is not limited to text-based documents such as meeting minutes. Text information can also be extracted from tabular checklist documents, and the evaluation criteria representing the checklist items can be extracted from those. Figure 30 shows the evaluation criteria acquisition prompt 640B (check content acquisition prompt) according to this embodiment. The evaluation criteria acquisition prompt 640B includes elements 641B and 642B.

[0092] Element 641B indicates a command / instruction to generate and retrieve check content (evaluation criteria) for each item (chapter) from the text information in element 642B. Element 642B shows text information extracted from the checklist document (the "Checklist for Functional Specifications" shown in Figure 30). Figure 31 shows the evaluation criteria 520B (check contents) when using the evaluation criteria acquisition prompt 640B according to this embodiment. The evaluation criteria are extracted together for each item (chapter).

[0093] As described above, the evaluation criteria extraction unit 115 creates a check content acquisition prompt (see evaluation criteria acquisition prompt 640B) that extracts the check content for each item (e.g., chapter) from a checklist document that contains the check content for each item of the document. The response acquisition unit 118 sends a check content acquisition prompt to the language model server 210 and acquires the check content directly or indirectly from the language model server 210. The evaluation criteria extraction unit 115 stores the check contents as evaluation criteria for the items in the storage unit 130 (document evaluation report rule database 150).

[0094] Document Evaluation Criteria Extraction Device: Evaluation Criteria Extraction Process Figure 32 is a flowchart of the evaluation criteria extraction process according to this embodiment. In step S51, the evaluation criteria extraction unit 115 retrieves the minutes of the meetings where the documents were evaluated / reviewed from the evaluation report database 160. In step S52, the evaluation criteria extraction unit 115 starts the process of repeating steps S53 to S54 for each set of meeting minutes.

[0095] In step S53, the evaluation criteria extraction unit 115 generates an evaluation criteria acquisition prompt 640A (see Figure 28). Here, element 642A is text information contained in the meeting minutes. Element 643A is an item (item identification information) contained in the format information of the document to be evaluated / reviewed in the meeting minutes.

[0096] In step S54, the evaluation criteria extraction unit 115 instructs the response acquisition unit 118 to send an evaluation criteria acquisition prompt 640A to the language model server 210. Next, the evaluation criteria extraction unit 115 stores the response evaluation criterion 520A (see Figure 29) in the document evaluation report rule database 150, associating it with the formal information items. Before storing the evaluation criterion 520A in the document evaluation report rule database 150, the evaluation criteria extraction unit 115 may inquire with the user of the document processing device 100 whether or not to adopt it as an evaluation criterion. Steps S55-S58 are the same process as steps S51-S54, but instead of processing meeting minutes, they process a checklist.

[0097] Features of the evaluation criteria extraction device The document processing device 100, acting as an evaluation criteria extraction device, extracts evaluation criteria for each document format (formal information) and for each item constituting the document from meeting minutes and checklists. This makes it possible to obtain evaluation criteria without manual intervention. By collecting evaluation criteria, organizing them manually as needed, and referring to them in the document evaluation report processing, it is expected that the quality of document evaluation by the document processing device 100 will improve.

[0098] ≪Document Processing Equipment: Document Proofreading Equipment≫ The document processing device 100, which functions as a document proofreading device, is described below. The document processing device 100 performs document proofreading, including resolving inconsistencies in wording that correspond to the document and the business related to the document.

[0099] Document proofreading device: Overview of processing Figure 33 is a diagram illustrating the outline of the document proofreading process (see Figure 40 below) performed by the document processing device 100 according to this embodiment. The dictionary generation unit 116 generates a synonym acquisition prompt 650 (see Figure 34 below) that refers to one or more related documents 540, such as documents related to the same business. The synonym acquisition prompt 650 is sent to the language model server 210, and a synonym dictionary 550 (see Figure 35 below) is returned. The synonym dictionary 550 is a synonym dictionary specialized for the context contained in the document 540 and the business related to the document 540.

[0100] The document proofreading unit 117 generates a proofreading document acquisition prompt 660 (see Figure 36 below) that refers to the synonym dictionary 550. The proofreading document acquisition prompt 660 is sent to the language model server 210, and the proofread document 560 is returned. The details of the document processing device 100 as a document proofreading device are described below.

[0101] Document proofreading device: Dictionary generation unit The dictionary generation unit 116 generates a synonym acquisition prompt 650 (see Figure 34 below) that commands / instructs the generation of a synonym dictionary. Subsequently, the dictionary generation unit 116 instructs the answer acquisition unit 118 to send the synonym acquisition prompt 650 to the language model server 210 and receives the answer, a synonym dictionary 550 (see Figure 35 below).

[0102] Document proofreading device: Synonym acquisition prompt Figure 34 shows the configuration of the synonym acquisition prompt 650 according to this embodiment. The synonym acquisition prompt 650 includes elements 651 to 652. Element 651 indicates a command / instruction to retrieve a thesaurus (group of synonyms) by referring to the text information of document 540 shown in element 652. The command / instruction includes criteria for selecting which word to use as the headword among the synonyms. Examples of criteria include prioritizing katakana spelling, kanji spelling, or abbreviations. Element 652 contains text information included in document 540.

[0103] Document proofreading device: Thesaurus Figure 35 shows the synonym dictionary 550 obtained according to this embodiment. The synonym dictionary 550 uses words that conform to the headword selection criteria commanded / instructed in element 651 as headwords (keys in JSON format).

[0104] Document proofreading equipment: Document proofreading department The document proofreading unit 117 generates a proofreading document acquisition prompt 660 (see Figure 36 below) that commands / instructs the proofreading of the document. Subsequently, the document proofreading unit 117 instructs the response acquisition unit 118 to send the proofreading document acquisition prompt 660 to the language model server 210 and receives the proofreading document as the response. The document proofreading unit 117 associates the proofreading document with the target document and stores it in the document database 131.

[0105] ≪Document Proofreading Device: Proofread Document Acquisition Prompt≫ Figure 36 shows the configuration of the calibration document acquisition prompt 660 according to this embodiment. The calibration document acquisition prompt 660 includes elements 661 to 663. Element 661 indicates a command / instruction to modify and proofread the document shown in element 663, referring to the thesaurus 550 shown in element 662. In addition to word variations, element 661 includes commands / instructions to correct tense variations and to change sentence endings to the polite "desu / masu" style. Element 662 shows the thesaurus 550. Element 663 contains the text information of the document to be proofread.

[0106] As described above, the document processing device 100 as a document proofreading device includes a dictionary generation unit 116 that creates a synonym acquisition prompt 650 which requests a synonym group, which is a collection of words in the text information (see element 652 shown in Figure 34) of a group of documents containing one or more documents, grouped by synonym. The document processing device 100, as a document proofreading device, includes a response acquisition unit 118 that sends a synonym acquisition prompt 650 to the language model server 210 and acquires a synonym group (see the synonym dictionary 550 shown in Figure 35) directly or indirectly from the language model server 210.

[0107] The document processing device 100, as a document proofreading device, includes a document proofreading unit 117 that creates a proofreading document acquisition prompt 660 (see Figure 36) that requests a proofreading document in which variations in word spelling in documents included in the document group have been corrected by referring to synonym groups. The response acquisition unit 118 sends a proofreading document acquisition prompt 660 to the language model server 210 and acquires the proofreading document directly or indirectly from the language model server 210.

[0108] The synonym acquisition prompt 650 includes a request for each synonym group to make a word that meets a predetermined criterion the headword (see element 651). The prescribed criteria are either prioritizing katakana notation, prioritizing kanji notation, or prioritizing abbreviations. The proofreading document acquisition prompt 660 requests a proofreading document that has corrected variations in word spelling to prioritize the headword (see element 661).

[0109] Proofreading document acquisition prompt 660 requests correction of inconsistencies in sentence endings, following correction of inconsistencies in word spelling (see element 661). Variations in sentence endings include at least one of the following: variations in tense, and variations in the use of polite / formal language (desu / masu) versus polite / formal language (da / dearu).

[0110] ≪Document proofreading device: Variations of synonym acquisition prompt and proofread document acquisition prompt≫ Figure 37 shows the synonym acquisition prompt 650A according to this embodiment. Compared to the synonym acquisition prompt 650 (see Figure 34), it emphasizes grouping synonyms while considering the context. Figure 38 shows a synonym dictionary 550A containing meaning, according to this embodiment. Figure 39 shows the proofreading document acquisition prompt 660A according to this embodiment. Compared to the proofreading document acquisition prompt 660 (see Figure 36), it emphasizes correcting word inconsistencies while considering the context.

[0111] As explained above, the synonym acquisition prompt 650A requests a group of synonyms that take into account their meaning in context. Synonym groups include meanings within the context. Proofreading document acquisition prompt 660A requests a proofreading document in which variations in word spelling have been corrected, taking into account the meaning in context.

[0112] Document proofreading device: Document proofreading process Figure 40 is a flowchart of the document proofreading process according to this embodiment. In step S61, the dictionary generation unit 116 retrieves relevant documents 540 from the document database 131.

[0113] In step S62, the dictionary generation unit 116 generates a synonym acquisition prompt 650 (see Figure 34). Here, element 652 is the entire text information contained in document 540. In step S63, the dictionary generation unit 116 instructs the answer acquisition unit 118 to send a synonym acquisition prompt 650 to the language model server 210 and receives the synonym dictionary 550, which is the answer.

[0114] In step S64, the document proofreading unit 117 starts the process of repeating steps S65 to S66 for each document acquired in step S61. In step S65, the document proofreading unit 117 generates a proofread document acquisition prompt 660 (see Figure 36). Here, element 662 is the text information contained in the document.

[0115] In step S66, the document proofreading unit 117 instructs the response acquisition unit 118 to send a proofreading document acquisition prompt 660 to the language model server 210. The document proofreading unit 117 receives the configured proofreading document, which is the response, associates it with the original document, and stores it in the document database 131.

[0116] Features of the document proofreading device The document processing device 100, acting as a document proofreading device, performs document proofreading to correct inconsistencies in synonyms, tenses, and sentence endings. For synonyms, corrections are made by referring to a synonym dictionary 550 generated based on related documents 540, such as documents related to the same task (see element 662). This makes it possible to correct words to correspond to meanings specific to the task or meanings according to the context in which the word appears.

[0117] <<Example of a document proofreading tool: Proofreading document acquisition prompt>> In the embodiment described above, variations in synonyms, variations in tense, and variations in polite / formal writing styles are corrected in a single proofreading document acquisition prompt 660, 660A. The corrections may be divided into multiple prompts. Dividing the corrections into multiple prompts is expected to improve accuracy.

[0118] ≪Variation: Language Model≫ The document processing device 100 described above sends a prompt to the language model server 210, which provides a language model service, and obtains a response. The document processing device 100 itself may also have a language model and query itself. Figure 41 is a functional block diagram of a document processing device 100A according to a modified example of this embodiment. The document processing device 100A includes a language model 132. The response acquisition unit 118A uses the language model 132 to take the prompt as input and output a response. By not using the language model server 210 of an external service, the risk of information leakage can be reduced. Furthermore, by using a language model 132 trained according to the organization's business, the accuracy of the response can be improved, and an improvement in the accuracy of document evaluation / review can be expected.

[0119] As described above, the response acquisition unit 118A of the document processing device 100A, which functions as a document analysis device, uses the language model 132 instead of the language model server 210 to acquire the item value acquisition result 430 for the item value acquisition prompt 610.

[0120] The response acquisition unit 118A of the document processing device 100A, which functions as a document evaluation device, uses the language model 132 instead of the language model server 210 to acquire evaluation results 450 for evaluation acquisition prompts 620 and reports 460 for report acquisition prompts 630.

[0121] The response acquisition unit 118A of the document processing device 100, which serves as a document evaluation criteria extraction device, uses the language model 132 instead of the language model server 210 to acquire the check content in response to the check content acquisition prompt (see evaluation criteria acquisition prompt 640B).

[0122] The answer acquisition unit 118A of the document processing device 100, which functions as a document proofreading device, uses the language model 132 instead of the language model server 210 to acquire a synonym group (synonym dictionary 550) for the synonym acquisition prompt 650 and acquire a proofread document for the proofread document acquisition prompt 660.

[0123] <<Other variations>> Although several embodiments and modifications of the present invention have been described above, these embodiments are merely illustrative and do not limit the technical scope of the present invention. The present invention can take many other embodiments, and various modifications such as omissions and substitutions can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention as described herein, and are also included in the scope of the invention and its equivalents as described in the claims.

[0124] Hardware Configuration The document processing devices 100 and 100A according to the above-described embodiment are implemented by a computer 900 having a configuration such as that shown in Figure 42. Figure 42 is a hardware configuration diagram showing an example of a computer 900 that implements the functions of the document processing devices 100 and 100A according to the above-described embodiment. The computer 900 includes a CPU 901, ROM 902, RAM 903, SSD 904, and an input / output interface 905 (labeled as input / output I / F (Interface) in Figure 42). Furthermore, the computer 900 includes a communication interface 906 (labeled as communication I / F in Figure 42) and a media interface 907 (labeled as media I / F in Figure 42). The computer 900 may be equipped with an HDD (Hard Disk Drive) instead of the SSD 904, or it may be equipped with an HDD in addition to the SSD 904.

[0125] The CPU 901 operates based on programs stored in the ROM 902 or SSD 904 and is controlled by the control unit 110 in Figure 1. The ROM 902 stores boot programs executed by the CPU 901 when the computer 900 starts up, as well as programs related to the computer 900's hardware.

[0126] The CPU 901 controls input devices 910, such as a mouse and keyboard, and output devices 911, such as a display and printer, via the input / output interface 905. The CPU 901 acquires data from the input devices 910 and outputs the generated data to the output devices 911 via the input / output interface 905.

[0127] SSD904 stores programs executed by CPU901 and data used by those programs. Communication interface906 receives data from other devices (e.g., language model server210) not shown via the communication network and outputs it to CPU901, and also transmits data generated by CPU901 to other devices via the communication network.

[0128] The media interface 907 reads a program or data stored in the recording medium 912 and outputs it to the CPU 901 via the RAM 903. The CPU 901 loads the program from the recording medium 912 onto the RAM 903 via the media interface 907 and executes the loaded program. The recording medium 912 can be an optical recording medium such as a DVD (Digital Versatile Disk), a magneto-optical recording medium such as an MO (Magneto Optical Disk), a magnetic recording medium, a conductive memory tape medium, or a semiconductor memory.

[0129] For example, when computer 900 functions as a document processing device 100, 100A according to the above embodiment, the CPU 901 of computer 900 realizes the functions of the document processing device 100, 100A by executing a program 138 (see Figure 1) loaded on RAM 903. The CPU 901 reads the program from the recording medium 912 and executes it. In addition, the CPU 901 may read the program from another device via a communication network, or it may install the program 138 from the recording medium 912 to the SSD 904 and execute it. Note that the document processing device 100, 100A is not limited to a hardware computer 900, but may also be in the form of a virtual machine. [Explanation of symbols]

[0130] 100 Document processing devices (document analysis devices, document evaluation devices, document evaluation criterion extraction devices, document proofreading devices) 111 Document Analysis Department 112 Document format generator 113 Document Evaluation Department 114 Evaluation Reporting Department 115 Evaluation Criteria Extraction Unit 116 Dictionary Generation Unit 117 Document Proofreading Department 118,118A Answer acquisition part 131 Document Databases 132 Language Models 140 Document Format Databases 150 Document Evaluation Reporting Rules Database 160 Evaluation Report Database 210 Language Model Servers 410 Tabular document 420, 420A Format Information Document Item value acquisition results for 430, 430A 450, 450A Evaluation Results 460 Report 510 documents 520, 520A Evaluation Criteria 520B Evaluation Criteria (Checklist) 540 documents 550, 550A Synonym Dictionary (Synonym Group) 560 Proofreading Documents 610 Item Value Retrieval Prompt 612 elements (format information) 613 elements (document information) 620, 620A, 620B, 620C Evaluation Acquisition Prompt 622 elements (evaluation criteria) 623 elements (information on documents to be evaluated) 624 elements (questions) 630, 630A, 630B, 630C Report acquisition prompt Element 632-635 (Example Report) 640, 640A Evaluation Criteria Acquisition Prompt 640B Evaluation Criteria Acquisition Prompt (Check Content Acquisition Prompt) 650, 650A Synonym acquisition prompt 660, 660A Calibration Document Acquisition Prompt

Claims

1. A dictionary generation unit creates a synonym retrieval prompt that requests synonym groups, which are created by grouping words contained in the text information of a document set containing one or more documents into synonym groups. A response acquisition unit that sends the aforementioned synonym acquisition prompt to a language model server and acquires the synonym group directly or indirectly from the language model server, The system includes a document proofreading unit that creates a proofreading document acquisition prompt requesting a proofreading document that corrects variations in word spelling in the documents included in the document group by referring to the aforementioned synonym group, The aforementioned response acquisition unit, The aforementioned proofreading document acquisition prompt is sent to the language model server, and the proofreading document is acquired directly or indirectly from the language model server. Document proofreading device.

2. The aforementioned synonym acquisition prompt is, For each of the aforementioned synonym groups, the requirement includes making a word that meets a predetermined criterion a headword, The aforementioned prompt for obtaining a proofreading document is: We request a proofread document in which the aforementioned variations in word spelling have been corrected to prioritize the aforementioned headword. The aforementioned prescribed standards are: Prioritize either katakana notation, kanji notation, or abbreviations. The document proofreading apparatus according to claim 1.

3. The aforementioned synonym acquisition prompt is, We request a group of synonyms that take into account their meaning in context. The aforementioned group of synonyms is Including meaning in context, The aforementioned prompt for obtaining a proofreading document is: We request a proofread document that corrects the inconsistencies in the aforementioned word spelling, taking into account their meaning in context. The document proofreading apparatus according to claim 1.

4. The aforementioned prompt for obtaining a proofreading document is: Following the correction of the aforementioned inconsistencies in word spelling, we request the correction of inconsistencies in sentence-ending spelling. The document proofreading apparatus according to claim 1.

5. The inconsistencies in the notation at the end of the sentence mentioned above are, It includes at least one of the following: fluctuations in tense, and fluctuations in the use of polite / formal language (desu / masu) or polite / formal language (da / dearu). The document proofreading apparatus according to claim 4.

6. The aforementioned response acquisition unit, Instead of the aforementioned language model server, use a language model: The synonym group is obtained in response to the synonym acquisition prompt. Acquire the calibration document in response to the calibration document acquisition prompt. The document proofreading apparatus according to claim 1.

7. The document proofreading device, The steps include creating a synonym retrieval prompt that requests synonym groups, which are created by grouping words contained in the text information of a set of documents containing one or more documents, by synonym; The steps include sending the synonym acquisition prompt to the language model server and acquiring the synonym group directly or indirectly from the language model server, The steps include creating a proofreading document acquisition prompt that requests proofreading documents that correct variations in word usage in the documents included in the document group, with reference to the aforementioned synonym group, The steps include: sending the aforementioned proofreading document acquisition prompt to the language model server and acquiring the proofreading document directly or indirectly from the language model server. Document proofreading methods.

Citation Information

Patent Citations

  • Log management device

    JP2022185696A