Information processing device, information processing method, and program

A machine learning-based system addresses personnel inconsistencies in document reviews by training on past documents to check for deviations and omissions, ensuring consistent adherence to document conventions and reducing human error.

JP7893490B2Active Publication Date: 2026-07-22NEC PLATFROMS LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC PLATFROMS LTD
Filing Date
2023-11-09
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing document review processes are inconsistent due to changes in personnel, leading to variations in checking depth and quality, and AI-powered systems are limited in checking adherence to specific document conventions.

Method used

A system utilizing machine learning to create a trained model that checks for deviations and omissions in new documents by learning from past documents and their feedback, categorizing issues, and outputting corrective actions.

Benefits of technology

Automates document checking, reducing human-dependent inconsistencies and ensuring adherence to established document standards, thereby maintaining document quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893490000001
    Figure 0007893490000001
  • Figure 0007893490000002
    Figure 0007893490000002
  • Figure 0007893490000003
    Figure 0007893490000003
Patent Text Reader

Abstract

To easily check deviations and writing omissions of a new document without depending on skills of a specific individual.SOLUTION: An information processing device acquires a past document, and extracts a described feature amount being a feature amount from described contents of the document. The information processing device creates a data set for described content learning on the basis of the described feature amount. The information processing device acquires indication contents of the document, and extracts an indication feature amount being a feature amount from the indication contents of the document. The information processing device creates a data set for indication content learning on the basis of the indication feature amount. When a document is inputted on the basis of the data set for described content learning and the data set for indication content learning, the information processing device creates a learned model learned so as to output a check result of the document.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus for checking the description content of a document.

Background Art

[0002] Documents such as specifications and design drawings for specific customers often follow the description methods and rules of documents created in the past. Therefore, when a document is created, a design review is conducted by experts such as the creators of past documents, customers, and salespersons, and the description content is carefully examined and corrected as necessary. Patent Document 1 describes a system that can efficiently perform correction while checking the finished image by displaying the finished screen and the editing screen.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, it is difficult to conduct the design review of a document with the same members every time due to changes in responsible persons, transfers, retirements, etc. For example, when one expert retires, a substitute expert will participate. However, since the design review depends heavily on the know-how of the experts, the checking range and depth of the document will change depending on the presence or absence of the know-how of the substitute expert. Therefore, depending on the members of the design review, many omissions and deficiencies in the description will occur, and the quality of the document will be affected by the checker, which has been a problem.

[0005] Furthermore, while AI-powered automated document checking currently exists in the market, it is limited to checking for typographical errors in general terminology and legal checks, and is unable to check new documents in accordance with the writing conventions and rules of past documents.

[0006] One of the objectives of this invention is to easily check for deviations and omissions in new documents without relying on the skills of a specific individual. [Means for solving the problem]

[0007] To solve the above problems, in one aspect of the present invention, the information processing device is: Document acquisition methods for obtaining documents, Based on the contents of the aforementioned document, The elements that make up the document A descriptive feature extraction means for extracting descriptive features, which are feature quantities, Based on the aforementioned features, The aforementioned document was used as input data, and the elements constituting that document were associated with it as ground truth data. A method for creating a data set for training on the content described, A means for obtaining the contents of the document mentioned above, Based on the points raised above, The findings are categorized based on the category of the issue, and correspond to the corrective actions that need to be taken and the actions that should be replaced with the corrective actions. A feature extraction method for extracting feature quantities, which are the pointed-out feature quantities, Based on the aforementioned features, The aforementioned document is used as input data, and category-specific error information, including the action items and the points raised for each category corresponding to the document, is associated with it as correct data. A means for creating a dataset for learning the content of the feedback, The aforementioned data set for learning and The aforementioned data set for learning the points raised. It includes both of the above as learning targets, uses the above document as input data, and is trained based on the correct answer data corresponding to the above content learning dataset or the above comment learning dataset. The system includes a pre-trained model learning means that creates a pre-trained model that outputs a check result when the aforementioned document is input.

[0008] In another aspect of the present invention, a computer-based information processing method is: Obtain the document, Based on the contents of the aforementioned document, The elements that make up the document Extract the descriptive features, which are the features, Based on the above-described feature amount, The aforementioned document was used as input data, and the elements constituting that document were associated with it as ground truth data. create a dataset for learning the description content, obtain the pointed-out content of the document, from the pointed-out content, The findings are categorized based on the category of the issue, and correspond to the corrective actions that need to be taken and the actions that should be replaced with the corrective actions. extract the pointed-out feature amount, which is a feature amount, Based on the pointed-out feature amount, The aforementioned document is used as input data, and category-specific error information, including the action items and the points raised for each category corresponding to the document, is associated with it as correct data. create a dataset for learning the pointed-out content, the dataset for learning the description content and the dataset for learning the pointed-out content It includes both of the above as learning targets, uses the above document as input data, and is trained based on the correct answer data corresponding to the above content learning dataset or the above comment learning dataset. Create a learned model that outputs a check result when the document is input.

[0009] In another aspect of the present invention, the program obtains a document, [[ID=2,7]]from the description content of the document, The elements that make up the document extract the description feature amount, which is a feature amount, Based on the description feature amount, The aforementioned document was used as input data, and the elements constituting that document were associated with it as ground truth data. create a dataset for learning the description content, obtain the pointed-out content of the document, from the pointed-out content, The findings are categorized based on the category of the issue, and correspond to the corrective actions that need to be taken and the actions that should be replaced with the corrective actions. extract the pointed-out feature amount, which is a feature amount, Based on the pointed-out feature amount, The aforementioned document is used as input data, and category-specific error information, including the action items and the points raised for each category corresponding to the document, is associated with it as correct data. create a dataset for learning the pointed-out content, the dataset for learning the description content and the dataset for learning the pointed-out content It includes both of the above as learning targets, uses the above document as input data, and is trained based on the correct answer data corresponding to the above content learning dataset or the above comment learning dataset. Cause the computer to execute a process of creating a learned model that outputs a check result when the document is input.

Advantages of the Invention

[0010] According to the present invention, it is possible to easily check for deviations and omissions in a new document without depending on the skills of a specific individual.

Brief Description of the Drawings

[0011] [Figure 1] Shows the configuration of the analysis system. [Figure 2] Shows the hardware configurations of the user terminal and the analysis server. [Figure 3] Shows the functional configuration of the analysis server during learning. [Figure 4] An example of the classification of past documents and their described contents. [Figure 5] An example of a checklist and a correction table by category. [Figure 6] A flowchart of the learning process. [Figure 7] Shows the functional configuration of the analysis server during checking. [Figure 8] A flowchart of the checking process.

Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [Hardware Configuration] FIG. 1 shows the configuration of an analysis system to which the analysis server of the present invention is applied. The analysis system 100 is a system that creates a learned model by learning the characteristics of past documents and checks for deviations and omissions in new documents (hereinafter also referred to as "new documents") using the learned model.

[0013] In the present embodiment, the document is a specification or a design document for a specific customer, and follows the description method and description rules of the documents created in the past. When a document is created, a design review is performed by knowledgeable persons such as the person in charge of creating the past document, the customer, and the salesperson, and it is checked whether the created document deviates from the description method and description rules. Then, a checklist listing the pointed-out contents pointed out and corrected by the knowledgeable person in the design review is created. The past documents and the checklist may be stored in the user terminal used by the user who created the document, or may be stored in a predetermined DB.

[0014] The analysis system 100 is connected to a user terminal 1 and an analysis server 20 via a network 5 such as the Internet, enabling communication between them. The user terminal 1 is an information processing device that processes, stores, and transmits various types of data, and is a personal computer (PC) or general-purpose tablet used by the user. The user terminal 1 is used to create documents such as specifications and design documents for specific customers, and to create checklists listing the issues pointed out in the created documents.

[0015] In this embodiment, for the sake of explanation, one user terminal 1 is connected to the analysis server 20. However, in reality, multiple user terminals, such as user terminals that store past documents and checklists, and user terminals that create new documents, are connected to the analysis server 20.

[0016] The analysis server 20 is an information processing device that processes, stores, and transmits various types of data, and can be, for example, a server device, a PC, or a general-purpose tablet. As will be described in detail later, the analysis server 20 has a document content table database (hereinafter, the database is also referred to as "DB") 27, a problem content table DB 28, and a trained model DB 29. The analysis server 20 creates a trained model that has learned the characteristics of past documents. The analysis server 20 also uses the created trained model to check for deviations and omissions in new documents and outputs the check results. Specifically, the analysis server 20 sends the check results to the user terminal 1 for display.

[0017] Figure 2(a) is a block diagram showing the hardware configuration of user terminal 1. As shown in the figure, user terminal 1 includes an interface 11, a processor 12, memory 13, a recording medium 14, a display unit 15, and an input unit 16.

[0018] Interface 11 exchanges data with the analysis server 20 via the network 5. Interface 11 is used to send various data to the analysis server 20 and to receive various data from the analysis server 20.

[0019] The processor 12 is a computer such as a CPU (Central Processing Unit) and controls the entire user terminal 1 by executing pre-prepared programs. The memory 13 consists of ROM (Read Only Memory), RAM (Random Access Memory), etc. The memory 13 stores programs executed by the processor 12. The memory 13 is also used as working memory while the processor 12 is executing various processes.

[0020] The recording medium 14 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the user terminal 1. The recording medium 14 stores various programs executed by the processor 12. The display unit 15 displays various information, for example, an LCD (Liquid Crystal Display). The input unit 16 is a touch panel or the like, used by the user to input various information.

[0021] Figure 2(b) is a block diagram showing the hardware configuration of the analysis server 20. As shown in the figure, the analysis server 20 comprises an interface 21, a processor 22, memory 23, a recording medium 24, a display unit 25, and an input unit 26. These components are interconnected with the description table DB 27, the issue description table DB 28, and the trained model DB 29 via a bus.

[0022] Interface 21 exchanges data with user terminal 1 via network 5. Processor 22 is a computer such as a CPU and controls the entire analysis server 20 by executing pre-prepared programs. Memory 23 is composed of ROM, RAM, etc. Memory 23 stores programs executed by processor 22. Memory 23 is also used as working memory while processor 22 is executing various processes.

[0023] The recording medium 24 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the analysis server 20. The recording medium 24 stores various programs executed by the processor 22.

[0024] The display unit 25 displays a predetermined image, for example, an LCD (Liquid Crystal Display). The input unit 26 is used by, for example, an operator managing the analysis server 20, and can be a keyboard, mouse, touch panel, etc.

[0025] The content table DB27 stores a data set for content learning, based on features extracted from the content of the document.

[0026] The DB28 table of issues stores a dataset for learning issues based on features extracted from a checklist that lists the issues found in documents.

[0027] The trained model DB29 stores the trained models created by the analysis server 20.

[0028] For the sake of explanation, the analysis server 20 has a description table DB27, a comment table DB28, and a trained model DB29. However, the present invention is not limited to these, and the type of DB and the data structure of the DB are arbitrary as long as the necessary data can be obtained.

[0029] [Functional configuration during learning] The analysis server 20 creates a trained model by training a machine learning model with the characteristics of past documents. Examples of machine learning methods include models utilizing neural networks. Figure 3 is a block diagram showing the functional configuration of the analysis server 20 during training. Functionally, the analysis server 20 comprises a document acquisition unit 41, a document content learning unit 42, a checklist acquisition unit 43, and a feedback content learning unit 44, and is realized by the processor 12 executing a program.

[0030] The document acquisition unit 41 acquires past documents from a designated database or the like.

[0031] The document content learning unit 42 extracts document features from the content of past documents and creates a data set for document content learning. Figure 4 shows an example of a past document and the classification of its content. As shown in Figure 4, the content of past documents can be classified into multiple elements such as customer name, title, chapter structure, and notes. The chapter structure can be further classified into sentence units and figure / table units based on the content of each chapter. In addition, notes can be further classified into recommendation units and standard units required by each customer. Note that notes can be classified not only into recommendation units and standard units, but also into special condition units for each customer.

[0032] The content learning unit 42 analyzes the content of past documents element by element and extracts description features. Based on the description features, the content learning unit 42 creates pairs of past documents and each element that constitutes the document as a content learning dataset and stores it in the content table DB27. Specifically, the content learning dataset is training data in which past documents are used as input data and the elements that constitute the document are used as ground truth data.

[0033] The document content learning unit 42 creates multiple document content learning datasets from multiple past documents and trains a machine learning model to create a trained model. The created trained model is stored in the trained model DB 29. Specifically, when a new document is input using the document content learning datasets, the document content learning unit 42 trains the trained model to output the elements that make up the content of the document. At this time, if there are any deviations in the content of the document that are not included in the elements, the trained model outputs those deviations as check results. As a result, the analysis server 20 can use the trained model to check whether the content of the new document deviates from the writing rules of past documents.

[0034] The checklist acquisition unit 43 acquires a checklist from a predetermined database or the like, which lists the contents of the documents acquired by the document acquisition unit 41.

[0035] The feedback content learning unit 44 extracts feedback features from feedback content of past documents and creates a feedback content learning dataset. Figure 5 shows an example of a checklist listing feedback content and a category-based errata. As shown in Figure 5, the checklist consists of feedback location, feedback item, and action item. Action item is a string of characters, etc., that requires appropriate action, which was pointed out by an expert during the design review of the document as being omitted or requiring correction. Feedback item is the correct string of characters, etc., that should be written in place of the action item, and feedback location is the location in the document where the action item and feedback item exist. In this embodiment, the chapter in the document where the action item and feedback item exist is used as the feedback location. Based on the checklist of past documents, the feedback content learning unit 44 classifies the feedback item and action item by category, such as symbols, diagrams, and text, and creates a category-based errata as shown in Figure 5.

[0036] The feedback learning unit 44 analyzes the category-specific errata and extracts feedback features. The feedback learning unit 44 creates a feedback learning dataset by pairing past documents with the category-specific errata of the document in question, and stores it in the feedback table DB28. Specifically, the feedback learning dataset is training data in which past documents are used as input data and the category-specific errata of the document in question is used as correct answer data.

[0037] The error detection learning unit 44 creates multiple error detection learning datasets from checklists of multiple past documents and trains the pre-trained model created by the document content learning unit 42. The trained model is stored in the pre-trained model DB 29. Specifically, the error detection learning unit 44 uses the error detection learning datasets to train the pre-trained model so that when a new document is input, it outputs a category-specific errata table for that document as a check result. This allows the analysis server 20 to use the pre-trained model to check for omissions, corrections, etc., in the content of the new document.

[0038] In this embodiment, the content learning unit 42 trains a trained model using the content learning dataset, and the feedback content learning unit 44 then trains the feedback content learning dataset on that trained model. However, the present invention is not limited to this, and the order of training can be arbitrarily set.

[0039] Furthermore, in the above configuration, the document acquisition unit 41 and the checklist acquisition unit 43 are examples of the document acquisition means and the feedback content acquisition means of the present invention, respectively. Also, the description content learning unit 42 is an example of the description feature extraction means, the description content learning dataset creation means and the trained model creation means of the present invention. Also, the feedback content learning unit 44 is an example of the feedback feature extraction means, the feedback content learning dataset creation means and the trained model creation means of the present invention.

[0040] [Learning Process] Next, the learning process performed by the analysis server 20 will be described. Figure 6 is a flowchart of the learning process performed by the analysis server 20. This process is achieved by the processor 22 shown in Figure 2(b) executing a pre-prepared program.

[0041] First, the analysis server 20 retrieves past documents from a designated database or user terminal 1 (step S101). Then, the analysis server 20 extracts descriptive features from the content of the past documents (step S102) and creates a data set for content learning (step S103).

[0042] Furthermore, the analysis server 20 obtains a checklist from a designated database or user terminal 1, which lists the content of the comments on past documents obtained in step S101 (step S104). Then, the analysis server 20 extracts the comment features from the checklist (step S105) and creates a dataset for learning the content of the comments (step S106).

[0043] The analysis server 20 creates a trained model by training it with the data set for content learning and the data set for feedback learning, and stores it in the trained model DB 29 (step S107). This completes the training process. When a document is input, the trained model created in this way can output a check result for parts of the document's content that deviate from the writing rules of past documents or that need correction.

[0044] [Functional configuration during check] The analysis server 20 checks the contents of a new document using the trained model that was created. Figure 7 is a block diagram showing the functional configuration of the analysis server 20 during the check. Functionally, the analysis server 20 comprises a document acquisition unit 51 and a check result output unit 52, and is realized by the processor 12 executing a program.

[0045] The document acquisition unit 51 acquires a new document from the user terminal 1.

[0046] The check result output unit 52 uses the trained models stored in the trained model DB 29 to check whether there are any deviations or omissions in the content of the new document, and outputs the check result.

[0047] Specifically, the check result output unit 52 inputs the new document into the trained model and sends the check result output from the trained model to the user terminal 1 for display. At this time, the trained model refers to the description table DB27 and determines whether or not there is a past document corresponding to the description of the new document. In other words, it determines whether or not there is a past document in the trained model's learned past documents that followed the description rules when the new document was created. If there is no past document in the description table DB27 that matches the new document, it is impossible for the trained model to perform the check, so the check result output unit 52 adds "Check Failed" to the check result.

[0048] Furthermore, the trained model refers to the issue details table DB28 to determine whether there is a past document corresponding to the content of the new document. If there is no relevant past document in the issue details table DB28, it is impossible for the trained model to perform the check, and the check result output unit 52 adds "Check Failed" to the check result.

[0049] In the above configuration, the document acquisition unit 51 and the check result output unit 52 are examples of the novel document acquisition means and check result output means of the present invention, respectively.

[0050] [Checking process] Next, the check process performed by the analysis server 20 will be explained. Figure 8 is a flowchart of the check process performed by the analysis server 20. This process is achieved by the processor 22 shown in Figure 2(b) executing a pre-prepared program.

[0051] First, the analysis server 20 obtains a new document from the user terminal 1 and inputs it into the trained model (step S201). Then, the analysis server 20 obtains the check result of the new document output by the trained model and sends it to the user terminal 1 (step S202). The user terminal 1 displays the check result. This completes the check process.

[0052] Thus, according to the analysis server 20, by extracting features from the content of past specifications and design documents for specific customers, as well as the feedback from design reviews of those documents, and using machine learning to train the model, it is possible to automate the checking of new documents using the trained model. In this way, by performing automated checks using AI that has learned the characteristics of past documents, it is possible to check whether there are any deviations in the new document from the writing rules and feedback of past documents. Furthermore, since automated checks by AI do not depend on specific individuals, it is possible to suppress check omissions and inconsistencies in checking perspectives that can occur due to human error, thereby resolving the problem of document quality being affected by individual differences.

[0053] [First variation] In the above embodiment, the description content learning unit 42 and the issue content learning unit 44 train a single machine learning model with a description content learning dataset and an issue content learning dataset, respectively. However, the present invention is not limited to this, and the description content learning unit 42 and the issue content learning unit 44 may be trained with different machine learning models to create two trained models. In this case, when a new document is input into one of the trained models, the output check results can identify areas that deviate from the description rules of past documents, and when a new document is input into the other trained model, the check results can identify areas that are missing or require correction.

[0054] [Second variation] The above embodiment applies to documents for specific customers that require adherence to past document formatting practices and unique documentation rules. However, the present invention is not limited to this. It is also possible to achieve automated checking by a trained model in documents such as internal design specifications, instruction manuals, and contracts by training the model with past documents and checklists from design reviews.

[0055] Some or all of the above embodiments may also be described as follows, but are not limited to the following:

[0056] (Note 1) Document acquisition methods for obtaining documents, A feature extraction means for extracting feature quantities, which are feature quantities, from the contents of the aforementioned document, A means for creating a data set for learning the content based on the aforementioned features, A means for obtaining the contents of the document mentioned above, A feature extraction means for extracting feature quantities, which are feature quantities, from the content of the aforementioned document, A means for creating a data set for learning the content of the feedback based on the aforementioned feedback features, A means for creating a trained model that, based on the aforementioned data set for learning the content described and the aforementioned data set for learning the content of the comments, creates a trained model that, when the document is input, outputs the check result of the document. An information processing device equipped with the following features.

[0057] (Note 2) The aforementioned feature extraction means is an information processing device according to Appendix 1, which classifies the described content into a plurality of elements and extracts a described feature for each element.

[0058] (Note 3) The aforementioned feature extraction means is an information processing device as described in Appendix 1, which creates an errata sheet classified into multiple categories based on the content of the criticism, and extracts the feature quantities of the criticism from the errata sheet.

[0059] (Note 4) The pre-trained model creation means is an information processing device according to Appendix 1, which, based on the content training dataset, trains the pre-trained model to output as a check result any parts of the document that deviate from the description rules of past documents when the document is input.

[0060] (Note 5) The pre-trained model creation means is an information processing device according to Appendix 1 that, based on the data set for learning the identified issues, trains the pre-trained model so that when the document is input, it outputs the parts of the document that need correction as a check result.

[0061] (Note 6) A method for obtaining a new document, The information processing apparatus according to Appendix 1, comprising: a check result output means that outputs a check result obtained by inputting a new document into the aforementioned trained model.

[0062] (Note 7) Obtain the document, From the contents of the aforementioned document, we extract the descriptive features, which are the feature quantities. Based on the aforementioned features, a dataset for learning the content is created. Obtain the points raised in the aforementioned document, From the points raised in the aforementioned document, we extract the feature quantities known as "pointed features," Based on the aforementioned feature characteristics, a dataset for learning the content of the comments is created. An information processing method for creating a trained model that, based on the aforementioned data set for learning the content described and the aforementioned data set for learning the content of the comments, outputs the result of checking the document when the document is input.

[0063] (Note 8) Obtain the document, From the contents of the aforementioned document, we extract the descriptive features, which are the feature quantities. Based on the aforementioned features, a dataset for learning the content is created. Obtain the points raised in the aforementioned document, From the points raised in the aforementioned document, we extract the feature quantities known as "pointed features," Based on the aforementioned feature characteristics, a dataset for learning the content of the comments is created. A program that causes a computer to execute a process to create a trained model that, based on the aforementioned data set for learning the content described and the aforementioned data set for learning the points raised, outputs the check results for the document when the document is input.

[0064] Although the present invention has been described above with reference to embodiments and examples, the present invention is not limited to the above embodiments and examples. Various modifications to the configuration and details of the present invention can be understood by those skilled in the art within the scope of the present invention. [Explanation of symbols]

[0065] 1 User terminal 5 Network 11, 21 Interfaces 12, 22 processors 13, 23 memory 14, 24 Recording media 15, 25 Display section 16, 26 Input section 20 Analysis Servers 27. Contents Table DB 28. Table of Issues (DB) 29. Pre-trained model database

Claims

1. Document acquisition methods for obtaining documents, A feature extraction means for extracting feature quantities that correspond to the elements constituting the document from the contents of the document, A means for creating a data set for learning the content of a document, which uses the document as input data and associates the elements constituting the document as ground truth data, based on the aforementioned features, A means for obtaining the contents of the document mentioned above, A feature extraction means for extracting feature quantities that correspond to the items requiring correction and the items that should be described in place of the items requiring correction, based on the category of the items being pointed out, from the aforementioned points of concern. A means for creating a data set for learning feedback content, which uses the document as input data based on the aforementioned feedback features, and associates the category-specific feedback content information, including the action items and feedback items for each category corresponding to the document, with the feedback content information as ground truth data. A trained model learning means that includes both the aforementioned description content learning dataset and the aforementioned criticism content learning dataset as learning targets, takes the document as input data, learns based on the correct answer data corresponding to the description content learning dataset or the criticism content learning dataset, and creates a trained model that outputs a check result when the document is input. An information processing device equipped with the following features.

2. The information processing apparatus according to Claim 1, wherein the check result includes at least one of the following: deviations in the content of the document that are not included in the elements constituting the document, or the category-specific content information corresponding to the document.

3. The category-specific information on the points raised is a category-specific errata sheet, The aforementioned category is a type of expression that includes one or more of the following: symbols, diagrams, and text. The aforementioned errata sheet by category is a table in which the action items requiring correction are marked as incorrect, and the points that should be listed in place of those action items are marked as correct. The information processing apparatus according to claim 1, wherein the aforementioned feature extraction means extracts the feature from the category-specific errata.

4. The information processing apparatus according to claim 1, wherein the feature extraction means extracts the described feature for each element.

5. A method for obtaining a new document, The information processing apparatus according to claim 1, further comprising: a check result output means for outputting check results obtained by inputting a new document into the trained model;

6. An information processing method performed by a computer, Obtain the document, From the contents of the aforementioned document, we extract the descriptive features, which are the feature quantities corresponding to the elements that constitute the document. Based on the features described above, a data set for learning the content of the document is created using the document as input data and associating the elements constituting the document as ground truth data. Obtain the points raised in the aforementioned document, Based on the aforementioned findings, we extract the "feedback features," which are feature quantities corresponding to the corrective actions that need to be corrected and the actions that should be replaced with the corrective actions, categorized according to the category of the issue. Based on the aforementioned feature characteristics, a dataset for learning the content of the comments is created by using the document as input data and associating the category-specific comment content information, including the action items and comments for each category corresponding to the document, with the comment content information, as ground truth data. An information processing method for creating a trained model that includes both the aforementioned description content training dataset and the aforementioned criticism content training dataset as training targets, takes the document as input data, trains based on the correct answer data corresponding to the description content training dataset or the criticism content training dataset, and outputs a check result when the document is input.

7. Obtain the document, From the contents of the aforementioned document, we extract the descriptive features, which are the feature quantities corresponding to the elements that constitute the document. Based on the features described above, a data set for learning the content of the document is created using the document as input data and associating the elements constituting the document as ground truth data. Obtain the points raised in the aforementioned document, Based on the aforementioned findings, we extract the "feedback features," which are feature quantities corresponding to the corrective actions that need to be corrected and the actions that should be replaced with the corrective actions, categorized according to the category of the issue. Based on the aforementioned feature characteristics, a dataset for learning the content of the comments is created by using the document as input data and associating the category-specific comment content information, including the action items and comments for each category corresponding to the document, with the comment content information, as ground truth data. A program that causes a computer to execute a process to create a trained model that includes both the aforementioned description content training dataset and the aforementioned criticism content training dataset as training targets, takes the document as input data, is trained based on the correct answer data corresponding to the description content training dataset or the criticism content training dataset, and outputs a check result when the document is input.