Information processing device, information processing method, and program
The system addresses inconsistent document reviews by using machine learning to create a trained model for automated document checking, ensuring adherence to established writing practices and reducing reliance on individual expertise.
Patent Information
- Application Number
- JP2023191283
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2043-11-09
AI Technical Summary
Existing document review processes are inconsistent due to changes in personnel, relying heavily on individual expertise, leading to potential omissions and deviations from established writing styles and rules, and current AI systems fail to check for these deviations effectively.
An information processing system that utilizes machine learning to create a trained model by analyzing past documents and checklists, extracting features and pointed out contents, enabling automated checking for deviations and omissions in new documents.
Enables consistent and automated document quality assurance by identifying deviations and omissions without relying on individual expertise, ensuring adherence to established writing practices.
Smart Images

Figure 2025078946000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus for checking the contents of a document. [Background technology]
[0002] Documents such as specifications and design documents for specific customers often follow the writing methods and writing rules of documents created in the past. Therefore, when a document is created, a design review is conducted by the person who created the previous document, customers, sales, and other knowledgeable people, and the contents are examined and proofread as necessary. Patent Document 1 describes a system that displays a finished screen and an editing screen, allowing efficient proofreading while checking the finished image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2017-117125 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, it is difficult to have the same members perform document design reviews every time due to changes in responsibilities, transfers, retirements, etc. For example, if one expert retires, a replacement expert will participate, but because design reviews rely heavily on the expertise of the expert, the scope and depth of document checks will change depending on whether or not a replacement expert has the know-how. As a result, depending on the design review members, there are many cases of omissions or insufficient information, and the quality of the document depends on the person who checks it, which has been an issue.
[0005] In addition, while there are currently automated document checks using AI on the market, they are limited to checking for typos in general terminology and legal checks, and are not able to check new documents in accordance with the writing style and rules of past documents.
[0006] One of the objects of the present invention is to easily check new documents for deviations and omissions without relying on the skills of a particular individual. [Means for solving the problem]
[0007] In order to solve the above problem, in one aspect of the present invention, an information processing device includes: A document acquisition means for acquiring a document; A description feature extraction means for extracting description feature values, which are feature values, from the description contents of the document; A description content learning data set creation means for creating a description content learning data set based on the description feature amount; A pointed out content acquisition means for acquiring pointed out contents of the document; A feature extraction unit for extracting a feature from the content of the document; A data set creation means for creating a data set for learning about suggested content based on the suggested feature amount; The system also includes a trained model creation means for creating a trained model that is trained to output a check result for a document when the document is input based on the description content learning dataset and the pointed out content learning dataset.
[0008] In another aspect of the present invention, an information processing method includes: Get the document, Extracting description features from the description content of the document; A data set for learning the description content is created based on the description feature amount, Obtain the contents of the document, Extracting a feature value from the content of the document; A dataset for learning about the indicated content is created based on the indicated feature amount; A trained model is created that is trained to output the check results for a document when the document is input based on the description content learning dataset and the pointed out content learning dataset.
[0009] In another aspect of the invention, a program includes: Get the document, Extracting description features from the description content of the document; A data set for learning the description content is created based on the description feature amount, Obtain the contents of the document, Extracting a feature value from the content of the document; A dataset for learning about the indicated content is created based on the indicated feature amount; The computer executes a process of creating a trained model that is trained to output a check result for a document when the document is input based on the description content learning dataset and the pointed out content learning dataset. Effect of the Invention
[0010] According to the present invention, deviations and omissions in new documents can be easily checked without relying on the skills of a particular individual. [Brief description of the drawings]
[0011] [Figure 1] The configuration of the analysis system is shown. [Diagram 2] The hardware configuration of the user terminal and analysis server is shown below. [Diagram 3] The functional configuration of the analysis server during learning is shown below. [Figure 4] This is an example of a classification of past documents and their contents. [Diagram 5]An example of a checklist and categorized errata. [Figure 6] 13 is a flowchart of a learning process. [Figure 7] The functional configuration of the analysis server at the time of checking is shown. [Figure 8] 13 is a flowchart of a check process. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [Hardware configuration] 1 shows the configuration of an analysis system to which the analysis server of the present invention is applied. The analysis system 100 is a system that creates a trained model that has learned the characteristics of past documents, and checks for deviations and omissions in new documents (hereinafter also referred to as "new documents") using the trained model.
[0013] In this embodiment, the document is a specification or design document for a specific customer, and follows the writing practices and rules of documents created in the past. When a document is created, a design review is performed by experts such as the person who created the previous document, the customer, and sales, to check whether the created document deviates from the writing practices and rules. Then, a checklist is created that lists the points pointed out and corrected by the experts during the design review. The past documents and checklists may be stored in a user terminal used by the user who created the document, or may be stored in a specified DB.
[0014] In the analysis system 100, a user terminal 1 and an analysis server 20 are communicatively connected via a network 5 such as the Internet. The user terminal 1 is an information processing device that processes, stores, and transmits / receives various data, and is a personal computer (PC) or a general-purpose tablet used by a user. The user terminal 1 is a terminal that creates documents such as specifications and design documents for a specific customer, and creates checklists that list points raised about the created documents.
[0015] In this embodiment, for ease of explanation, one user terminal 1 is connected to the analysis server 20, but in reality, multiple user terminals are connected to the analysis server 20, such as a user terminal that stores past documents and checklists and a user terminal that has created a new document.
[0016] The analysis server 20 is an information processing device that processes, stores, and transmits / receives various data, and is, for example, a server device, a PC, or a general-purpose tablet. Although details will be described later, the analysis server 20 has a description content table database (hereinafter, the database is also referred to as "DB") 27, a pointed out content table DB 28, and a trained model DB 29. The analysis server 20 creates a trained model that has learned the characteristics of past documents. In addition, the analysis server 20 uses the created trained model to check deviations and omissions in new documents, and outputs the check results. Specifically, the analysis server 20 transmits the check results to the user terminal 1 to display them.
[0017] 2(a) is a block diagram showing the hardware configuration of the user terminal 1. As shown in the figure, the user terminal 1 includes an interface 11, a processor 12, a memory 13, a recording medium 14, a display unit 15, and an input unit 16.
[0018] The interface 11 transmits and receives data to and from the analysis server 20 via the network 5. The interface 11 is used when transmitting various data to the analysis server 20 and receiving various data from the analysis server 20.
[0019] The processor 12 is a computer such as a CPU (Central Processing Unit), and controls the entire user terminal 1 by executing a prepared program. The memory 13 is composed of a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The memory 13 stores the programs executed by the processor 12. The memory 13 is also used as a working memory while the processor 12 is executing various processes.
[0020] The recording medium 14 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or a semiconductor memory, and is configured to be detachable from the user terminal 1. The recording medium 14 records various programs executed by the processor 12. The display unit 15 is, for example, an LCD (Liquid Crystal Display) or the like, and displays various information. The input unit 16 is, for example, a touch panel, and is used by the user to input various information.
[0021] 2(b) is a block diagram showing the hardware configuration of the analysis server 20. As shown in the figure, the analysis server 20 includes an interface 21, a processor 22, a memory 23, a recording medium 24, a display unit 25, and an input unit 26. These components, a description content table DB 27, a pointed out content table DB 28, and a trained model DB 29 are interconnected via a bus.
[0022] The interface 21 transmits and receives data to and from the user terminal 1 via the network 5. The processor 22 is a computer such as a CPU, and controls the entire analysis server 20 by executing a program prepared in advance. The memory 23 is composed of a ROM, a RAM, etc. The memory 23 stores the programs executed by the processor 22. The memory 23 is also used as a working memory while the processor 22 is executing various processes.
[0023] The recording medium 24 is a non-volatile, non-transient recording medium such as a disk-shaped recording medium or a semiconductor memory, and is configured to be detachable from the analysis server 20. The recording medium 24 records various programs executed by the processor 22.
[0024] The display unit 25 is, for example, an LCD (Liquid Crystal Display) that displays a predetermined image. The input unit 26 is, for example, a keyboard, a mouse, a touch panel, etc., and is used by, for example, an operator who manages the analysis server 20.
[0025] The description content table DB27 stores a description content learning data set based on feature amounts extracted from the description content of a document.
[0026] The suggested content table DB28 stores a data set for learning suggested content based on feature amounts extracted from a checklist that lists suggested content of a document.
[0027] The trained model DB29 stores trained models created by the analysis server 20.
[0028] For ease of explanation, the analysis server 20 has a description content table DB27, a pointed out content table DB28, and a trained model DB29, but the present invention is not limited to this, and the type of DB and the data structure of the DB are arbitrary as long as the necessary data can be obtained.
[0029] [Function configuration during learning] The analysis server 20 creates a trained model by having a machine learning model learn the features of past documents. As a machine learning method, for example, a model using a neural network can be mentioned. FIG. 3 is a block diagram showing the functional configuration of the analysis server 20 during training. Functionally, the analysis server 20 includes a document acquisition unit 41, a description content learning unit 42, a checklist acquisition unit 43, and a pointed out content learning unit 44, and is realized by the processor 12 executing a program.
[0030] The document acquisition unit 41 acquires past documents from a predetermined DB or the like.
[0031] The description content learning unit 42 extracts description features, which are feature quantities, from the description contents of past documents, and creates a data set for learning the description contents. FIG. 4 is an example of classification of past documents and their description contents. As shown in FIG. 4, the description contents of past documents can be classified into multiple elements such as customer name, title, chapter structure, and notes. The chapter structure can be further classified into sentence units and figures and tables based on the contents described in each chapter. Also, the notes can be further classified into recommendation units and standard units required by each customer. Note that the notes can be classified not only by recommendation units and standard units, but also by special condition units for each customer.
[0032] The description content learning unit 42 analyzes the description content of past documents for each element and extracts description features. The description content learning unit 42 creates pairs of past documents and elements constituting the documents as a description content learning dataset based on the description features, and stores the pairs in the description content table DB 27. Specifically, the description content learning dataset is teacher data in which past documents are associated as input data and elements constituting the documents as correct answer data.
[0033] The description content learning unit 42 creates a plurality of description content learning data sets from a plurality of past documents, and creates a learned model by having a machine learning model learn the data sets. The created learned model is stored in the learned model DB 29. Specifically, the description content learning unit 42 trains the learned model so that when a new document is input, the described content learning unit 42 uses the description content learning data set to output elements that constitute the description content of the document. At this time, if there is a deviant part not included in the elements in the description content of the document, the learned model outputs the part as a check result. This allows the analysis server 20 to use the learned model to check whether the description content of the new document deviates from the description rules of past documents.
[0034] The checklist acquisition unit 43 acquires a checklist that lists the contents of the issues pointed out in the document acquired by the document acquisition unit 41 from a predetermined DB or the like.
[0035] The pointed out content learning unit 44 extracts pointed out feature quantities, which are feature quantities, from pointed out contents of past documents, and creates a pointed out content learning data set. FIG. 5 is an example of a checklist in which pointed out contents are listed, and an errata by category. As shown in FIG. 5, the checklist is composed of a pointed out location, a pointed out matter, and a treatment matter. The treatment matter is a character string that requires appropriate treatment, etc., that is pointed out by an expert in a design review of the document as being omitted or corrected. The pointed out matter is a correct character string that should be written instead of the treatment matter, and the pointed out location is a location in the document where the treatment matter and the pointed out matter exist. In this embodiment, the chapter in the document where the treatment matter and the pointed out matter exist is the pointed out location. The pointed out content learning unit 44 classifies the pointed out matters and treatment matters by category, such as symbols, diagrams, and sentences, based on the checklist of the past documents, and creates an errata by category as shown in FIG. 5.
[0036] The pointed out contents learning unit 44 analyzes the errata by category and extracts pointed out features. The pointed out contents learning unit 44 creates a pair of a past document and its errata by category as a pointed out contents learning data set, and stores it in the pointed out contents table DB 28. Specifically, the pointed out contents learning data set is training data in which the past document is associated as input data with the errata by category of the document as correct answer data.
[0037] The pointed out content learning unit 44 creates multiple pointed out content learning data sets from checklists of multiple past documents, and trains the trained model created by the description content learning unit 42. The trained trained model is stored in the trained model DB 29. Specifically, the pointed out content learning unit 44 trains the trained model using the pointed out content learning data sets so that when a new document is input, the trained model outputs a category-specific errata for the document as a check result. This allows the analysis server 20 to use the trained model to check whether there are any omissions or corrections in the description content of the new document.
[0038] In this embodiment, the description content learning unit 42 trains a trained model using a data set for learning description content, and then the comment content learning unit 44 trains the trained model using a data set for learning comment content. However, the present invention is not limited to this, and the order of learning can be set arbitrarily.
[0039] In the above configuration, the document acquisition unit 41 and the checklist acquisition unit 43 are examples of the document acquisition means and the pointed out content acquisition means of the present invention, respectively. The written content learning unit 42 is an example of the written feature extraction means, the written content learning dataset creation means, and the trained model creation means of the present invention. The pointed out content learning unit 44 is an example of the pointed out feature extraction means, the pointed out content learning dataset creation means, and the trained model creation means of the present invention.
[0040] [Learning process] Next, a description will be given of the learning process performed by the analysis server 20. Fig. 6 is a flowchart of the learning process performed by the analysis server 20. This process is realized by the processor 22 shown in Fig. 2(b) executing a program prepared in advance.
[0041] First, the analysis server 20 acquires past documents from a predetermined DB or the user terminal 1 (step S101). Then, the analysis server 20 extracts description features from the contents of the past documents (step S102) and creates a data set for learning the contents (step S103).
[0042] The analysis server 20 also acquires a checklist that lists the issues raised in the past documents acquired in step S101 from a predetermined DB or the user terminal 1 (step S104).Then, the analysis server 20 extracts issue features from the checklist (step S105) and creates a dataset for learning the issues raised (step S106).
[0043] The analysis server 20 creates a trained model by training the description content learning dataset and the indication content learning dataset, and stores the trained model in the trained model DB 29 (step S107). This ends the training process. When a document is input, the trained model created in this way can output, as a check result, any part of the description of the document that deviates from the description rules of past documents or that requires correction.
[0044] [Function configuration when checked] The analysis server 20 checks the contents of new documents using the created trained model. Fig. 7 is a block diagram showing the functional configuration of the analysis server 20 at the time of checking. The analysis server 20 functionally includes a document acquisition unit 51 and a check result output unit 52, and is realized by the processor 12 executing a program.
[0045] The document acquisition unit 51 acquires a new document from the user terminal 1 .
[0046] The check result output unit 52 uses the trained model stored in the trained model DB 29 to check whether there are any deviations or omissions in the contents of the new document, and outputs the check result.
[0047] Specifically, the check result output unit 52 inputs the new document into the trained model, and transmits the check result output from the trained model to the user terminal 1 for display. At this time, the trained model refers to the description content table DB27 and determines whether or not there is a past document corresponding to the description content of the new document. In other words, it determines whether or not a past document that followed the description rules when the new document was created is included in the past documents learned by the trained model. If the target past document is not included in the description content table DB27, it is impossible to perform a check using the trained model, and therefore the check result output unit 52 adds "check NG" to the check result.
[0048] In addition, the trained model refers to the pointed out content table DB 28 and determines whether there is a past document that corresponds to the description content of the new document. If there is no target past document in the pointed out content table DB 28, it is impossible to perform a check using the trained model, so the check result output unit 52 adds "check NG" to the check result.
[0049] In the above configuration, the document acquisition unit 51 and the check result output unit 52 are examples of a new document acquisition means and a check result output means, respectively, of the present invention.
[0050] [Check process] Next, a description will be given of the check processing by the analysis server 20. Figure 8 is a flowchart of the check processing by the analysis server 20. This processing is realized by the processor 22 shown in Figure 2(b) executing a program prepared in advance.
[0051] First, the analysis server 20 acquires a new document from the user terminal 1 and inputs it to the trained model (step S201). Then, the analysis server 20 acquires the check result of the new document output by the trained model and transmits it to the user terminal 1 (step S202). The user terminal 1 displays the check result. This ends the check process.
[0052] In this way, the analysis server 20 can extract features from the contents of documents such as past specifications and design documents for a specific customer and the comments made in design reviews of the documents, and machine-learns the features to realize automatic checking of new documents using a trained model. In this way, automatic checking by AI that has learned the features of past documents can be performed to check whether there are any deviations in the new document from the description rules and comments made in past documents. In addition, automatic checking by AI does not depend on a specific individual, so it is possible to prevent missed checks and inconsistencies in the check perspective due to personal checking, and to solve the problem of document quality being dependent on people.
[0053] [First Modification] In the above embodiment, the description content learning unit 42 and the pointed out content learning unit 44 have one machine learning model learn the description content learning data set and the pointed out content learning data set, respectively, but the present invention is not limited to this, and the description content learning unit 42 and the pointed out content learning unit 44 may learn different machine learning models to create two trained models. In this case, when a new document is input into one trained model, it is possible to check from the output check result which parts deviate from the description rules of past documents, and when a new document is input into the other trained model, it is possible to check from the check result which parts are missing or require correction, etc.
[0054] [Second modified example] In the above embodiment, the present invention is applied to documents for specific customers that require following the writing style of past documents and unique writing rules, but the present invention is not limited to this. It is also possible to realize automatic checking using a trained model for documents such as in-house design specifications, instruction manuals, contracts, etc. by training the model on checklists from past documents and design reviews.
[0055] A part or all of the above-described embodiments can be described as, but is not limited to, the following supplementary notes.
[0056] (Appendix 1) A document acquisition means for acquiring a document; A description feature extraction means for extracting description feature values, which are feature values, from the description contents of the document; A description content learning data set creation means for creating a description content learning data set based on the description feature amount; A pointed out content acquisition means for acquiring pointed out contents of the document; A feature extraction unit for extracting a feature from the content of the document; A data set creation means for creating a data set for learning about suggested content based on the suggested feature amount; A trained model creation means for creating a trained model that is trained to output a check result of the document when the document is input based on the description content learning dataset and the indication content learning dataset; An information processing device comprising:
[0057] (Appendix 2) 2. The information processing device according to claim 1, wherein the description feature extraction means classifies the description content into a plurality of elements and extracts description features for each element.
[0058] (Appendix 3) The information processing device according to claim 1, wherein the pointed out feature extraction means creates an errata sheet classified into a plurality of categories based on the pointed out contents, and extracts the pointed out feature from the errata sheet.
[0059] (Appendix 4) The information processing device described in Appendix 1, wherein the trained model creation means trains a trained model so that when the document is input, any part of the document that deviates from the description rules of past documents is output as a check result based on the description content learning dataset.
[0060] (Appendix 5) The information processing device described in Appendix 1, wherein the trained model creation means trains a trained model so that when the document is input, parts of the document that need to be corrected are output as check results based on the dataset for learning the indicated content.
[0061] (Appendix 6) A new document acquisition means for acquiring a new document; and a check result output means for outputting a check result output by inputting a new document into the trained model.
[0062] (Appendix 7) Get the document, Extracting description features from the description content of the document; A data set for learning the description content is created based on the description feature amount, Obtain the contents of the document, Extracting a feature value from the content of the document; A dataset for learning about the indicated content is created based on the indicated feature amount; An information processing method for creating a trained model that is trained to output a check result for a document when the document is input based on the description content learning dataset and the pointed out content learning dataset.
[0063] (Appendix 8) Get the document, Extracting description features from the description content of the document; A data set for learning the description content is created based on the description feature amount, Obtain the contents of the document, Extracting a feature value from the content of the document; A dataset for learning about the indicated content is created based on the indicated feature amount; A program that causes a computer to execute a process of creating a trained model that has been trained to output a check result for a document when the document is input based on the description content training dataset and the pointed out content training dataset.
[0064] Although the present invention has been described above with reference to the embodiments and examples, the present invention is not limited to the above-mentioned embodiments and examples. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. [Explanation of symbols]
[0065] 1 User terminal 5. Network 11, 21 Interface 12, 22 processors 13, 23 Memory 14, 24 Recording media 15, 25 Display section 16, 26 Input section 20 Analysis Server 27 Contents table DB 28 Point out contents table DB 29 Trained Model DB
Claims
1. A document acquisition means for acquiring a document; A description feature extraction means for extracting description feature values, which are feature values, from the description contents of the document; A description content learning data set creation means for creating a description content learning data set based on the description feature amount; A pointed out content acquisition means for acquiring pointed out contents of the document; A feature extraction unit for extracting a feature from the content of the document; A data set creation means for creating a data set for learning about suggested content based on the suggested feature amount; A trained model creation means for creating a trained model that is trained to output a check result of the document when the document is input based on the description content learning dataset and the indication content learning dataset; An information processing device comprising:
2. The information processing apparatus according to claim 1 , wherein the description feature extracting means classifies the description contents into a plurality of elements and extracts the description feature for each element.
3. 2 . The information processing apparatus according to claim 1 , wherein the pointed out feature extracting means creates an errata sheet classified into a plurality of categories based on the pointed out contents, and extracts the pointed out feature from the errata sheet.
4. The information processing device according to claim 1, wherein the trained model creation means trains a trained model so that when the document is input, any part of the document that deviates from the description rules of past documents is output as a check result based on the description content learning dataset.
5. The information processing device according to claim 1, wherein the trained model creation means trains a trained model so that when the document is input, parts of the document that need to be corrected are output as check results based on the indicated content learning dataset.
6. A new document acquisition means for acquiring a new document; The information processing device according to claim 1 , further comprising: a check result output means for outputting a check result output by inputting a new document to the trained model.
7. Get the document, Extracting description features from the description content of the document; A data set for learning the description content is created based on the description feature amount, Obtain the contents of the document, Extracting a feature value from the content of the document; A dataset for learning about the indicated content is created based on the indicated feature amount; An information processing method for creating a trained model that is trained to output a check result for a document when the document is input based on the description content learning dataset and the pointed out content learning dataset.
8. Get the document, Extracting description features from the description content of the document; A data set for learning the description content is created based on the description feature amount, Obtain the contents of the document, Extracting a feature value from the content of the document; A dataset for learning about the indicated content is created based on the indicated feature amount; A program that causes a computer to execute a process of creating a trained model that has been trained to output a check result for a document when the document is input based on the description content training dataset and the pointed out content training dataset.
Citation Information
Patent Citations
Inspection device, inspection method, program, and learning device
JP2019212115A
Learning apparatus, determination apparatus, learning method, determination method, learning program, and determination program
JP2022035703A
Information processing apparatus, information processing method, terminal program, server program, and contract correction support system
JP2022169992A
Ceramic particle mixtures containing coal combustion fly ash.
JP2022504880A
Information processing device, information processing method, and program
JP2023083926A