Information processing device, information processing method, and program
The information processing device efficiently checks document content conformity by analyzing and decomposing documents using a trained model, addressing inefficiencies in comparing non-standardized documents and reducing human error.
Patent Information
- Application Number
- JP2024056096
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-29
- Estimated Expiration
- 2044-03-29
AI Technical Summary
Existing systems struggle to efficiently determine whether the content of non-standardized documents conforms to required conditions, leading to potential human errors and inefficiencies in document comparisons.
An information processing device utilizing a trained model to analyze documents, perform task decomposition, and conduct matching and conformance determination, including search and code execution to verify content against predefined conditions.
Enhances the efficiency and accuracy of checking document content conformity to required conditions, reducing human error and improving the processing of non-standardized documents.
Smart Images

Figure 0007762249000001 
Figure 0007762249000002 
Figure 0007762249000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] Traditionally, document comparisons were performed visually by humans, but in recent years, as large volumes of electronic documents have been handled, the task has become more efficient using computers. However, when large volumes of documents are reviewed by humans, human error is likely to occur. Furthermore, while computer-based comparison of electronic documents is easy if the documents are in a standardized format, it is not easy to compare documents that do not have a standardized format. For example, such comparisons are not suitable for tasks such as checking whether the contents of specifications, which do not necessarily have a standardized format, comply with standards.
[0003] Patent Document 1 proposes a technology that calculates the semantic identity of each of multiple regulation contents in an organization's regulations for each standard content in a standard criteria, calculates an identity score for each of the regulation contents with the highest identity score, averages the identity scores between each standard content and its associated regulation content to calculate a score of agreement between the standard criteria and the organization's regulations, and displays to the user the agreement score for each standard content and the regulation content associated with the standard content.
[0004] Patent document 2 proposes a technology that decomposes a query document containing multiple data types into elements of different data types, performs a data type similarity search on the decomposed data type elements, and finds documents similar to the query document. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2020-86737 [Patent Document 2] Special Publication No. 2002-537604 Summary of the Invention [Problem to be solved by the invention]
[0006] There is a need for a system that can use an information processing device such as a computer to determine whether the content of a target document conforms to required conditions, regardless of its format.The present invention aims to provide an information processing device, an information processing method, and a program that can efficiently check whether the content of a target document conforms to required conditions. [Means for solving the problem]
[0007] The information processing device according to the present invention includes an acquisition means for acquiring a first document, an analysis means for extracting contents to be checked from the first document and generating a task for each of the contents, a search means for searching a second document for conditions related to the contents to be checked in the task, and a search means for searching the second document for conditions related to the contents to be checked in the task. and a condition regarding the content to be confirmed in the task searched by the search means. The present invention has a processing means for sequentially inputting the inputted contents to a trained model, and causing the trained model to infer and answer whether the contents to be confirmed in the inputted task conform to the conditions searched from the second document, and an output means for outputting the processing result by the processing means. When performing a numerical comparison according to the condition searched from the second document, the processing means executes a code related to the numerical comparison generated by the trained model and inputs the execution result to the trained model. It is characterized by: [Effects of the Invention]
[0008] According to the present invention, it is possible to provide an information processing device, an information processing method, and a program that can improve the efficiency of the task of checking whether the contents of a target document conform to required conditions. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an example of a hardware configuration of an information processing device. [Figure 2] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing device according to the first embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of processing performed by the information processing device according to the first embodiment. [Figure 4] FIG. 10 is a diagram illustrating an example of dividing a specification. [Figure 5] FIG. 10 is a diagram illustrating an example of task decomposition. [Figure 6] FIG. 10 is a diagram illustrating an example of a prompt in task decomposition. [Figure 7] FIG. 10 is a diagram illustrating an example of task decomposition. [Figure 8] FIG. 10 is a diagram illustrating an example of a prompt in task decomposition. [Figure 9] FIG. 10 is a diagram illustrating an example of a prompt in extracting prerequisite information. [Figure 10] FIG. 10 is a diagram illustrating an example of a standards document DB. [Figure 11] FIG. 10 is a diagram illustrating an example of a search using a search tool. [Figure 12] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 13] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 14] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 15] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 16A] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 16B] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 17A] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 17B] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 18] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 19] FIG. 10 is a diagram illustrating an example of matching processing. [Figure 20] FIG. 10 is a diagram illustrating an example of an output of a processing result. [Figure 21] FIG. 10 is a diagram illustrating an example of a functional configuration of an information processing device according to a second embodiment. [Figure 22] FIG. 10 is a diagram illustrating an example of processing performed by an information processing device according to a second embodiment. [Figure 23] FIG. 10 is a diagram illustrating an example of attribute information extraction. [Figure 24] FIG. 10 is a diagram illustrating an example of rule index construction. [Figure 25] FIG. 10 is a diagram illustrating processing in the second embodiment. [Figure 26A] FIG. 10 is a diagram illustrating an example of a sample relating to matching processing. [Figure 26B] FIG. 10 is a diagram illustrating an example of a sample relating to matching processing. [Figure 26C] FIG. 10 is a diagram illustrating an example of a sample relating to matching processing. [Figure 26D] FIG. 10 is a diagram illustrating an example of a sample relating to matching processing. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the following description, when a "document" is referred to, it is assumed that an electronic document (digital document) made up of digital data is also included.
[0011] First Embodiment 1 is a diagram showing an example of the hardware configuration of an information processing device 100 according to one embodiment of the present invention. The information processing device 100 includes a CPU 101, a ROM 102, a RAM 103, an auxiliary storage device 104, an output device 105, an input device 106, and a network I / F 107. The CPU 101, the ROM 102, the RAM 103, the auxiliary storage device 104, the output device 105, the input device 106, and the network I / F 107 are communicatively connected via a system bus 108.
[0012] The CPU (Central Processing Unit) 101 is a central processing unit that controls various operations of the information processing device 100. For example, the CPU 101 may control the operation of the entire information processing device 100. The ROM (Read Only Memory) 102 stores control programs, boot programs, etc. that can be executed by the CPU 101. The RAM (Random Access Memory) 103 is the main storage memory of the CPU 101 and is used as a work area or a temporary storage area for expanding various programs.
[0013] The auxiliary storage device 104 stores various data, various programs, etc. The auxiliary storage device 104 is realized by a storage device capable of temporarily or permanently storing various data, such as a non-volatile memory represented by an HDD (Hard Disk Drive) or an SSD (Solid State Drive).
[0014] The output device 105 is a device that outputs various types of information and is used to present various types of information to a user. For example, the output device 105 is realized by a display device such as a display. The output device 105 presents information to a user by displaying various types of display information. As another example, the output device 105 may be realized by an audio output device that outputs sounds such as voices and electronic sounds. In this case, the output device 105 presents information to a user by outputting sounds such as voices and electronic sounds. Furthermore, the device used as the output device 105 may be changed as appropriate depending on the medium used to present information to a user.
[0015] The input device 106 is used to receive various instructions from the user. For example, the input device 106 includes input devices such as a mouse, a keyboard, and a touch panel. As another example, the input device 106 may include a sound collection device such as a microphone, and collect the voice uttered by the user. In this case, various analysis processes such as acoustic analysis and natural language processing are performed on the collected voice, and the content indicated by this voice is recognized as an instruction from the user. Furthermore, the device used as the input device 106 may be changed as appropriate depending on the method for recognizing an instruction from the user. Furthermore, multiple types of devices may be used as the input device 106.
[0016] The network I / F 107 is used for communication with external devices etc. via a network. The device used as the network I / F 107 may be changed as appropriate depending on the type of communication path and the communication method used.
[0017] The CPU 101 loads a program stored in the ROM 102 or the auxiliary storage device 104 into the RAM 103 and executes the program, thereby realizing various functions and processes of the information processing device described below.
[0018] 2 is a diagram showing an example of the functional configuration of the information processing device 100 according to the first embodiment. The information processing device 100 includes an acquisition unit 201, an analysis unit 202, a database (DB) creation unit 203, an information extraction unit 204, a conformance determination unit 205, an output unit 209, and a storage unit 210.
[0019] Here, in the information processing device 100 of this embodiment, some of the functions are realized using a trained model 211 that is generated by machine learning (deep learning) and stored in the storage unit 210. In the following, the trained model 211 will be described as a large-scale language model (LLM). The large-scale language model (LLM) is a language model constructed using a large amount of text data (such as a large-scale corpus) and deep learning technology. When text data called a prompt indicating an instruction, command, or the like is input, the LLM performs inference based on the prompt, and generates and outputs text data corresponding to the input prompt. Note that the trained model 211 is not limited to the large-scale language model (LLM), and for example, a large-scale multimodal model (LMM) that can handle text data may also be applied.
[0020] The acquisition unit 201 acquires documents related to the conformity confirmation work. The acquisition unit 201 acquires a first document (document to be confirmed) to be confirmed whether it conforms to required conditions, and a second document indicating the conditions to be confirmed corresponding to the first document. The acquisition unit 201 is an example of an acquisition means. Examples of the first document include specifications for a product or service, a contract regarding the contents of the contract, a proposal, etc. Examples of the second document include documents such as specifications and regulations indicating standard rules, and an RFP (Request for Proposal). Note that the examples of the first document and second document described above are merely examples and are not limited to these.
[0021] The analysis unit 202 analyzes the first document acquired by the acquisition unit 201, extracts content from the first document to be checked for conformance with required conditions, and generates a matching task (hereinafter also simply referred to as a "task") for each extracted content. The analysis unit 202 is an example of an analysis means. The analysis unit 202 generates a task related to the checking of the first document by inputting the first document and a prompt including an instruction to extract content that needs to be checked and output a task to the trained model 211. The generated matching task (task) is stored in the storage unit 210 as task data 213. Note that the analysis unit 202 is not limited to generating a matching task (task) using the trained model 211. For example, the analysis unit 202 may extract content that needs to be checked from a document (first document) having a predetermined format by mechanical processing using an algorithm, and generate a matching task (task).
[0022] The DB creation unit 203 creates a database (retriever) related to documents based on the documents acquired by the acquisition unit 201. The database (retriever) created by the DB creation unit 203 is stored as a document DB (retriever) in the storage unit 210. That is, the DB creation unit 203 creates a first document DB (retriever) 215 based on the first document acquired by the acquisition unit 201, and creates a second document DB (retriever) 216 based on the second document acquired by the acquisition unit 201. These document DBs (retriever) 215, 216 are DBs (retrievers) as search tools used when determining relevance of document contents using the trained model 211. The document DBs (retriever) 215, 216 are, for example, full-text search engines, and searches are performed using BM25. The DB creation unit 203 is an example of a creation means.
[0023] The information extraction unit 204 extracts predetermined information from the first document acquired by the acquisition unit 201. The information extraction unit 204 is an example of information extraction means. In the first embodiment, the information extraction unit 204 extracts premise information from the first document acquired by the acquisition unit 201. Here, the premise information is information common throughout the first document, such as information related to a standard (rule) referenced in the first document. The information extraction unit 204 extracts premise information related to the first document by inputting the first document and a prompt including an instruction to extract common information throughout the first document to the trained model 211. The extracted premise information is stored in the storage unit 210 as premise information 214 and is provided to the trained model 211 as hint information in a conformance determination process for the content of the document using the trained model 211.
[0024] The conformance determination unit 205 performs conformance determination on the first document to be verified. Based on a task related to the first document stored as task data 213 in the storage unit 210, the conformance determination unit 205 confirms whether the content of the first document to be verified in the task conforms to required conditions. Specifically, the conformance determination unit 205 inputs the task and premise information related to the first document as a prompt into the trained model 211, thereby determining whether the content of the first document indicated in the task conforms to the conditions indicated in the second document. Furthermore, if the task to be verified in the conformance determination is not a single task that can be verified by the matching processing unit 206 (described later) but is composed of multiple tasks, i.e., multiple contents, the conformance determination unit 205 divides the task into verifiable tasks. The conformance determination unit 205 includes a matching processing unit 206, an information search unit 207, and a code execution unit 208.
[0025] The matching processor 206 performs matching processing between the first document and the second document, and determines whether the contents of the first document conform to required conditions. The matching processor 206 is an example of a processing means. The matching processor 206 inputs a prompt including a task and premise information related to the first document into the trained model 211, thereby inferring whether the contents of the first document indicated in the task conform to the conditions indicated in the second document. The matching processor 206 also outputs to the information search unit 207 and the code execution unit 208 a search request generated by inputting the task and premise information related to the first document as a prompt into the trained model 211, and a code (program) for performing calculation processing, etc.
[0026] The information searching unit 207 searches the first document and the second document in response to a search request from the matching processing unit 206, and outputs the search results to the matching processing unit 206. The information searching unit 207 is an example of a searching means. The information searching unit 207 receives a search query input from the matching processing unit 206, performs search processing on the first document DB 215 and the second document DB 216 using the trained model 211, and returns search results according to the input search query to the matching processing unit 206. Note that the information searching unit 207 is not limited to performing search processing using the trained model 211 and outputting search results. For example, the information searching unit 207 may receive a search query input from the matching processing unit 206, perform rule-based search processing on the first document DB 215 and the second document DB 216, and return the acquired search results to the matching processing unit 206. For example, in the search processing shown in FIG. 11, which will be described later, the search results 1106 obtained by the search processing 1104 may be output as the final search results 1108 as they are.
[0027] The code execution unit 208 executes the code (program) and outputs the execution result. The code execution unit 208 executes the code (program) input from the matching unit 206 and returns the execution result to the matching unit 206.
[0028] The output unit 209 outputs the result of the conformance determination for the first document made by the conformance determination unit 205, i.e., the determination result as to whether the content of the first document extracted as a task for the first document conforms to the required conditions. The output unit 209 may output the determination result as to whether the content of the first document conforms to the required conditions, together with the reason for the determination result.
[0029] The storage unit 210 stores various data and the like used when processing is performed in the information processing device 100. The storage unit 210 stores, for example, a trained model 211, a setting file 212, task data 213, premise information 214, a first document DB 215, and a second document DB 216. The trained model 211 is a model on which machine learning (deep learning) has been performed and is used for various inferences performed in the information processing device 100. The setting file 212 is a file in which prompts and the like to be input to the trained model 211 are stored. The setting file 212 may also store training samples to be input as prompts. A plurality of training samples may be prepared in the setting file 212, etc., and one may be selected and provided to the trained model 211 depending on the function to be realized. The task data 213 is a task related to a confirmation task generated by the analysis unit 202, and the premise information 214 is premise information extracted by the information extraction unit 204. The first document DB 215 is a DB for the first document created by the DB creating unit 203, and the second document DB 216 is a DB for the second document created by the DB creating unit 203.
[0030] Processing in the information processing device 100 in the first embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram illustrating an example of processing by the information processing device 100 in the first embodiment. In the following, as an example, a case will be described in which a first document to be checked is a product specification 301, a second document indicating the required conditions to be checked is a standard document 302 relating to the product standards, and whether or not the contents described in the specification 301 conform to the standard of the standard document 302 is checked.
[0031] The information processing device 100 acquires a specification 301 as a first document, and also acquires a standards document 302 as a second document.
[0032] (Processes 303 and 305: Dividing specifications and decomposing tasks) The information processing device 100 performs a specification division process 303 on the acquired specification 301, dividing the entire specification 301 into multiple documents 304 that become parts of the specification. In the specification division process 303, the information processing device 100 divides the specification 301 into multiple documents by mechanical processing using an algorithm so that each divided document has a size that can be processed by the trained model 211. For example, the information processing device 100 divides the text data of the specification 301 read by an OCR or PDF parser, etc., using a predetermined pattern based on numerical items, etc., to divide the specification 301 into multiple documents 304. In the example of dividing a specification shown in FIG. 4, the information processing device 100 divides the specification 401 into text data 402 to 405 for each item number (item) based on the numerical items. Tables in the specification are converted into data in CSV format, for example.
[0033] Furthermore, the information processing device 100 performs task decomposition processing 305 on the multiple documents 304 after division, and generates matching tasks (tasks) 306 related to content confirmation of the specifications 301. In the task decomposition processing 305, the information processing device 100 generates tasks 306 related to content confirmation of the specifications 301 by inputting the divided documents 304 and a prompt including an instruction to extract content that needs to be confirmed for compliance with standards and output a task to the trained model 211.
[0034] An example of task decomposition in the task decomposition process 305 will be described with reference to FIGS. An example of generating tasks for checking quantitative items will be described with reference to Figures 5 and 6. Figure 5(A) shows a description in a specification 301, in which the values of carbon (C), manganese (Mn), silicon (Si), and phosphorus (P) are listed for each product as chemical components of the product. Data 304 after division of the description shown in Figure 5(A) is the data shown in Figure 5(B). The information processing device 100 inputs the data shown in Figure 5(B) and the prompt 600 shown in Figure 6 into the trained model 211, thereby breaking down the contents that need to be checked into tasks and generating tasks as shown in Figure 5(C).
[0035] As shown in FIG. 6 , prompt 600 includes prompt 601, which instructs the user to comprehensively output values related to the standard and therefore need to confirm the consistency of the described values with the standard, and to break down the items to be checked into as small verifiable tasks as possible. Prompt 600 may also include training sample 602 for tabular data or training sample 603 for cases where the desired information to be extracted is not available. By using such prompts, as shown in FIG. 5(C), it is possible to control the granularity of the information (tasks) to be extracted and break them down into small tasks that can be verified by the matching process described below, thereby generating tasks that confirm whether the values described in specification 301 conform to the standard. Furthermore, by including training samples such as those shown in the example in prompt 600, it is possible to obtain task output in a desired format by taking advantage of the characteristics of large-scale language models, which have the high accuracy of returning output adapted to the training samples.
[0036] Next, an example of generating a task (matching task) for checking qualitative matters will be described with reference to FIGS. 7 and 8. FIG. 7(A) shows the description in the specification 301, which describes product quality control standards, product identification and traceability, and surface finish. The information processing device 100 inputs the segmented data for the quality control standards shown in FIG. 7(A) and the prompt 800 shown in FIG. 8 to the trained model 211, thereby breaking down the contents that need to be checked into tasks and generating a task (matching task) as shown in FIG. 7(B). Similarly, the trained model 211 inputs the segmented data for product identification and traceability shown in FIG. 7(A) and the prompt 800 shown in FIG. 8 to generate a task (matching task) as shown in FIG. 7(C), and the trained model 211 inputs the segmented data for surface finish shown in FIG. 7(A) and the prompt 800 shown in FIG. 8 to generate a task (matching task) as shown in FIG. 7(D).
[0037] As shown in FIG. 8, prompt 800 includes prompt 801, which instructs the user to output content related to the standard and therefore to check for consistency with the standard, and to break down the items to be checked into as small, verifiable tasks as possible. Prompt 800 may also include training samples 802 and 803 for extracting qualitative information, or training samples for cases where the desired information is not available. By using such prompts, as shown in FIGS. 7(B) to 7(D), it is possible to control the granularity of the information (tasks) to be extracted and break them down into small tasks that can be verified by the matching process described below, thereby generating tasks that verify whether the content described in specification 301 conforms to the standard. Furthermore, by including training samples such as those shown in the example in prompt 800, it is possible to obtain task output in a desired format by taking advantage of the characteristics of large-scale language models, which have the ability to accurately return output adapted to the training samples.
[0038] (Process 307: Extraction of prerequisite information) The information processing device 100 performs premise information extraction processing 307 on the acquired specification 301, and extracts premise information 308, which is common information, through the specification 301. In the premise information extraction processing 307, the information processing device 100 inputs the specification 301 and a prompt including an instruction to extract common information through the specification to the trained model 211, thereby generating premise information 308 related to the specification 301. In this example, as shown in FIG. 9 , by inputting a prompt 900 instructing to extract the standards and product types specified in the specification from the specification 301, information on the standards referenced in the specification, the product types described in the specification, etc. are extracted as premise information 308.
[0039] (Process 309, 311: DB creation) The information processing device 100 performs specification DB creation processing 309 based on the acquired specification 301, and creates specification DB (retriever) 310. Furthermore, the information processing device 100 performs standard document DB creation processing 311 based on the acquired standard document 302, and creates standard document DB (retriever) 312. The specification DB (retriever) 310 corresponds to the first document DB (retriever) 215 shown in Fig. 2, and the standard document DB (retriever) 312 corresponds to the second document DB (retriever) 216 shown in Fig. 2. In DB creation processing 309 and 311, the information processing device 100 processes the specification 301 and the standard document 302 into a searchable form by mechanical processing using an algorithm, and creates a database.
[0040] For example, as shown in FIG. 10 , the information processing device 100 creates a record for each page of the standard document 302, including a file name such as “Iron Standard 2023,” a document name such as “Steel Plate Standard,” a main text such as “1. Introduction This document is...,” and a page number such as “p1,” to construct the standard document DB (retriever) 312. For example, by including a page number in the record, it becomes possible to refer to a different page, such as the next page or the previous page, when performing a search using the trained model 211 described below. The specification DB (retriever) 310 may be configured in a similar manner. Note that the information contained in each record is not limited to the above example, and other information may also be included. Furthermore, while a record is created for each page of the document, this is merely an example, and a record may be created for each other structural unit.
[0041] Here, the standards document DB (retriever) 312 may be created in advance, or may be created when checking whether the contents of the specification 301 conform to the standard. For example, if a standards document DB (retriever) 312 is created and maintained for standards documents that are used repeatedly or over a long period of time, it is not necessary to create it each time, thereby reducing the amount of processing. Furthermore, for standards documents that cannot be permanently stored, creating the standards document DB (retriever) 312 when checking whether the contents of the specification 301 conform to the standard can prevent the occurrence of information leaks, etc.
[0042] (Process 320: Matching determination) For the tasks generated in the task decomposition process 305, the information processing device 100 performs conformance determination process 320 for each task using the trained model 211 to determine whether the contents described in the specification 301 conform to the standards in the standard document 302. Tasks 306 generated by the task decomposition process 305 are stored in a queue format in a task list 321. In the conformance determination process 320, the information processing device 100 performs matching process 323 using the trained model 211 for tasks 322 extracted one by one from the task list 321, and outputs a processing result 340. The matching process 323 is repeated until there are no more tasks in the task list 321.
[0043] In the matching process 323, the trained model 211 is made to simulate the human workflow when matching the contents of documents, a tool is used to collect information necessary for the matching, and semantic inference is performed based on the collected information, etc. to obtain processing results. In addition, when making the inference, a small number or a large number of workflow samples of a specific domain are trained (few-shot learning, etc.) in the trained model 211 to adapt it to the workflow domain.
[0044] The matching process 323 in this embodiment performs processing in accordance with a framework called ReAct (Reasoning and Acting) and outputs the processing results. That is, in accordance with the ReAct framework, the matching process 323 generates text for each phase of Thought, Action, and Observation in response to a Question, and derives a Final Answer by going through this Thought → Action → Observation flow one or more times.
[0045] In this example, Task 322 corresponds to Question. In Thought, a plan based on the current process to achieve the goal is considered using the trained model 211, and in Action, a tool to execute the plan is selected using the trained model 211. Then, input is given to execute the tool, and the execution results of the executed tool are observed in Observation. Based on the execution results of the tool obtained in Observation, a plan is considered in Thought, and if a plan is found, it proceeds to Action, and if a conclusion has been reached, a Final Answer is generated.
[0046] In this embodiment, the tools used by the trained model 211 in the matching process 323 include a task division tool 324, a specification search tool 326, a standards search tool 329, and a code execution tool 332. The task division tool 324 is realized by the conformance determination unit 205, the specification search tool 326 and the standards search tool 329 are realized by the information search unit 207, and the code execution tool 332 is realized by the code execution unit 208.
[0047] The task division tool 324 has a function of adding the tasks 325 divided by the matching process 323 to the task list 321. As will be described later, if the task 322 extracted from the task list 321 and input to the matching process 323 is not a single verifiable task but is composed of multiple tasks, i.e., multiple contents, the matching process 323 divides the task 322 into verifiable tasks 325. The task division tool 324 adds these multiple divided tasks to the task list 321.
[0048] The specification search tool 326 has a function of searching the specification 301 for detailed information necessary for verifying the task 322. The specification search tool 326 accepts a specification search query 327, searches the specification DB (Retriever) 310, and returns the information obtained by the search as specification information 328. The specification search tool 326 also has a function of extracting and summarizing relevant parts from the search results of the specification DB (Retriever) 310, and a function of adding an explanation of the extracted content. This function of summarizing the search results makes it possible, for example, to reduce the number of characters so that the search results can be given to the trained model 211 as a prompt.
[0049] The standards search tool 329 has a function of searching the standards document 302 for details of standards corresponding to the content described in the specification 301. The standards document search tool 329 accepts a standards document search query 330, searches the standards document DB (Retriever) 312, and returns the information obtained by the search as standards information 331. The standards document search tool 329 also has a function of extracting and summarizing relevant parts from the search results of the standards document DB (Retriever) 312, and a function of adding an explanation of the extracted content. This function of summarizing the search results makes it possible, for example, to reduce the number of characters so that the search results can be given to the trained model 211 as a prompt.
[0050] The code execution tool 332 has a function of executing code (program) generated by the trained model 211 for performing calculations such as numerical comparison (including unit conversion and arithmetic operations). The code execution tool 332 accepts code (program) 333 generated by the trained model 211, executes the code (program), and returns an execution result 334.
[0051] Here, a description will be given of searches using the search tools 326 and 329. Below, a search using the standard search tool 329 will be described as an example, but the same applies to searches using the specification search tool 326. Fig. 11 is a diagram for explaining an example of a search using the standard search tool 329.
[0052] When searching for specifications, the information processing device 100 generates a search intention 1101 and a search query 1103 based on the task to be confirmed and premise information. In the example shown in Fig. 11, the information processing device 100 generates a search intention 1101 of "I want to know the fatigue strength of IS360B-SSB" and a search query 1103 including "IronStandard2023", "steel plate", "IS360B-SSB", and "fatigue strength".
[0053] Next, the information processing device 100 executes a search process 1104 using the generated search query 1103. In the search process 1104, a rule-based search process is performed on standard document data 1105 using the search query 1103, and search results 1106 corresponding to the search query 1103 are output.
[0054] In the matching process described below, search results from the search tools 326 and 329 are added to the prompt and used for inference by the trained model 211. For example, if information on a specific row of a multi-row table is needed without an explanation and only the relevant portion is extracted, header information and the like will be lost, and only information listing words will be extracted. Even if only such information is input as search results to the trained model 211, it is difficult for the trained model 211 to interpret it, which may result in incorrect inferences or plans regarding the matching process, or inferences that lack planning. This is unlikely to occur when extracting information from documents containing only normal text, but can occur when extracting information from documents with layouts, such as specifications and standards documents, because they often use charts and tables.
[0055] Therefore, in this embodiment, the information processing device 100 executes a search result summarization process 1107 for summarizing the search results 1106 obtained by the search process 1104 based on the search intention 1101, and outputs the processing result as a final search result 1108. The search result summarization process 1107 is performed using the trained model 211. The search result summarization process 1107 extracts relevant sections (portions related to the search intention 1101) from the search results 1106 obtained by the search process 1104 and provides an explanation of the extracted content. For example, if target information is obtained with respect to the search intention 1101, the relevant sections and explanations are output; if target information is not obtained but reference information such as a table to be referenced is obtained, the information and explanations are output; and if target information is not obtained, the explanation "NO_OUTPUT" is output. The explanation is generated, for example, taking into consideration sentences written on one or more pages from which the information was obtained. For example, in the example shown in Figure 11, the first search result 1106 includes information about the fatigue strength of IS360A, IS360B, and IS340, but the information processing device 100 extracts only information about the fatigue strength of IS360B from the first search result 1106 in accordance with the search intent 1101 and outputs it as the final search result 1108 together with an explanation.
[0056] In this way, by performing search result summary processing 1107 on search results 1106 from search tools 326, 329 and outputting final search results 1108, appropriate information can be input into the trained model 211 by providing the search results and their explanatory information to the trained model 211 in the matching process, thereby improving the accuracy of inference and planning.
[0057] An example of matching processing will be described below. In the inference process of matching processing shown in Figures 12 to 19, Question is the task to be confirmed (input information for matching processing), Thought, Action, Action Input, and Final Answer are inferences made by the trained model 211, and Observation is the execution result of each tool.
[0058] As described above, in this embodiment, when performing inference in matching processing, the trained model 211 is trained (by few-shot learning, etc.) with workflow samples related to matching processing, and adapted to the domain related to matching processing. For example, training is performed by providing the trained model 211 with prompts including samples such as those shown in FIGS. 26A to 26D. When performing inference in matching processing, elements 2600A to 2600D shown in FIGS. 26A to 26D are combined, and a task is added to the end as a Question to perform inference. In FIG. 26A, 2601 is a rule related to matching processing, and 2602 is a sample relating to an example of searching for and matching standards and an example of matching involving numerical comparison. In FIG. 26B, 2603 is a sample relating to an example of searching for and matching standards and specifications and an example of matching involving numerical comparison. In FIG. 26C, 2604 is a sample relating to an example of matching using appropriate information in a multi-stage search. 26D, 2606 is a sample relating to an example of handling when a search fails, and 2607 is a sample relating to an example of dividing a task consisting of multiple tasks into single tasks. By providing a prompt including such a sample to the trained model 211, it is possible to execute processing related to the task in an appropriate flow according to the task.
[0059] FIG. 12 is a diagram illustrating an example of matching processing. FIG. 12 shows an inference process for an example in which standards are searched for and matching processing is performed to confirm quantitative items. In the example shown in FIG. 12, as shown in Question, a task 322 is input to confirm, "Does the description of the applicable size "10m" for product model number IS380A(-SSP) meet the Iron Standard restrictions?" In addition, information on standard candidates and steel type extracted as premise information 308 from specification 301 is input as additional information.
[0060] By inputting information including the task and additional information described in the Question as a prompt into the trained model 211, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, the information processing device 100 plans to investigate the applicable size of the relevant product in the standard document 302 (Thought), selects the standard search tool 329 (SearchStandard) as the tool to carry out this plan (Action), and generates information (arguments) to be input into the standard search tool 329 (Action Input). Here, standard_no, title, and sentence are information to be input into the search tool as a search query, and question is information to be input into the search tool as the search intent.
[0061] The standards search tool 329 searches the standards document DB 312 based on the information written in the Action Input, and the search results by the standards search tool 329, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). The standards search tool 329 performs a summary process of the search results as described above, and outputs the final search results, which are composed of, for example, related parts extracted from the search results for the standards document 302 and an explanation of the extracted content, as the search results by the standards search tool 329 (the same applies to other inference processes).
[0062] Next, by inputting the execution result of the tool described in Observation (search result of the standard document) as an additional prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, it plans to verify with the code (program) whether it satisfies the applicable size limit of the relevant product inferred from the description of the standard document 302 (Thought), selects the code execution tool 332 (RunCode) as the tool to execute it (Action), and generates code to be executed by the code execution tool 332 (Action Input).
[0063] The code (program) written in the Action Input is executed by the code execution tool 332, and the execution result (True) by the code execution tool 332, which is the execution result of the tool, is returned to the matching processing unit 206 (Observation). By additionally inputting the execution result of the tool (execution result of the code) written in the Observation as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. From this inference, it is determined from the verification result using the code that the applicable size of product model number IS380A satisfies the applicable size restrictions inferred from the description in the standard document 302 (Thought), and an answer indicating that various information related to the confirmation and the confirmation result are OK (compliant) is generated as a matching processing result (Final Answer).
[0064] Fig. 13 is a diagram for explaining an example of matching processing. Fig. 13 shows an inference process for an example in which standards are searched for and matching processing is performed to confirm quantitative items. In the example shown in Fig. 13, as shown in the Question, matching processing is performed for task 322, which confirms, "Does the description of the applicable size "9m" for product model number IS360B(-SSS) meet the restrictions of the Iron Standard?"
[0065] 13, in the same manner as in the example shown in FIG. 12, inference is performed using the trained model 211, processing is performed using the standards search tool 329 and the code execution tool 332, and the execution result (False) by the code execution tool 332 is returned to the matching processing unit 206. By additionally inputting the execution result by the code execution tool 332 as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. From this inference, it is determined from the verification result by the code that the applicable size of product model number IS360B does not satisfy the applicable size restrictions inferred from the description in the standards document 302 (Thought), and an answer indicating that various information related to the confirmation and the confirmation result are NG (not compliant) is generated as the matching processing result (Final Answer).
[0066] Fig. 14 is a diagram illustrating an example of matching processing. Fig. 14 shows an inference process for an example in which standards are searched for and matching processing is performed to confirm qualitative matters. In the example shown in Fig. 14, as shown in Question, a task 322 is input to confirm, "Is the surface pretreatment in the standard a pickling treatment?" In addition, information on standard candidates and steel type extracted as premise information 308 from the specification 301 is input as additional information.
[0067] By inputting information including the task and additional information described in these Questions as a prompt into the trained model 211, the information processing device 100 performs inference using the trained model 211 based on the information input as a prompt. Through this inference, the information processing device 100 plans to investigate whether the surface pretreatment in the standard document 302 is a pickling treatment (Thought), selects the standard search tool 329 (SearchStandard) as a tool to execute this (Action), and generates information (arguments) to be input into the standard search tool 329 (Action Input).
[0068] The standards search tool 329 searches the standards document DB 312 based on the information described in the Action Input, and the search results by the standards search tool 329, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). By additionally inputting the execution results of the tool described in Observation (search results of the standards document) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, it is determined that the surface pretreatment inferred from the description in the standards document 302 is pickling treatment and conforms to the standards (Thought), and an answer indicating that various information related to the confirmation and the confirmation result are OK (conforms) is generated as the matching processing result (Final Answer).
[0069] Fig. 15 is a diagram for explaining an example of matching processing. Fig. 15 shows an inference process for an example in which standards are searched for and matching processing is performed to confirm qualitative matters. In the example shown in Fig. 15, matching processing is performed for task 322, which confirms, as shown in the Question, "Is the final surface finish according to the standards a brush finish?"
[0070] 15, in the same manner as in the example shown in Fig. 14, inference is performed using the trained model 211, processing is performed using the standard search tool 329, and the execution result (search result) by the standard search tool 329 is returned to the matching processing unit 206. By additionally inputting the execution result by the standard search tool 329 as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, it is determined that the final finish of the surface inferred from the description in the standard document 302 is not a brush finish and does not conform to the standard (Thought), and various information related to the confirmation and an answer indicating that the confirmation result is NG (not conforming) are generated as the matching processing result (Final Answer).
[0071] 16A and 16B are diagrams illustrating an example of matching processing. FIGS. 16A and 16B show an inference process for an example in which quantitative items are searched for and checked against standards and specifications. In this example, as shown in the Question in FIG. 16A, a task 322 is input to confirm, "Does the description of the fatigue strength of IS380A (-SSP) '280≦F≦290' meet the Iron Standard's limits?" Additionally, information on the candidate standards and steel type extracted from the specification 301 as premise information 308 is input as additional information.
[0072] By inputting information including the tasks and additional information described in these Questions as prompts into the trained model 211, the information processing device 100 performs inference using the trained model 211 based on the information input as prompts. Through this inference, it plans to investigate the fatigue strength of the relevant product in the standard document 302 (Thought), selects the standard search tool 329 (SearchStandard) as a tool to carry out this plan (Action), and generates information (arguments) to be input into the standard search tool 329 (Action Input).
[0073] The standards search tool 329 searches the standards document DB 312 based on the information described in the Action Input, and the search results by the standards search tool 329, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). By additionally inputting the execution results of the tool described in Observation (search results of the standards document) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Because this inference indicates that whether the fatigue strength of the relevant product, estimated from the description in the standards document 302, is satisfied depends on the length of the product, it is planned to check the length of the product from the specification 301 (Thought), the specification search tool 326 (SearchDesignDocument) is selected as the tool to execute this (Action), and information (arguments) to be input to the specification search tool 326 is generated (Action Input).
[0074] The specification search tool 326 searches the specification DB 310 based on the information described in the Action Input, and as shown in FIG. 16B, the search results by the specification search tool 326, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). By additionally inputting the execution results of the tool described in Observation (specification search results) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, it is planned (Thought) to verify with code (program) whether the description of fatigue strength in the specification 301 satisfies the fatigue strength limit of the relevant product inferred from the description in the standard document 302, and the code execution tool 332 (RunCode) is selected as the tool to execute this (Action), and code to be executed by the code execution tool 332 is generated (Action Input).
[0075] The code (program) written in the Action Input is executed by the code execution tool 332, and the execution result (True) by the code execution tool 332, which is the execution result of the tool, is returned to the matching processing unit 206 (Observation). By additionally inputting the execution result of the tool (execution result of the code) written in the Observation as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. From this inference, it is determined from the verification result using the code that the fatigue strength of IS380A satisfies the fatigue strength limit estimated from the description in the standard document 302 (Thought), and an answer indicating that various information related to the confirmation and the confirmation result are OK (compliant) is generated as a matching processing result (Final Answer). In the example shown in FIGS. 16A and 16B, a task for checking quantitative items is exemplified, but the same processing is also performed for a task for checking qualitative items.
[0076] 17A and 17B are diagrams illustrating an example of matching processing. FIGS. 17A and 17B show an inference process for an example of performing matching processing and confirmation using appropriate information in a multi-stage search for quantitative items. In this example, as shown in the Question in FIG. 17A, a task 322 is input to confirm, "Does the carbon value of 0.20% in the chemical composition of IS380A (-SSP) meet the Iron Standard's restrictions?" Additionally, information on the candidate standard and type of steel extracted as premise information 308 from the specification 301 is input as additional information.
[0077] By inputting information including the task and additional information described in these Questions as a prompt into the trained model 211, the information processing device 100 performs inference using the trained model 211 based on the information input as a prompt. Through this inference, the information processing device 100 plans to look up the carbon value in the chemical composition of the relevant product in the standard document 302 (Thought), selects the standard search tool 329 (SearchStandard) as a tool to carry out this plan (Action), and generates information (arguments) to be input into the standard search tool 329 (Action Input).
[0078] The standards search tool 329 searches the standards document DB 312 based on the information described in the Action Input, and the search results by the standards search tool 329, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). By additionally inputting the execution results of the tool described in Observation (search results from the standards document) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, the carbon value in the chemical composition of the product in question could not be directly obtained from the standards document 302, but it is inferred from the description in the standards document 302 that the chemical composition is listed in Table 2. Therefore, a plan is made to search for the chemical composition in Table 2 of the standards document 302 (Thought), the standards search tool 329 (SearchStandard) is selected as the tool to execute this (Action), and information (arguments) to be input to the standards search tool 329 is generated (Action Input).
[0079] The specifications DB 310 is searched by the standards search tool 329 based on the information described in the Action Input, and as shown in FIG. 17B, the search results by the standards document search tool 329, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). By additionally inputting the execution results of the tool described in Observation (search results of the standards document) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, a plan is made to verify with code (program) whether the carbon value limit in the chemical composition of the relevant product inferred from the description in the standards document 302 is met (Thought), the code execution tool 332 (RunCode) is selected as the tool to execute this (Action), and code to be executed by the code execution tool 332 is generated (Action Input).
[0080] The code (program) described in the Action Input is executed by the code execution tool 332, and the execution result (True) by the code execution tool 332, which is the execution result of the tool, is returned to the matching processing unit 206 (Observation). By additionally inputting the execution result of the tool described in the Observation (the execution result of the code) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. From this inference, it is determined from the verification result using the code that the carbon value in the chemical composition of IS380A satisfies the restriction on the carbon value in the chemical composition estimated from the description in the standard document 302 (Thought), and an answer indicating that various information related to the confirmation and the confirmation result are OK (compliant) is generated as the matching processing result (Final Answer). In the example shown in FIGS. 17A and 17B, a task for checking quantitative items is exemplified, but the same processing is also performed for a task for checking qualitative items.
[0081] FIG. 18 is a diagram illustrating an example of matching processing. FIG. 18 shows an inference process for an example in which an input task containing multiple tasks is broken down into single tasks. In the example shown in FIG. 18, as shown in Question, task 322 is input to confirm whether the fatigue strengths of IS380A(-SSA) and IS380B(-SSS), "270≦F≦280" and "260≦F≦308," satisfy the Iron Standard restrictions. In addition, information on the candidate standards and steel type extracted from specification 301 as premise information 308 is input as additional information.
[0082] By inputting information including the tasks and additional information described in these Questions as prompts into the trained model 211, the information processing device 100 performs inference using the trained model 211 based on the information input as prompts. This inference determines that task decomposition is necessary, so the information processing device 100 plans to decompose the task and add it to the task list (Thought), selects the task division tool 324 (AddTask) as a tool to execute this (Action), and generates information to be input into the task division tool 324 (Action Input).
[0083] Based on the information written in the Action Input, the task division tool 324 adds each decomposed task to the task list 321, and the processing results by the task division tool 324, which are the execution results of the tool, are returned to the matching processor 206 (Observation). For example, the task shown in this example, which checks whether "the fatigue strengths '270≦F≦280' and '260≦F≦308' of IS380A(-SSA) and IS380B(-SSS) satisfy the limits of the Iron Standard," is decomposed into tasks which check whether "the fatigue strength '270≦F≦280' of IS380A(-SSA) satisfies the limits of the Iron Standard" and whether "the fatigue strength '260≦F≦308' of IS380B(-SSS) satisfies the limits of the Iron Standard," and these are added to the task list 321.
[0084] By inputting the execution result of the tool described in Observation as an additional prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. From this inference, it is determined that the task division tool 324 has successfully added the decomposed task to the task list 321 (Thought), and an answer indicating skipping is generated as a matching processing result (Final Answer).
[0085] FIG. 19 is a diagram illustrating an example of matching processing. FIG. 19 shows an inference process for an example of processing to suppress hallucination when a search fails. In the example shown in FIG. 19, as shown in Question, a task 322 is input to confirm whether the electrical resistance of IS380A(-SSA) "between 1.0×10^-7 and 1.5×10^-7 Ω·m" satisfies the Iron Standard's restrictions. In addition, information on the candidate standard and type of steel material extracted as premise information 308 from the specification 301 is input as additional information.
[0086] By inputting information including the task and additional information described in these Questions as prompts into the trained model 211, the information processing device 100 performs inference using the trained model 211 based on the information input as prompts. Through this inference, the information processing device 100 plans to check the electrical resistance of the relevant product in the standard document 302 (Thought), selects the standard search tool 329 (SearchStandard) as a tool to carry out this plan (Action), and generates information (arguments) to be input into the standard search tool 329 (Action Input).
[0087] The standards search tool 329 searches the standards document DB 312 based on the information described in the Action Input, and the search results by the standards search tool 329, which are the execution results of the tool, are returned to the matching processing unit 206 (Observation). By additionally inputting the execution results of the tool described in Observation (search results from the standards document) as a prompt, the information processing device 100 performs inference using the trained model 211 based on the information input as the prompt. Through this inference, it is determined that information about the electrical resistance of the relevant product cannot be obtained from the standards document 302, and it is not possible to confirm whether the description of the electrical resistance in the specification 301 satisfies the standard restrictions (Thought). As a result of the matching processing, various pieces of information related to the confirmation and an answer indicating that the confirmation result is NG (not compliant) are generated (Final Answer). In this way, when the target information cannot be obtained, "NO_OUTPUT" is output, and by explicitly indicating that the target information could not be obtained, it is possible to prevent hallucination from occurring.
[0088] 20(A) and 20(B) are diagrams showing examples of output of processing results in this embodiment. The processing results in the information processing device 100 described above are displayed as a list of verification contents and verification results for each verified task, for example, as in list 2001 shown in Fig. 20(A) and list 2011 shown in Fig. 20(B). Although list 2001 shown in Fig. 20(A) and list 2011 shown in Fig. 20(B) are configured to display all tasks that have been verified, it is also possible to display only tasks for which the verification result is NG (for which it was not determined that the description in the specification complies with the standard).
[0089] Furthermore, when a task for which the search result is NG is selected from the tasks displayed in the list by a user operation on the input device 106, the reason for the NG verification result may be displayed, for example, as in detailed report 2002 shown in FIG. 20(A) or detailed report 2012 shown in FIG. 20(B). In this example, detailed reports 2002 and 2012 are composed of a problem portion and a verified process portion, and can be generated based on an inference process in a matching process. Note that a GUI such as a button for instructing the display of a detailed report may be provided for a task for which the search result is NG, and the detailed report may be displayed in response to an operation on the GUI.
[0090] A detailed report 2002 shown in Figure 20(A) is a detailed report on the task whose inference process is shown in Figure 13, and a detailed report 2012 shown in Figure 20(B) is a detailed report on the task whose inference process is shown in Figure 19. For example, in the detailed report, the problem part is generated based on the Final Answer in the inference process, and the verified process part is generated based on the Thought, Action Input, Observation, etc. in the inference process. Also, in the verified process part, for example, a query is generated based on the Action Input in the inference process, a search result is generated based on the Observation in the inference process, and a consideration is generated based on the Thought in the inference process.
[0091] In the above example, the list of results and the detailed report showing the reasons why the verification result was NG are displayed separately, but the list of results and the detailed report may be displayed together. Furthermore, the information shown in the detailed report is not limited to the above example, and may include, for example, information about the tools used in the verification.
[0092] According to the first embodiment, the information processing device 100 extracts contents to be checked from a first document (e.g., a specification), generates a task for each content, sequentially inputs the generated tasks to the trained model 211, and causes the trained model 211 to respond as to whether the contents to be checked in the tasks conform to conditions retrieved from a second document (e.g., a standard document), and outputs the response as a processing result. This makes it possible to efficiently check whether the contents of the first document to be checked conform to the required conditions indicated in the second document.
[0093] In addition, hallucination generally occurs when generating information that requires knowledge outside the domain that has not been learned in the trained model. However, in this embodiment, when performing verification, information is searched for using an external tool (search tools 326, 329), and the search results are added to the prompt and given to the trained model that performs the matching process, thereby making it possible to prevent hallucination from occurring.
[0094] <Second embodiment> In the first embodiment, an example was shown in which a second document indicating the conditions to be confirmed can be identified from a first document to be confirmed, but there are cases in which it is difficult to identify a second document from a first document. In the second embodiment, an example in which it is difficult to identify a second document indicating the conditions to be confirmed from a first document to be confirmed will be described. The hardware configuration of the information processing device 2100 in the second embodiment is the same as the hardware configuration of the information processing device 100 in the first embodiment shown in FIG. 1, and therefore a description thereof will be omitted.
[0095] Fig. 21 is a diagram showing an example of the functional configuration of an information processing device 2100 according to the second embodiment. In Fig. 21, components having the same functions as those shown in Fig. 2 are denoted by the same reference numerals, and duplicated descriptions will be omitted. The information processing device 2100 includes an acquisition unit 201, an analysis unit 202, a DB creation unit 203, an information extraction unit 204, a match determination unit 205, an output unit 209, a storage unit 210, and a candidate selection unit 2101. The match determination unit 205 also includes a matching processing unit 206, an information search unit 207, and a code execution unit 208.
[0096] In the information processing device 2100, similarly to the information processing device 100 in the first embodiment, some of the functions are realized using a trained model 211 that is generated by machine learning (deep learning) and stored in the storage unit 210. The trained model 211 is, for example, a large-scale language model (LLM) or a large-scale multimodal model (LMM) that can handle text data.
[0097] In the second embodiment, the analysis unit 202 generates a matching task (task) related to the confirmation work of the contents of the first document, similarly to the first embodiment. The analysis unit 202 also analyzes one or more second documents acquired by the acquisition unit 201, extracts information representing the characteristics of the criteria (rules) described in each second document, and constructs a rule index for the second document. The analysis unit 202 extracts information representing the characteristics of the criteria (rules) from each second document by inputting the second document and a prompt including an instruction to extract and output the information representing the characteristics of the criteria (rules) to the trained model 211. The extracted information representing the characteristics of the criteria (rules) is linked to the second document from which it was extracted and stored in the storage unit 210 as a rule index 2103 for the second document. The second document may be divided into multiple documents in a predetermined pattern based on chapters or the like and input to the trained model 211.
[0098] In the second embodiment, the information extraction unit 204 extracts attribute information from a first document acquired by the acquisition unit 201. Here, the attribute information is information indicating characteristics of the first document, such as information that contributes to identifying a standard (rule) referenced in the first document. The information extraction unit 204 extracts attribute information about the first document by inputting the first document and a prompt including an instruction to extract information indicating the characteristics of the first document to the trained model 211. The extracted attribute information is stored as attribute information 2102 in the storage unit 210. The information extraction unit 204 divides the first document into multiple documents based on, for example, chapters, and inputs the documents to the trained model 211. Therefore, attribute information is extracted for each of the multiple documents divided based on chapters, etc.
[0099] The candidate selection unit 2101 selects a second document (candidate document) to be used to confirm the content of the first document from one or more second documents based on the attribute information 2102 extracted from the first document by the information extraction unit 204 and the rule index 2103 of the second document constructed by the analysis unit 202. The candidate selection unit 2101 is an example of a selection means. The candidate selection unit 2101 selects a second document to be used to confirm the content of the first document based on the similarity between the features of the first document and the features of the second document. For example, the candidate selection unit 2101 searches the rule index 2103 of the second document for each attribute information 2102 extracted from the first document, adds up the scores (feature similarity) related to the search results obtained for each attribute information 2102 for each second document, and selects the second document with the highest score as the candidate document. Note that the candidate selection unit 2101 is not limited to selecting one second document, and may select, for example, multiple second documents whose scores (feature similarity) exceed a threshold. For the second document selected by the candidate selection unit 2101 , the DB creation unit 203 creates a database (retriever) relating to the document, and stores it in the storage unit 210 as a second document DB (retriever) 216 .
[0100] Processing in the information processing device 2100 in the second embodiment will be described with reference to Fig. 22. Fig. 22 is a diagram illustrating an example of processing in the information processing device 2100 in the second embodiment. In the following, as an example, a case will be described in which the first document is a product specification 2201, the second document is a standard document 2202 relating to the product standard, and it is confirmed whether the contents described in the specification document 2201 conform to the standard of the standard document 2202.
[0101] The information processing device 2100 acquires a specification 2201 as a first document, and also acquires one or more standard documents 2202 as second documents.
[0102] (Processes 2203, 2205: Dividing specifications and decomposing tasks) The information processing device 2100 performs specification division processing 2203 on the acquired specification 2201, and divides the entire specification 2201 into multiple documents 2204 that become parts of the specification. In the specification division processing 2203, similar to the specification division processing 303 in the first embodiment, the information processing device 2100 divides the specification 2201 into multiple documents by mechanical processing using an algorithm so that each document after division has a size that can be processed by the trained model 211.
[0103] Furthermore, the information processing device 2100 performs task decomposition processing 2205 on the plurality of documents 2204 after division, and generates matching tasks (tasks) 2206 related to content confirmation of the specifications 2201. In the task decomposition processing 2205, similar to the task decomposition processing 305 in the first embodiment, the information processing device 2100 generates tasks 2206 related to content confirmation of the specifications 2201 by inputting the divided documents 2204 and a prompt including an instruction to extract content that needs to be confirmed for compliance with the standard and output a task to the trained model 211.
[0104] (Process 2207: Extract attribute information) The information processing device 2100 performs attribute information extraction processing 2207 on the acquired specification 2201, extracting attribute information indicating features of the first document that contribute to identifying the standard referenced in the first document. In the attribute information extraction processing 2207, the information processing device 2100 inputs, to the trained model 211, the specification 2201 divided into multiple documents based on chapters or the like, and a prompt including an instruction to extract information indicating features in the specification 2201. This generates attribute information 2208 related to the specification 2201. In this example, by inputting a prompt 2300 that instructs extracting information indicating the features of a product described in the specification from the specification 2201 as shown in FIG. 23(A), information that contributes to identifying the standard referenced in the specification as shown in FIG. 23(B) is extracted as attribute information 2208.
[0105] (Process 2209: Create DB of specifications) The information processing device 2100 performs DB creation processing 2209 for the specifications based on the acquired specifications 2201, and creates a specification DB (retriever) 2210. The specification DB (retriever) 2210 corresponds to the first document DB (retriever) 215 shown in Fig. 21. The information processing device 2100 processes the specifications 2201 into a searchable form by mechanical processing using an algorithm, and creates a database.
[0106] (Process 2211: Rule index construction) The information processing device 2100 performs a standard document rule index construction process 2211 on the acquired standard document 2202, extracts information representing the characteristics of the standard described in each standard document, and constructs a standard document rule index 2212. In the standard document rule index construction process 2211, the information processing device 2100 inputs, to the trained model 211, the standard document divided into multiple documents based on chapters or the like, and a prompt including an instruction to extract and output information representing the characteristics of the standard, thereby extracting information representing the characteristics of the standard from each standard document. For example, the information processing device 2100 inputs the divided standard document and a prompt 2400 shown in FIG. 24(A) to the trained model 211, thereby extracting information representing the characteristics of the standard from the standard document, as shown in FIG. 24(B). The information processing device 2100 constructs a standard document rule index 2212 by linking the extracted information representing the characteristics of the standard with the standard document. Prompt 2400 shown in FIG. 24(A) includes prompt 2401 that instructs the user to concisely output information representing the characteristics of the standard from the standard document, and training samples 2402 and 2403 that are to be trained by trained model 211.
[0107] (Process 2213: Extraction of candidate documents for standard documents) The information processing device 2100 performs candidate document extraction processing 2213 for standard documents to select standard documents 2202 as candidate documents to be used for confirming the contents of the specifications 2201. In the candidate document extraction processing 2213 for standard documents, the information processing device 2100 searches the rule index 2212 for the standard documents based on the attribute information 2208 extracted in the attribute information extraction processing 2207, and selects the standard documents 2202 with high scores (feature similarities) related to the search results as candidate documents to be used for confirming the contents of the specifications 2201. In addition, the information processing device 2100 creates a standard document DB (retriever) 2214 based on the selected candidate documents (selected standard documents 2202). The standard document DB (retriever) 2214 corresponds to the second document DB (retriever) 216 shown in FIG. 21 .
[0108] (Process 2220: Matching determination) For the tasks generated by the task decomposition process 2205, the information processing device 2100 performs conformance determination process 2220 for each task using the trained model 211, as in the first embodiment, to determine whether the contents of the specification 2201 conform to the standards of the standard document 2202. In the conformance determination process 220, the information processing device 2100 performs matching process 2223 using the trained model 211 for tasks 2222 extracted one by one from a task list 2221 in which the tasks 2206 generated by the task decomposition process 2205 are stored in a queue format, and outputs a processing result 2240. The matching process 2223 is performed until there are no more tasks in the task list 2221.
[0109] In the matching process 2223, as in the first embodiment, the trained model 211 is made to simulate the human workflow when matching the contents of documents, a tool is used to collect information necessary for matching, and semantic inference is performed based on the collected information, etc., to obtain processing results. In this embodiment, the matching process 2223 also performs processing following a framework called ReAct, and outputs processing results.
[0110] In this embodiment, the tools used by the trained model 211 in the matching process 2223 include a task division tool 2224, a specification search tool 2226, a standards search tool 2229, and a code execution tool 2232. The task division tool 2224 is realized by the conformance determination unit 205, the specification search tool 2226 and the standards search tool 2229 are realized by the information search unit 207, and the code execution tool 2232 is realized by the code execution unit 208. The task division tool 2224, the specification search tool 2226, the standards search tool 2229, and the code execution tool 2232 are similar to the task division tool 324, the specification search tool 326, the standards search tool 329, and the code execution tool 332 in the first embodiment, respectively, and therefore description thereof will be omitted.
[0111] 25, the information processing device 2100 performs extraction processing 2501 to extract content that needs to be checked for conformance with standards in the specification 2201, and generates a matching task (task) 2502. The information processing device 2100 performs extraction processing 2503 to extract attribute information that indicates characteristics of the specification 2201, and generates attribute information 2504 related to the specification 2201.
[0112] Furthermore, the information processing device 2100 performs extraction processing 2505 to extract information representing the characteristics of the standard from the standard document 2202, and constructs a standard document rule index 2506. In the standard document rule index 2506, the information representing the characteristics of the standard (conditional phrase) is associated with the standard document in which it is described (the rule to be followed).
[0113] The information processing device 2100 performs a search process 2507 on a rule index 2506 of the standard document based on the attribute information 2504, and selects a candidate document 2508 of the standard document to be used for confirming the contents of the specification 2201 based on the similarity of features obtained from the search results. Then, the information processing device 2100 performs a conformance determination process 2509 for each generated task 2502 using the trained model 211, determines whether the contents described in the specification 2201 conform to the standard of the standard document 2202 selected as the candidate document 2508, and outputs a processing result 2240. Note that if, as a result of the search for the standard document in this conformance determination process 2509, it is necessary to refer to another standard not included in the candidate document 2508, the search may be performed on other documents (for example, all standard documents) in addition to the selected candidate document 2508.
[0114] For example, if it is difficult to identify the second document from the first document to be verified, it is possible to perform a keyword search or embedded search by inputting the original text directly into a search engine. Keyword searches tend to produce noisy search results because they do not take semantics into account. Embedded searches are also expected to produce similar results based on meaning and context, but it can be difficult to understand why a particular search result was selected.
[0115] In contrast, in this embodiment, the information processing device 2100 selects a second document to be used for content confirmation of the first document based on the similarity between the features of the first document and the features of the second document obtained from search results using attribute information extracted from the first document and a rule index for the second document constructed by extracting information representing the features of criteria (rules) from the second document. By selecting a second document to be used for content confirmation of the first document in this manner, the selection of unrelated second documents as search results is reduced, making it possible to select second documents that are highly relevant as search results. Furthermore, by using words (technical terms) specialized for a specific domain as prompts provided to the trained model, it is possible to respond to domain-specific areas and enable specialized searches.
[0116] Furthermore, according to the second embodiment, even if it is difficult to identify the second document to be used for content confirmation from the first document to be confirmed, it is possible to select the second document according to the characteristics of the first document, thereby making it possible to efficiently confirm whether the content of the first document meets the required conditions.
[0117] In the first and second embodiments described above, the trained model 211 is configured to be stored in the storage unit 210 of the information processing device 100 (2100), but it may also be configured to be provided in another information processing device (server device, etc.) that can communicate with the information processing device 100 (2100) via a network. In this case, the information processing device 100 (2100) may transmit an input to the trained model 211 via the network I / F 107 or the like to the other information processing device that has the trained model 211, and receive the processing result for that input.
[0118] Furthermore, in the first and second embodiments described above, a prompt including a learning sample is given to train the trained model 211 (few-shot learning), but it is also possible to use a trained model that has been fine-tuned to suit each function.
[0119] It should be noted that the above-described embodiments are merely examples of specific embodiments of the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be embodied in various forms without departing from its technical concept or main features. [Explanation of symbols]
[0120] 100, 2100 Information processing equipment 101 CPU 102 ROM 103 RAM 104 Auxiliary storage 105 Output Device 106 Input Device 107 Network I / F 201 Acquisition Department 202 Analysis Department 203 DB creation department 204 Information extraction section 205 Compliance Determination Department 206 Matching Processing Section 207 Information Search Department 208 Code Execution Unit 209 Output section 210 Storage section 211 trained models 212 Configuration File 213 Task Data 214 Prerequisite information 215 First Document DB (Retriever) 216 Second Document DB (Retriever) 2101 Candidate Selection Unit 2102 Attribute information 2103 Rule Index
Claims
1. an acquisition means for acquiring a first document; an analysis means for extracting contents to be confirmed from the first document and generating a task for each of the contents; a search means for searching a second document for conditions related to the content to be confirmed in the task; a processing means for sequentially inputting the task generated by the analysis means and conditions related to the content to be confirmed in the task searched for by the search means into a trained model, and for causing the trained model to infer and respond as to whether or not the content to be confirmed in the input task conforms to the conditions searched for in the second document; output means for outputting a processing result by the processing means; The information processing device is characterized in that, when performing a numerical comparison according to the conditions searched from the second document, the processing means executes code related to the numerical comparison generated by the trained model and inputs the execution result into the trained model.
2. an information extraction means for generating premise information related to the first document by inputting a prompt including an instruction to extract premise information, which is information common throughout the first document, to a trained model, and extracting the premise information from the first document as the premise information referenced in the first document by inputting a prompt instructing to extract the premise information from the first document; 2. The information processing apparatus according to claim 1, wherein the search means searches the second document for the condition based on the content to be confirmed in the task and the premise information.
3. the search means searches the first document based on the condition searched from the second document; The information processing device according to claim 1, characterized in that, when the first document is searched for by the search means, the trained model responds as to whether the content to be confirmed in the task conforms to the conditions searched for in the second document, based on the content to be confirmed in the task, the search results of the first document by the search means, and the conditions searched for in the second document.
4. The information processing device according to claim 1, characterized in that, if the second document does not contain any conditions regarding the content to be confirmed in the task, the trained model responds that it cannot determine whether the content to be confirmed in the task conforms to the conditions retrieved from the second document.
5. a creating means for creating a database relating to the documents acquired by the acquiring means, the acquiring means acquires the first document and the second document; 2. The information processing apparatus according to claim 1, wherein said search means searches the database created by said creation means.
6. 2. The information processing apparatus according to claim 1, wherein said search means searches a database created based on said second document.
7. The information processing device according to claim 1 , wherein the trained model is a large-scale language model.
8. The information processing device according to claim 1 , wherein the trained model is a large-scale multimodal model.
9. the analysis means generates tasks for each of the contents by inputting the first document into a trained model; The information processing device according to claim 1 , wherein the search means searches for conditions related to the content to be confirmed in the task by inputting the content to be confirmed in the task into a trained model.
10. An information processing method executed by an information processing device, an acquisition step of acquiring a first document; an analysis step of extracting content to be confirmed from the first document and generating a task for each of the content; a search step of searching a second document for conditions related to the content to be confirmed in the task; a processing step of sequentially inputting the task generated in the analysis step and conditions related to the content to be confirmed in the task searched in the search step into a trained model, and having the trained model infer and respond as to whether the content to be confirmed in the input task conforms to the conditions searched from the second document; an output step of outputting a processing result in the processing step, An information processing method characterized in that, in the processing step, when a numerical comparison is performed according to the conditions searched from the second document, code related to the numerical comparison generated by the trained model is executed and the execution result is input into the trained model.
11. On the computer, an acquisition step of acquiring a first document; an analysis step of extracting content to be confirmed from the first document and generating a task for each of the content; a searching step of searching a second document for conditions related to the content to be confirmed in the task; a processing step of sequentially inputting the task generated in the analysis step and conditions related to the content to be confirmed in the task searched in the search step into a trained model, and having the trained model infer and respond as to whether or not the content to be confirmed in the input task conforms to the conditions searched from the second document; an output step for outputting a processing result in the processing step; A program for executing a process in which, when a numerical comparison is performed according to the conditions searched from the second document in the processing step, the code related to the numerical comparison generated by the trained model is executed and the execution results are input into the trained model.
Citation Information
Patent Citations
Sparse data processing method and device of neural network processor
CN115809683A
Project document review system and method based on artificial intelligence technology
CN116703337A
Document similarity search
JP2002537604A
Standard conformity support device and method thereof
JP2020086737A
Computer program, task generation device, and task generation method
JP2024160599A