Search result consultation apparatus, risk assessment recommendation information consultation apparatus, and search result consultation method

The optimization of language model combinations and prompts through a conferencing and model optimization unit addresses the inefficiencies in using multiple large-scale language models, enhancing accuracy and reducing costs in search result and risk assessment processes.

JP2026007047APending Publication Date: 2026-01-16HITACHI LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
JP2024106509
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

The use of multiple large-scale language models for improved output quality increases calculation time, and there is a lack of consideration for optimizing the roles and prompts to achieve both answer accuracy and resource cost efficiency.

Method used

A conferencing processing unit and model optimization unit are employed to optimize the combination of language models and prompts, evaluating and optimizing the number of models and prompts used for search results and risk assessment recommendation information.

Benefits of technology

This optimization allows for efficient and accurate integration of multiple language models, reducing calculation time and resource costs while maintaining high-quality output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026007047000001_ABST
    Figure 2026007047000001_ABST
Patent Text Reader

Abstract

To optimize a combination of a plurality of language models to be discussed and their prompts.SOLUTION: The recommended information consultation device 1 for risk assessment includes an answer generation unit 131 that answers a search result obtained by extracting and summarizing specific information based on a plurality of cases, a consultation processing unit 132 that consults the search result answered by the answer generation unit 131 with a predetermined number of language models and a predetermined prompt, and a model optimization unit 133 that evaluates a consultation result by the consultation processing unit 132 and optimizes the number of language models used by the consultation processing unit 132 and a combination of prompts.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a search result collaborating device, a risk assessment recommendation information collaborating device, and a search result collaborating method. [Background technology]

[0002] Large language models (LLMs) are language models built using extremely large datasets and deep learning technology. The name LLM comes from the fact that they are built using significantly increased computational effort, data volume, and number of parameters compared to conventional natural language models. LLMs are a type of generative AI and an artificial intelligence used in natural language processing, capable of learning large amounts of text data to perform tasks such as sentence generation and question answering. LLMs are attracting attention worldwide because they are capable of fluent, human-like conversation and can perform a variety of natural language processing tasks with high accuracy.

[0003] Non-Patent Document 1 describes that the quality of output data from a large-scale language model is improved by discussing multiple large-scale language models. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Mingxu Tao, Dongyan Zhao, Yansong Feng,"Chain-of-Discussion: A Multi-Model Framework for Complex Evidence-Based Question Answering", arXiv, :2402.16313,Internet<URL: https: / / arxiv.org / abs / 2402.16313> Summary of the Invention [Problem to be solved by the invention]

[0005] According to the invention described in Non-Patent Document 1, the quality of output data from a large-scale language model can be improved by using multiple large-scale language models for discussion. However, using multiple large-scale language models increases the amount of calculation, which may delay the time it takes to obtain an answer. Furthermore, no consideration has been given to how to combine the roles of respondents in large-scale language models to achieve both answer accuracy and resource cost scores.

[0006] Therefore, an object of the present invention is to optimize the combination of multiple collaborating language models and their prompts. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems, the search result conferencing device of the present invention is characterized by comprising: a conferencing processing unit that confer on search results obtained from a plurality of cases using a predetermined number of language models and predetermined prompts; and a model optimization unit that evaluates the conferencing results by the conferencing processing unit and optimizes the number of language models and combinations of prompts used by the conferencing processing unit.

[0008] The risk assessment recommendation information consultation device of the present invention is characterized by comprising a consultation processing unit that consults on recommendation information obtained from multiple risk-related cases using a predetermined number of language models and predetermined prompts, and a model optimization unit that evaluates the consultation results by the consultation processing unit and optimizes the number of language models and combinations of prompts used by the consultation processing unit.

[0009] The search result collaborating method of the present invention is characterized by comprising the steps of: a collaborating processor collaborating on search results obtained from a plurality of cases using a predetermined number of language models and predetermined prompts; and a model optimization unit evaluating the collaborating results by the collaborating processor and optimizing the number of language models and combinations of prompts used by the collaborating processor. Other means will be described in the detailed description of the invention. [Effects of the Invention]

[0010] The present invention allows for the optimization of the combination of collaborating language models and their prompts. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a configuration diagram of a recommendation information consultation device for risk assessment according to an embodiment of the present invention. [Figure 2A] 10 is a flowchart of a data conversion process. [Figure 2B] This is a prompt in the data conversion process. [Figure 3] FIG. 10 is a diagram illustrating data to be input into the data conversion process. [Figure 4] FIG. 10 is a diagram illustrating data output by the data conversion process. [Figure 5] 10 is a flowchart of a work information output process. [Figure 6] 10 is a flowchart of an answer generation process. [Figure 7] 10A and 10B are diagrams illustrating data selected by a user through work information input processing. [Figure 8] FIG. 10 is a diagram illustrating a prompt to be handed over to an information extraction unit. [Figure 9] FIG. 2 is a configuration diagram of a response generation unit. [Figure 10A] 10 is a flowchart of a process executed by an information extraction unit. [Figure 10B] This is a prompt that the information extraction unit issues to the language model. [Figure 11A] 10 is a flowchart of a process executed by a model optimization unit. [Figure 11B] This is a prompt that the model optimizer gives to the language model. [Figure 11C] This is a prompt that the model optimizer gives to the language model. [Figure 11D] This is a prompt that the model optimizer gives to the language model. [Figure 12] FIG. 10 is a diagram showing verification data. [Figure 13] FIG. 10 is a diagram showing the correct risk assessment result for the first task. [Figure 14A] 10 is a flowchart of a role distribution process executed by a collegial discussion processing unit. [Figure 14B] This is a prompt that the collegial processing unit issues to multiple language models. [Figure 14C] This is a prompt that the collegial processing unit instructs the first language model. [Figure 14D] This is a prompt that the collegial processing unit instructs the Nth language model. [Figure 15A] 10 is a flowchart of a role distribution process executed by a collegial discussion processing unit. [Figure 15B] This is a prompt issued by the collegial processing section in the role distribution process. [Figure 16] FIG. 10 is a diagram showing examples of risk assessment within a group. [Figure 17] FIG. 10 is a diagram showing examples of accidents outside the group. [Figure 18] FIG. 10 is a diagram showing a recommendation result display screen. [Figure 19] FIG. 10 is a diagram showing a recommendation result display screen. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. FIG. 1 is a configuration diagram of a risk assessment recommendation information consultation device 1 according to this embodiment. This recommendation information conferencing device 1 is an interactive computer device specialized in providing recommended information for risk assessment using a language model. The recommendation information conferencing device 1 includes a data conversion unit 11, an input unit 12, a search conferencing unit 13, and an output unit 14. MES (Manufacturing Execution System) information 51, process control sheet information 52, work procedure manual information 53, facility usage information 54, tool information 55, and site photo information 56 are input to the recommendation information conferencing device 1. Based on these inputs, the recommendation information conferencing device 1 outputs recommended information for risk assessment.

[0013] The data conversion unit 11 converts the MES information 51, the process management sheet information 52, the work procedure manual information 53, the facility information 54, the tool information 55, and the site photo information 56 into a data format that can be input to the language model.

[0014] The input unit 12 includes a work information database 121, a work information selection unit 122, and a work information output unit 123. The work information output unit 123 displays a work information selection screen to allow the user to select relevant information.

[0015] The search collaborating unit 13 includes a response generating unit 131 , information extracting units 136 and 137 , a collaborating processing unit 132 , a model optimizing unit 133 , an evaluating unit 134 , and an optimizing unit 135 .

[0016] The answer generation unit 131 extracts specific information from the question and answers the question based on the internal risk assessment case 21 via the information extraction unit 136. The answer generation unit 131 extracts search results for recommended information from the question based on multiple risk-related cases and answers the question. Here, the answer generation unit 131 presents hazard sources, risk estimation results, and countermeasure contents that are assumed from the input process, work name, work content, facilities used, and tools used, based on the in-group case information or out-of-group case information.

[0017] The information extraction unit 136 searches for internal risk assessment cases 21 using the language model 31. The information extraction unit 136 also includes a Retrieval-Augmented Generation (RAG) unit that uses the language model to extract specific information and return a summarized result. RAG is also known as retrieval augmentation generation. RAG is a model used in the field of natural language processing, and is a method for extracting related sentences through a sentence search and passing them as prompts to the language model.

[0018] In this embodiment, the use of RAG is not limited to this, and when the cases are in a table format, the information extraction unit 136 may execute only a search process that does not involve a language model, such as extracting a specific column from multiple cases and passing it to the answer generation unit 131, and presenting hazard sources, risk estimation results, and countermeasure contents that are assumed from the input process, work name, work content, facilities used, and tools used. The answer generation unit 131 searches the external accident cases 22 via the information extraction unit 137, and presents the hazards assumed from the input process, work name, work content, facilities used, and tools used, the risk estimation results, and countermeasures. The information extraction unit 137 uses the language model 32 to search the external accident cases 22, and presents the hazards assumed from the input process, work name, work content, facilities used, and tools used, the risk estimation results, and countermeasures.

[0019] The collegial processing unit 132 assigns roles to each language model 33 based on information about the roles. Furthermore, the collegial processing unit 132 collates the recommendation information, which is the answer generated by the answer generation unit 131, using multiple language models 33 and predetermined prompts. The model optimization unit 133 evaluates the collegial result using the evaluation unit 134 and optimizes the collegial result using the optimization unit 135. Then, the model optimization unit 133 determines whether to output the result. If a good answer has been obtained and it is determined that the result should be output, the optimized number of roles, prompts, and other results are output.

[0020] If a good answer is not obtained, the optimal number of LLM instances is calculated again, and the process returns to the role definition process. If the consultation result is not optimized, the model optimization unit 133 returns the number of roles N, a flag indicating whether or not to handle N roles with one instance, and prompts for N roles to the consultation processing unit 132, and the allocation of roles and consultation are repeated again.

[0021] The model optimization unit 133 evaluates the results of the discussion by the discussion processing unit 132 using an evaluation tool, and optimizes the number of language models 33 and the combination of prompts used by the discussion processing unit 132. Based on this, the model optimization unit 133 determines whether to output a result. If it determines that an answer with an evaluation value equal to or greater than a predetermined value has been obtained, the model optimization unit 133 outputs the optimized number of roles and the optimized prompts as results. If the evaluation value of the answer is less than the predetermined value, the process returns to calculating the optimal number of LLM instances and defining the roles.

[0022] The model optimization unit 133 includes an evaluation unit 134 that evaluates the results of the discussion by the discussion processing unit 132 based on verification data that matches questions with answers to those questions. Based on the evaluation of the discussion results by the evaluation unit 134, the optimization unit 135 optimizes the number of language models 33 that the discussion processing unit 132 will next use or the content of the prompt.

[0023] If the evaluation result by the evaluation unit 134 does not reach a predetermined level, the model optimization unit 133 causes the collegial processing unit 132 to collate using a number of language models 33 different from the predetermined number and / or a prompt different from the predetermined prompt. The predetermined prompt includes a description that defines the role of each language model 33. The collegial processing unit 132 may further use multiple types of language models, and is not limited to this. The collegial processing unit 132 and the model optimization unit 133 repeatedly optimize the number of language models 33 and the prompts until an output result of a predetermined quality is obtained with a predetermined amount of calculation.

[0024] The output unit 14 outputs the results of optimization by the model optimization unit 133 on the screen, and also displays the details of the discussion by the discussion processing unit 132.

[0025] FIG. 2A is a flowchart of the data conversion process. First, the data conversion unit 11 acquires documents of a given domain (step S10). Then, the data conversion unit 11 uses a large-scale language model for the documents of the given domain and extracts specified data from each piece of information using the prompt 61 in FIG. 2B (step S11). The data conversion unit 11 extracts and formats data from work information using, for example, a language model. Note that the data conversion unit 11 may extract and format data from work information using rule-based processing, and is not limited to this method. The formatted data is categorized into items such as process, work name, work content, facilities used, tools used, and energy parameters. The given domain documents and examples are converted using techniques such as embedding vectors and knowledge graphs, and then used for search, etc. Note that the embodiments of the present invention are not limited to these. Thereafter, the data conversion unit 11 stores the extracted data in the work information database 121 (step S12), and the processing of FIG. 2A ends.

[0026] FIG. 2B is a prompt 61 in the data conversion process. This prompt 61 includes a heading for "Sample_Original Data", a heading for "Sample_Converted Data", a heading for "Question", and a heading for "Original Data". Below the heading "Sample_Original Data" is a sample of the original data. Specifically, the original data is listed, separated by commas, such as "Task name, work details, start and end dates, person in charge, and work procedure manual to be used."

[0027] A sample of the converted data is listed under the heading "Sample_Converted Data." Specifically, a colon is written immediately after each item, followed by the corresponding content, such as "Process: xx, Work Name: xx, Work Content: xx, Facility Used: xx, Tool Used: xx, Energy Parameter: xx."

[0028] Below the question heading is the following string: "Please convert the following original data into the specified data based on the sample that will be converted from {#sample_original_data} to {#sample_converted_data}." Below the heading of the original data, the following strings are written: "Task name, work details, work start and end dates, person in charge, and work procedure manual to be used."

[0029] FIG. 3 is a diagram illustrating data to be input to the data conversion process. The first row of the table stores the process control table of the process control system, including the task name, task details, start and end dates, person in charge, and the work procedure manual to be used. The second row of the table stores the work procedure manual, which includes the work name, tools and facilities used, and the work target.

[0030] The third row of the table stores the facility document, which includes the facility name, facility layout, and energy parameters, including the platform height, voltage, temperature, weight, and speed.

[0031] The fourth row of the table stores the tool information document, which includes the tool name, size, weight, purpose, usage, and parts. The fifth row of the table stores a site photo, which includes the name and location of the object. The data conversion unit 11 utilizes a large-scale language model for documents in a given domain to extract specified data from each piece of information.

[0032] FIG. 4 is a diagram illustrating data output by the data conversion process. The extracted data of the process control table of the process control system is composed of segments of the process, the work name, the work content, the facilities used, the tools used, and the energy parameters. The extracted data of the work procedure manual is composed of segments each including the work name, the work content, the facilities used, the tools used, and the energy parameters.

[0033] The extracted data of the facility in use is composed of segments for the facility in use and the energy parameters. The extracted data of the tool information is configured by segmenting the tool used and the energy parameters. The extracted data of the site photographs is composed of segments of the facilities used, the tools used, and the energy parameters.

[0034] FIG. 5 is a flowchart of the work information output process. First, the work information output unit 123 lists the work information stored in the work information database 121 and displays it on the work information selection screen (step S20), thereby allowing the user to make a selection. When the task information selected by the user is passed to the answer generating unit 131 (step S21), the processing of FIG. 5 ends.

[0035] FIG. 6 is a flowchart of the answer generation process. First, answer generation unit 131 extracts K pieces of question-and-answer pair data from verification data 57 (step S30). Answer generation unit 131 passes prompts to information extraction units 136 and 137 based on verification data 57 and the task information passed from task information output unit 123 (step S31). An example of this prompt is shown in FIG. 8, which will be described later. When the responses from the information extraction units 136 and 137 are passed to the collegial processing unit 132 (step S32), the processing in FIG. 6 ends.

[0036] FIG. 7 is a diagram illustrating data selected by the user through the work information input process. The table in Fig. 7 has a selection column on the far left side of the table shown in Fig. 4. Check marks are placed in the second, third, and fourth lines of this selection column, which indicates that the user has selected the work procedure manual data, facility data, and tool data.

[0037] FIG. 8 is a diagram illustrating the prompt 62 that is passed to the information extraction units 136 and 137. As shown in FIG. The prompt 62 includes the referenceable information, the question, and the heading of the verification data. Below the heading of the available information, the following string is written: "Work name: xx, Work content: xx, Facility used: xx, Tool used: xx, Energy parameters: xx."

[0038] Below the question heading is the following string: "Please answer the following questions based on the above information. Based on the information on cases within the group and cases outside the group, please provide the anticipated hazards, risk estimates, and countermeasures based on the process, work name, work content, facilities used, and tools used."

[0039] Under the heading of verification data, verification data items 1 through K are listed. Each verification data item contains the specified work information, multiple assumed hazards, risk estimation results, countermeasures, and the final answer, which is the correct risk assessment result.

[0040] The work information output unit 123 displays the work information on the screen and allows the user to select it. Based on the selected work information, the answer generation unit 131 uses the information extraction units 136 and 137 to present the hazards assumed from the case information, the risk estimation results, and the details of countermeasures. Note that the answer generation unit 131 extracts K items of verification data from the verification data file to be used in the subsequent consultation process and passes them to the information extraction units 136 and 137.

[0041] FIG. 9 is a configuration diagram of the answer generation unit 131. The answer generation unit 131 includes an internal case search processing unit 1311 and an external case search processing unit 1312. The internal case search processing unit 1311 and the external case search processing unit 1312 selectively search for cases from the input verification data 57.

[0042] The answer generation engine of the answer generation unit 131 acquires K pairs of questions and answers and passes them to the internal case search processing unit 1311 or the external case search processing unit 1312. The prompt input unit 1310 inputs a prompt to the internal case search processing unit 1311 or the external case search processing unit 1312. The prompt contains, along with items of referable information, a question item, a character string such as "From the information on the cases within the group / outside the group, please provide the process, work name, work content, facilities used, tools used, and hazard sources assumed from energy parameters, risk estimation results, and countermeasure content," as well as items of verification data in the form of a character string.

[0043] Specifically, the internal case search processing unit 1311 includes an information extraction unit 136, an internal risk assessment case 21, and a search unit (not shown). The search unit searches for and acquires context related to the input from the internal risk assessment case 21 based on the verification data 57, and passes the search results to the language model 31. The language model 31 extracts cases from each reference destination using the input and context. As a result, the answer generation unit 131 outputs an answer to a subsequent processing unit.

[0044] The external case search processing unit 1312 includes an information extraction unit 137, external accident cases 22, and a search unit (not shown). The search unit searches for external accident cases 22 based on the verification data 57 and passes the search results to the language model 31. The language model 31 extracts cases from each reference destination using the input and context. As a result, the answer generation unit 131 outputs an answer to a subsequent processing unit.

[0045] FIG. 10A is a flowchart of the process executed by information extraction units 136 and 137. A case where the information extraction unit 136 is the main subject will be described. First, the information extraction unit 136 searches for and acquires relevant context from each case based on the referable information in the prompt passed from the answer generation unit 131 (step S40). Next, the information extraction unit 136 adds the acquired context to the prompt, passes the prompt to the language model 31 that extracts cases, and causes cases to be extracted (step S41). Finally, the information extraction unit 136 passes the answer from the language model 31 that extracts cases to the answer generation unit 131.

[0046] A case where the information extraction unit 137 is the main subject will be described. First, the information extraction unit 137 searches for and acquires relevant context from each case based on the referable information in the prompt passed from the answer generation unit 131 (step S40). Next, the information extraction unit 137 adds the acquired context to the prompt, passes the prompt to the language model 32 that extracts cases, and causes cases to be extracted (step S41). Finally, the information extraction unit 137 passes the answer from the language model 32 that extracts cases to the answer generation unit 131.

[0047] FIG. 10B shows a prompt 63 that information extraction units 136 and 137 instruct the language model. This prompt 63 includes the work information selected by the user through the work information input process shown in FIG. 7, the context acquired from the case, the question, and the verification data. The items of information that can be referenced include, for example, the type of information obtained from the verification data and the information in question. The item of context obtained from the case includes the type of information extracted from each case and the information. The question field contains the instructions to be given to the language model as a question. In this case, it is the following string: "Based on the information on cases within the group and cases outside the group, please provide the hazards assumed from the process, work name, work content, facilities used, and tools used, as well as the risk estimation results and countermeasures." The verification data items, from item 1 to item K, include the specified work information, multiple assumed hazard sources, risk estimation results, countermeasures, and final answers.

[0048] FIG. 11A is a flowchart of the process executed by the model optimization unit 133. First, the model optimization unit 133 provides an initial value for the number of roles, N (step S50). Then, the model optimization unit 133 checks a flag indicating that N roles are performed by one instance (step S51). Note that the initial value of the flag is False, and N roles are performed by N instances.

[0049] The model optimization unit 133 passes the prompt 65 shown in FIG. 11B to the language model 34, and the language model 34 generates N system prompts for assigning N roles to the language model 34 (step S52).

[0050] In step S52, if the loop processing is the second or subsequent time and the number of roles is reduced from the previous processing, the model optimization unit 133 stores the prompt for the previous role and randomly deletes one role. If the response when that role is deleted is inappropriate (no output is obtained), the deleted role is restored and another role is randomly deleted, and this process is repeated. The collegial processing unit 132 executes collegiality for all combination patterns when N roles are reduced to N-1 roles and calculates the Shapley value, which represents the contribution of each role. The role with the lowest Shapley value is deleted and subsequent processing continues. The Shapley value is an index that evaluates the importance of features in a machine learning model. The contribution of each feature is calculated and interpreted taking into account interactions and correlations. However, if N is large, the number of combinations of patterns resulting from changing the number of LLM instances becomes enormous. Therefore, it is necessary to start the optimization loop with a relatively small initial value, such as N=2.

[0051] Furthermore, in step S52, if the number of roles is increased from the previous processing, the N previously created prompts are included in the prompt of the instruction text when creating N+1. For the increased number of roles, an instruction is included in the prompt so that the roles become other than the N previously created prompts. An example of a prompt in this case is, "Roles for N roles have already been defined below. Please generate one system prompt to give one additional role to the LLM so that the roles do not overlap."

[0052] Returning to the flowchart, the explanation will be continued. The model optimization unit 133 provides information about roles to the collegial processing unit 132, and the collegial processing unit 132 passes a role prompt to the language model 33 (step S53).

[0053] Thereafter, the collegial processing unit 132 passes the prompt to the language model 33, and generates opinions for N roles for K verification questions (step S54). The model optimization unit 133 passes the prompt to the language model 34, and creates an answer that integrates the N opinions (step S55).

[0054] In step S55, the method for creating a response differs depending on the flag indicating whether or not N roles are performed in one instance. When N roles are performed by one instance, the process of integrating them into a final answer passes a prompt 66 shown in FIG. 11C to the language model 34 to create an answer. When N roles are performed by N instances, the process of integrating them into the final answer is (1) to (3) described below, but any other method that achieves the same effect can also be used.

[0055] (1) Take a majority vote on the opinions of the N roles and have them create an answer. (2) When N roles are being performed by N instances, another instance of a large-scale language model is prepared whose role is to create the final answer, and this instance is used to create the final answer. (3) After showing other views to the language model that created N views, it creates N sets of views ranked in order of support for these views, scores them using a rank fusion algorithm such as Reciprocal Rank Fusion, and creates a final answer from the best view. Reciprocal Rank Fusion is based on the concept of inverse rank, which is the inverse of the rank of the first relevant document in a list of search results. The goal of this technique is to give more importance to items that rank highly in multiple lists, taking into account the position of the item in the original ranking. This improves the overall quality and reliability of the final ranking and is useful for the task of fusing multiple ordered search results.

[0056] 11A, the model optimization unit 133 passes the prompt to the language model 34, and scores the answers to the K questions (step S56). An LLM instance is prepared to evaluate how closely the answers correspond to the correct risk assessment results, and scoring is performed automatically. The model optimization unit 133 then uses the LLM to evaluate the following elements. Evaluation is expected to be performed using RAGAS or similar. RAGAS is a framework for quantitatively evaluating the output results of RAG. The optimization engine is expected to be Optuna or similar.

[0057] When the model optimization unit 133 performs scoring using RAGAS, the score reflects the semantic similarity between the model answer and the answer generated by the AI. It is assumed that the semantic similarity of the answer (Answer Semantic Similarity) metrics will be used. Furthermore, the model optimization unit 133 may reflect the total costs of the collegial processing time, LLM usage fees, communication costs, machine resource usage fees, etc. in the score. In the case of the collegial processing time, the shorter the time, the higher the score should be. The lower the total costs of the LLM usage fees, communication costs, machine resource usage fees, etc., the higher the score should be.

[0058] ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is an evaluation metric for evaluating the quality of automatic summarization. The model optimization unit 133 calculates the overlap between the generated summary and the reference summary and evaluates the quality of the summary based on the recall. When the model optimization unit 133 performs scoring using ROUGE, it evaluates the similarity of the text between the model answer and the answer from the large-scale language model. During evaluation, accuracy is optimized and settings that are compatible with the resource cost score are selected based on the history. The resource cost score is ignored during the optimization process, and only the answer accuracy score is optimized. When the resource cost score is a constraint, the settings such as the number of LLMs and prompts that provide the best accuracy among those that satisfy the cost from the history of the optimization loop are output as the optimization result.

[0059] Here, the answer accuracy score can be the average of the values ​​of the proximity between the model answer and the answer generated by the AI ​​in N pieces of validation data. The model optimization unit 133 may optimize accuracy and select settings that are compatible with the resource cost score based on the history. Specifically, the model optimization unit 133 ignores resource considerations during the optimization process and optimizes only the accuracy score. Then, when the resource cost score is a constraint, the model optimization unit 133 may output, as the optimization result, the settings of the number of LLMs and prompts that provide the best accuracy among those that satisfy the cost requirement from the history of the optimization process loop.

[0060] For example, consider a case where the average accuracy in the first case is 0.80 and the average calculation time is 2 minutes, and the average accuracy in the second case is 0.75 and the average calculation time is 0.5 minutes. If it is desired to keep the average calculation time within 1 minute, the model optimization unit 133 adopts the number of LLMs and prompts in the second case.

[0061] 11A, the model optimization unit 133 determines the number N of roles to be tried next and a flag indicating whether each role should be performed with one instance, based on the average value of the K scores (step S57). The optimization unit 135 here is assumed to have an architecture that sequentially determines the parameters to be tried next from the transition of the objective variable corresponding to the parameters up to now, like Optuna's TPE (tree-structured Parzen estimator) algorithm.

[0062] Then, the model optimization unit 133 determines whether N and the flag have converged (step S88). If N and the flag have converged (Yes), the process proceeds to step S59. If N and the flag have not converged (No), the process returns to step S51.

[0063] One condition for convergence in step S88 is that the variance of the scores of the K questions and answers in the most recent Y loops is less than a threshold, but this is not limited to this in actual operation. In step S88, the optimization unit 135 determines whether to output the results. If a good answer has been obtained and it is determined that the results should be output, the model optimization unit 133 outputs the optimized number of roles, prompts, and other results. If a good answer has not been obtained and it is determined that the results should not be output, the model optimization unit 133 calculates the optimal number of LLM instances again and defines the roles, so the process returns to step S51. The process of automatically calculating the number of LLM instances to be discussed refers to the process of automatically calculating the number of roles N and whether the N roles should be performed by N different instances or by the same LLM instance. If the conventional optimum value no longer provides the accuracy of the answer expected by the user, new pairs of questions from the user and their answers may be stored in the verification data and the optimum value may be recalculated.

[0064] When the model optimization unit 133 outputs N, the flag, N prompts, the work content designated by the user, and the risk assessment results for the work content on the screen (step S59), the processing of FIG. 11 ends.

[0065] The model optimization unit 133 determines convergence based on a general convergence criterion in machine learning and mathematical optimization. For example, the model optimization unit 133 determines convergence when none of the values ​​of the score formula in the most recent predetermined number of trials have been able to update the best score. This is a standard method for determining convergence, and is often used in machine learning as a criterion for stopping learning earlier than the planned number of trials.

[0066] The model optimization unit 133 may further set a convergence criterion such that the first first predetermined number of attempts to calculate the score are always performed, and then a second predetermined number of attempts are performed, but the median score of the first predetermined number of attempts immediately before that is not exceeded when the convergence is determined. This is called Median Pruner, and is a convergence criterion implemented by default in Optuna.

[0067] To measure the proximity between the model answer and the answer generated by the language model, the model optimization unit 133 may implement a Rouge-based method that matches the words that appear in both documents and their order of appearance. The model optimization unit 133 may also implement a method that calculates the embedding vectors of both documents and finds the cosine value or distance between them. When finding the distance between them, the model optimization unit 133 corrects the original vector length by converting it to 1.

[0068] The model optimization unit 133 may measure the similarity between the model answer and the answer generated by the language model by inputting both documents into a large-scale language model and instructing the user to rate the similarity of the content on a scale of 1 to 10. In this case, it is common for the language model to show specific examples with scores ranging from the lowest to the highest.

[0069] FIG. 11B shows a prompt 65 that the model optimization unit 133 issues to the language model. Prompt 65 states, "For the specified task, multiple hazards, risk estimates, and countermeasures are presented, based on accident cases within the company and accident cases from the Ministry of Health, Labor, and Welfare. You must rank these to determine which hazards, risk estimates, and countermeasures are most appropriate. In order to hold this discussion, we would like to prepare N LLMs with different roles. Please generate N system prompts to assign N roles to the LLMs."

[0070] FIG. 11C shows a prompt 66 that the model optimization unit 133 issues to the language model. Prompt 66 states, "Based on the views reached during the discussion, please rank the most appropriate hazards, risk estimates, and countermeasures to create your final answer."

[0071] FIG. 11D shows a prompt 67 that the model optimization unit 133 issues to the language model. Prompt 67 contains the text "Please score the Final Answer using RAGAS using the following factors." Next, the heading "Scoring Factors" is written, and the scoring factors are listed.

[0072] 12 is a diagram showing the verification data 57. The user prepares data for which the correct risk assessment results for the task are known. This allows the model optimization unit 133 to obtain the optimal number of LLMs and the optimal prompt. The verification data 57 includes a number column, a task information column, a list of multiple assumed hazards, risk estimation results, and a list of countermeasure information. Based on the verification data 57, the prompt input unit 1310 creates a prompt and inputs it to the information extraction units 136 and 137.

[0073] FIG. 13 is a diagram showing the correct risk assessment result for the first task. The risk assessment result is composed of a priority column, a hazard column, a risk estimation result column, a countermeasure content column, and a reference document column. Using this risk assessment result as training data, the number of LLMs and prompts are optimized so that the recommendation results by the recommendation information consultation device 1 are close to the risk assessment result and are suitable.

[0074] FIG. 14A is a flowchart of the role allocation process executed by the collegial discussion processing unit 132. First, the collegial processing unit 132 obtains the number N of roles, a flag indicating whether or not the N roles are handled in one instance, and prompts for the N roles (step S60). Then, the collegial processing unit 132 passes the role prompts to the language models 33 to be collaborating based on the information from process 1 (step S61), and at that time, it determines whether to pass the prompts to one language model 33 or multiple language models 33 based on the contents of the flag. When the processing of step S61 ends, the processing of Fig. 14A ends.

[0075] FIG. 14B shows a prompt 681 that the collegial processing section issues to a plurality of language models. This prompt 681 includes a role prompt heading and a question heading. Under the role prompt heading is "Role 1:xxx, ..., Role N:xxx." The name of each role is listed after the role and a colon. Under the question heading it says, "You have the role listed in {#role}. Please assume that role and continue with the following steps."

[0076] FIG. 14C shows a prompt 682 that the collegial processing unit 132 instructs the first language model. This prompt 682 includes a role prompt heading and a question heading. Under the role prompt heading is "Role 1:xxx." The name of the role is listed after Role 1 and a colon. Under the question heading it says, "Please assume the role listed in {#role} and continue with the following steps."

[0077] FIG. 14D shows a prompt 683 that the collegial processing unit 132 instructs the Nth language model. This prompt 683 includes a role prompt heading and a question heading. Under the role prompt heading is "role N:xxx," where the name of the role is listed after role N and a colon. Under the question heading it says, "Please assume the role listed in {#role} and continue with the following steps."

[0078] FIG. 15A is a flowchart of the role allocation process executed by the collegial discussion processing unit 132. Here, the collegial processing unit 132 defines the viewpoint of the discussion. Specifically, the collegial processing unit 132 passes a prompt to the language model 33 and defines the viewpoint of the discussion when collaborating (step S70). Then, when the collegial processing unit 132 passes a prompt to the language model 33 and defines the viewpoint of the discussion when collaborating (step S71), the processing of FIG. 15A ends.

[0079] FIG. 15B shows a prompt 69 given by the collegial discussion processing unit 132 in the role distribution process. Prompt 69 reads, "Assuming you are in the role you were given, please discuss which is the best risk assessment example for the specified task. Then, please output the suggestions you have gained from the discussion in bullet points." This is followed by the headings of the verification data, the heading of the role prompt, the heading of the task information entered by the user, the heading of "hazards, risk estimation results, and countermeasures presented by the RAG application," the heading of the viewpoint of the discussion, and the heading of the constraints for executing Process 2.

[0080] FIG. 16 is a diagram showing an internal risk assessment example 21 within a group. This internal risk assessment example 21 includes a work name column, a work location classification column, a process classification column, an equipment classification column, a work target product column, a work location column, a high-risk work classification column, a hazard column, a hazard energy parameter column, a countermeasure column, and a risk assessment column. The answer generation unit 131 uses the information extraction unit 136 and the language model 31 to present the hazards, risk estimation results, and countermeasure contents assumed based on the internal risk assessment example 21 from the input process, work name, work content, facilities used, and tools used.

[0081] FIG. 17 is a diagram showing an external accident case 22 outside the group. The external accident case 22 includes an industry column, a column for the type of machinery, equipment, or hazardous substance, a column for the type of disaster, a column for the number of victims, a column for the cause of human casualties, a column for the cause of material casualties, a column for the occurrence status, a column for the cause, and a column for countermeasures. The answer generation unit 131 uses the information extraction unit 137 and the language model 31 to present the assumed hazard sources, risk estimation results, and countermeasures based on the external accident case 22 from the input process, work name, work content, facilities used, and tools used.

[0082] FIG. 18 is a diagram showing the recommendation result display screen 40. As shown in FIG. The recommendation result display screen 40 includes an optimum number text box 401, the number of LLM instances 402, and a role details table 403. The optimal number of roles for the collaborative process is displayed in the optimal number text box 401. The optimal number of LLM instances is displayed in the number of LLM instances 402. The optimal combination of roles assigned to the LLM is displayed in the role details table 403. This allows the user to confirm the optimized number of LLM instances and the optimal combination of roles.

[0083] FIG. 19 is a diagram showing the recommendation result display screen 41. As shown in FIG. The recommendation result display screen 41 includes a work content table 411 and a risk assessment result table 412 .

[0084] The recommendation information consultation device 1 optimizes the number of LLM instances, role prompts, etc. using the verification data 57 shown in Figure 9, then extracts cases using the content entered by the user under the optimized conditions, then consults with them to create a final answer, and displays the final answer together with the optimization conditions on the recommendation result display screen 41 in Figure 19.

[0085] The configuration and effects of the present invention will be described below.

[0086] [1] a collation processing unit (132) that collates search results obtained from a plurality of cases using a predetermined number of language models and a predetermined prompt; a model optimization unit (133) that evaluates a result of the collaboration by the collaboration processing unit (132) and optimizes the number of language models and combinations of prompts used by the collaboration processing unit (132); A search result collaborating device (1) comprising:

[0087] [2] an answer generation unit (131) that extracts and summarizes specific information based on multiple cases; 2. The search result collaborating device according to claim 1, further comprising:

[0088] This allows the search result collaborating device to optimize the combination of multiple collaborating language models and their prompts.

[0089] [3] the answer generation unit (131) extracts relevant cases from the plurality of case information using a language model and returns a summarized search result; 3. The search result collaborating device according to claim 2.

[0090] This makes it possible to easily extract relevant cases from multiple pieces of case information.

[0091] [4] The model optimization unit (133) includes an evaluation unit (134) that evaluates a result of the discussion by the discussion processing unit (132) based on verification data in which questions and answers to the questions correspond to each other. 2. The search result collaborating device according to claim 1.

[0092] This makes it possible to evaluate highly rated discussion results and extract conditions that will yield highly rated discussion results.

[0093] [5] The model optimization unit (133) determines an optimized number of language models and prompts when the evaluation result by the evaluation unit (134) reaches a predetermined level. 5. The search result collaborating device according to claim 4.

[0094] This makes it possible to evaluate highly rated discussion results and extract conditions that will yield highly rated discussion results.

[0095] [6] If the evaluation result by the evaluation unit (134) does not reach a predetermined level, the model optimization unit (133) causes the collaborating processing unit (132) to collaborate using a number of language models different from the predetermined number and / or a prompt different from the predetermined prompt. 5. The search result collaborating device according to claim 4.

[0096] This makes it possible to evaluate the results of discussions based on a large number of language model patterns and prompt patterns.

[0097] [7] the predetermined prompt includes a description defining a role of each of the language models; 2. The search result collaborating device according to claim 1.

[0098] This allows for defining suitable roles for the language model.

[0099] [8] The collegial processing unit (132) further uses a plurality of types of language models. 2. The search result collaborating device according to claim 1.

[0100] This makes it possible to select from a plurality of language models one that will provide a suitable answer.

[0101] [9] a collation processing unit (132) that collates recommendation information obtained from a plurality of risk-related cases using a predetermined number of language models and a predetermined prompt; a model optimization unit (133) that evaluates a result of the collaboration by the collaboration processing unit (132) and optimizes the number of language models and combinations of prompts used by the collaboration processing unit (132); A risk assessment recommendation information consultation device (1) comprising:

[0102] This enables the risk assessment recommendation information conferencing device to optimize the combination of multiple language models to be conferred and their prompts.

[0103]

[10] A step in which a collation processing unit (132) collates search results obtained from a plurality of cases using a predetermined number of language models and a predetermined prompt; a model optimization unit (133) evaluating a result of the collaboration by the collaboration processing unit (132) and optimizing the number of language models and combinations of prompts used by the collaboration processing unit (132); A search result collaborating method comprising:

[0104] As a result, the search result collaborating method makes it possible to optimize the combination of collaborating language models and their prompts.

[0105] <<Variation>> The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. It is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is also possible to add, delete, or replace part of the configuration of each embodiment with other configurations.

[0106] The above-described configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware such as an integrated circuit. The above-described configurations, functions, etc. may be realized by software by a processor interpreting and executing a program that realizes each function. Information such as the programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or on a storage medium such as a flash memory card or a DVD (Digital Versatile Disk).

[0107] In each embodiment, the control lines and information lines shown are those that are considered necessary for the explanation, and not all control lines and information lines in the product are necessarily shown. In reality, it can be considered that almost all components are interconnected. As modified examples of the present invention, for example, the following (a) and (b) are available.

[0108] (a) The large-scale language model called by the recommendation information consultation apparatus is not limited to one built in an on-premise environment, and may be one built in a cloud environment. (b) The large-scale language models called by the recommendation information consultation device are not limited to assigning different roles to the same type of model, but different types may be selected, and different roles may be assigned to different types of large-scale language models. [Explanation of symbols]

[0109] 1. Recommendation information conferencing device (search result conferencing device) 11 Data conversion section 12 Input section 13 Search Council 14 Output section 51 MES Information 52 Process control table information 53 Work procedure information 54 Facility Information 55 Tool information 56 On-site photo information 121 Work Information Database 122 Work information selection section 123 Work information output section 131 Answer generation part 136 Information extraction part 137 Information extraction part 132 Collegial Processing Section 133 Model Optimization Department 134 Evaluation Department 135 Optimization Department 21 Internal Risk Assessment Examples 31 Language Models 22 External Accident Cases 32 language models 33 Language Models 1311 Internal case search processing unit 1312 External case search processing unit 57 Verification data

Claims

1. a collation processing unit that collates search results obtained from a plurality of cases using a predetermined number of language models and a predetermined prompt; a model optimization unit that evaluates a result of the collaboration by the collaboration processing unit and optimizes the number of language models and combinations of prompts used by the collaboration processing unit; A search result collaborating device comprising:

2. an answer generation unit that extracts and summarizes specific information based on multiple cases; The search result collaborating device according to claim 1, further comprising:

3. the answer generation unit extracts relevant cases from a plurality of case information using a language model and returns a summarized search result; 3. The search result collaborating device according to claim 2.

4. the model optimization unit includes an evaluation unit that evaluates a result of the discussion by the discussion processing unit based on verification data in which questions and answers to the questions correspond to each other; 2. The search result collaborating device according to claim 1.

5. the model optimization unit determines an optimized number of language models and prompts when the evaluation result by the evaluation unit reaches a predetermined level.

5. The search result collaborating device according to claim 4.

6. If the evaluation result by the evaluation unit does not reach a predetermined level, the model optimization unit causes the collaborating processing unit to collaborate using a number of language models different from the predetermined number and / or a prompt different from the predetermined prompt.

5. The search result collaborating device according to claim 4.

7. the predetermined prompt includes a description defining a role of each of the language models; 2. The search result collaborating device according to claim 1.

8. The collegial processing unit further uses a plurality of types of language models.

2. The search result collaborating device according to claim 1.

9. a collation processing unit that collates recommendation information obtained from a plurality of risk-related cases using a predetermined number of language models and a predetermined prompt; a model optimization unit that evaluates a result of the collaboration by the collaboration processing unit and optimizes the number of language models and combinations of prompts used by the collaboration processing unit; A risk assessment recommendation information consultation device comprising:

10. a step in which a collation processing unit collates search results obtained from a plurality of cases using a predetermined number of language models and a predetermined prompt; a model optimization unit evaluating a result of the collaboration by the collaboration processing unit and optimizing the number of language models and combinations of prompts used by the collaboration processing unit; A search result collaborating method comprising:

Citation Information

Cited By

  • Natural language processing system and natural language processing method

    JP2026046302A

  • Natural language processing system and natural language processing method

    JP7906813B2