Information processing device, assistance method, and assistance program
The information processing apparatus evaluates the contribution of each pre-trained model to the final answer, addressing the challenge of reward distribution and model improvement in multi-model systems.
Patent Information
- Application Number
- PCT/JP2024/023745
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-07-01
- Publication Date
- 2025-06-19
AI Technical Summary
Current technologies lack the ability to evaluate the contribution of each pre-trained model to the final answer generated when using multiple models in combination, making it difficult to provide reasonable rewards and reflect user feedback effectively.
An information processing apparatus and method that includes an answer generation mechanism using multiple learned models and a contribution degree calculation mechanism to determine the contribution of each model to the generated answer, allowing for evaluation and appropriate reward distribution.
Enables accurate evaluation of each model's contribution, facilitating fair reward distribution and improving model performance through targeted re-training, thereby enhancing the overall answer generation process.
Smart Images

Figure JP2024023745_19062025_PF_FP_ABST
Abstract
Description
Information processing device, support method, and support program
[0001] The present disclosure relates to an information processing device, a support method, and a support program.
[0002] In recent years, with the development of large language models (LLMs), question-answering systems using natural language processing technology have become widespread. LLMs acquire a wide range of knowledge by learning large amounts of text data, and can generate appropriate answers to questions entered in natural language. For example, Patent Document 1 listed below discloses an information providing device that generates answers to questions about labor management using a language model.
[0003] Japanese Patent No. 7353695
[0004] In technologies that generate answers to queries input by users using trained models such as language models, including the information providing device described in Patent Literature 1, there is room for further improvement in answer generation performance by using multiple trained models in combination. For example, by using multiple trained models in combination, it becomes possible to generate highly accurate answers that take advantage of the characteristics of each trained model, or to run multiple trained models in parallel to generate answers at high speed.
[0005] When multiple trained models are used in combination, it is expected that it will be necessary to evaluate the degree to which each trained model contributed to the finally generated answer. Without knowing the degree of contribution of each trained model, it is difficult to provide an appropriate reward to the provider of each trained model, and it is also difficult to appropriately reflect user feedback on the quality of the final answer in the retraining of each trained model.
[0006] However, currently, no technology is known that evaluates the degree to which each trained model contributed to a finally generated answer, including the information providing device of Patent Document 1. An exemplary objective of the present disclosure is to provide a technology that enables evaluation of the degree to which each trained model contributed to a generated answer when an answer to a query is generated using multiple trained models.
[0007] An information processing device according to an exemplary aspect of the present disclosure includes an answer generation means that generates an answer to a query input by a user using a plurality of trained models generated by machine learning, and a contribution calculation means that calculates the contribution of each trained model to the answer generated by the answer generation means.
[0008] An assistance method according to an exemplary aspect of the present disclosure includes: at least one processor executing an answer generation process that generates an answer to a query entered by a user using multiple trained models generated by machine learning; and a contribution calculation process that calculates the contribution of each trained model to the answer generated in the answer generation process.
[0009] An assistance program according to an exemplary aspect of the present disclosure causes a computer to function as an answer generation means that generates an answer to a query entered by a user using multiple trained models generated by machine learning, and a contribution calculation means that calculates the contribution of each trained model to the answer generated by the answer generation means.
[0010] According to one exemplary aspect of the present disclosure, one exemplary effect is achieved in that it is possible to evaluate the extent to which each trained model contributed to the generated answer.
[0011] 1 is a block diagram showing a configuration of an information processing device according to the present disclosure; FIG. 2 is a flow diagram showing a flow of a support method according to the present disclosure; FIG. 3 is a diagram showing a configuration example of a response system according to the present disclosure; FIG. 4 is a block diagram showing a configuration of an information processing device according to the present disclosure; FIG. 5 is a flow diagram showing a flow of processing executed by the information processing device shown in FIG. 4; FIG. 6 is a block diagram showing a configuration of a computer that functions as an information processing device according to the present disclosure.
[0012] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technologies (part or all of the products or methods) employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technologies employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.
[0013] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described. Note that the scope of application of each technique employed in this exemplary embodiment is not limited to this exemplary embodiment. In other words, each technique employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, as long as no particular technical obstacles arise.
[0014] (Key Points of the Invention) Results from multiple LLMs (i.e., answers to multiple decomposed queries obtained by decomposing an input query) are aggregated to produce a good result. In other words, a single consolidated answer is generated from the multiple answers, and each LLM is evaluated for how much it contributed to that result. Cash from customers is then distributed according to their contribution.
[0015] (Configuration of information processing device) The information processing device according to this exemplary embodiment includes an aggregation means that aggregates multiple answers generated by decomposing an input query, inputting the generated decomposed queries into language models corresponding to the decomposed queries, and generating an answer, for each of multiple decomposed queries, and generates an aggregated answer; and a contribution calculation means that calculates the contribution of each of the multiple language models to the aggregated answer.
[0016] According to the above configuration, it is possible to provide appropriate feedback to the providers of each language model.
[0017] The aggregator aggregates answers using a language model trained to aggregate answers, and if the aggregator records dependencies between decomposed queries (e.g., an answer to one decomposed query is part of another decomposed query), the aggregator aggregates according to the recorded dependencies.
[0018] The contribution calculation means calculates a Shapley value in game theory for each language model. This is the contribution of each language model. For example, if the aggregation means aggregates answers from language models A to C to generate an aggregated answer, the contribution calculation means calculates the following contributions and calculates the Shapley value from these contributions: 1. Contribution of language model A alone: Similarity between the answer from language model A and the aggregated answer. 2. Contribution of language model B alone: Similarity between the answer from language model B and the aggregated answer. 3. Contribution of language model C alone: Similarity between the answer from language model C and the aggregated answer. 4. Contribution of the combination of language models A and B: Similarity between the answer aggregating answers from language models A and B and the aggregated answer. 5. Contribution of the combination of language models A and C: Similarity between the answer aggregating answers from language models A and C and the aggregated answer. 6. Contribution of the combination of language models B and C: Similarity between the answer aggregating answers from language models B and C and the aggregated answer.
[0019] It is also possible to consider the difference between the answer of language model A and the aggregated answer as the contribution of the combination of language models B and C. Similarly, it is also possible to calculate the contribution of the combination of language models A and B and the combination of A and C.
[0020] Furthermore, language models may include not only those that generate text content but also those that generate numerical sequences (e.g., graphs and tables). In this case, by using the insight derivation model to convert the numerical sequences into text, they can be treated in the same way as other texts.
[0021] In addition to allocating rewards, the calculated contribution level can also be used to allocate user feedback. For example, if the user scores an aggregated answer in which the contribution levels of language models A to C were 5:3:2, respectively, and the scores for language models A to C were 40, 24, and 16 points, respectively. These scores can be used, for example, when determining which language model will generate the next answer.
[0022] In addition to using the Shapley value, it is also possible to calculate the contribution for a combination of an answer from a single language model and an aggregated answer using a contribution calculation model generated using training data in which the contribution of the single language model is associated as correct answer data.
[0023] (Flowchart) 1. Obtain answers from each LLM 2. Aggregate 3. Present 4. Calculate contribution level 5. Distribute rewards according to contribution level [Second Exemplary Embodiment] A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is the basis for each of the exemplary embodiments described below. Note that the scope of application of each technology employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology employed in this exemplary embodiment can also be employed in other exemplary embodiments included in this disclosure, to the extent that no significant technical obstacles arise. Furthermore, each technology shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in this disclosure, to the extent that no significant technical obstacles arise. These matters also apply to the third exemplary embodiment.
[0024] (Configuration of Information Processing Device 1) The configuration of the information processing device 1 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1. As shown in Fig. 1, the information processing device 1 includes an answer generation unit 101 and a contribution calculation unit 102.
[0025] The answer generation unit 101 generates an answer to an input query using multiple trained models. Here, a "query" refers to an inquiry to the information processing device 1. For example, a question or request in natural language input by a user seeking information is a typical query. The content of the query is not particularly limited, and may be, for example, a question such as "What's the weather like in Tokyo tomorrow?" or a command such as "Please create a summary of the input text." Therefore, the term "query" in the following embodiments may also be interpreted as an inquiry, question, command, request, or the like. A query may be written in, for example, a natural language or an artificial language such as a programming language.
[0026] The query may be input as text data or as voice data. In the latter case, the input voice data can be processed in the same way as the former by converting it into text data through voice recognition processing.
[0027] Furthermore, the "trained model" may be one that has been machine-learned to output information that can be used to generate an answer to an input query. It may generate part or all of an answer, or it may output information that is referenced in generating an answer. For example, when a natural language query is used, a language model such as an LLM may be used as the trained model. For example, the language model may be a Generative Pre-Trained Transformer (GPT) that predicts a character string that is likely to follow an input character string and outputs a sentence containing the input character string. Other language models that may be used include a Text-to-Text Transfer Transformer (T5), a Bidirectional Encoder Representations from Transformers (BERT), a Robustly optimized BERT approach (RoBERTa), and an Efficiently Learning an Encoder that Classifies Token Replacements Accurately (ELECTRA).
[0028] The answer generated by the trained model may be text data, or may be data in other formats such as image data, audio data, or numerical values. It is preferable to apply the multiple trained models with different characteristics, such as different training data used for training or different formats of output data. This makes it possible to generate highly accurate answers that take advantage of the characteristics of each trained model. It is also possible to apply the multiple trained models with common characteristics. In this case, too, it is possible to expect the effects of shortening the time required to complete answer generation and distributing the load.
[0029] Note that some or all of the plurality of trained models may be provided in the information processing device 1, or all of the plurality of trained models may be provided in a device external to the information processing device 1. When using a trained model provided in an external device, the answer generation unit 101 transmits a query to the external device to generate an answer.
[0030] The contribution calculation unit 102 calculates the contribution of each trained model to the answer generated by the answer generation unit 101. Here, the contribution is a numerical value indicating the degree of contribution to the generation of the answer. The contribution calculated by the contribution calculation unit 102 can be used for various purposes. For example, the calculated contribution can also be used as an evaluation index for the performance of each trained model. In this case, in the next and subsequent answer generation, it becomes possible to select a trained model that can generate a better answer based on the previously calculated contribution.
[0031] As described above, the information processing device 1 according to this exemplary embodiment is configured to include an answer generation unit 101 that generates an answer to an input query using multiple trained models, and a contribution calculation unit 102 that calculates the contribution of each trained model to the answer generated by the answer generation unit 101.
[0032] According to the above configuration, an answer to a query is generated using multiple trained models, and the contribution of each trained model to the generated answer is calculated, thereby making it possible to evaluate the degree to which each trained model contributed to the generated answer. Furthermore, according to the above configuration, an answer to a query is generated using multiple trained models, making it possible to optimize the answer to the query. Furthermore, according to the above configuration, the performance of the trained model can be improved by retraining the trained model using the calculated contribution. Furthermore, since the contribution of each trained model can also provide clues for the development of new AI (artificial intelligence) systems, it can be said that the information processing device 1 contributes to the overall development of AI systems that use trained models.
[0033] (Assistance Program) The functions of the information processing device 1 described above can also be realized by a program. The assistance program according to this exemplary embodiment is a program that supports the use of trained models, and causes a computer to function as answer generation means that generates an answer to a query entered by a user using multiple trained models generated by machine learning, and as contribution calculation means that calculates the contribution of each trained model to the answer generated by the answer generation means. This assistance program has the effect of making it possible to evaluate the degree to which each trained model contributed to the generated answer.
[0034] (Flow of the Support Method) The flow of the support method according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the support method. Note that the execution entity of each step in this support method may be a processor provided in the information processing device 1, or a processor provided in another device, or each step may be executed by a processor provided in a different device.
[0035] In S1 (answer generation process), at least one processor generates an answer to a query entered by a user using multiple trained models generated by machine learning.
[0036] In S2 (contribution calculation process), at least one processor calculates the contribution of each trained model to the answer generated in S1.
[0037] As described above, the assistance method according to this exemplary embodiment is an assistance method for assisting in the utilization of trained models, in which at least one processor executes an answer generation process that generates an answer to a query entered by a user using multiple trained models generated by machine learning, and a contribution calculation process that calculates the contribution of each trained model to the answer generated in the answer generation process. This assistance method has the effect of enabling the degree to which each trained model contributed to the generated answer to be evaluated.
[0038] Third Exemplary Embodiment (Configuration of Response System 5A) A response system 5A according to this exemplary embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example configuration of the response system 5A. As shown in the figure, the response system 5A includes an information processing device 1A, trained models 21A and 22A, and a terminal device 3A. These components included in the response system 5A are connected to each other so as to be able to communicate with each other via a network.
[0039] The answer system 5A is a system that has the function of accepting query input from a user via the terminal device 3A, causing the trained models 21A and 22A to generate an answer to the query, and presenting the generated answer to the user.
[0040] The information processing device 1A is a device that has a function of causing the trained models 21A and 22A to generate answers to queries input by a user. For example, the information processing device 1A may be a platform server provided on a cloud. As will be described in detail later, the information processing device 1A decomposes a query input by a user to generate multiple decomposed queries and allocates each of the generated multiple decomposed queries to one of the trained models 21A and 22A. The information processing device 1A then transmits the decomposed queries to the trained models 21A and 22A according to the determined allocation, causes them to generate answers, and aggregates and presents the generated answers to the user. This enables improved answer generation performance compared to using only one of the trained models 21A and 22A.
[0041] The trained models 21A and 22A are trained models that have been machine-learned to generate answers to queries. For example, when accepting input of a query written in natural language, general-purpose language models that have been machine-learned to learn the arrangement of components (such as words) of sentences written in natural language or the arrangement of sentences in a text may be used as the trained models 21A and 22A. Note that the trained models used by the information processing device 1A are not limited to language models; the information processing device 1A can also use inference models such as prediction models and classification models. It is also possible to use a model that combines a language model and an inference model and is capable of generating an answer in natural language based on the inference results of the inference model.
[0042] It is preferable that the trained model 21A and the trained model 22A have different characteristics. For example, trained models 21A and 22A with different numbers of trained parameters may be applied. Generally, the greater the number of trained parameters, the higher the accuracy of the answer, but the longer it takes to generate the answer. In the answering system 5A, answers to decomposed queries for which priority should be placed on answer speed can be generated by a trained model with a smaller number of trained parameters, and answers to decomposed queries for which priority should be placed on answer accuracy can be generated by a trained model with a larger number of trained parameters. This makes it possible to achieve both answer speed and answer accuracy.
[0043] Furthermore, for example, by using different training data for at least one of machine learning and fine tuning, trained models 21A and 22A with different characteristics can be generated. For example, a general-purpose language model may be used as is as the trained model 21A, and a general-purpose language model may be fine-tuned as the trained model 22A using combinations of questions and their answers in a specific technical field as training data. In this case, answers to decomposed queries including general questions may be generated by the trained model 21A, and answers to decomposed queries including questions related to a specific technical field may be generated by the trained model 22A.
[0044] The terminal device 3A is a device that serves as an interface with the user in the answer system 5A. Specifically, the terminal device 3A has a function of accepting a query input by the user and presenting the user with an answer to the query (generated under the control of the information processing device 1A). While Fig. 3 shows an example in which the terminal device 3A is a smartphone, a stationary personal computer or the like can also be used as the terminal device 3A.
[0045] When a general-purpose computer such as a smartphone or a personal computer is used as the terminal device 3A, a predetermined client application may be installed on the terminal device 3A to enable the use of the service provided by the response system 5A. In this case, when using the service provided by the response system 5A, the user simply operates the terminal device 3A to launch the client application. When starting to use the service, the user may be prompted to enter pre-registered authentication information to log in to the service.
[0046] In the example of Fig. 3, a query such as "Create a training report for this month and a menu for next month" is input to the terminal device 3A. In this manner, the reply system 5A can also be used for healthcare purposes. The query input to the terminal device 3A is transmitted to the information processing device 1A. Note that the query is preferably transmitted in an encrypted state using a secure communication protocol such as HTTPS (HyperText Transfer Protocol Secure).
[0047] The information processing device 1A decomposes the received query to generate a plurality of decomposed queries. For example, the information processing device 1A decomposes the input query into two queries: "Create a training report for this month" and "Create a training menu for next month."
[0048] Next, the information processing device 1A assigns each query generated by the decomposition, i.e., the decomposed queries, to the trained model 21A or 22A. Here, for example, the trained model 21A is provided on a server that stores various data indicating the user's training results and is capable of generating answers using the data, while the trained model 22A is a model trained to be able to create a training menu. In this case, the information processing device 1A assigns the decomposed query "Create a report on this month's training" to the trained model 21A, and assigns the decomposed query "Create a training menu for next month" to the trained model 22A.
[0049] Next, the information processing device 1A transmits each decomposition query to the trained model 21A or 22A. The transmission of the decomposition query is performed, for example, via an API (Application Programming Interface). Then, the answers to the decomposition queries generated by the trained model 21A or 22A are also transmitted to the information processing device 1A, for example, via the API. The information processing device 1A aggregates the answers generated by the trained models 21A and 22A, formats them as necessary, and transmits them to the terminal device 3A. It is preferable that these answers be transmitted in an encrypted state, similar to the queries.
[0050] The terminal device 3A that receives the answer decodes and outputs the received answer as necessary. In the example of Fig. 3, a "This Month's Report" and a "Recommended Menu for Next Month" are displayed on the terminal device 3A as answers to the above query. In this way, the answering system 5A can present answers that include a "This Month's Report" with appropriate content using various data indicating the user's training results and a "Recommended Menu for Next Month" with appropriate content based on the user's prior learning results.
[0051] As described above, a user can use an advanced question-answering service that uses multiple trained models, trained models 21A and 22A, simply by inputting a query in natural language into the user's own terminal device 3A. This service provides efficient responses by switching between trained models 21A and 22A transparently to the user.
[0052] The query may be input as text or as voice. The answer to the query may be presented as text as in the example of FIG. 3 or as voice. The answer system 5A may include three or more trained models. The trained models included in the answer system 5A may be stored inside the information processing device 1A or may be stored in another device. For example, the terminal device 3A may include the trained models.
[0053] Here, conventional answering systems that answer queries using a single model may not be able to provide sufficient answers to complex questions, such as questions that span multiple fields or questions that require advanced specialized knowledge. Furthermore, even if the content of the answer is appropriate, it may take a long time to generate it. In this regard, answering system 5A is capable of providing appropriate and prompt answers even to complex questions.
[0054] The information processing device 1A also calculates the contribution of each of the trained models 21A and 22A to the answer generated as described above. This makes it possible to pay, for example, a reward based on the contribution, in other words, a usage fee for the trained models 21A and 22A, to the provider of the trained models 21A and 22A.
[0055] (Configuration of Information Processing Device 1A) The configuration of the information processing device 1A according to this exemplary embodiment will be described with reference to FIG. 4. FIG. 4 is a block diagram showing the configuration of the information processing device 1A. Note that the information processing device 1A may be a device whose main function is to control the generation of answers to queries, or may be a general-purpose device that also has other functions. Furthermore, the information processing device 1A may be a stationary device as shown in FIG. 3, or may be a portable device. For example, it is also possible to provide the functions of the information processing device 1A to a terminal device 3A shown in FIG. 3.
[0056] 4, the information processing device 1A includes a control unit 10A that controls the various units of the information processing device 1A and a storage unit 11A that stores various data used by the information processing device 1A. The information processing device 1A also includes a communication unit 12A that enables the information processing device 1A to communicate with other devices, an input unit 13A that accepts input to the information processing device 1A, and an output unit 14A that enables the information processing device 1A to output data. The control unit 10A includes an answer generation unit 101A, a contribution calculation unit 102A, a query decomposition unit 103A, an allocation unit 104A, a presentation control unit 105A, a reward granting unit 106A, a verification unit 107A, a related information acquisition unit 108A, and a reliability determination unit 109A.
[0057] The answer generation unit 101A generates an answer to an input query using multiple trained models, similar to the answer generation unit 101 described in exemplary embodiment 2. For example, the answer generation unit 101A generates one answer to one query using the trained models 21A and 22A shown in FIG. 3 . That is, the answer generation unit 101A aggregates multiple answers generated by multiple trained models to generate one aggregated answer. Details of the answer aggregation method will be described later.
[0058] As described in exemplary embodiment 1, the trained model used to generate an answer may be one that has been trained by machine learning so as to be able to output information that can be used to generate an answer to an input query. For example, a generative model (which may also be referred to as generative AI) such as a language model may be used, or an inference model that performs inference such as classification or prediction may be used.
[0059] Furthermore, as will be described in detail later, the input query is decomposed into multiple decomposed queries by the query decomposing unit 103A, and the allocation unit 104A allocates each decomposed query to one of multiple trained models. Therefore, the answer generating unit 101A transmits the decomposed queries to the multiple trained models in accordance with the allocation determined by the allocation unit 104A, causing them to generate answers.
[0060] The contribution calculation unit 102A calculates the contribution of each trained model to the answer generated by the answer generation unit 101A, similar to the contribution calculation unit 102 described in exemplary embodiment 2. For example, when one answer to one query is generated using the trained models 21A and 22A, the contribution calculation unit 102A calculates the contribution of the trained models 21A and 22A to the answer, respectively. The method of calculating the contribution will be described later.
[0061] As described above, the information processing device 1A includes the answer generation unit 101A that generates an answer to an input query using multiple trained models, and the contribution calculation unit 102A that calculates the contribution of each trained model to the answer generated by the answer generation unit 101A. Thus, similar to the information processing device 1 described in exemplary embodiment 2, the information processing device 1A has the effect of being able to evaluate the degree to which each trained model contributed to the generated answer.
[0062] (Query Decomposition) The query decomposition unit 103A decomposes a query input by a user to generate multiple decomposed queries. The method for generating the decomposed queries is not particularly limited. For example, the query decomposition unit 103A may generate the decomposed queries using a question decomposition model that has been machine-learned to decompose a question into multiple queries. As the question decomposition model, for example, a known model such as DecompRC may be applied. Furthermore, for example, if the query input by the user is in text format, the query decomposition unit 103A may decompose the query at the positions of periods and commas. Furthermore, for example, the query decomposition unit 103A may perform morphological analysis on the query input by the user and decompose the query based on the results of the morphological analysis.
[0063] Here, an example will be described in which the query decomposition unit 103A decomposes queries having dependencies using a model such as DecompRC. When the query input by the user is "Which companies in industries where sales have decreased this year are expected to recover next year?", the query decomposition unit 103A generates, for example, the following three decomposed queries Q1, Q2, and Q3. Q1: In which industry are sales decreasing this year? Q2: Which companies in [Answer A1] are expected to recover next year? Q3: What is the specific company name in [Answer A2]? In this case, the decomposed query Q2 has a format in which the answer A1 of the decomposed query Q1 is substituted for "[Answer A1]," making it a decomposed query that depends on the decomposed query Q1. Similarly, the decomposed query Q3 has a format in which the answer A2 of the decomposed query Q2 is substituted for "[Answer A2]," making it a decomposed query that depends on the decomposed query Q2.
[0064] In this way, the query decomposition unit 103A decomposes a query input by a user into multiple decomposed queries, and the answer generation unit 101A can acquire appropriate related documents and related tables for each decomposed query, thereby efficiently deriving a final answer. In particular, by generating answers to each decomposed query in consideration of the dependency relationships between the decomposed queries, answers can be obtained in stages, and an appropriate answer can be generated even for complex queries.
[0065] For example, the query decomposing unit 103A may decompose a query input by a user into a language model. In this case, the query decomposing unit 103A may generate a command statement that includes the query input by the user and commands the query to be divided into multiple sentences according to its content, and input the generated command statement to the language model. This causes the decomposed query to be output from the language model. If at least one of the trained models 21A and 22A is a language model, the query decomposing unit 103A may cause the trained model 21A or 22A, which is a language model, to generate the decomposed query.
[0066] When the query decomposition unit 103A decomposes an input query into decomposed queries having dependencies in this manner, the answer generation unit 101A generates an answer to one decomposed query and then generates answers to other decomposed queries by referring to the answer. This makes it possible to generate appropriate answers that take into account the dependencies of each decomposed query. For example, if the answer to the decomposed query Q1 described above is "food service industry," the answer generation unit 101A may generate an answer by inputting this answer and the decomposed query Q2 into the trained model. Similarly, the answer generation unit 101A may generate a final answer, i.e., an aggregate answer, by inputting the answer to the decomposed query Q2 and the decomposed query Q3 into the trained model.
[0067] It is also possible to generate answers to decomposed queries that have a dependency relationship using a prediction model and a language model. For example, suppose a query is input saying, "Predict store X's sales this month, and plan next month's ordering policy for product A based on the sales." In this case, the query decomposition unit 103A decomposes this query into a decomposed query Q4 saying, "Predict store X's sales this month," and a decomposed query Q5 saying, "Plan next month's ordering policy for product A based on the answer to Q4."
[0068] Here, the decomposition query Q4 instructs a prediction of sales for a specific store, and it is difficult to generate an accurate answer to such a decomposition query Q4 using a general language model. Therefore, a prediction model for predicting sales for a specific store may be prepared in advance, and the allocation unit 104A may allocate the prediction model to a decomposition query instructing a prediction of sales for a specific store. This makes it possible to generate an accurate answer based on accurate prediction results. Note that the prediction model can be generated by machine learning using training data that is correlated with the prediction target (sales in the above example) (data correlated with sales at store X in the above example). Furthermore, in predictions using a prediction model, the answer generation unit 101A may acquire data correlated with the prediction target and input it into the prediction model. The same applies when a classification model or the like is applied.
[0069] As described above, the information processing device 1A includes a query decomposition unit 103A that decomposes a query to generate multiple decomposed queries. The answer generation unit 101A causes one of multiple trained models to generate an answer to one of the multiple decomposed queries, and also causes one of multiple trained models to generate an answer to another of the multiple decomposed queries by referring to the generated answer. This provides the effect of enabling the generation of appropriate answers that take into account the dependency relationships of the decomposed queries, in addition to the effect provided by the information processing device 1.
[0070] (Regarding Allocation of Decomposed Queries) The allocation unit 104A allocates each of a plurality of decomposed queries to one of a plurality of trained models that have been machine-learned to generate an answer to the query. For example, as described with reference to FIG. 3, the allocation unit 104A may allocate the decomposed queries to the trained models 21A and 22A.
[0071] The method of allocating decomposition queries to trained models is arbitrary. For example, the allocation unit 104A may allocate each decomposition query to a trained model that is highly compatible with the decomposition query based on the evaluation result of the compatibility between each decomposition query and each trained model. This makes it possible to perform appropriate allocation taking into account the compatibility of each combination of decomposition query and trained model. Note that the compatibility evaluation method will be described in detail later.
[0072] Furthermore, for example, the allocation unit 104A may allocate decomposed queries using load information indicating the load status in answer generation using each trained model. This allows allocation to be performed taking into account the load status in answer generation using each trained model. For example, the allocation unit 104A may assign a decomposed query to a trained model with a smaller load indicated in the load information, which is expected to result in a faster answer.
[0073] The load information may be generated by the information processing device 1A, or may be load information generated by another device and acquired via the communication unit 12A or the input unit 13A. It is preferable that the load information be updated in real time. The load information may be discrete information, such as high or low load, or may indicate the magnitude of the load using a continuous value (e.g., a numerical value between 0 and 1). The load in answer generation using a trained model can also be referred to as the load on the device that causes the trained model to generate an answer (the device equipped with the trained model). Any method, including known methods, can be used to calculate the load. For example, the load information may include the number of queries that have already been input to the target trained model but for which no answer has been generated, or the predicted time from when the trained model is instructed to generate an answer until answer generation is completed.
[0074] The allocating unit 104A may allocate decomposition queries using both the suitability evaluation results and the load information, thereby distributing the load in the generation using each trained model and allocating the decomposition queries to trained models that are suitable for the decomposition queries.
[0075] For example, both the compatibility evaluation result and the load information may be expressed as numerical values. In this case, the allocation unit 104A may calculate, for each combination of a decomposed query to be allocated, the difference between the numerical value indicating the compatibility evaluation result and the numerical value indicating the load information (hereinafter referred to as the overall evaluation value) with respect to each combination with multiple trained models. The allocation unit 104A may then allocate the trained model with the highest overall evaluation value to the decomposed query. For example, if the compatibility evaluation result between a certain decomposed query X and the trained model 21A is 0.8 (the larger this numerical value, the higher the compatibility) and the load information of the trained model 21A is 0.2 (the larger this numerical value, the higher the load), the overall evaluation value of the combination of the decomposed query X and the trained model 21A is 0.6. In this case, if the compatibility evaluation result between the decomposed query X and the trained model 22A is 0.6 and the load information of the trained model 22A is 0.1, the overall evaluation value of the combination of the decomposed query X and the trained model 22A is 0.5. Therefore, in this example, the allocation unit 104A allocates the decomposed query X to the trained model 21A that has a larger overall evaluation value.
[0076] The allocating unit 104A may also allocate decomposed queries based on the contribution calculated by the contribution calculation unit 102A. That is, the allocating unit 104A may allocate decomposed queries to trained models with high contribution calculated by the contribution calculation unit 102A, preferentially over trained models with low contribution. This makes it possible to generate highly accurate answers using trained models that are likely to contribute to the answer.
[0077] The allocation unit 104A may allocate decomposed queries using an allocation model for allocating decomposed queries. The allocation model may be generated by machine learning the relationship between various information related to allocation (e.g., features of the decomposed query to be allocated, attribute information of the user to whom the answer is to be presented, load information, features of each trained model, and statistical values of contributions previously calculated by the contribution calculation unit 102A for each trained model) as explanatory variables and the trained model to which the decomposed query should be allocated as the objective variable. Furthermore, when the contribution calculation unit 102A calculates a new contribution, the allocation unit 104A may update the allocation model by relearning using training data including the calculated contribution. This makes it possible to continuously allocate decomposed queries to trained models that are likely to contribute to generating an answer.
[0078] (Method for Evaluating Compatibility Between Decomposed Query and Trained Model) Any method can be applied as a method for evaluating compatibility between a decomposed query and a trained model, as long as it can produce valid evaluation results. For example, the allocation unit 104A may evaluate the compatibility between a decomposed query and a trained model using query feature information indicating the features of the decomposed query and model feature information indicating the features of the trained model. This has the effect of enabling allocation to be performed while taking into account the features of both the decomposed query and the trained model.
[0079] The method for generating the query feature information is not particularly limited. For example, the allocating unit 104A may generate the query feature information of the decomposed query using a feature extraction model that has been machine-learned to generate query feature information (e.g., a feature vector) that indicates the features of the input query.
[0080] On the other hand, the model feature information may be generated in advance and stored in the storage unit 11A or the like. The method for generating the model feature information is not particularly limited. For example, the model feature information may be information generated using training data used in machine learning of the trained model. This provides the effect of enabling allocation that takes into account the features of both the decomposition query and the trained model by using the training data.
[0081] For example, if the trained model is a large-scale language model, a large-scale text corpus is used for training. By analyzing sentences and words contained in such a text corpus, it is possible to generate model feature information that indicates the features of the trained model (large-scale language model) generated by training using such a text corpus.
[0082] For example, characteristic words may be extracted from the text corpus used for training using a technique such as TF-IDF (Term Frequency - Inverse Document Frequency), and some or all of the extracted words may be used as model feature information. Alternatively, some or all of the extracted words may be synthesized to generate model feature information. For example, some or all of the extracted words may be input to a language model, and words or sentences derived from those words may be output, and the output words or sentences may be used as model feature information.
[0083] Furthermore, for example, multiple trained models may be artificially classified, in which case the artificial classification serves as model feature information. For example, classifications such as "good at answering mathematical questions," "good at everyday conversation," and "fast processing" may serve as model feature information. Sentences contained in the above-described text corpus may be input to a language model to determine which classification they fall into or their degree of conformance to each classification. Furthermore, when a fine-tuned language model is used as the trained model, model feature information may be generated from the training data used for fine-tuning.
[0084] When a feature vector generated using a feature extraction model is used as the query feature information, the model feature information may also be a feature vector generated using the feature extraction model. This makes it possible to easily determine the similarity between the query feature information and the model feature information within the same feature space. This also makes it possible to assign a trained model having similar features to the decomposed query to the decomposed query.
[0085] For example, the model feature information may be a feature vector obtained by inputting words and sentences extracted from a text corpus or the like as described above and metadata of the trained model (indicating the application field and characteristics (e.g., response speed and accuracy) of the trained model) into a feature extraction model. The allocation unit 104A may then calculate the similarity (e.g., cosine similarity) between the query feature information of the decomposed query and the model feature information of each trained model as a value indicating the compatibility between the decomposed query and the trained model. In this case, the allocation unit 104A may allocate the decomposed query to the trained model with the highest similarity. In this case, queries with similar meanings are allocated to the same trained model.
[0086] Furthermore, queries previously input to a trained model and answers previously generated by the trained model can also be considered to indicate the characteristics of the trained model. For this reason, model characteristic information may be generated using at least one of queries previously input to the trained model and answers generated in response to those queries.
[0087] For example, the model feature information may be words or sentences extracted from a query previously input to the trained model and the answer generated in response to that query, or a feature vector obtained by inputting the words or sentences into a feature extraction model. Furthermore, for example, an answer that best represents the features of the trained model may be extracted from answers previously output by the trained model, the similarity between that answer and other answers may be calculated, and the number or percentage of answers whose similarity is equal to or greater than a threshold may be used as the conformance to the features. For example, suppose that 100 answers generated by the trained model 22A include one that indicates the recommended exercise menu shown in FIG. 3, and that 59 of the other 99 answers have a similarity to this answer that is equal to or greater than a threshold. In this case, the model feature information of the trained model 22A may indicate a conformance to the "exercise menu generation" of 60%. In this method, the model feature information is generated based on answers actually generated by the trained model, so the results of directly evaluating the trained model are reflected in the model feature information.
[0088] (Presenting Various Information) The presentation control unit 105A presents information related to the answer system 5A to the user. For example, the presentation control unit 105A presents to the user an aggregated answer that aggregates answers generated by multiple trained models. In addition to the effects of the information processing device 1, the information processing device 1A equipped with the presentation control unit 105A has the effect of being able to present multiple generated answers in a form that is easy for the user to understand. Details of the method of aggregating answers will be described later.
[0089] (Regarding the answer aggregation method) As described above, the answer generation unit 101A aggregates answers generated by each of the multiple trained models. For example, the answer generation unit 101A may generate an aggregated answer by inputting the answers generated by each of the multiple trained models into a language model that has been machine-learned to aggregate multiple answers. This provides, in addition to the effects of the information processing device 1, the effect of being able to generate natural and complete answers that take into account the relevance of each answer generated by each of the multiple trained models and eliminate redundancy. The language model used to aggregate answers may be, for example, a general-purpose language model, a general-purpose language model that has been fine-tuned for answer aggregation, or a language model dedicated to answer aggregation that has been generated for answer aggregation. Furthermore, if a language model is included in the trained model used to generate an answer, the language model may also be used to aggregate answers.
[0090] Note that the method of aggregating answers to each decomposition query is arbitrary and is not limited to the above example. For example, in the example of FIG. 3, the answer displayed on the terminal device 3A is a parallel listing of the answer generated by the trained model 21A and the answer generated by the trained model 22A. In this way, combining multiple generated answers without processing them also falls within the category of "aggregation." Furthermore, as described above, sequentially generating answers to multiple dependent decomposition queries to generate a final answer also falls within the category of "aggregation."
[0091] Furthermore, the multiple trained models used to generate an answer may include multiple language models. In this case, the allocation unit 104A may allocate one decomposed query to multiple language models. In this case, the answer generation unit 101A may generate an answer by combining candidate outputs predicted by each of the multiple language models for one decomposed query. This provides the effect of enabling the generation of more diverse and flexible answers in addition to the effect achieved by the information processing device 1.
[0092] Generally, a language model repeats the process of predicting a probability distribution for each of the candidate outputs that may follow an input string, indicating the probability that each candidate output follows the input string, and selecting a candidate output based on the predicted probability distribution, thereby generating an answer consisting of an array of multiple candidate outputs.
[0093] Therefore, when multiple language models are caused to generate answers for one decomposed query, each language model predicts a probability distribution of candidate outputs. Therefore, the answer generation unit 101A can sample candidate outputs to be used as components of an aggregated answer from among the candidate outputs of each language model based on these probability distributions. The answer generation unit 101A then causes each language model to generate a probability distribution of the candidate output that follows the sampled candidate output. Similarly, the answer generation unit 101A repeats sampling and generation of probability distributions to generate aggregated answers.
[0094] For example, suppose a decomposed query, "Please tell me what pets are popular in City X," is input to language model A and language model B. Language model A infers that the probability that the candidate output of "cat" will follow the decomposed query is 0.4, while language model B infers that the probability that the candidate output of "dog" will follow the decomposed query is 0.3. In this case, answer generation unit 101A may determine the candidate output of "cat," which has a higher probability, as the output following the decomposed query. By repeating this process, an aggregate answer that combines the outputs of language models A and B can be generated.
[0095] Any sampling method, including known techniques, can be applied. For example, a greedy method that always selects the most probable candidate output, a method that randomly selects a candidate output based on probability, or top-k sampling that selects a candidate output from the top k most probable texts can be applied.
[0096] Furthermore, for example, the answer generation unit 101A may generate an aggregated answer using beam search. In beam search, the answer generation unit 101A searches multiple candidate outputs generated by each language model and outputs the most likely output sequence as an answer. Specifically, the answer generation unit 101A stores the candidate outputs generated by each language model in multiple candidate sequences called beams. For example, if the beam width is 3, three candidate outputs are stored as candidate sequences at each step. In the next step, the answer generation unit 101A adds the candidate outputs generated by the language model to each candidate output stored as a candidate sequence and calculates a score (likelihood). The top three candidate outputs with the highest scores are then selected as a new beam, and the process proceeds to the next step. The answer generation unit 101A repeats the above process and finally selects the sequence of candidate outputs with the highest score as an output. In beam search, the outputs of language models are combined to generate an optimal answer that is in line with the context. Furthermore, when a priority or reliability is set for each language model, the answer generation unit 101A may preferentially select a candidate output generated by a language model with a high priority or reliability. Similarly, the answer generation unit 101A may preferentially select a candidate output generated by a language model with a large sum of contributions previously calculated by the contribution calculation unit 102A.
[0097] Furthermore, the answer generation unit 101A may aggregate numerical values included in each answer generated by multiple trained models for one decomposed query. For example, suppose that in response to a decomposed query such as "What is the probability of rain tomorrow in Tokyo?", trained model A generates the answer "60%," and trained model B generates the answer "80%." In this case, the answer generation unit 101A may calculate the arithmetic mean of the numerical values included in each answer to generate the answer "70%." Furthermore, if reliability is set for trained models A and B, the answer generation unit 101A may calculate a weighted average value by applying a weight according to the reliability, and generate an answer including the calculated weighted average value.
[0098] (Regarding Reward Granting) The reward granting unit 106A grants a reward to each provider of each trained model used to generate an answer, according to the contribution calculated by the contribution calculation unit 102A. For example, the reward granting unit 106A may allocate the reward paid by the user for the generated and presented answer to each provider of each trained model according to the contribution. The reward is typically money or points equivalent thereto, but is not limited to these examples, and any reward can be applied.
[0099] (Regarding Answer Verification) When a contradiction occurs between answers generated by each of a plurality of trained models, which are language models, the verification unit 107A causes each trained model that generated each of the contradictory answers to provide the basis for that answer. In addition to the effects achieved by the information processing device 1, the information processing device 1A equipped with the verification unit 107A has the effect of preventing contradictory answers from being aggregated and presented, and making it possible to obtain information for determining which answer is appropriate. Note that the term "contradiction" here can also be rephrased as "discrepancy" or "inconsistency," etc.
[0100] Furthermore, the presentation control unit 105A may present the rationale generated by the verification unit 107A to the user. Then, the verification unit 107A may have the user select one of the contradictory answers based on the presented rationale. In this case, the answer selected by the user is used to generate the aggregated answer, and the answers not selected are not used to generate the aggregated answer. This makes it possible to generate and present a valid aggregated answer even if each trained model generates a contradictory answer. Furthermore, the verification unit 107A may perform at least one of a process of increasing the reliability or priority of a trained model that generated an answer selected by the user and a process of decreasing the reliability or priority of a trained model that generated an answer not selected by the user.
[0101] Whether or not a contradiction occurs between the answers can be determined, for example, by a language model. When determining whether or not a contradiction occurs using a language model, the verification unit 107A may, for example, input to the language model the answers generated by each of the multiple trained models, along with a query asking whether any of the answers are contradictory in content. This causes the language model to output information indicating whether or not there are contradictory answers. The same applies when providing the basis for an answer; the verification unit 107A may input to the language model, along with the answer, a query asking about the basis for the answer. When a language model is included in the trained model used to generate an answer, the language model may also be used to determine whether or not there is a contradiction and to generate the basis.
[0102] (Regarding Use of Related Information) The related information acquisition unit 108A acquires related information, which is information used to generate an aggregated answer. The related information may be any information that can be used to generate an aggregated answer. For example, the related information acquisition unit 108A may acquire related information that indicates at least one of the reliability or priority set for each of the multiple trained models, the user attributes, and the answer generation environment.
[0103] The reliability and priority may be set in advance by a user or a service provider, or may be determined by automatically evaluating the answers generated by each trained model. For example, the public nature of the information source (whether it is a public or private institution), the recency and comprehensiveness of the information contained in the answer, and other such indicators may be used for the automatic evaluation of the answer. Examples of user attributes include personal characteristics (age, gender, height, weight, medical history, etc.) previously registered by the user. Furthermore, the answer generation environment may include information indicating the hardware environment, such as whether the terminal device used by the user is a mobile terminal or a fixed terminal. Furthermore, the answer generation environment may also include information indicating the user's surrounding environment, such as the user's location information acquired by a Ground Positioning System (GPS) or meteorological information, such as the weather in the area identified by the location information.
[0104] The answer generation unit 101A may generate an aggregated answer based on the related information acquired by the related information acquisition unit 108A. As described above, the related information may indicate, for example, at least one of the reliability or priority set for each of the multiple trained models, the user's attributes, and the answer generation environment. This provides the effect of enabling the generation of an aggregated answer that reflects the related information, in addition to the effect achieved by the information processing device 1.
[0105] A method for reflecting related information in the aggregated answer may be applied according to the method for generating the aggregated answer. For example, when generating an aggregated answer using a language model, the answer generation unit 101A may generate a query including related information and generate the aggregated answer using the query. For example, the answer generation unit 101A may generate the aggregated answer using a query that includes the answers generated by each trained model and the reliability or priority of the trained model that generated the answer, and instructs the answers to be aggregated taking into account the reliability or priority. Furthermore, for example, the answer generation unit 101A may generate the aggregated answer using a query that includes user attributes and instructs the system to generate an aggregated answer adapted to users with the attributes.
[0106] Furthermore, the answer generation unit 101A may input the aggregated answer and related information once generated into a language model and arrange the input aggregated answer to match the related information. For example, the answer generation unit 101A may generate a query that includes the user's age as related information and instructs the generation of an answer that rewrites the aggregated answer for users of that age, and input this query into the language model together with the aggregated answer once generated. This makes it possible to generate an aggregated answer that matches a user of the age indicated in the related information.
[0107] (Regarding Determining the Reliability of an Aggregated Answer) The reliability determination unit 109A determines the reliability of an aggregated answer that aggregates the answers based on at least one of the answers generated by each of a plurality of trained models. In addition to the effects of the information processing device 1, the information processing device 1A equipped with the reliability determination unit 109A has the effect of making it possible to grasp the reliability of the aggregated answer. Furthermore, by presenting the reliability determination result to the user together with the aggregated answer, it is possible to provide the user with information to use when making a decision based on the aggregated answer.
[0108] The method for determining the reliability is not particularly limited. For example, if a prediction model is included among the multiple trained models used to generate an answer, the reliability determination unit 109A may use the confidence level of the prediction result of the prediction model as the reliability of the aggregated answer. For example, suppose an input query is, "Please plan a sales strategy for the new store based on next month's sales forecast for the new store." From this query, decomposition queries are generated: "Next month's sales forecast for the new store" and "Please plan a sales strategy for the new store based on the sales forecast." The answer to the former is generated by the sales prediction model for the new store, and the answer to the latter is generated by the language model by referencing the prediction results of the sales prediction model. An aggregated answer is generated: "Sales are expected to increase significantly from this month, so we recommend securing sufficient personnel and conducting active sales promotion." In this case, the reliability determination unit 109A may use the confidence level of the prediction result of the sales prediction model (a numerical value indicating the likelihood of the prediction result, generated together with the prediction result) as the reliability of the aggregated answer.
[0109] Furthermore, in the case of a classification model or a language model, when an inference result is calculated, a numerical value indicating the confidence level of the inference result is calculated, and the reliability determination unit 109A can use such a numerical value as the confidence level of the aggregate answer. Furthermore, the reliability determination unit 109A may specify the confidence level of the inference result for each of a plurality of trained models, and use a statistical value (e.g., an average value, a minimum value, a maximum value, or the like) calculated from the confidence levels as the confidence level.
[0110] The reliability determination unit 109A may determine the reliability using a determination model for determining the reliability of the aggregated answer. Such a determination model can be generated by machine learning the relationship between various answers and the reliability of each answer.
[0111] (Method for calculating contribution degree) The contribution degree calculation unit 102A may calculate the contribution degree of each trained model based on the similarity between an aggregated answer that aggregates answers generated by multiple trained models and each answer generated by each of the multiple trained models. If the similarity between an answer generated by a certain trained model and the aggregated answer is high, it can be said that the answer contributed greatly to generating the aggregated answer, so this configuration makes it possible to calculate an appropriate contribution degree.
[0112] The method for calculating the similarity is not particularly limited. For example, the contribution calculation unit 102A may calculate the cosine similarity or the Jaccard coefficient, which are used as a measure of similarity in natural language processing.
[0113] Furthermore, the method for calculating the contribution based on the similarity is not particularly limited. For example, as described above, the contribution calculation unit 102A may calculate a Shapley value for each trained model and use the calculated Shapley value as the contribution of each trained model. Calculating the Shapley value is an example of a method for calculating the contribution based on the similarity between the aggregated answer and each answer generated by each of the multiple trained models.
[0114] The Shapley value is known as a method for fairly distributing the contributions of multiple players when playing a game cooperatively. When the Shapley value is used as the contribution, each trained model is associated with a player, and the generation of an aggregated answer is associated with a cooperative game. Specifically, as described in the exemplary embodiment 1, the contribution calculation unit 102A calculates the similarity between the aggregated answer and an answer obtained by aggregating the answers of each combination of multiple trained models used to generate an answer (e.g., single, pair, trio, etc.), and calculates the Shapley value from these similarities.
[0115] Furthermore, the contribution calculation unit 102A may calculate the contribution using a contribution calculation model generated by machine learning. Such a contribution calculation model can be generated by machine learning the relationship between an answer generated by one trained model and an aggregated answer that aggregates multiple answers generated by multiple trained models including the one trained model, and the contribution of the one trained model to the aggregated answer. Even when a configuration using a contribution calculation model is adopted, it is possible to calculate an appropriate contribution. Another advantage of this configuration is that it can calculate the contribution more directly than when calculating the contribution based on similarity.
[0116] (Processing Flow) The processing flow executed by the information processing device 1A will be described with reference to Fig. 5. Fig. 5 is a flow diagram showing the processing flow executed by the information processing device 1A. Note that the flow in Fig. 5 includes each step of the support method according to this exemplary embodiment.
[0117] In S11, the query decomposition unit 103A accepts a query input by a user. Next, in S12, the query decomposition unit 103A decomposes the query accepted in S11 to generate multiple decomposed queries. Next, in S13, the allocation unit 104A allocates the multiple decomposed queries generated in S12 to multiple trained models.
[0118] In S14, the answer generation unit 101A generates an answer to the query input in S11 using the multiple trained models generated by machine learning. More specifically, the answer generation unit 101A transmits the decomposed query to each trained model in accordance with the allocation determined by the processing in S13, causing the trained models to generate an answer.
[0119] In S15, the verification unit 107A determines whether there are any contradictions among the multiple answers generated in S 14. If the determination in S15 is YES, the process proceeds to S19, and if the determination in S15 is NO, the process proceeds to S16.
[0120] In S16, the verification unit 107A causes each trained model that generated each contradictory answer to provide the basis for that answer. Next, in S17, the presentation control unit 105A presents the basis for each answer generated in S16 to the user along with each contradictory answer. Next, in S18, the verification unit 107A accepts the user's selection of one of the answers presented in S17. Note that the answer selected by the user from among the contradictory answers is used to generate an aggregated answer in S20. After S18 is completed, the process proceeds to S19.
[0121] In S19, the related information acquisition unit 108A acquires related information to be used for generating the aggregated answer. As described above, the related information acquisition unit 108A may acquire related information indicating at least one of the reliability or priority set for each of the multiple trained models, the user attributes, and the answer generation environment, for example.
[0122] In S20, the answer generation unit 101A aggregates the answers generated in S14 to generate an aggregated answer. For example, the answer generation unit 101A may input the answers generated in S14 into a language model to generate an aggregated answer. At this time, the answer generation unit 101A may also input the related information acquired in S19 into the language model to generate an aggregated answer that takes the related information into consideration. Note that in the flow of FIG. 5, the processes of S14 and S20 correspond to an answer generation process that generates an answer to a query input by a user using multiple trained models generated by machine learning.
[0123] In S21 (contribution calculation process), the contribution calculation unit 102A calculates the contribution of each trained model to the aggregated answer generated by the processes of S14 and S20. For example, the contribution calculation unit 102A may calculate a Shapley value for each trained model and use the calculated Shapley value as the contribution of each trained model.
[0124] In S22, the reward granting unit 106A grants a reward according to the contribution calculated in S21 to each provider of each trained model used to generate the answer in S14. Note that the reward does not necessarily have to be granted immediately after the contribution is calculated. For example, the reward granting unit 106A may grant a lump sum of rewards generated within a predetermined period.
[0125] In S23, the reliability determination unit 109A determines the reliability of the aggregated response generated in S20. The process of S23 may be performed at any timing after the process of S20. For example, the process of S23 may be performed before the processes of S21 to S22, or the process of S23 may be performed in parallel with the processes of S21 to S22.
[0126] In S24, the presentation control unit 105A presents the aggregated answer generated by the processing of S20 to the user, and the processing of Fig. 5 is thereby terminated. Note that in S24, the reliability determined in S23 may also be presented together with the aggregated answer. Furthermore, after the processing of S24, the processing may return to S11 and accept input of a new query.
[0127] (Example of not generating a decomposed query) Providing the query decomposition unit 103A is not essential, and an answer to an input query may be generated using multiple trained models without decomposing the query. For example, in response to a disaster, a query may be sent to language models established by a local government, a country, or an individual, and the collected answers may be aggregated to present useful information. Note that the language model used in this case may be one that has learned various information about the entity that established the language model, or may be one that can refer to various information about the local government, or may satisfy both of these requirements.
[0128] For example, suppose a query such as "Please tell me the damage situation in District X" is input to the information processing device 1A when an earthquake occurs. In this case, the answer generation unit 101A sends the query to a language model established by the local government that has jurisdiction over District X to generate an answer. The answer generation unit 101A also sends the same query to a language model established by the national government and a language model established by residents of District X to generate answers. The answer generation unit 101A then aggregates these answers to generate an aggregated answer, and the presentation control unit 105A presents the aggregated answer. This makes it possible to quickly grasp the overall picture of the damage and to plan and implement effective disaster prevention measures. Furthermore, if there is a discrepancy in the answers generated by each language model, the verification unit 107A can detect the discrepancy, thereby preventing decisions from being made based on incorrect information.
[0129] Furthermore, when multiple language models are used to generate answers to one query, the answer generation unit 101A may sample candidate outputs to be components of an aggregated answer from among the candidate outputs of each language model based on the probability distribution of the candidate outputs predicted by each language model. The answer generation unit 101A then causes each language model to generate a probability distribution of the candidate output that follows the sampled candidate output. Similarly, the answer generation unit 101A may repeat sampling and generation of probability distributions to generate aggregated answers. Furthermore, as described in the section "About Answer Aggregation Method," the answer generation unit 101A may aggregate numerical values included in each answer generated by multiple trained models for one query.
[0130] [Example of Software Implementation] Some or all of the functions of the information processing device 1, 1A may be implemented by hardware such as an integrated circuit (IC chip), or may be implemented by software.
[0131] In the latter case, the information processing device 1 or 1A is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Fig. 6. Fig. 6 is a block diagram showing the hardware configuration of computer C that functions as information processing device 1 or 1A.
[0132] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program (assistance program) P for causing the computer C to function as the information processing device 1 or 1A. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of the information processing device 1 or 1A.
[0133] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0134] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0135] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0136] Furthermore, each of the above functions of the information processing device 1 or 1A may be realized by a single processor provided in a single computer, by multiple processors provided in a single computer working together, or by multiple processors provided in each of multiple computers working together. Furthermore, the program for causing the information processing device 1 or 1A to realize each of the above functions may be stored in a single memory provided in a single computer, or may be distributed and stored in multiple memories provided in a single computer, or may be distributed and stored in multiple memories provided in each of multiple computers.
[0137] [Additional Notes] This disclosure includes the technologies described in the following supplementary notes. However, the present invention is not limited to the technologies described in the following supplementary notes, and various modifications are possible within the scope of the claims.
[0138] (Appendix A1) An information processing device comprising: an answer generation means that generates an answer to a query input by a user using a plurality of trained models generated by machine learning; and a contribution calculation means that calculates the contribution of each trained model to the answer generated by the answer generation means.
[0139] (Supplementary Note A2) The information processing device according to Supplementary Note A1, further comprising: a presentation control unit that presents to the user an aggregated answer that aggregates answers generated by the plurality of trained models.
[0140] (Supplementary Note A3) The information processing device according to Supplementary Note A2, wherein the answer generation means inputs answers generated by each of the plurality of trained models into a language model that has been machine-learned to aggregate the plurality of answers, thereby generating the aggregated answer.
[0141] (Appendix A4) The information processing device according to any one of Appendices A1 to A3, wherein the contribution calculation means calculates the contribution of each trained model based on the similarity between an aggregated answer in which each answer generated by the plurality of trained models is aggregated and each answer generated by each of the plurality of trained models.
[0142] (Appendix A5) An information processing device described in any of Appendices A1 to A3, wherein the contribution calculation means calculates the contribution using a contribution calculation model generated by machine learning the relationship between an answer generated by one trained model and an aggregated answer that aggregates multiple answers generated by multiple trained models including the one trained model, and the contribution of the one trained model to the aggregated answer.
[0143] (Appendix A6) The information processing device according to any one of Appendices A1 to A5, further comprising: a query decomposition means that decomposes the query to generate a plurality of decomposed queries, wherein the answer generation means causes one of the plurality of trained models to generate an answer to one of the plurality of decomposed queries, and causes another of the plurality of trained models to generate an answer to another of the plurality of decomposed queries by referring to the generated answer.
[0144] (Appendix A7) An information processing device according to any one of Appendices A1 to A6, comprising a verification means for, when a contradiction occurs between answers generated by each of the plurality of trained models which are language models, making each of the trained models which generated each answer in which a contradiction occurs provide a basis for the answer.
[0145] (Supplementary Note A8) The information processing device according to any one of Supplementary Notes A1 to A7, further comprising a reliability determination means configured to determine, based on at least any of the answers generated by each of the plurality of trained models, a reliability of an aggregated answer that aggregates the answers.
[0146] (Appendix A9) The information processing device according to any one of Appendices A1 to A8, wherein the answer generation means generates an aggregated answer by aggregating answers generated by each of the plurality of trained models based on at least one of a reliability or priority set for each of the plurality of trained models, attributes of the user, and an environment in which the answer is generated.
[0147] (Appendix B1) An assistance method in which at least one processor executes an answer generation process that generates an answer to a query entered by a user using multiple trained models generated by machine learning, and a contribution calculation process that calculates the contribution of each trained model to the answer generated in the answer generation process.
[0148] (Supplementary Note B2) The assistance method according to Supplementary Note B1, including a presentation control unit that presents to the user an aggregated answer that aggregates answers generated by the plurality of trained models.
[0149] (Appendix B3) The support method described in Appendix B2, wherein, in the answer generation process, the at least one processor inputs answers generated by each of the plurality of trained models into a language model that has been machine-learned to aggregate multiple answers, to generate the aggregated answer.
[0150] (Appendix B4) An assistance method described in any of Appendices B1 to B3, wherein in the contribution calculation process, the at least one processor calculates the contribution of each trained model based on the similarity between an aggregated answer that aggregates the answers generated by the multiple trained models and each answer generated by each of the multiple trained models.
[0151] (Appendix B5) An assistance method described in any of Appendices B1 to B3, wherein in the contribution calculation process, the at least one processor calculates the contribution using a contribution calculation model generated by machine learning the relationship between an answer generated by one trained model and an aggregated answer that aggregates multiple answers generated by multiple trained models including the one trained model, and the contribution of the one trained model to the aggregated answer.
[0152] (Appendix B6) The support method according to any one of Appendices B1 to B5, wherein the at least one processor includes a query decomposition process that decomposes the query to generate a plurality of decomposed queries, and in the answer generation process, the at least one processor causes one of the plurality of trained models to generate an answer to one of the plurality of decomposed queries, and also causes another of the plurality of trained models to generate an answer to another of the plurality of decomposed queries by referring to the generated answer.
[0153] (Appendix B7) An assistance method described in any of Appendices B1 to B6, including a verification process in which, when a contradiction occurs between answers generated by each of the multiple trained models, which are language models, the at least one processor causes each of the trained models that generated each answer in which a contradiction occurs to provide the basis for that answer.
[0154] (Appendix B8) The support method according to any one of Appendices B1 to B7, further comprising a reliability determination process in which the at least one processor determines the reliability of an aggregated answer that aggregates the answers generated by at least one of the plurality of trained models, based on the answers.
[0155] (Appendix B9) An assistance method described in any of Appendices B1 to B8, wherein in the answer generation process, the at least one processor aggregates answers generated by each of the plurality of trained models based on at least one of the reliability or priority set for each of the plurality of trained models, the attributes of the user, and the answer generation environment to generate an aggregated answer.
[0156] (Appendix C1) An assistance program that causes a computer to function as an answer generation means that generates an answer to a query entered by a user using multiple trained models generated by machine learning, and a contribution calculation means that calculates the contribution of each trained model to the answer generated by the answer generation means.
[0157] (Appendix C2) The assistance program according to Appendix C1, which causes the computer to function as a presentation control means that presents to the user an aggregated answer that aggregates each answer generated by the plurality of trained models.
[0158] (Appendix C3) The assistance program according to Appendix C2, wherein the answer generation means inputs answers generated by each of the plurality of trained models into a language model that has been machine-learned to aggregate the plurality of answers, thereby generating the aggregated answer.
[0159] (Appendix C4) An assistance program described in any of Appendices C1 to C3, wherein the contribution calculation means calculates the contribution of each trained model based on the similarity between an aggregated answer that aggregates each answer generated by the multiple trained models and each answer generated by each of the multiple trained models.
[0160] (Appendix C5) An assistance program described in any of Appendices C1 to C3, wherein the contribution calculation means calculates the contribution using a contribution calculation model generated by machine learning the relationship between an answer generated by one trained model and an aggregated answer that aggregates multiple answers generated by multiple trained models including the one trained model, and the contribution of the one trained model to the aggregated answer.
[0161] (Appendix C6) The assistance program according to any one of Appendices C1 to C5, wherein the computer is caused to function as a query decomposition process that decomposes the query to generate a plurality of decomposed queries, and the answer generation means causes one of the plurality of trained models to generate an answer to one of the plurality of decomposed queries, and also causes another of the plurality of trained models to generate an answer to another of the plurality of decomposed queries by referring to the generated answer.
[0162] (Appendix C7) An assistance program described in any of Appendices C1 to C6, which causes the computer to function as a verification means that, when a contradiction occurs between answers generated by each of the multiple trained models, which are language models, requires each trained model that generated each answer in which a contradiction occurs to provide the basis for that answer.
[0163] (Appendix C8) The assistance program according to any one of Appendices C1 to C7, which causes the computer to function as a reliability determination means that determines the reliability of an aggregated answer that aggregates answers generated by at least one of the answers generated by each of the plurality of trained models.
[0164] (Appendix C9) The assistance program described in any of Appendices C1 to C8, wherein the answer generation means aggregates answers generated by each of the plurality of trained models based on at least one of a reliability or priority set for each of the plurality of trained models, attributes of the user, and an environment in which the answer is generated to generate an aggregated answer.
[0165] (Appendix D1) An information processing device comprising at least one processor, the at least one processor executing an answer generation process that generates an answer to a query input by a user using a plurality of trained models generated by machine learning, and a contribution calculation process that calculates the contribution of each trained model to the answer generated by the answer generation process.
[0166] The information processing device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.
[0167] (Supplementary Note D2) The information processing device according to Supplementary Note D1, wherein the at least one processor executes a presentation control process to present to the user an aggregated answer that aggregates answers generated by the plurality of trained models.
[0168] (Appendix D3) The information processing device described in Appendix D2, wherein in the answer generation process, the at least one processor inputs answers generated by each of the plurality of trained models into a language model that has been machine-learned to aggregate multiple answers, to generate the aggregated answer.
[0169] (Appendix D4) An information processing device described in any of Appendices D1 to D3, wherein in the contribution calculation process, the at least one processor calculates the contribution of each trained model based on the similarity between an aggregated answer that aggregates each answer generated by the multiple trained models and each answer generated by each of the multiple trained models.
[0170] (Appendix D5) An information processing device described in any of Appendices D1 to D3, wherein in the contribution calculation process, the at least one processor calculates the contribution using a contribution calculation model generated by machine learning the relationship between an answer generated by one trained model and an aggregated answer that aggregates multiple answers generated by multiple trained models including the one trained model, and the contribution of the one trained model to the aggregated answer.
[0171] (Appendix D6) The information processing device of any of Appendices D1 to D5, wherein the at least one processor executes a query decomposition process that decomposes the query to generate a plurality of decomposed queries, and in the answer generation process, the at least one processor causes one of the plurality of trained models to generate an answer to one of the plurality of decomposed queries, and also causes another of the plurality of trained models to generate an answer to another of the plurality of decomposed queries by referring to the generated answer.
[0172] (Appendix D7) An information processing device described in any of Appendices D1 to D6, wherein, when a contradiction occurs between answers generated by each of the multiple trained models, which are language models, the at least one processor performs a verification process to have each of the trained models that generated each answer in which a contradiction occurs provide the basis for the answer.
[0173] (Appendix D8) The information processing device according to any one of appendices D1 to D7, wherein the at least one processor executes a reliability determination process to determine the reliability of an aggregated answer that aggregates answers generated by at least any of the answers generated by each of the plurality of trained models.
[0174] (Appendix D9) An information processing device described in any of Appendices D1 to D8, wherein in the answer generation process, the at least one processor aggregates answers generated by each of the plurality of trained models based on at least one of the reliability or priority set for each of the plurality of trained models, the attributes of the user, and the answer generation environment to generate an aggregated answer.
[0175] (Appendix E) A non-transient recording medium having recorded thereon an assistance program that causes a computer to execute an answer generation process that generates an answer to a query entered by a user using multiple trained models generated by machine learning, and a contribution calculation process that calculates the contribution of each trained model to the answer generated by the answer generation process.
[0176] REFERENCE SIGNS LIST 1 Information processing device 101 Answer generation unit (answer generation means) 102 Contribution calculation unit (contribution calculation means) 1A Information processing device 101A Answer generation unit (answer generation means) 102A Contribution calculation unit (contribution calculation means) 103A Query decomposition unit (query decomposition means) 105A Presentation control unit (presentation control means) 107A Verification unit (verification means) 109A Reliability determination unit (reliability determination means) 21A, 22A Trained model
Claims
1. An information processing device comprising: an answer generation means for generating an answer to a query input by a user using a plurality of trained models generated by machine learning; and a contribution calculation means for calculating the contribution of each trained model to the answer generated by the answer generation means.
2. The information processing device according to claim 1, further comprising a presentation control unit that presents to the user an aggregated answer that aggregates each answer generated by the multiple trained models.
3. The information processing device according to claim 2, wherein the answer generation means inputs answers generated by each of the plurality of trained models into a language model that has been machine-learned to aggregate a plurality of answers, thereby generating the aggregated answer.
4. An information processing device as described in claim 1 or 2, wherein the contribution calculation means calculates the contribution of each trained model based on the similarity between an aggregated answer in which each answer generated by the multiple trained models is aggregated and each answer generated by each of the multiple trained models.
5. An information processing device as described in claim 1 or 2, wherein the contribution calculation means calculates the contribution using a contribution calculation model generated by machine learning the relationship between an answer generated by one trained model and an aggregated answer that aggregates multiple answers generated by multiple trained models including the one trained model, and the contribution of the one trained model to the aggregated answer.
6. An information processing device as described in claim 1 or 2, further comprising a query decomposition means for decomposing the query to generate a plurality of decomposed queries, wherein the answer generation means causes one of the plurality of trained models to generate an answer to one of the plurality of decomposed queries, and causes one of the plurality of trained models to generate an answer to another of the plurality of decomposed queries by referring to the generated answer.
7. An information processing device as described in claim 1 or 2, further comprising a verification means for, when a contradiction occurs between answers generated by each of the multiple trained models which are language models, having each of the trained models which generated each answer with a contradiction provide a reason for that answer.
8. An information processing device as described in claim 1 or 2, further comprising a reliability determination means for determining the reliability of an aggregated answer that aggregates the answers based on at least one of the answers generated by each of the multiple trained models.
9. An information processing device as described in claim 1 or 2, wherein the answer generation means aggregates answers generated by each of the plurality of trained models based on at least one of a reliability or priority set for each of the plurality of trained models, attributes of the user, and an environment in which the answer is generated, to generate an aggregated answer.
10. An assistance method in which at least one processor executes: an answer generation process that generates an answer to a query input by a user using multiple trained models generated by machine learning; and a contribution calculation process that calculates the contribution of each trained model to the answer generated in the answer generation process.
11. An assistance program that causes a computer to function as an answer generation means that generates an answer to a query input by a user using a plurality of trained models generated by machine learning, and a contribution calculation means that calculates the contribution of each trained model to the answer generated by the answer generation means.
Citation Information
Patent Citations
Method, device and equipment for constructing drug sensitivity prediction model sample
CN115985413A
Medical imaging system, medical image processing apparatus, and program
JP2013102851A
A system that collects and identifies skin conditions using images and expert knowledge
JP2022549433A
Cited By
Method for constructing intelligent question-answering service and related device
TWI940259B